Hajime Nagahara

dblp:78/5843 · also Hajime Magahara · DBLP profile ↗
← Back
104ranked-venue papers
12as first author
39since 2021 · last 2026
0000-0003-1579-8767ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 69 · 8 first-author · 23 since 2021Artificial intelligence and machine learning · 65 · 6 first-author · 23 since 2021Systems, architecture and hardware · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Japanese Dataset for Aspect-based Sentiment Polarity Classification and Emotion Intensity Estimation
Kentaro Hanafusa, Kota Manabe, Yuki Maeda, Daisuke Maekawa, Tomoyuki Kajiwara, Hideaki Hayashi, Yuta Nakashima, Hajime Nagahara
LREC8
2026 SMILE: Sensor Data-driven Detection and Counterfactual Pattern Analysis for Loneliness in Older Adults
abstract
Prolonged loneliness in older adults increases the risk of dementia, cardiovascular diseases, and premature death. However, timely detection and behavioral pattern analysis are challenging due to the slow progression of loneliness and diverse individual lifestyles. To address these barriers, we propose SMILE, an end-to-end framework for sensor-driven detection and counterfactual pattern analysis of loneliness. Our approach uses non-intrusive sensors to passively observe behavior and build a personalized profile over time. When SMILE detects meaningful changes in activity patterns, it employs generative models to produce counterfactual sensor profiles—synthetic representations of behavior predicted to result in lower loneliness scores. Suggested pattern changes are identified as the difference between observed and counterfactual profiles, highlighting the behaviors most associated with loneliness. We conducted a 6-month feasibility study with 18 participants from the United States and 15 from Japan to develop and evaluate SMILE. Our framework integrates a Time-series Transformer (TST) model for loneliness detection, which achieved 96.25% accuracy, and a diffusion model for generating behavioral explanations, which outperformed a baseline by up to 76.69%. These findings highlight the potential of SMILE to support sensor-driven, personalized loneliness detection and behavioral explanation in aging populations.
Xiayan Ji, Ahhyun Yuh, Viktor Erdélyi, Teruhiro Mizumoto, Hyonyoung Choi, Sean Lee Harrison, Emma Cho, Takashi Suehiro, Takeshi Nakagawa, Yutian Cheng, Yasuyuki Gondo, Hajime Nagahara, Teruo Higashino, George Demiris, Oleg Sokolsky, Insup Lee 0001
ACM Trans. Comput. Heal.12
2025 NeISF++: Neural Incident Stokes Field for Polarized Inverse Rendering of Conductors and Dielectrics
abstract
Recent inverse rendering methods have improved shape, material, and illumination reconstruction using polarization cues. However, they only support dielectrics, ignoring conductors, which are common in everyday life. Since conductors and dielectrics have different reflection properties, using previous dielectrics-based methods will lead to obvious errors. In addition, conductors are glossy, which may cause strong specular reflection and is hard to reconstruct. To solve the above issues, we propose NeISF++, an inverse rendering pipeline that supports conductors and dielectrics. The key ingredient for our proposal is a general pBRDF that describes both conductors and dielectrics. As for the strong specular reflection problem, we propose a novel geometry initialization method using DoLP images. This physical cue is invariant to intensities and thus robust to strong specular reflections. Experimental results on our synthetic and real datasets show that our method surpasses the existing polarized inverse rendering methods for geometry and material decomposition as well as downstream tasks like relighting.
Taishi Ono, Takeshi Uemori, Sho Nitta, Hajime Mihara, Alexander Gatto, Hajime Nagahara, Yusuke Moriuchi
CVPR7
2025 Gaussian-Based Instance-Adaptive Intensity Modeling for Point-Supervised Facial Expression Spotting
abstract
Point-supervised facial expression spotting (P-FES) aims to localize facial expression instances in untrimmed videos, requiring only a single timestamp label for each instance during training. To address label sparsity, hard pseudo-labeling is often employed to propagate point labels to unlabeled frames; however, this approach can lead to confusion when distinguishing between neutral and expression frames with various intensities, which can negatively impact model performance. In this paper, we propose a two-branch framework for P-FES that incorporates a Gaussian-based instance-adaptive Intensity Modeling (GIM) module for soft pseudo-labeling. GIM models the expression intensity distribution for each instance. Specifically, we detect the pseudo-apex frame around each point label, estimate the duration, and construct a Gaussian distribution for each expression instance. We then assign soft pseudo-labels to pseudo-expression frames as intensity values based on the Gaussian distribution. Additionally, we introduce an Intensity-Aware Contrastive (IAC) loss to enhance discriminative feature learning and suppress neutral noise by contrasting neutral frames with expression frames of various intensities. Extensive experiments on the SAMM-LV and CAS(ME)$^2$ datasets demonstrate the effectiveness of our proposed framework. Code is available at https://github.com/KinopioIsAllIn/GIM.
Yicheng Deng, Hideaki Hayashi, Hajime Nagahara
ICLR3
2025 Multi-task Learning of Classification and Generation for Set-structured Data
abstract
In this study, we propose a multi-task learning model of classification and generation for set-structured data. The proposed model learns data generation and classification in a single neural network by integrating a classification layer into a variational autoencoder while maintaining permutation invariance and equivariance nature, which are charac-teristics of set-structured data. The proposed model allows for semi-supervised learning in set-structured data classifi-cation and can also be applied to confidence calibration using the input data distribution estimated by the generative model. In the experiments, we evaluated the performance of the proposed model in a semi-supervised classification task on set-structured datasets and compared it with a baseline model consisting only of a classifier. The results demon-strated that simultaneous learning of the classification and generation effectively improves the classification accuracy and confidence reliability for set-structured data, even with a limited number of labeled data.
Fumioki Sato, Hideaki Hayashi, Hajime Nagahara
WACV3
2025 Simultaneous acquisition of geometry and material for translucent objects
abstract
Reconstructing the geometry and material properties of translucent objects from images is a challenging problem due to the complex light propagation of translucent media and the inherent ambiguity of inverse rendering. Therefore, previous works often make the assumption that the objects are opaque or use a simplified model to describe translucent objects, which significantly affects the reconstruction quality and limits the downstream tasks such as relighting or material editing. We present a novel framework that tackles this challenge through a combination of physically grounded and data-driven strategies. At the core of our approach is a hybrid rendering supervision scheme that fuses a differentiable physical renderer with a learned neural renderer to guide reconstruction. To further enhance supervision, we introduce an augmented loss tailored to the neural renderer. Our system takes as input a flash/no-flash image pair, enabling it to disambiguate complex light propagation that happens inside translucent objects. We train our model on a large-scale synthetic dataset of 117 K scenes and evaluate across both synthetic benchmarks and real-world captures. To mitigate the domain gap between synthetic and real data, we contribute a new real-world dataset with ground-truth surface normals and fine-tune our model accordingly. Extensive experiments validate the robustness and accuracy of our method across diverse scenarios. • We tackle the inverse rendering of translucent objects by jointly estimating shape, spatially-varying surface reflectance, SSS parameters, and illumination from a single-view flash/no-flash pair. • We propose a hybrid rendering architecture combining physically accurate surface modeling with a learned neural renderer for SSS. • We introduce a novel augmented loss strategy to improve the supervision of SSS effects. • We build a large-scale synthetic dataset and a real-world benchmark to train and validate our model. • We demonstrate state-of-the-art performance on both synthetic and real translucent objects, supported by fine-tuning to address domain shift.
Trung Ngo Thanh, Hajime Nagahara
Image Vis. Comput.3
2025 Built year prediction of buddha face with heterogeneous label modeled as probabilistic distribution
abstract
Abstract Analysis of cultural heritages, particularly their construction years, provides new insights into human history. However, due to natural disasters, wars, material deterioration, and human errors, records documenting the construction years of many artifacts have often been lost. Historians and experts can estimate construction years within specific ranges using chemical-based analysis technologies or extensive historical research. Given the vast number of collected artifacts, applying these conventional methods to every artifact is impractical. To address this challenge, we developed a deep neural network model designed for Buddha statues to estimate an artifact’s construction year from its image. One major challenge in this task is the heterogeneity of the labels: the training samples include both precise construction years and possible ranges (e.g., a dynasty or a century) estimated by historians. To unify these heterogeneous labels during training, we represent them as probabilistic distributions. In our previous work Qian et al. (2021), we assumed that the ambiguity in heterogeneous construction year labels followed a Gaussian distribution, assigning the highest likelihood to the midpoint of the designated time range. However, this assumption does not always hold. In this paper, we propose representing heterogeneous construction year labels as a uniform distribution, assigning equal probability to all points within the designated time range. Based on this label representation, we designed a semi-supervised learning loss function to leverage both labeled and unlabeled samples during training. Our experimental results demonstrate that our method achieves a mean absolute error of 34.3 years on a test set consisting of Buddha statues constructed between 400 and 1403. These results are further analyzed in two ways. First, we compared our model’s performance to the image quality BRISQUE score, revealing a correlation between higher image quality and lower prediction error rates. Second, we validated our predictions with experts, assessing the level of agreement with our model, the challenges in determining construction years, and identifying features of interest in the artifacts.
Yiming Qian, Cheikh Brahim El Vaigh, Yuta Nakashima, Benjamin Renoust, Hajime Nagahara, Yutaka Fujioka
Multim. Tools Appl.5
2025 GNNBoost: boosting artwork classification with graph embeddings
abstract
Abstract The use of AI systems for managing large-scale cultural heritage artifacts has become possible due to the rise of digitization. To classify such content, machine learning is typically used, where contextual information is important to structure the data. One way to capture context is through a knowledge graph. In this study, we propose a newd graph neural networks, we can improve artwork classification by utilizing the relationships between entities of the knowledge graph. Our experiments demonstrate that this approach achieves state-of-the-art results on multiple classification tasks across three datasets (SemArt paintings, Buddha statues, and Ukiyo-e woodblock prints). Moreover, our approach is effective in dealing with unbalanced data and we explore the use of both graph attention mechanisms and focal loss functions.
Cheikh Brahim El Vaigh, Noa Garcia, Benjamin Renoust, Chenhui Chu, Yuta Nakashima, Yiming Qian, Hajime Nagahara
Multim. Tools Appl.7
2024 Time-Efficient Light-Field Acquisition Using Coded Aperture and Events
abstract
We propose a computational imaging method for time-efficient light-field acquisition that combines a coded aperture with an event-based camera. Differentfrom the conventional coded-aperture imaging method, our method applies a sequence of coding patterns during a single exposure for an image frame. The parallax information, which is related to the differences in coding patterns, is recorded as events. The image frame and events, all of which are measured in a single exposure, are jointly used to computationally reconstruct a light field. We also designed an algorithm pipeline for our method that is end-to-end trainable on the basis of deep optics and compatible with real camera hardware. We experimentally showed that our method can achieve more accurate reconstruction than several other imaging methods with a single exposure. We also developed a hardware prototype with the potential to complete the measurement on the camera within 22 msec and demonstrated that light fields from real 3-D scenes can be obtained with convincing visual quality. Our software and supplementary video are available from our project website1:
Shuji Habuchi, Keita Takahashi 0001, Chihiro Tsutake, Toshiaki Fujii, Hajime Nagahara
CVPR5
2024 NeISF: Neural Incident Stokes Field for Geometry and Material Estimation
abstract
Multi-view inverse rendering is the problem of estimating the scene parameters such as shapes, materials, or il-luminations from a sequence of images captured under dif-ferent viewpoints. Many approaches, however, assume single light bounce and thus fail to recover challenging sce-narios like inter-reflections. On the other hand, simply ex-tending those methods to consider multi-bounced light re-quires more assumptions to alleviate the ambiguity. To address this problem, we propose Neural Incident Stokes Fields (NeISF), a multi-view inverse rendering framework that reduces ambiguities using polarization cues. The pri-mary motivation for using polarization cues is that it is the accumulation of multi-bounced light, providing rich infor-mation about geometry and material. Based on this knowl-edge, the proposed incident Stokes field efficiently models the accumulated polarization effect with the aid of an orig-inal physically-based differentiable polarimetric renderer. Lastly, experimental results show that our method outper-forms the existing works in synthetic and real scenarios.
Taishi Ono, Takeshi Uemori, Hajime Mihara, Alexander Gatto, Hajime Nagahara, Yusuke Moriuchi
CVPR6
2024 Deep Polarization Cues for Single-Shot Shape and Subsurface Scattering Estimation
Trung Ngo Thanh, Hajime Nagahara
ECCV (67)3
2024 Multi-Scale Spatio-Temporal Graph Convolutional Network for Facial Expression Spotting
abstract
Facial expression spotting is a significant but challenging task in facial expression analysis. The accuracy of expression spotting is affected not only by irrelevant facial movements but also by the difficulty of perceiving subtle motions in micro-expressions. In this paper, we propose a Multi-Scale Spatio-Temporal Graph Convolutional Network (SpoT-GCN) for facial expression spotting. To extract more robust motion features, we track both short- and long-term motion of facial muscles in compact sliding windows whose window length adapts to the temporal receptive field of the network. This strategy, termed the receptive field adaptive sliding window strategy, effectively magnifies the motion features while alleviating the problem of severe head movement. The subtle motion features are then converted to a facial graph representation, whose spatio-temporal graph patterns are learned by a graph convolutional network. This network learns both local and global features from multiple scales of facial graph structures using our proposed facial local graph pooling (FLGP). Furthermore, we introduce supervised contrastive learning to enhance the discriminative capability of our model for difficult-to-classify frames. The experimental results on the SAMM-LV and CAS(ME)2datasets demonstrate that our method achieves state-of-the-art performance, particularly in micro-expression spotting. Ablation studies further verify the effectiveness of our proposed modules.
Yicheng Deng, Hideaki Hayashi, Hajime Nagahara
FG3
2024 CALICO: Confident Active Learning with Integrated Calibration
Lorenzo S. Querol, Hajime Nagahara, Hideaki Hayashi
ICANN (1)2
2024 Is Internal State Feedback in an E-Learning Environment Acceptable to People?
abstract
In on-demand e-learning environments, the lack of direct intervention can lead to a decline in learners' engagement. To address this issue, systems that estimate the learners' attitudes and provide feedback have been proposed. However, the acceptability of such systems has not been sufficiently researched. In this study, we investigated the acceptability by people to an e-learning system with internal state feedback, for future personalized learning support. To this end, we developed a system that estimates and visualizes the learner's internal state in real-time. The system was exhibited in a public space for free use, and users' impressions were analyzed. To estimate the learners' internal state, we developed a machine-learning model that recognizes learners' alertness from facial videos. The system was deployed in an exhibition space, and 131 responses were collected. These responses were coded and analyzed using a co-occurrence network. The result indicated that learners tend to dislike the system due to feelings of being observed by supervisors. In contrast, instructors expressed favorable options toward the introduction of the system.
Atsushi Ashida, Ryosuke Kawamura, Shizuka Shirai, Noriko Takemura, Mehrasa Alizadeh, Hideaki Hayashi, Hajime Nagahara
ICCE7
2024 LoHoSC: Low Order High Order Style Consistency for Syn-to-Real Domain Generalized Semantic Segmentation
Sudhakar Kumawat, Hajime Nagahara
ICPR (3)2
2024 Deep Hardware Modality Fusion for Image Segmentation
Sudhakar Kumawat, Hajime Nagahara
ICPR (7)3
2024 Deep Volume Reconstruction from Multi-focus Microscopic Images
Caio Azevedo, Sanchayan Santra, Sudhakar Kumawat, Hajime Nagahara, Ken'ichi Morooka
MICCAI (4)4
2024 DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models
abstract
Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GPT-4 excel in medical question answering but may face challenges in the lack of interpretability when handling complex tasks in real clinical settings. We thus introduce the diagnostic reasoning dataset for clinical notes (DiReCT), aiming at evaluating the reasoning ability and interpretability of LLMs compared to human doctors. It contains 511 clinical notes, each meticulously annotated by physicians, detailing the diagnostic reasoning process from observations in a clinical note to the final diagnosis. Additionally, a diagnostic knowledge graph is provided to offer essential knowledge for reasoning, which may not be covered in the training data of existing LLMs. Evaluations of leading LLMs on DiReCT bring out a significant gap between their reasoning ability and that of human doctors, highlighting the critical need for models that can reason effectively in real-world clinical scenarios.
Bowen Wang 0002, Jiuyang Chang, Yiming Qian, Guoxin Chen, Zhouqiang Jiang, Yuta Nakashima, Hajime Nagahara
NeurIPS9
2024 MIDAS: Mixing Ambiguous Data with Soft Labels for Dynamic Facial Expression Recognition
abstract
Dynamic facial expression recognition (DFER) is an important task in the field of computer vision. To apply automatic DFER in practice, it is necessary to accurately recognize ambiguous facial expressions, which often appear in data in the wild. In this paper, we propose MIDAS, a data augmentation method for DFER, which augments ambiguous facial expression data with soft labels consisting of probabilities for multiple emotion classes. In MIDAS, the training data are augmented by convexly combining pairs of video frames and their corresponding emotion class labels, which can also be regarded as an extension of mixup to soft-labeled video data. This simple extension is remarkably effective in DFER with ambiguous facial expression data. To evaluate MIDAS, we conducted experiments on the DFEW dataset. The results demonstrate that the model trained on the data augmented by MIDAS outperforms the existing state-of-the-art method trained on the original dataset.
Ryosuke Kawamura, Hideaki Hayashi, Noriko Takemura, Hajime Nagahara
WACV4
2024 Revisiting Pixel-Level Contrastive Pre-Training on Scene Images
abstract
Contrastive image representation learning through instance discrimination has shown impressive transfer performance. Recent strategies have focused on pushing the limit of their transfer performance for dense prediction tasks, particularly when conducting pre-training on scene images with complex structures. Initial approaches employ pixel-level contrastive pre-training to optimize dense spatial features, while subsequent methods utilize region-mining algorithms to capture holistic regional semantics and address the issue of semantically inconsistent scene image crops. In this paper, we revisit pixel-level contrastive pre-training on scene images. Contrary to the assumption that pixel-level learning falls short in achieving these objectives, we demonstrate its under-explored potentials: (1) it can effectively learn holistic regional semantics more simply compared to region-level methods, and (2) it intrinsically provides tools to mitigate the impact of semantically inconsistent views involved with scene-level training images. We propose PixCon, a pixel-level contrastive learning framework, and explore two variants with different positive matching strategies to investigate the potential of pixel-level learning. Additionally, when PixCon incorporates a novel semantic reweighting approach tailored for scene image pre-training, it outperforms or matches the performance of previous region-level methods in object detection and semantic segmentation tasks across multiple benchmarks.1
Zongshang Pang, Yuta Nakashima, Mayu Otani, Hajime Nagahara
WACV4
2024 Instruct Me More! Random Prompting for Visual In-Context Learning
abstract
Large-scale models trained on extensive datasets, have emerged as the preferred approach due to their high generalizability across various tasks. In-context learning (ICL), a popular strategy in natural language processing, uses such models for different tasks by providing instructive prompts but without updating model parameters. This idea is now being explored in computer vision, where an input-output image pair (called an in-context pair) is supplied to the model with a query image as a prompt to exemplify the desired output. The efficacy of visual ICL often depends on the quality of the prompts. We thus introduce a method coined Instruct Me More (InMeMo), which augments in-context pairs with a learnable perturbation (prompt), to explore its potential. Our experiments on mainstream tasks reveal that InMeMo surpasses the current state-of-the-art performance. Specifically, compared to the baseline without learnable prompt, InMeMo boosts mIoU scores by 7.35 and 15.13 for foreground segmentation and single object detection tasks, respectively. Our findings suggest that InMeMo offers a versatile and efficient way to enhance the performance of visual ICL with lightweight training. Code is available at https://github.com/Jackieam/InMeMo.
Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Hajime Nagahara
WACV5
2023 Inverse Rendering of Translucent Objects using Physical and Neural Renderers
abstract
In this work, we propose an inverse rendering model that estimates 3D shape, spatially-varying reflectance, homogeneous subsurface scattering parameters, and an environment illumination jointly from only a pair of captured images of a translucent object. In order to solve the ambiguity problem of inverse rendering, we use a physically-based renderer and a neural renderer for scene reconstruction and material editing. Because two renderers are differentiable, we can compute a reconstruction loss to assist parameter estimation. To enhance the supervision of the proposed neural renderer, we also propose an augmented loss. In addition, we use a flash and no-flash image pair as the input. To supervise the training, we constructed a large-scale synthetic dataset of translucent objects, which consists of 117K scenes. Qualitative and quantitative results on both synthetic and real-world datasets demonstrated the effectiveness of the proposed model. Code and Data are available at https://github.com/ligoudaner377/homo_translucent
Trung Ngo Thanh, Hajime Nagahara
CVPR3
2023 Learning Bottleneck Concepts in Image Classification
abstract
Interpreting and explaining the behavior of deep neural networks is critical for many tasks. Explainable AI provides a way to address this challenge, mostly by providing per-pixel relevance to the decision. Yet, interpreting such explanations may require expert knowledge. Some recent attempts toward interpretability adopt a concept-based framework, giving a higher-level relationship between some concepts and model decisions. This paper proposes Bottleneck Concept Learner (BotCL), which represents an image solely by the presence/absence of concepts learned through training over the target task without explicit supervision over the concepts. It uses self-supervision and tailored regularizers so that learned concepts can be human-understandable. Using some image classification tasks as our testbed, we demonstrate BotCL's potential to rebuild neural networks for better interpretability11Code is avaliable at https://github.com/wbw520/BotCL and a simple demo is available at https://botcl.liangzhili.com/.
Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Hajime Nagahara
CVPR4
2023 Automated Detection of Students' Gaze Interactions in Collaborative Learning Videos: A Novel Approach
Qi Zhou 0011, Amartya Bhattacharya, Wannapon Suraworachet, Hajime Nagahara, Mutlu Cukurova
EC-TEL4
2023 A Compact BRDF Scanner with Multi-conjugate Optics
abstract
Reflectance, represented as Bidirectional reflectance distribution functions (BRDFs), is an important scene property that we wish to obtain from the real world together with the scene’s 3D shape. BRDF acquisition, however, remains a difficult task because it is time-consuming and requires huge and costly devices. To make the BRDF acquisition easier and accessible to everyone, we develop a compact projector-camera-based BRDF scanner with coaxial multi-conjugate optics, where the sensor/source is conjugate with the Fourier transform plane of the objective lens, and the camera/projector pupil is conjugate with the sample plane. It consists of only off-the-shelf components, and the carefully designed conjugate optics make the entire system compact. BRDFs of outdoor objects are scanned to show the validity of our device.
Kensuke Uchida, Hajime Nagahara, Yasuyuki Matsushita
ICCP2
2023 Contrastive Losses Are Natural Criteria for Unsupervised Video Summarization
abstract
Video summarization aims to select the most informative subset of frames in a video to facilitate efficient video browsing. Unsupervised methods usually rely on heuristic training objectives such as diversity and representativeness. However, such methods need to bootstrap the online-generated summaries to compute the objectives for importance score regression. We consider such a pipeline inefficient and seek to directly quantify the frame-level importance with the help of contrastive losses in the representation learning literature. Leveraging the contrastive losses, we propose three metrics featuring a desirable key frame: local dissimilarity, global consistency, and uniqueness. With features pre-trained on the image classification task, the metrics can already yield high-quality importance scores, demonstrating competitive or better performance than past heavily-trained methods. We show that by refining the pre-trained features with a lightweight contrastively learned projection module, the frame-level importance scores can be further improved, and the model can also leverage a large number of random videos and generalize to test videos with decent performance.
Zongshang Pang, Yuta Nakashima, Mayu Otani, Hajime Nagahara
WACV4
2023 Cross-language font style transfer
abstract
Abstract In this paper, we propose a cross-language font style transfer system that can synthesize a new font by observing only a few samples from another language. Automatic font synthesis is a challenging task and has attracted much research interest. Most previous works addressed this problem by transferring the style of the given subset to the content of unseen ones. Nevertheless, they only focused on the font style transfer in the same language. In many cases, we need to learn font style from one language and then apply it to other languages. Existing methods make this difficult to accomplish because of the abstraction of style and language differences. To address this problem, we specifically designed the network into a multi-level attention form to capture both local and global features of the font style. To validate the generative ability of our model, we constructed an experimental font dataset of 847 fonts, each containing English and Chinese characters with the same style. Results show that our model generates 80.3% of users’ preferred images compared with state-of-the-art models.
Yuta Taniguchi, Min Lu 0003, Shin'ichi Konomi, Hajime Nagahara
Appl. Intell.5
2023 Match them up: visually explainable few-shot image classification
abstract
Abstract Few-shot learning (FSL) approaches, mostly neural network-based, assume that pre-trained knowledge can be obtained from base (seen) classes and transferred to novel (unseen) classes. However, the black-box nature of neural networks makes it difficult to understand what is actually transferred, which may hamper FSL application in some risk-sensitive areas. In this paper, we reveal a new way to perform FSL for image classification, using a visual representation from the backbone model and patterns generated by a self-attention based explainable module. The representation weighted by patterns only includes a minimum number of distinguishable features and the visualized patterns can serve as an informative hint on the transferred knowledge. On three mainstream datasets, experimental results prove that the proposed method can enable satisfying explainability and achieve high classification results. Code is available at https://github.com/wbw520/MTUNet .
Bowen Wang 0002, Liangzhi Li 0001, Manisha Verma, Yuta Nakashima, Ryo Kawasaki, Hajime Nagahara
Appl. Intell.6
2023 Action Recognition From a Single Coded Image
abstract
The unprecedented success of deep convolutional neural networks (CNN) on the task of video-based human action recognition assumes the availability of good resolution videos and resources to develop and deploy complex models. Unfortunately, certain budgetary and environmental constraints on the camera system and the recognition model may not be able to accommodate these assumptions and require reducing their complexity. To alleviate these issues, we introduce a deep sensing solution to directly recognize human actions from coded exposure images. Our deep sensing solution consists of a binary CNN-based encoder network that emulates the capturing of a coded exposure image of a dynamic scene using a coded exposure camera, followed by a 2D CNN for recognizing human action in the captured coded exposure image. Furthermore, we propose a novel knowledge distillation framework to jointly train the encoder and the action recognition model and show that the proposed training approach improves the action recognition accuracy by an absolute margin of 6.2%, 2.9%, and 7.9% on Something$^{2}$-v2, Kinetics-400, and UCF-101 datasets, respectively, in comparison to our previous approach. Finally, we built a prototype coded exposure camera using LCoS to validate the feasibility of our deep sensing solution. Our evaluation of the prototype camera show results that are consistent with the simulation results.
Sudhakar Kumawat, Tadashi Okawara, Michitaka Yoshida, Hajime Nagahara, Yasushi Yagi
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Multi-label Disengagement and Behavior Prediction in Online Learning
Manisha Verma, Yuta Nakashima, Noriko Takemura, Hajime Nagahara
AIED (1)4
2022 Acquiring a Dynamic Light Field through a Single-Shot Coded Image
abstract
We propose a method for compressively acquiring a dynamic light field (a 5-D volume) through a single-shot coded image (a 2-D measurement). We designed an imaging model that synchronously applies aperture coding and pixel-wise exposure coding within a single exposure time. This coding scheme enables us to effectively embed the original information into a single observed image. The observed image is then fed to a convolutional neural network (CNN) for light-field reconstruction, which is jointly trained with the camera-side coding patterns. We also developed a hardware prototype to capture a real 3-D scene moving over time. We succeeded in acquiring a dynamic light field with 5x5 viewpoints over 4 temporal sub-frames (100 views in total)from a single observed image. Repeating capture and reconstruction processes over time, we can acquire a dynamic light field at 4x the frame rate of the camera. To our knowledge, our method is the first to achieve a finer temporal resolution than the camera itself in compressive light-field acquisition. Our software is available from our project webpage.11https://www.fujii.nuee.nagoya-u.ac.jp/Research/CompCam2
Ryoya Mizuno, Keita Takahashi 0001, Michitaka Yoshida, Chihiro Tsutake, Toshiaki Fujii, Hajime Nagahara
CVPR6
2022 Privacy-Preserving Action Recognition via Motion Difference Quantization
Sudhakar Kumawat, Hajime Nagahara
ECCV (13)2
2022 Blockwise Feature-Based Registration of Deformable Medical Images
Su Wai Tun, Takashi Komuro, Hajime Nagahara
ICIC (1)3
2022 A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog Domain
abstract
We annotate 35,000 SNS posts with both the writer’s subjective sentiment polarity labels and the reader’s objective ones to construct a Japanese sentiment analysis dataset. Our dataset includes intensity labels (none, weak, medium, and strong) for each of the eight basic emotions by Plutchik (joy, sadness, anticipation, surprise, anger, fear, disgust, and trust) as well as sentiment polarity labels (strong positive, positive, neutral, negative, and strong negative). Previous studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently. In other words, there are few corpora that are annotated with both basic emotions and sentiment polarity. Our dataset is the first large-scale corpus to annotate both of these emotion labels, and from both the writer’s and reader’s perspectives. In this paper, we analyze the relationship between basic emotion intensity and sentiment polarity on our dataset and report the results of benchmarking sentiment polarity classification.
Haruya Suzuki, Yuto Miyauchi, Kazuki Akiyama, Tomoyuki Kajiwara, Takashi Ninomiya, Noriko Takemura, Yuta Nakashima, Hajime Nagahara
LREC8
2022 Surface Normals and Light Directions From Shading and Polarization
abstract
We introduce a method of recovering the shape of a smooth dielectric object using diffuse polarization images taken with different directional light sources. We present two constraints on shading and polarization and use both in a single optimization scheme. This integration is motivated by photometric stereo and polarization-based methods having complementary abilities. Polarization gives strong cues for the surface orientation and refractive index, which are independent of the light direction. However, employing polarization leads to ambiguities in selecting between two ambiguous choices of the surface orientation, in the relationship between the refractive index and zenith angle (observing angle). Moreover, polarization-based methods for surface points with small zenith angles perform poorly owing to the weak polarization. In contrast, the photometric stereo method with multiple light sources disambiguates the surface normals and gives a strong relationship between surface normals and light directions. However, the method has limited performance for large zenith angles and refractive index estimation and faces strong ambiguity when light directions are unknown. Taking the advantages of these methods, our proposed method recovers surface normals for small and large zenith angles, light directions, and refractive indexes of the object. The proposed method is positively evaluated in simulations and real-world experiments.
Trung Ngo Thanh, Hajime Nagahara, Rin-Ichiro Taniguchi
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 SCOUTER: Slot Attention-based Classifier for Explainable Image Recognition
abstract
Explainable artificial intelligence has been gaining attention in the past few years. However, most existing methods are based on gradients or intermediate features, which are not directly involved in the decision-making process of the classifier. In this paper, we propose a slot attention-based classifier called SCOUTER for transparent yet accurate classification. Two major differences from other attention-based methods include: (a) SCOUTER’s explanation is involved in the final confidence for each category, offering more intuitive interpretation, and (b) all the categories have their corresponding positive or negative explanation, which tells "why the image is of a certain category" or "why the image is not of a certain category." We design a new loss tailored for SCOUTER that controls the model’s behavior to switch between positive and negative explanations, as well as the size of explanatory regions. Experimental results show that SCOUTER can give better visual explanations in terms of various metrics while keeping good accuracy on small and medium-sized datasets. Code is available1.
Liangzhi Li 0001, Bowen Wang 0002, Manisha Verma, Yuta Nakashima, Ryo Kawasaki, Hajime Nagahara
ICCV6
2021 Learners' Efficiency Prediction Using Facial Behavior Analysis
abstract
In the e-learning context, how much the learner is concentrated and engaged, or the learners’ efficiency, is essential for providing adaptive and flexible materials, timely suggestions, etc., which can lead to efficient learning. In this work, we explore to predict learners’ efficiency with a realistic configuration, in which we use a webcam or a laptop PC’s built-in camera. Specifically, we first provide a feasible definition of the learners’ efficiency, and based on this definition, we predict one’s efficiency from facial behavior. We predict the learners’ efficiency using various convolutional neural networks. Results are discussed using different evaluation metrics.
Manisha Verma, Yuta Nakashima, Hirokazu Kobori, Ryota Takaoka, Noriko Takemura, Tsukasa Kimura, Hajime Nagahara, Masayuki Numao, Kazumitsu Shinohara
ICIP7
2021 GCNBoost: Artwork Classification by Label Propagation through a Knowledge Graph
abstract
The rise of digitization of cultural documents offers large-scale contents, opening the road for development of AI systems in order to preserve, search, and deliver cultural heritage. To organize such cultural content also means to classify them, a task that is very familiar to modern computer science. Contextual information is often the key to structure such real world data, and we propose to use it in form of a knowledge graph. Such a knowledge graph, combined with content analysis, enhances the notion of proximity between artworks so it improves the performances in classification tasks. In this paper, we propose a novel use of a knowledge graph, that is constructed on annotated data and pseudo-labeled data. With label propagation, we boost artwork classification by training a model using a graph convolutional network, relying on the relationships between entities of the knowledge graph. Following a transductive learning framework, our experiments show that relying on a knowledge graph modeling the relations between labeled data and unlabeled data allows to achieve state-of-the-art results on multiple classification tasks on a dataset of paintings, and on a dataset of Buddha statues. Additionally, we show state-of-the-art results for the difficult case of dealing with unbalanced data, with the limitation of disregarding classes with extremely low degrees in the knowledge graph.
Cheikh Brahim El Vaigh, Noa Garcia, Benjamin Renoust, Chenhui Chu, Yuta Nakashima, Hajime Nagahara
ICMR6
2021 WRIME: A New Dataset for Emotional Intensity Estimation with Subjective and Objective Annotations
abstract
Tomoyuki Kajiwara, Chenhui Chu, Noriko Takemura, Yuta Nakashima, Hajime Nagahara. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tomoyuki Kajiwara, Chenhui Chu, Noriko Takemura, Yuta Nakashima, Hajime Nagahara
NAACL-HLT5
2020 Acquiring Dynamic Light Fields Through Coded Aperture Camera
Kohei Sakai, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara
ECCV (19)4
2020 YOLO in the Dark - Domain Adaptation Method for Merging Multiple Models
Yukihiro Sasagawa, Hajime Nagahara
ECCV (21)2
2020 Action Recognition from a Single Coded Image
abstract
Cameras are prevalent in society at the present time, for example, surveillance cameras, and smartphones equipped with cameras and smart speakers. There is an increasing demand to analyze human actions from these cameras to detect unusual behavior or within a man-machine interface for Internet of Things (IoT) devices. For a camera, there is a trade-off between spatial resolution and frame rate. A feasible approach to overcome this trade-off is compressive video sensing. Compressive video sensing uses random coded exposure and reconstructs higher than read out of sensor frame rate video from a single coded image. It is possible to recognize an action in a scene from a single coded image because the image contains multiple temporal information for reconstructing a video. In this paper, we propose reconstruction-free action recognition from a single coded exposure image. We also proposed deep sensing framework which models camera sensing and classification models into convolutional neural network (CNN) and jointly optimize the coded exposure and classification model simultaneously. We demonstrated that the proposed method can recognize human actions from only a single coded image. We also compared it with competitive inputs, such as low-resolution video with a high frame rate and high-resolution video with a single frame in simulation and real experiments.
Tadashi Okawara, Michitaka Yoshida, Hajime Nagahara, Yasushi Yagi
ICCP3
2020 5D Light Field Synthesis from a Monocular Video
abstract
Commercially available light field cameras have difficulty in capturing 5D (4D + time) light field videos. They can only capture still light field images or are excessively expensive for normal users to capture the light field video. To tackle this problem, we propose a deep learning-based method for synthesizing a light field video from a monocular video. We propose a new synthetic light field video dataset that renders photorealistic scenes using Unreal Engine because no light field video dataset is available. The proposed deep learning framework synthesizes the light field video with a full set (9 × 9) of sub-aperture images from a normal monocular video. The proposed network consists of three sub-networks, namely, feature extraction, 5D light field video synthesis, and temporal consistency refinement. Experimental results show that our model can successfully synthesize the light field video for synthetic and real scenes and outperforms the previous frame-by-frame method quantitatively and qualitatively.
Kyuho Bae, Andre Ivan, Hajime Nagahara, In Kyu Park
ICPR3
2020 Constructing a Public Meeting Corpus
abstract
In this paper, we propose a full pipeline of analysis of a large corpus about a century of public meeting in historical Australian news papers, from construction to visual exploration. The corpus construction method is based on image processing and OCR. We digitize and transcribe texts of the specific topic of public meeting. Experiments show that our proposed method achieves a F-score of 87.8% for corpus construction. As a result, we built a content search tool for temporal and semantic content analysis.
Koji Tanaka, Chenhui Chu, Haolin Ren, Benjamin Renoust, Yuta Nakashima, Noriko Takemura, Hajime Nagahara, Takao Fujikawa
LREC7
2020 IterNet: Retinal Image Segmentation Utilizing Structural Redundancy in Vessel Networks
abstract
Retinal vessel segmentation is of great interest for diagnosis of retinal vascular diseases. To further improve the performance of vessel segmentation, we propose IterNet, a new model based on UNet [1], with the ability to find obscured details of the vessel from the segmented vessel image itself, rather than the raw input image. IterNet consists of multiple iterations of a mini-UNet, which can be 4× deeper than the common UNet. IterNet also adopts the weight-sharing and skip-connection features to facilitate training; therefore, even with such a large architecture, IterNet can still learn from merely 10~20 labeled images, without pre-training or any prior knowledge. IterNet achieves AUCs of 0.9816, 0.9851, and 0.9881 on three mainstream datasets, namely DRIVE, CHASE-DB1, and STARE, respectively, which currently are the best scores in the literature. The source code is available1.
Liangzhi Li 0001, Manisha Verma, Yuta Nakashima, Hajime Nagahara, Ryo Kawasaki
WACV4
2019 A 3-D Display Pipeline from Coded-Aperture Camera to Tensor Light-Field Display Through CNN
abstract
We propose an efficient pipeline from input to output for a tensor light-field display. Conventionally, a dense light field (i.e., tens of images taken with narrow viewpoint intervals) is required as an input in such displays. However, obtaining dense light fields is a challenging task for real scenes. To make the acquisition process more efficient, we adopted a coded-aperture camera as an input device, which is suitable for acquiring dense light fields in a compressive manner. Moreover, we modeled the entire process from acquisition to display using a convolutional neural network. As a result of training the network on a massive light field data, we can reproduce the whole light field on the display from only a few images taken with the camera. Both simulative and real experiments were conducted to show the effectiveness of our method.
Keita Maruyama, Yasutaka Inagaki, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara
ICIP5
2019 Facial Expression Recognition with Skip-Connection to Leverage Low-Level Features
abstract
Deep convolutional neural networks (CNNs) have established their feet in the ground of computer vision and machine learning, used in various applications. In this work, an attempt is made to learn a CNN for a task of facial expression recognition (FER). Our network has convolutional layers linked with an FC layer with a skip-connection to the classification layer. Motivation behind this design is that lower layers of a CNN are responsible for lower level features, and facial expressions can be mainly encoded in low-to-mid level features. Hence, in order to leverage the responses from lower layers, all convo-lutional layers are integrated via FC layers. Moreover, a network with shared parameters is used to extract landmark motion trajectory features. These visual and landmark features are fused to improve the performance. Our method is evaluated on the CK+ and Oulu-CASIA facial expression datasets.
Manisha Verma, Hirokazu Kobori, Yuta Nakashima, Noriko Takemura, Hajime Nagahara
ICIP5
2019 BUDA.ART: A Multimodal Content Based Analysis and Retrieval System for Buddha Statues
abstract
We introduce BUDA.ART, a system designed to assist researchers in Art History, to explore and analyze an archive of pictures of Buddha statues. The system combines different CBIR and classical retrieval techniques to assemble 2D pictures, 3D statue scans and meta-data, that is focused on the Buddha facial characteristics. We build the system from an archive of 50,000 Buddhism pictures, identify unique Buddha statues, extract contextual information, and provide specific facial embedding to first index the archive. The system allows for mobile, on-site search, and to explore similarities of statues in the archive. In addition, we provide search visualization and 3D analysis of the statues.
Benjamin Renoust, Matheus Oliveira Franca, Jacob Chan, Van Le, Ayaka Uesaka, Yuta Nakashima, Hajime Nagahara, Jueren Wang, Yutaka Fujioka
ACM Multimedia7
2019 Reflectance and Shape Estimation with a Light Field Camera Under Natural Illumination
Trung Ngo Thanh, Hajime Nagahara, Ko Nishino, Rin-Ichiro Taniguchi, Yasushi Yagi
Int. J. Comput. Vis.2
2018 A Coded Aperture for Watermark Extraction from Defocused Images
Hiroki Hamasaki, Shingo Takeshita, Kentaro Nakai, Toshiki Sonoda, Hiroshi Kawasaki, Hajime Nagahara, Satoshi Ono
ACCV (6)6
2018 Learning to Capture Light Fields Through a Coded Aperture Camera
Yasutaka Inagaki, Yuto Kobayashi, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara
ECCV (7)5
2018 Joint Optimization for Compressive Video Sensing and Reconstruction Under Hardware Constraints
Michitaka Yoshida, Akihiko Torii, Masatoshi Okutomi, Kenta Endo, Yukinobu Sugiyama, Rin-Ichiro Taniguchi, Hajime Nagahara
ECCV (10)7
2017 Reflectance and Shape Estimation with a Light Field Camera under Natural Illumination
Trung Ngo Thanh, Hajime Nagahara, Ko Nishino, Rin-Ichiro Taniguchi, Yasushi Yagi
BMVC2
2017 PCA-coded aperture for light field photography
abstract
A light field, which is often understood as a set of dense multi-view images, has been utilized in various 2D/3D applications. Efficient light field acquisition using a coded aperture camera is the target problem considered in this paper. Specifically, the entire light field, which consists of many images, should be reconstructed from only a few images that are captured through different aperture patterns. In previous work, this problem has often been discussed from the context of compressed sensing (CS). In contrast, we formulated this problem from the perspective of principal component analysis (PCA) to derive optimal non-negative aperture patterns and a straight-forward reconstruction algorithm. Even though it is based on a conventional technique, our method has proven to be more accurate and much faster than a state-of-the-art CS-based method.
Yusuke Yagi, Keita Takahashi 0001, Toshiaki Fujii, Toshiki Sonoda, Hajime Nagahara
ICIP5
2017 Adaptive background model registration for moving cameras
Tsubasa Minematsu, Hideaki Uchiyama, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
Pattern Recognit. Lett.4
2016 Real-Time Surface of Revolution Reconstruction on Dense SLAM
abstract
We present a fast and accurate method for reconstructing surfaces of revolution (SoR) on 3D data and its application to structural modeling of a cluttered scene in real-time. To estimate a SoR axis, we derive an approximately linear cost function for fast convergence. Also, we design a framework for reconstructing SoR on dense SLAM. In the experiment results, we show our method is accurate, robust to noise and runs in real-time.
Hideaki Uchiyama, Jean-Marie Normand, Guillaume Moreau, Hajime Nagahara, Rin-Ichiro Taniguchi
3DV5
2016 4D light field segmentation with spatial and angular consistencies
abstract
In this paper, we describe a supervised four-dimensional (4D) light field segmentation method that uses a graph-cut algorithm. Since 4D light field data has implicit depth information and contains redundancy, it differs from simple 4D hyper-volume. In order to preserve redundancy, we define two neighboring ray types (spatial and angular) in light field data. To obtain higher segmentation accuracy, we also design a learning-based likelihood, called objectness, which utilizes appearance and disparity cues. We show the effectiveness of our method via numerical evaluation and some light field editing applications using both synthetic and real-world light fields.
Hajime Mihara, Takuya Funatomi, Kenichiro Tanaka, Hiroyuki Kubo, Yasuhiro Mukaigawa, Hajime Nagahara
ICCP6
2016 High-speed imaging using CMOS image sensor with quasi pixel-wise exposure
abstract
Several recent studies in compressive video sensing have realized scene capture beyond the fundamental trade-off limit between spatial resolution and temporal resolution using random space-time sampling. However, most of these studies showed results for higher frame rate video that were produced by simulation experiments or using an optically simulated random sampling camera, because there are currently no commercially available image sensors with random exposure or sampling capabilities. We fabricated a prototype complementary metal oxide semiconductor (CMOS) image sensor with quasi pixel-wise exposure timing that can realize nonuniform space-time sampling. The prototype sensor can reset exposures independently by columns and fix these amount of exposure by rows for each 8×8 pixel block. This CMOS sensor is not fully controllable via the pixels, and has line-dependent controls, but it offers flexibility when compared with regular CMOS or charge-coupled device sensors with global or rolling shutters. We propose a method to realize pseudo-random sampling for high-speed video acquisition that uses the flexibility of the CMOS sensor. We reconstruct the high-speed video sequence from the images produced by pseudo-random sampling using an over-complete dictionary. The proposed method also removes the rolling shutter effect from the reconstructed video.
Hajime Nagahara, Toshiki Sonoda, Kenta Endo, Yukinobu Sugiyama, Rin-Ichiro Taniguchi
ICCP1
2016 Dynamic photometric stereo method using multi-tap CMOS image sensor
abstract
The photometric stereo method enables estimation of surface normals from images that have been captured using different but known lighting directions. The classical photometric stereo method requires at least three images to determine the normals in a given scene. However, this method cannot be applied to dynamic scenes because it is assumed that the scene remains static while the required images are captured. In this work, we present a dynamic photometric stereo method for estimation of the surface normals in a dynamic scene. We use a multi-tap complementary metal-oxide-semiconductor (CMOS) image sensor to capture the input images required for the proposed photometric stereo method. This image sensor can divide the electrons from the photodiode from a single pixel into the different taps of the exposures and can thus capture multiple images under different lighting conditions with almost identical timing. We implemented a camera lighting system and created a software application to enable estimation of the normal map in real time. We also evaluated the accuracy of the estimated surface normals and demonstrated that our proposed method can estimate the surface normals of dynamic scenes.
Takuya Yoda, Hajime Nagahara, Rin-Ichiro Taniguchi, Keiichiro Kagawa, Keita Yasutomi, Shoji Kawahito
ICPR2
2016 Background light ray modeling for change detection
Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
J. Vis. Commun. Image Represent.2
2016 Learning multi-task local metrics for image annotation
Xing Xu 0001, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
Multim. Tools Appl.3
2015 Estimating Surface Normals with Depth Image Gradients for Fast and Accurate Registration
abstract
We present a fast registration framework with estimating surface normals from depth images. The key component in the framework is to utilize adjacent pixels and compute the normal at each pixel on a depth image by following three steps. First, image gradients on a depth image are computed with a 2D differential filtering. Next, two 3D gradient vectors are computed from horizontal and vertical depth image gradients. Finally, the normal vector is obtained from the cross product of the 3D gradient vectors. Since horizontal and vertical adjacent pixels at each pixel are considered composing a local 3D plane, the 3D gradient vectors are equivalent to tangent vectors of the plane. Compared with existing normal estimation based on fitting a plane to a point cloud, our depth image gradients based normal estimation is extremely faster because it needs only a few mathematical operations. We apply it to normal space sampling based 3D registration and validate the effectiveness of our registration framework by evaluating its accuracy and computational cost with a public dataset.
Yosuke Nakagawa, Hideaki Uchiyama, Hajime Nagahara, Rin-Ichiro Taniguchi
3DV3
2015 Person re-identification visualization tool for object tracking across non-overlapping cameras
abstract
In this paper, we present a visualization tool for person re-identification when tracking objects across non-overlapping cameras. Tracking objects across non-overlapping cameras is challenging because the observations from different cameras are widely separated in both time and space. Hence, these systems need a large amount of labeled training data. Commonly, this training data is constructed manually at significant human cost. We support this process efficiently by visualizing the correspondences of objects across multiple cameras. Our tool facilitates the construction of a database for person re-identification with ease. Moreover, the accuracy of person re-identification can be increased using the generated database because the amount of training data is increased. In the experiments, we apply the proposed tool to real world situations to verify the validity of the proposed system.
Etienne Pot, Maiya Hori, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
AVSS4
2015 Change detection on light field for active video surveillance
abstract
Existing background model based change detection methods have difficulty in distinguishing between foreground and background changes when both changes are caused by the same factors. We explore the possibility of using a light field camera to resolve the problem of existing single-view camera-based approaches. We present a new change detection strategy that processes light rays captured by the light field camera. The light rays are used for three purposes: 1) generating an active surveillance field (ASF) to determine in-focus and out-focus areas, 2) evaluating focusness to determine whether the light rays come from the ASF, and 3) creating and updating light-ray background models to capture temporal changes in light rays. To investigate the effectiveness of the proposed approach, we evaluated several video sequences captured by a light field camera. Experimental results show that our change detection scheme can robustly handle challenging situations that cannot be resolved by existing single-view approaches.
Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
AVSS2
2015 Shape and light directions from shading and polarization
abstract
We introduce a method to recover the shape of a smooth dielectric object from polarization images taken with a light source from different directions. We present two constraints on shading and polarization and use both in a single optimization scheme. This integration is motivated by the fact that photometric stereo and polarization-based methods have complementary abilities. The polarization-based method can give strong cues for the surface orientation and refractive index, which are independent of the light direction. However, it has ambiguities in selecting between two ambiguous choices of the surface orientation, in the relationship between refractive index and zenith angle (observing angle), and limited performance for surface points with small zenith angles, where the polarization effect is weak. In contrast, photometric stereo method with multiple light sources can disambiguate the surface orientation and give a strong relationship between the surface normals and light directions. However, it has limited performance for large zenith angles, refractive index estimation, and faces the ambiguity in case the light direction is unknown. Taking their advantages, our proposed method can recover the surface normals for both small and large zenith angles, the light directions, and the refractive indexes of the object. The proposed method is successfully evaluated by simulation and real-world experiments.
Trung Ngo Thanh, Hajime Nagahara, Rin-Ichiro Taniguchi
CVPR2
2015 TransCut: Transparent Object Segmentation from a Light-Field Image
abstract
The segmentation of transparent objects can be very useful in computer vision applications. However, because they borrow texture from their background and have a similar appearance to their surroundings, transparent objects are not handled well by regular image segmentation methods. We propose a method that overcomes these problems using the consistency and distortion properties of a light-field image. Graph-cut optimization is applied for the pixel labeling problem. The light-field linearity is used to estimate the likelihood of a pixel belonging to the transparent object or Lambertian background, and the occlusion detector is used to find the occlusion boundary. We acquire a light field dataset for the transparent object, and use this dataset to evaluate our method. The results demonstrate that the proposed method successfully segments transparent objects from the background.
Yichao Xu, Hajime Nagahara, Atsushi Shimada 0001, Rin-Ichiro Taniguchi
ICCV2
2015 Adaptive search of background models for object detection in images taken by moving cameras
abstract
We propose a strategy of background subtraction for an image sequence captured by a moving camera. To adapt for camera motion, it is necessary to estimate the relation between consecutive frames in background subtraction. However, simple background subtraction using the relation between consecutive frames results in many false detections. We use re-projection error to handle this problem. The re-projection error has a low value in a background region. According to re-projection error, our method searches neighboring background models and tunes a threshold value for detection in order to reduce false detections. We evaluated the accuracy of detection of our method in experiments. Our method provided better detection than a method that does not search neighboring background models. Our method thus reduced the number of false detections.
Tsubasa Minematsu, Hideaki Uchiyama, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
ICIP4
2015 Tutorial 2: Computational Imaging and Projection
abstract
Summary form only given. In this tutorial, we will introduce emerging technologies on computational imaging and light field projection to AR/MR researchers.Light is the most important medium in AR/VR technologies to not only obtain information but also show and modify visual cue in the real scenes. Therefore in this area, latest techniques on optics, imaging and lighting have played an important role to make a next step toward the sophisticated experiences. Computational photography is one of the most influential technology in computer vision and optical engineering areas, and we think most techniques in computational imaging and projection can be applied to common problems in mixed reality, such as scene modeling, modification of the appearances of actual objects and user interactions.
Shinsaku Hiura, Hajime Nagahara, Daisuke Iwai, Toshiyuki Amano
ISMAR2
2015 Light field distortion feature for transparent object classification
Yichao Xu, Kazuki Maeno, Hajime Nagahara, Atsushi Shimada 0001, Rin-Ichiro Taniguchi
Comput. Vis. Image Underst.3
2015 Camera array calibration for light field acquisition
Yichao Xu, Kazuki Maeno, Hajime Nagahara, Rin-Ichiro Taniguchi
Frontiers Comput. Sci.3
2015 Similar gait action recognition using an inertial sensor
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
Pattern Recognit.3
2014 Light Transport Refocusing for Unknown Scattering Medium
abstract
In this paper we propose a new light transport refocusing method for depth estimation as well as for investigation inside scattering media with unknown scattering properties. Propagated visible light rays through scattering media are utilized in our proposed refocusing method. We use 2D light source to illuminate the scattering media and 2D image sensor for capturing transported rays. The proposed method that uses 4D light transport can clearly visualize shallow depth, as well as deep depth plane of the medium. We apply our light transport refocusing method for depth estimation using conventional depth-from-focus method and for clear visualization by descattering the light rays passing through the medium. To evaluate the effectiveness we have done experiments using acrylic and milk-water type scattering medium in various optical and geometrical conditions. Finally, we show up the results of depth estimation and clear visualization, as well as with numeric evaluation.
Md. Abdul Mannan, Seiichi Tagawa, Toru Tamaki, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
ICPR4
2014 Anonymous Camera for Privacy Protection
abstract
Privacy protection in the surveillance video data has received great attention. Although tremendous works have been proposed to provide effective privacy protection techniques, most of the algorithms are based on post-processing that deletes, obscures or encrypts the privacy information after privacy-included raw data are recorded. Consequently, they are vulnerable to raw data leak out, which may lead to unauthorized use. Therefore, it is imperative to develop a new privacy protection scheme which is capable of excluding any privacy information at the video recording phase. In this paper, we propose an anonymous camera aiming to protect the privacy of individuals at the video capturing phase by optical masking technique. It effectively reduces the risk of raw data leakage because no privacy information will be recorded by the camera. We implemented a prototype camera, which consists of an infrared camera, a RGB camera and a liquid crystal on silicon (LCoS) device. We introduce optical design and performance of the anonymous camera, the masking algorithm as well as the calibration methodology. Experimental results demonstrate that our prototype anonymous camera can perform accurate real time masking of the face for privacy protection.
Yuheng Lu, Hajime Nagahara, Rin-Ichiro Taniguchi
ICPR3
2014 Object detection based on spatiotemporal background models
Satoshi Yoshinaga, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
Comput. Vis. Image Underst.3
2014 Half-sweep imaging for depth from defocus
Shuhei Matsui, Hajime Nagahara, Rin-Ichiro Taniguchi
Image Vis. Comput.2
2014 Case-based background modeling: associative background database towards low-cost and high-performance change detection
Atsushi Shimada 0001, Yosuke Nonaka, Hajime Nagahara, Rin-Ichiro Taniguchi
Mach. Vis. Appl.3
2014 The largest inertial sensor-based gait database and performance evaluation of gait-based personal authentication
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
Pattern Recognit.3
2013 Light Field Distortion Feature for Transparent Object Recognition
abstract
Current object-recognition algorithms use local features, such as scale-invariant feature transform (SIFT) and speeded-up robust features (SURF), for visually learning to recognize objects. These approaches though cannot apply to transparent objects made of glass or plastic, as such objects take on the visual features of background objects, and the appearance of such objects dramatically varies with changes in scene background. Indeed, in transmitting light, transparent objects have the unique characteristic of distorting the background by refraction. In this paper, we use a single-shot light field image as an input and model the distortion of the light field caused by the refractive property of a transparent object. We propose a new feature, called the light field distortion (LFD) feature, for identifying a transparent object. The proposal incorporates this LFD feature into the bag-of-features approach for recognizing transparent objects. We evaluated its performance in laboratory and real settings.
Kazuki Maeno, Hajime Nagahara, Atsushi Shimada 0001, Rin-Ichiro Taniguchi
CVPR2
2013 Background Modeling Based on Bidirectional Analysis
abstract
Background modeling and subtraction is an essential task in video surveillance applications. Most traditional studies use information observed in past frames to create and update a background model. To adapt to background changes, the background model has been enhanced by introducing various forms of information including spatial consistency and temporal tendency. In this paper, we propose a new framework that leverages information from a future period. Our proposed approach realizes a low-cost and highly accurate background model. The proposed framework is called bidirectional background modeling, and performs background subtraction based on bidirectional analysis, i.e., analysis from past to present and analysis from future to present. Although a result will be output with some delay because information is taken from a future period, our proposed approach improves the accuracy by about 30% if only a 33-millisecond of delay is acceptable. Furthermore, the memory cost can be reduced by about 65% relative to typical background modeling.
Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
CVPR2
2012 Motion-Invariant Coding Using a Programmable Aperture Camera
Toshiki Sonoda, Hajime Nagahara, Rin-Ichiro Taniguchi
ACCV (4)2
2012 Inertial-sensor-based walking action recognition using robust step detection and inter-class relationships
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
ICPR3
2011 Phase registration in a gallery improving gait authentication
abstract
In this paper, we propose a method of inertial sensor-based gait authentication by inter-period phase registration of an owner's gallery. In spite of the importance for gait authentication of constructing a gallery of phase-registered gait patterns, previous implementations just relied on simple methods of period detection based on heuristic knowledge such as local peaks/valleys or local auto-correlation of the gait signals. Consequently, we propose to improve a gait gallery by incorporating a phase registration technique which globally optimizes inter-period phase consistency in an energy minimization framework. However, the previous phase registration technique suffers from a phase distortion problem due to ambiguities in the combination of a periodic signal function and a phase evolution function. We present a linear phase evolution prior to constructing an undistorted gait signal for better matching performance. Experiments using real gait signals from 32 subjects show that the proposed methods outperform the latest methods in the field.
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Yasushi Yagi
IJCB3
2011 Half-Sweep Imaging for Depth from Defocus
Shuhei Matsui, Hajime Nagahara, Rin-Ichiro Taniguchi
PSIVT (1)2
2011 Flexible Depth of Field Photography
abstract
The range of scene depths that appear focused in an image is known as the depth of field (DOF). Conventional cameras are limited by a fundamental trade-off between depth of field and signal-to-noise ratio (SNR). For a dark scene, the aperture of the lens must be opened up to maintain SNR, which causes the DOF to reduce. Also, today's cameras have DOFs that correspond to a single slab that is perpendicular to the optical axis. In this paper, we present an imaging system that enables one to control the DOF in new and powerful ways. Our approach is to vary the position and/or orientation of the image detector during the integration time of a single photograph. Even when the detector motion is very small (tens of microns), a large range of scene depths (several meters) is captured, both in and out of focus. Our prototype camera uses a micro-actuator to translate the detector along the optical axis during image integration. Using this device, we demonstrate four applications of flexible DOF. First, we describe extended DOF where a large depth range is captured with a very wide aperture (low noise) but with nearly depth-independent defocus blur. Deconvolving a captured image with a single blur kernel gives an image with extended DOF and high SNR. Next, we show the capture of images with discontinuous DOFs. For instance, near and far objects can be imaged with sharpness, while objects in between are severely blurred. Third, we show that our camera can capture images with tilted DOFs (Scheimpflug imaging) without tilting the image detector. Finally, we demonstrate how our camera can be used to realize nonplanar DOFs. We believe flexible DOF imaging can open a new creative dimension in photography and lead to new capabilities in scientific imaging, vision, and graphics.
Sujit Kuthirummal, Hajime Nagahara, Changyin Zhou, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Phase Registration of a Single Quasi-Periodic Signal Using Self Dynamic Time Warping
Yasushi Makihara, Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Yasushi Yagi
ACCV (3)3
2010 Object Detection Using Local Difference Patterns
Satoshi Yoshinaga, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
ACCV (4)3
2010 Programmable Aperture Camera Using LCoS
Hajime Nagahara, Changyin Zhou, Hiroshi Ishiguro, Shree K. Nayar
ECCV (6)1
2009 Adaptive-Scale Robust Estimator Using Distribution Model Fitting
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ACCV (3)2
2009 An adaptive-scale robust estimator for motion estimation
abstract
Although RANSAC is the most widely used robust estimator in computer vision, it has certain limitations making it ineffective in some situations, such as the motion estimation problem, in which uncertainty on the image features changes according to the capturing conditions. The greatest problem is that the threshold used by RANSAC to detect inliers cannot be changed adaptively; instead it is fixed by the user. An adaptive scale algorithm must therefore be applied in such cases. In this paper, we propose a new adaptive scale robust estimator that adaptively finds the best solution with the best scale to fit the inliers, without the need for predefined information. Our new adaptive scale estimator matches the residual probability density from an estimate and the standard Gaussian probability density function to find the best inlier scale. Our algorithm is evaluated in several motion estimation experiments under varying conditions and the results are compared with several of the latest adaptive-scale robust estimators.
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ICRA2
2008 Flexible Depth of Field Photography
Hajime Nagahara, Sujit Kuthirummal, Changyin Zhou, Shree K. Nayar
ECCV (4)1
2008 Robust and real-time egomotion estimation using a compound omnidirectional sensor
abstract
We propose a new egomotion estimation algorithm for a compound omnidirectional camera. Image features are detected by a conventional feature detector and then quickly classified into near and far features by checking infinity on the omnidirectional image of the compound omnidirectional sensor. Egomotion estimation is performed in two steps: first, rotation is recovered using far features; then translation is estimated from near features using the estimated rotation. RANSAC is used for estimations of both rotation and translation. Experiments in various environments show that our approach is robust and provides good accuracy in real-time for large motions.
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ICRA2
2008 Depth estimation from the color drift of a route panorama
abstract
Image based modeling methods has been well studied for generating a 3D model from an image sequence. Most of them require redundant and huge spatio-temporal images for estimating a scene depth. It is not good characteristic for taking a higher resolution of texture. A route panorama is a continuous panoramic image along a path. It is suitable for modeling large environments such as a city or town. The panorama captured by a line scan sensor also has advantage for capturing higher resolution easily. In this paper, we propose a method for depth estimation from the panorama. The route panorama has color drifts that correspond to the distances of captured objects. We use these color drifts to estimate the depth of an image. The proposed method detects the color drift by window matching using belief propagation. It also uses a Gaussian pyramid to stabilize the estimation and decrease its computation cost. We confirmed that the proposed method estimated depth maps from a single high-resolution panorama in experiments.
Hajime Nagahara, Atsushi Ichikawa, Masahiko Yachida
IROS1
2007 An Omnidirectional Vision Sensor with Single View and Constant Resolution
abstract
Many omnidirectional vision sensors based on convex mirrors have been proposed for observing a 360 degrees field of view. Some have a single viewpoint using a hyperboloidal or parabolic mirror, others have a constant resolution property. Until now, there have not been any that have had both these properties. In this paper, we propose an omnidirectional vision sensor that has both properties; that is, a single viewpoint and a constant resolution. The proposed omnidirectional sensor uses two mirrors, which improves the degree of freedom of the design for satisfying each property. We discuss optimization of the design in terms of geometry and optics.
Hajime Nagahara, Koji Yoshida, Masahiko Yachida
ICCV1
2007 Robust and Real-time Rotation Estimation of Compound Omnidirectional Sensor
abstract
Camera ego-motion consists of translation and rotation, in which rotation can be described simply by distant features. We present a robust rotation estimation using distant features given by our compound omnidirectional sensor. Features are detected by a conventional feature detector, and then distant features are identified by checking the infinity on the omnidirectional image of the compound sensor. The rotation matrix is estimated between consecutive video frames using RANSAC with only distant features. Experiments with various environments show that our approach is robust and also gives reasonable accuracy in real-time.
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ICRA2
2006 Calibration of Rotating Line Camera for Spherical Imaging
Tomoyuki Hirota, Hajime Nagahara, Masahiko Yachida
ACCV (1)2
2006 Video Synthesis with High Spatio-Temporal Resolution Using Motion Compensation and Image Fusion in Wavelet Domain
Kiyotaka Watanabe, Yoshio Iwai, Hajime Nagahara, Masahiko Yachida, Toshiya Suzuki
ACCV (1)3
2006 An Omnidirectional Vision Sensor with Single Viewpoint and Constant Resolution
abstract
Many omnidirectional vision sensors using a convex mirror have been proposed for observing a 360 degrees field of view. Some of them have a single viewpoint using a hyperboloidal or parabolic mirror, some of the others have a constant resolution property. Until now, there have not been any that have had both these properties. In this paper, we propose an omnidirectional vision sensor that has both a single viewpoint and a constant resolution. The proposed omnidirectional sensor uses two mirrors, which improves the degree of freedom of the design for satisfying each property. We introduce three practical settings of constant resolution; angular constant, horizontal constant and vertical constant resolutions
Koji Yoshida, Hajime Nagahara, Masahiko Yachida
IROS2
2005 Dual-sensor camera for acquiring image sequences with different spatio-temporal resolution
abstract
In accordance with advances in camera technology, the requirement of high-quality video has greatly increased. Among the factors required for high-quality video are a high-resolution and a high frame rate. However, the limitation of the pixel transfer rate restricts compatibility of a high-resolution and a high frame rate in commercial camera. We propose a dual-sensor camera that consists of two cameras: one with a high-resolution and a low frame rate and the other with a low-resolution and a high frame rate. The system is capable of capturing two different image sequences: high-resolution images and high-frame-rate images. A sensor calibration for the dual-sensor camera is also proposed.
Hajime Nagahara, Akira Hoshikawa, Tomohiro Shigemoto, Yoshio Iwai, Masahiko Yachida, Hiroyuki Tanaka
AVSS1
2004 Super wide viewer using catadioptrical optics
abstract
Many applications have used a head-mounted display (HMD), such as in virtual and mixed realities and telepresence. However, the field of view (FOV) of commercial HMD systems is too narrow for feeling immersion. In this paper, we propose a super-wide field of view head-mounted display consisting of an ellipsoidal mirror and a hyperboloidal curved mirror. The horizontal FOV of the proposed HMD is 180 degrees and includes the peripheral vision of humans. It increases the reality and immersion for users.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
ACM Trans. Graph.1
2003 Wide field of view catadioptrical head-mounted display
abstract
Many applications have used a Head-Mounted Display (HMD), such as in virtual and mixed realities, and tele-presence. The advantage of HMD systems is the ease of feeling a 3D world in the display of animation or movies. However, the field of view (FOV) of commercial HMD systems is too narrow for feeling immersion. The horizontal FOV of many commercial HMDs is around 60 degrees, significantly narrower than that of humans. In this paper, we propose a super wide field of view catadioptrical head-mounted display consisting of an ellipsoidal and a hyperboloidal curved mirror. The horizontal FOV of the proposed HMD is 180 degrees and includes the peripheral view of humans. It increases reality and immersion of users. As well, the central region (60 degrees) of the FOV can measure 3D distances using stereoscopics.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
IROS1
2003 Super wide viewer using catadioptrical optics
abstract
Many applications have used a Head-Mounted Display (HMD), such as in virtual and mixed realities, and tele-presence. The advantage of HMD systems is the ease of feeling a 3D world in the display of animation or movies. However, the field of view (FOV) of commercial HMD systems is too narrow for feeling immersion. The horizontal FOV of many commercial HMDs is around 60 degrees, significantly narrower than that of humans. In this paper, we propose a super wide field of view head-mounted display consisting of an ellipsoidal and a hyperboloidal curved mirror. The horizontal FOV of the proposed HMD is 180 degrees and includes the peripheral view of humans. It increases reality and immersion of users. As well, the central region (60 degrees) of the FOV can measure 3D distances using stereoscopics.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
VRST1
2003 Superresolution modeling using an omnidirectional image sensor
abstract
Recently, many virtual reality and robotics applications have been called on to create virtual environments from real scenes. A catadioptric omnidirectional image sensor composed of a convex mirror can simultaneously observe a 360-degree field of view making it useful for modeling man-made environments such as rooms, corridors, and buildings, because any landmarks around the sensor can be taken in and tracked in its large field of view. However, the angular resolution of the omnidirectional image is low because of the large field of view captured. Hence, the resolution of surface texture patterns on the three-dimensional (3-D) scene model generated is not sufficient for monitoring details. To overcome this, we propose a high resolution scene texture generation method that combines an omnidirectional image sequence using image mosaic and superresolution techniques.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
IEEE Trans. Syst. Man Cybern. Part B1
2002 Resolution Improving Method for a 3D Environment Modeling using Omnidirectional Image Sensor
abstract
Recently, many applications in virtual reality and robotics require to create virtual environment from a real scene. A catadioptric omnidirectional image sensor composed of a convex mirror can observe a 360-degree field of view at once. It is useful for modeling a man-made environment such as a room, a corridor and a building, because the landmarks around the sensor can be taken and tracked by its large field of view. However, the angular resolution of the omnidirectional image is low owing to capturing the large field of view. Therefore, the resolution of texture patterns of each surface on the generated 3D scene model is not enough for monitoring details. To solve this problem, we propose a high-resolution scene texture generation method that combines an omnidirectional image sequence using image mosaic and super-resolution techniques.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
ICRA1
2001 Resolution improving method from multi-focal omnidirectional images
abstract
The omnidirectional image sensor named HyperOmniVision, is composed of a hyperboloidal mirror and conventional video camera. It can observe a 360 degree field of view and can transform an input image to a perspective image. However, it has an intrinsic problem where the image resolution of HyperOmniVision is lower than that of an ordinary video camera, because it has a structure whereby only one CCD captures whole periphery scene. We propose a resolution improvement method using sub-pixel displaced and multi-focused images.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
ICIP (1)1