EDBT 2026 Demo / reviewers in the wild / expert
Kazuaki Nakamura
dblp:48/2611
· DBLP profile ↗
15ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorSecurity and privacy · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Label-Only Model Inversion Attacks with Query-Free Training of Conditional Diffusion-Based Attack ModelabstractAs DNN-based AI models have rapidly developed, the risk of cyber-attacks against them has also increased. This paper focuses on Model Inversion Attacks (MIA), which is the attack to reveal a victim model’s training data by estimating input samples that causes the victim model to produce adversaries’ designated output. Modern MIA methods often tackle the task of Label-Only MIA, where a black-box image classification model is assumed as a victim model whose network structure and parameters are not disclosed to adversaries. A key technique for attacking such a model is to exploit an image generation model (called "attack model") trained by the adversaries’ own image set and its corresponding class label set provided by the victim model. However, to get the class label set, existing methods have to send a massive amount of images to the victim model as queries, which is highly suspicious behavior. To solve this unpracticality, we propose a Label-Only MIA method that requires no queries to train the attack model. The proposed method employs a Conditional Diffusion Model (CDM) as the attack model to obtain plausible attack results. Specifically, we first train an image feature extractor as an auxiliary module and use it to embed all images in the adversaries’ own dataset to a feature space. Then, we train a CDM-based attack model using a set of pairs of an image and its embedded feature. These training procedures are performed independently from the victim model. We experimentally confirmed that the proposed method achieves an attack performance comparable to white-box-oriented MIA methods, demonstrating that MIA risks are significant even for practical image classification models that are capable of blocking malicious users who send too many queries. Hikaru Mori, Kazuaki Nakamura |
AVSS | 2 |
| 2025 | Detoxification of Unlabeled Dataset: Reducing Implicit Class Imbalance Using Pseudo-Jacobian of GAN's Generator
Kosei Suyama, Kazuaki Nakamura |
MMM (1) | 2 |
| 2023 | Social IoT Approach to Cyber Defense of a Deep-Learning-Based Recognition System in Front of Media Clones Generated by Model Inversion AttackabstractModel inversion attack (MIA) is a cyber threat with an increasing alert even for deep-learning-based recognition systems (DLRSs). By targeting a DLRS under a scenario of attacker access to the model structure and parameters, MIA generates a data clone for a certain targeted class label. To avoid the possible threats of such MIA-generated data clones, this research work proposes a social IoT approach to a collaborative cyber-defense among the online recognition systems (RSs) sharing the targeted class label. Since, the generation of an MIA-clone is by targeting an RS model and using its structure, parameters, and class labels output scores in an iterative optimization process, the generated clone is partially inherent to the targeted model. Thus, it is expected for an MIA-clone to show a different performance on a secondary RS wherein the same targeted class label is included. It is because, in the MIA generation of the clone, not only the targeted class label but also other class labels, and model parameters and structure affect the process, while the second model has just the targeted class label in common with the target model. Deploying the Social Internet of Recognition Systems (SIoRS), the proposed technique utilizes a collaborative recognition by SIoRC which plays the role of a complementary recognition besides the targeted RS. The recognition output by the targeted RS is further verified by the SIoRS complementary recognition result. To avoid the MIA-targeted data clones, the verification of recognition is by the log-likelihood ratio test between the targeted RS and the SIoRS complementary recognition confidence scores. The proposed technique is evaluated by statistical analysis on deep face RSs in 10000 Monte Carlo runs for each of the conventional, dc-generative adversarial network (GAN) and$\alpha $-GAN integrated MIA techniques in targeting two different user identities. The$Z$scores of the fitted normal distribution of the log-likelihood ratios indicate almost 100% detection rate of clones generated by conventional MIA and 95.23% and 86% of clones, respectively, generated by DC-GAN and$\alpha $-GAN integrated deep MIA techniques. Mahdi Khosravy, Kazuaki Nakamura, Naoko Nitta, Nilanjan Dey, Rubén González Crespo, Enrique Herrera-Viedma, Noboru Babaguchi |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Anonymization of Human Gait in Video Based on Silhouette Deformation and Texture TransferabstractThese days, a lot of videos are uploaded onto web-based video sharing services such as YouTube. These videos can be freely accessed from all over the world. On the other hand, they often contain the appearance of walking private people, which could be identified by silhouette-based gait recognition techniques rapidly developed in recent years. This causes a serious privacy issue. To avoid it, this paper proposes a method for anonymizing the appearance of walking people, namely human gait, in video. In the proposed method, we first crop human regions from all frames in an input video and binarize them to get their silhouettes. Next, we slightly deform the silhouettes from the aspects of static body shape and dynamic walking rhythm so that the person in the input video cannot be correctly identified by gait recognition techniques. After that, the textures of the original human regions are transferred onto the deformed silhouettes. We achieve this by a displacement field-based approach, which is training-free and thus robust to a variety of clothes. Finally, the anonymized human regions with the transferred textures are filled back into the input video. In the results of our experiments, we successfully degraded the accuracy of CNN-based gait recognition systems from 100% to 1.57% in the lowest case without yielding serious distortion in the appearance of the human regions, which demonstrated the effectiveness of the proposed method. Yuki Hirose, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Model Inversion Attack by Integration of Deep Generative Models: Privacy-Sensitive Face Generation From a Face Recognition SystemabstractCybersecurity in front of attacks to a face recognition system is an emerging issue in the cloud era, especially due to its strong bonds with the privacy of the users registered to the system. A possible attack is the model inversion attack (MIA) which aims to reveal the identity of a targeted user by generating the most proper datapoint input to the system with maximum corresponding confidence score at the output. The generated data of a registered user can be maliciously used as a serious invasion of the user privacy. In literature, MIA processes are categorized into white-box and black-box scenarios which are respectively with and without information about the system structure, parameters, and partially about the users. This research work assumes the MIA under semi-white box scenario of availability of system model structure and parameters but not any user data information, and verifies it as a severe threat even for a deep-learning-based face recognition system despite its complex structure and the diversity of registered user data. The alert state is promoted by Deep MIA which is the integration of deep generative models in MIA, and$\alpha $-GAN integrated MIA-initilized by a face based seed ($\alpha $-GAN-MIA-FS) is proposed. As a novel MIA search strategy, a pre-trained deep generative model with capability of generating a face image from a random feature vector is used for narrowing down the image search space to the feature vectors space, which has much lower dimensions. This allows the MIA process to efficiently search for a low-dimensional feature vector whose corresponding face image maximizes the confidence score. We have experimentally evaluated the proposed method by two objective criteria and three subjective criteria in comparison to$\alpha $-GAN-integrated MIA initialized with a random seed ($\alpha $-GAN-MIA-RS), DCGAN-integrated MIA (DCGAN-MIA), and the conventional MIA. The evaluation results approve the efficiency and superiority of the proposed technique in generating natural looking face clones with high recognizability as the targeted users. Mahdi Khosravy, Kazuaki Nakamura, Yuki Hirose, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Semi-Supervised Outdoor Image Generation Conditioned on Weather SignalsabstractIn recent years, various types of sensors observe the real world. Especially, weather sensors are densely installed all over the world to observe current weather situations at various places. However, weather signals such as the temperature or humidity obtained by weather sensors are intuitively difficult for humans to understand. On the other hand, images captured by typical RGB cameras can tell weather situations at the captured places in a more comprehensible way for humans; however, cameras are only installed at limited places and are not necessarily open to public due to privacy issues. In order to solve this problem, the goal of our work is to generate images which can tell weather situations at arbitrary time and locations. This can be realized by using a conditional generative adversarial network architecture that takes an image and a condition to transform the image accordingly to the condition. Training such image generator requires a large number of image and condition pairs as the training data. Although weather signals can be easily collected from weather sensors, collecting their spatially and temporally synchronized outdoor images is not easy. Thus, we propose a semi-supervised method for training the image generator. A relatively small number of pairs of an outdoor image and weather signals is collected, each from different web services, by considering their semantic consistency. The collected pairs are used to train a predictor for predicting weather signals from a given outdoor image. Then, the image generator is trained by using a large number of pairs of an outdoor image and pseudo weather signals predicted by the predictor as the training data. Sota Kawakami, Kei Okada, Naoko Nitta, Kazuaki Nakamura, Noboru Babaguchi |
ICPR | 4 |
| 2019 | Training-Free Method for Generating Motion Video Clones From A Still Image Considering Self-Occlusion of Human BodyabstractIn this paper, we propose a method for generating photo-realistic video in which a person virtually performs a motion that is not performed in the real world. we refer to such video as motion video clones (MVCs). The proposed method requires only two kinds of source information as input data: a reference video, in which a person A performs some motion, and a target image, which includes the whole body of another person B. Using these data, our method generates a MVC in which the person A's motion is re-enacted by the person B's body. Since our method does not need 3D human body model nor any training phase, it is suitable to MVC-based entertainment systems. To handle the self-occlusion of the human body in the reference video, we employ a part-based approach. For each part such as the right arm and the left leg, we first extract its skeleton from the target image and move it so that the motion represented by the reference video is reenacted. Next, we compute a 2D affine transform between the original and moved positions of the skeleton. This transform is used to map the texture of the target image onto each frame of a resultant MVC. Finally, we extend the part-wise affine transforms to pixel-wise ones by computing their linear combination for each pixel, whose combination weights are computed according to the geodesic distance between the pixel and the center of each part. This allows us to avoid unnatural appearance around the joints of the body parts. In our experiments, the proposed method generated much visually-natural MVCs than existing methods. Teppei Tsutsumi, Kazuaki Nakamura, Seiko Myojin, Naoko Nitta, Noboru Babaguchi |
ICIP | 2 |
| 2019 | Encryption-Free Framework of Privacy-Preserving Image Recognition for Photo-Based Information ServicesabstractNowadays, mobile devices, such as smartphones, have been widely used all over the world. In addition, the performance of image recognition has drastically increased with deep learning technologies. From these backgrounds, some photo-based information services provided in a client-server architecture are getting popular: client users take a photo of a certain spot and send it to a server, while the server identifies the spot with an image recognizer and returns its related information to the users. However, this kind of client-server image recognition can cause a privacy issue because image recognition results are sometimes privacy-sensitive. To tackle the privacy issue, in this paper, we propose a framework of privacy-preserving image recognition called EnfPire, in which the server cannot uniquely determine the recognition result but client users can do so. An overview of EnfPire is as follows. First, client users extract a visual feature from their taken photo and transform it so that the server cannot uniquely determine the recognition result. Then, the users send the transformed feature to the server that returns a set of candidates of the recognition result to the users. Finally, the users compare the candidates to the original visual feature for obtaining the final result. Our experimental results demonstrate that EnfPire successfully degrades the server's spot-recognition accuracy from 99.8% to 41.4% while keeping 86.9% of the spot-recognition accuracy on the user side. Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Generating Handwritten Character Clones from an Incomplete Seed Character Set using Collaborative FilteringabstractIn this paper, we propose a method for generating clones of a target writer's handwritten character images (called handwritten character clones or HCCs) using an incomplete seed character set, which consists of at most one or no example of his/her actual handwriting per character. In HCC generation, not a single HCC but its distribution should be created for each character because humans' actual handwriting images differ from each other even if the same writer writes the same character. However, it is difficult to achieve this from the incomplete seed character set. To solve the problem, in the proposed method, we first create a number of HCC distributions for each character by clustering a set of handwritten character images offered by other writers. Next, for each character contained in the seed character set, we choose the distribution best fit to its example. Finally, for the other characters, we estimate the best distribution for them employing collaborative filtering. We conducted pilot experiments focusing on Japanese character images, in which the proposed method successfully generated various HCCs with a certain level of quality for each character. Kazuaki Nakamura, Eiji Miyazaki, Naoko Nitta, Noboru Babaguchi |
ICFHR | 1 |
| 2017 | A Framework of Privacy-Preserving Image Recognition for Image-Based Information Services
Kojiro Fujii, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
MMM (1) | 2 |
| 2017 | Effect of Junk Images on Inter-concept Distance Measurement: Positive or Negative?
Yusuke Nagasawa, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
MMM (2) | 2 |
| 2016 | Detection of groups in crowd considering their activity stateabstractIn this paper, we focus on the problem of group detection in crowd, which is a task of partitioning a set of pedestrians in a scene into small subsets called groups based on their trajectories. Most of previous methods use only a single model for representing a relationship between trajectories of pedestrians who belong to the same group. However, such relationship would vary depending on the activity state (e.g. walking together, approaching, splitting, and so on) of the group. In this paper, we propose a novel group detection method which can cope with a variation of groups' activity state. The proposed method constructs different models for each activity state in order to appropriately evaluate the relationship of pedestrians' trajectories. In addition, our method regards groups' activity state as hidden variables and estimates their probability distributions, which is used for integrating the constructed models. The proposed method outperforms existing methods in the experiment on the public dataset. Kazuaki Nakamura, Tsukasa Ono, Noboru Babaguchi |
ICPR | 1 |
| 2015 | Temporal spotting of human actions from videos containing actor's unintentional motionsabstractThis paper proposes a method for temporal action spotting: the temporal segmentation and classification of human actions in videos. Naturally performed human actions often involve actor's unintentional motions. These unintentional motions yield false visual evidences in the videos, which are not related to the performed actions and degrade the performance of temporal action spotting. To deal with this problem, our proposed method empolys a voting-based approach in which the temporal relation between each action and its visual evidence is probabilistically modeled as a voting score function. Due to the approach, our method can robustly spot the target actions even when the actions involve several unintentional motions, because the effect of the false visual evidences yielded by the unintentional motions can be canceled by other visual evidences observed with the target actions. Experimental results showed that the proposed method is highly robust to the unintentional motions. Keita Hara, Kazuaki Nakamura, Noboru Babaguchi |
ICME | 2 |
| 2013 | People counting across spatially disjoint cameras by flow estimation between foreground regionsabstractOur goal is to develop a method for counting the number of people traveling over a wide area monitored by spatially disjoint multiple cameras with non-overlapping fields of view. The proposed method counts the number of people traversing across each pair of cameras' fields of view by estimating the flows between the foreground regions which have disappeared from and appeared in the camera views within a short time interval. The approach aims at resolving two problems: people in a foreground region can split and merge outside the cameras' fields of view, and the appearance variance of the same person and persons with similar appearance can lead to errors in person re-identification across cameras. The average errors of 0.184 persons in counting the number of people traveling over an area in a university campus monitored by four virtual cameras for 100 minutes have demonstrated the effectiveness of the proposed method. Naoko Nitta, Takayuki Nakazaki, Kazuaki Nakamura, Ryota Akai, Noboru Babaguchi |
AVSS | 3 |
| 2012 | Investigation of a Method to Estimate Learners' Interest Level for Agent-Based Conversational e-Learning
Kazuaki Nakamura, Koh Kakusho, Tetsuo Shoji, Michihiko Minoh |
IPMU (2) | 1 |