EDBT 2026 Demo / reviewers in the wild / expert
Koichiro Niinuma
dblp:35/214
· DBLP profile ↗
23ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-8367-3988ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Security and privacy · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolving Enactions of Expertise: Software Engineers' Evaluation and Demonstration of Coding Expertise with AI Coding AssistantsabstractAI coding assistants are changing how software engineers engage in coding work. This shift raises a key question: does the changing of coding work also alter how software engineers evaluate and demonstrate coding expertise? We explore this question through a simulated live coding interview involving two software engineers, one as evaluator and the other as candidate, with AI tools allowed. Participants continued to rely on familiar criteria but adjusted the evidence they sought, as AI assistants both introduced new forms of demonstrating expertise and obscured some established workflows. The importance of these evolving enactions varied with evaluators’ emphasis on implementation versus planning. Lacking a clear link to expertise, heightened productivity expectations created additional tensions around these evolving enactions. We conclude by discussing how extended enactions can be supported through AI-focused tools and training, and how tensions between diminished enactions and productivity call for collaborative attention. Yeonju Jang, Mose Sakashita, Koichiro Niinuma, Aakar Gupta |
CHI | 3 |
| 2026 | DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion WorkflowsabstractIn data-driven systems, integrating disparate data sources becomes challenging when incoming data does not conform to the system’s specifications. Despite advances in automated schema matching systems, data integration tasks involving complex semantic interrelationships still require users to manually identify and define transformations between datasets, which can be cognitively demanding and time-consuming. We present DataSpeck, an end-to-end system that automates the conversion of disparate data sources to fit any pre-existing data specification. DataSpeck employs an AI-driven human-in-the-loop design, using LLMs to analyze semantic relationships and generate step-by-step transformation pipelines autonomously, while only requesting user attention to resolve semantic ambiguities. In our technical evaluation, DataSpeck successfully automated ~86% of varied data transformations while generating interpretable strategies with confidence scores and targeted clarification requests. In a user study (N=12), participants completed data conversion tasks ~53% faster with significantly reduced cognitive load using DataSpeck compared to Microsoft Excel with Copilot. Adil Rahman, Koichiro Niinuma, Aakar Gupta |
CHI | 2 |
| 2026 | FuzzySeek: Multimodal Refinement of Imprecise Video Queries for Moment RetrievalabstractRecent AI advances have made it possible to retrieve specific moments from long-form videos using natural language queries. However, existing systems can struggle to align retrieval results with user intent due to the lack of means for users to express their intents in simple natural language text. Moreover, there is limited support for helping users express or refine their intents interactively. We present FuzzySeek, a video moment retrieval interface that supports the expression and specification of imprecise or broad exploratory queries through multimodal interaction. FuzzySeek proposes three key components (1) Multimodality-blended text querying to improve expressivity, enabling users to directly anchor multimodal content within their textual queries, (2) Proactive Multimodal Guidance, which identifies imprecise/broad terms and phrases and surfaces targeted clarifications across modalities to improve query specificity and, (3) Query rollback to enable iterative back and forth exploration to enable direct or exploratory searches. Through a technical evaluation, multiple illustrative use cases and a user study with 11 participants, we show that FuzzySeek improves clarification efficiency, reduces cognitive load, and better supports video moment retrieval for imprecise queries compared to a baseline system without such support. Aditi Mishra, Koichiro Niinuma, Aakar Gupta |
IUI | 2 |
| 2025 | AdaptiveSliders: User-aligned Semantic Slider-based Editing of Text-to-Image Model Output
Rahul Jain 0018, Amit Goel, Koichiro Niinuma, Aakar Gupta |
CHI | 3 |
| 2025 | Custom Condition Generation for Zero-Shot Human-Scene Interactions SynthesisabstractExisting methods for creating human interactions within scenes show promise for common interactions, but often fail with less frequent ones. To overcome this, we introduce a new approach that creates tailored conditions for generating these interactions without previously seen examples. This method leverages the strengths of both large language models (LLMs) and vision-language models (VLMs). Unlike the GenZI, the current state-of-the-art approach, which struggles with rare interactions due to its reliance on VLM inpainting, our method follows a three-step process: first, we generate a preliminary human posture using VLMs and then estimate this posture in three dimensions. Next, we refine the conditions to fit the specific scene and interaction by analyzing the inputs with both LLMs and VLMs. Finally, we fine-tune the placement, orientation, and posture of the human figure using specific optimization techniques. Our experimental results show that this method performs well across a wide range of interactions, including those that are less common. Ryosuke Kawamura, Zoltán Ádám Milacski, Fernando De la Torre, László A. Jeni, Koichiro Niinuma |
FG | 5 |
| 2025 | RN-Sam: Road Network-Aided Sam Optimization for Road Segmentation In Satellite ImageryabstractRoad segmentation in satellite imagery is critical for various applications, and the Segment Anything Model (SAM) has recently been applied to this task, as with other remote sensing applications. However, despite its advancements, applying SAM to road segmentation poses notable challenges. First, inaccurate prompts can degrade SAM’s performance. Second, existing methods often lack practical evaluation in cross-region scenario. To address these issues, this paper introduces RN-SAM, a novel framework that leverages OpenStreetMap (OSM) as a reliable source of auxiliary road network information to enhance SAM for road segmentation and incorporates a new dataset designed for robust cross-region evaluations. The proposed framework comprises two phases: fine-tuning SAM with road network-based prompts and applying test-time adaptation using OSM road network data. Experimental results on our dataset demonstrate that the proposed framework significantly enhances SAM’s performance for road segmentation and outperforms existing methods in both same-region and cross-region scenarios. Ryosuke Kawamura, Pablo Guarda, Pradeep Narwade, Koichiro Niinuma |
ICIP | 5 |
| 2025 | Gaussian Splatting Lucas-KanadeabstractGaussian Splatting and its dynamic extensions are effective for reconstructing 3D scenes from 2D images when there is significant camera movement to facilitate motion parallax and when scene objects remain relatively static. However, in many real-world scenarios, these conditions are not met. As a consequence, data-driven semantic and geometric priors have been favored as regularizers, despite their bias toward training data and their neglect of broader movement dynamics.
Departing from this practice, we propose a novel analytical approach that adapts the classical Lucas-Kanade method to dynamic Gaussian splatting. By leveraging the intrinsic properties of the forward warp field network, we derive an analytical velocity field that, through time integration, facilitates accurate scene flow computation. This enables the precise enforcement of motion constraints on warp fields, thus constraining both 2D motion and 3D positions of the Gaussians. Our method excels in reconstructing highly dynamic scenes with minimal camera movement, as demonstrated through experiments on both synthetic and real-world scenes. Liuyue Xie, Joel Julin, Koichiro Niinuma, László A. Jeni |
ICLR | 3 |
| 2025 | GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text ContextsabstractThe connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities makes capturing this connection challenging with a fixed set of descriptors. Specifically, closed vocabulary scene encoders, which require learning text-scene associations from scratch, have been favored in the literature, often resulting in inaccurate motion grounding. In this paper, we propose a method that integrates an open vocabulary scene encoder into the architecture, establishing a robust connection between text and scene. Our two-step approach starts with pretraining the scene encoder through knowledge distillation from an existing open vocabulary semantic image segmentation model, ensuring a shared text-scene feature space. Subsequently, the scene encoder is fine-tuned for conditional motion generation, incorporating two novel regularization losses that regress the category and size of the goal object. Our methodology achieves up to a 30% reduction in the goal object distance metric compared to the prior state-of-the-art baseline model on the HUMANISE dataset. This improvement is demonstrated through evaluations conducted using three implementations of our framework, a perceptual study, and an open vocabulary experiment. Additionally, our method is designed to accommodate future 2D open vocabulary segmentation methods for distillation in a plug-and-play manner. Zoltán Ádám Milacski, Koichiro Niinuma, Ryosuke Kawamura, Fernando De la Torre, László A. Jeni |
WACV | 2 |
| 2024 | CoGS: Controllable Gaussian SplattingabstractCapturing and re-animating the 3D structure of artic-ulated objects present significant barriers. On one hand, methods requiring extensively calibrated multi-view setups are prohibitively complex and resource-intensive, limiting their practical applicability. On the other hand, while single-camera Neural Radiance Fields (NeRFs) offer a more streamlined approach, they have excessive training and rendering costs. 3D Gaussian Splatting would be a suitable alternative but for two reasons. Firstly, existing methods for 3D dynamic Gaussians require synchronized multi- view cameras, and secondly, the lack of controllability in dynamic scenarios. We present CoGS, a methodfor Controllable Gaussian Splatting, that enables the direct ma-nipulation of scene elements, offering real-time control of dynamic scenes without the prerequisite of pre-computing control signals. We evaluated CoGS using both synthetic and real-world datasets that include dynamic objects that differ in degree of difficulty. In our evaluations, CoGS con-sistently outperformed existing dynamic and controllable neural representations in terms of visual fidelity. Joel Julin, Zoltán Ádám Milacski, Koichiro Niinuma, László A. Jeni |
CVPR | 4 |
| 2024 | Video Question Answering with Procedural Programs
Rohan Choudhury, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni |
ECCV (38) | 2 |
| 2024 | Synthetic Video Generation for Weakly Supervised Cross-Domain Video Anomaly Detection
Pradeep Narwade, Ryosuke Kawamura, Gaurav Gajbhiye, Koichiro Niinuma |
ICPR (15) | 4 |
| 2024 | Enhanced Product Classification Using Learned Prompt Ensembling and Dual Interpolation with CLIP-Based ModelabstractRetail product classification is a crucial technology due to its market size and potential. Several algorithms have been proposed; however, these approaches are not suitable due to their inability to accommodate new data without retraining or their insufficient performance caused by class names reflecting product-specific names. In this study, we adopt a CLIP-based model for retail product classification to overcome the challenges associated with model retraining for new products and the issue of unique class names. Our approach, learned prompt ensembling and dual interpolation (LPEDI), combines prompt learning and its ensembling with encoder fine-tuning, and employs dual interpolation for coefficient adjustment. The method outperforms existing solutions on two retail product datasets, achieving a 5.9% improvement for in-distribution data and a 4.5% gain for out-of-distribution data. These results establish LPEDI as a practical and effective solution for retail product classification. Takahisa Yamamoto, Koichiro Niinuma, László A. Jeni |
MMSP | 2 |
| 2024 | Don't Look Twice: Faster Video Transformers with Run-Length TokenizationabstractVideo transformers are slow to train due to extremely large numbers of input tokens, even though many video tokens are repeated over time. Existing methods to remove uninformative tokens either have significant overhead, negating any speedup, or require tuning for different datasets and examples. We present Run-Length Tokenization (RLT), a simple approach to speed up video transformers inspired by run-length encoding for data compression. RLT efficiently finds and removes `runs' of patches that are repeated over time before model inference, then replaces them with a single patch and a positional encoding to represent the resulting token's new length.
Our method is content-aware, requiring no tuning for different datasets, and fast, incurring negligible overhead.
RLT yields a large speedup in training, reducing the wall-clock time to fine-tune a video transformer by 30% while matching baseline model performance. RLT also works without training, increasing model throughput by 35% with only 0.1% drop in accuracy.
RLT speeds up training at 30 FPS by more than 100%, and on longer video datasets, can reduce the token count by up to 80\%. Our project page is at rccchoudhury.github.io/projects/rlt. Rohan Choudhury, Guanglei Zhu, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni |
NeurIPS | 4 |
| 2024 | Occlusion Sensitivity Analysis with Augmentation Subspace Perturbation in Deep Feature SpaceabstractDeep Learning of neural networks has gained prominence in multiple life-critical applications like medical diagnoses and autonomous vehicle accident investigations. However, concerns about model transparency and biases persist. Explainable methods are viewed as the solution to address these challenges. In this study, we introduce the Occlusion Sensitivity Analysis with Deep Feature Augmentation Subspace (OSA-DAS), a novel perturbation-based interpretability approach for computer vision. While traditional perturbation methods make only use of occlusions to explain the model predictions, OSA-DAS extends standard occlusion sensitivity analysis by enabling the integration with diverse image augmentations. Distinctly, our method utilizes the output vector of a DNN to build low-dimensional subspaces within the deep feature vector space, offering a more precise explanation of the model prediction. The structural similarity between these subspaces encompasses the influence of diverse augmentations and occlusions. We test extensively on the ImageNet-1k, and our class- and model-agnostic approach outperforms commonly used interpreters, setting it apart in the realm of explainable AI. Pedro H. V. Valois, Koichiro Niinuma, Kazuhiro Fukui |
WACV | 2 |
| 2023 | DyLiN: Making Light Field Networks DynamicabstractLight Field Networks, the re-formulations of radiance fields to oriented rays, are magnitudes faster than their coordinate network counterparts, and provide higher fidelity with respect to representing 3D structures from 2D observations. They would be well suited for generic scene representation and manipulation, but suffer from one problem: they are limited to holistic and static scenes. In this paper, we propose the Dynamic Light Field Network (DyLiN) method that can handle non-rigid deformations, including topological changes. We learn a deformation field from input rays to canonical rays, and lift them into a higher dimensional space to handle discontinuities. We further introduce CoDyLiN, which augments DyLiN with controllable attribute inputs. We train both models via knowledge distillation from pretrained dynamic radiance fields. We evaluated DyLiN using both synthetic and real world datasets that include various non-rigid deformations. DyLiN qualitatively outperformed and quantitatively matched state-of-the-art methods in terms of visual fidelity, while being 25 – 71× computationally faster. We also tested CoDyLiN on attribute annotated data and it surpassed its teacher model. Project page: https://dylin2023.github.io. Joel Julin, Zoltán Ádám Milacski, Koichiro Niinuma, László A. Jeni |
CVPR | 4 |
| 2023 | CoNFies: Controllable Neural Face AvatarsabstractNeural Radiance Fields (NeRF) are compelling techniques for modeling dynamic 3D scenes from 2D image collections. These volumetric representations would be well suited for synthesizing novel facial expressions but for two problems. First, deformable NeRFs are object agnostic and model holistic movement of the scene: they can replay how the motion changes over time, but they cannot alter it in an interpretable way. Second, controllable volumetric representations typically require either time-consuming manual annotations or 3D supervision to provide semantic meaning to the scene. We propose a controllable neural representation for face self-portraits (CoNFies), that solves both of these problems within a common framework, and it can rely on automated processing. We use automated facial action recognition (AFAR) to characterize facial expressions as a combination of action units (AU) and their intensities. AUs provide both the semantic locations and control labels for the system. CoNFies outperformed competing methods for novel view and expression synthesis in terms of visual and anatomic fidelity of expressions. Koichiro Niinuma, László A. Jeni |
FG | 2 |
| 2023 | Visually explaining 3D-CNN predictions for video classification with an adaptive occlusion sensitivity analysisabstractThis paper proposes a method for visually explaining the decision-making process of 3D convolutional neural networks (CNN) with a temporal extension of occlusion sensitivity analysis. The key idea here is to occlude a specific volume of data by a 3D mask in an input 3D temporalspatial data space and then measure the change degree in the output score. The occluded volume data that produces a larger change degree is regarded as a more critical element for classification. However, while the occlusion sensitivity analysis is commonly used to analyze single image classification, it is not so straightforward to apply this idea to video classification as a simple fixed cuboid cannot deal with the motions. To this end, we adapt the shape of a 3D occlusion mask to complicated motions of target objects. Our flexible mask adaptation is performed by considering the temporal continuity and spatial co-occurrence of the optical flows extracted from the input video data. We further propose to approximate our method by using the first-order partial derivative of the score with respect to an input image to reduce its computational cost. We demonstrate the effectiveness of our method through various and extensive comparisons with the conventional methods in terms of the deletion/insertion metric and the pointing metric on the UCF101. The code is available at: https://github.com/uchiyama33/AOSA. Tomoki Uchiyama, Naoya Sogi, Koichiro Niinuma, Kazuhiro Fukui |
WACV | 3 |
| 2021 | Synthetic Expressions are Better Than Real for Learning to Detect Facial ActionsabstractCritical obstacles in training classifiers to detect facial actions are the limited sizes of annotated video databases and the relatively low frequencies of occurrence of many actions. To address these problems, we propose an approach that makes use of facial expression generation. Our approach reconstructs the 3D shape of the face from each video frame, aligns the 3D mesh to a canonical view, and then trains a GAN-based network to synthesize novel images with facial action units of interest. To evaluate this approach, a deep neural network was trained on two separate datasets: One network was trained on video of synthesized facial expressions generated from FERA17; the other network was trained on unaltered video from the same database. Both networks used the same train and validation partitions and were tested on the test partition of actual video from FERA17. The network trained on synthesized facial expressions outperformed the one trained on actual facial expressions and surpassed current state-of-the-art approaches. Koichiro Niinuma, Itir Önal, Jeffrey F. Cohn, László A. Jeni |
WACV | 1 |
| 2019 | Unmasking the Devil in the Details: What Works for Deep Facial Action Coding?
Koichiro Niinuma, László A. Jeni, Itir Önal, Jeffrey F. Cohn |
BMVC | 1 |
| 2011 | Erratum to "Soft Biometric Traits for Continuous User Authentication"abstractIn the Acknowledgment for the above paper (ibid., vol. 5, no. 4, pp. 771-780, Dec 2010), due to a production error, the corresponding author's name was spelled incorrectly. The correct spelling is Unsang Park. Koichiro Niinuma, Unsang Park, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2010 | Soft Biometric Traits for Continuous User AuthenticationabstractMost existing computer and network systems authenticate a user only at the initial login session. This could be a critical security weakness, especially for high-security systems because it enables an impostor to access the system resources until the initial user logs out. This situation is encountered when the logged in user takes a short break without logging out or an impostor coerces the valid user to allow access to the system. To address this security flaw, we propose a continuous authentication scheme that continuously monitors and authenticates the logged in user. Previous methods for continuous authentication primarily used hard biometric traits, specifically fingerprint and face to continuously authenticate the initial logged in user. However, the use of these biometric traits is not only inconvenient to the user, but is also not always feasible due to the user's posture in front of the sensor. To mitigate this problem, we propose a new framework for continuous user authentication that primarily uses soft biometric traits (e.g., color of user's clothing and facial skin). The proposed framework automatically registers (enrolls) soft biometric traits every time the user logs in and fuses soft biometric matching with the conventional authentication schemes, namely password and face biometric. The proposed scheme has high tolerance to the user's posture in front of the computer system. Experimental results show the effectiveness of the proposed method for continuous user authentication. Koichiro Niinuma, Unsang Park, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2009 | Difference sphere: An approach to near light source estimation
Takeshi Takai, Atsuto Maki, Koichiro Niinuma, Takashi Matsuyama |
Comput. Vis. Image Underst. | 3 |
| 2004 | Difference Sphere: An Approach to Near Light Source Estimation
Takeshi Takai, Koichiro Niinuma, Atsuto Maki, Takashi Matsuyama |
CVPR (1) | 2 |