EDBT 2026 Demo / reviewers in the wild / expert
Shin'ya Nishida
dblp:83/1571
· DBLP profile ↗
17ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-5098-4752ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
4 papers |
Computational photography and imaging · 53% Computer animation and physical simulation · 24% Image and video processing · 12% | |
| Artificial intelligence
2 papers |
3D vision · 89% Autonomous driving · 11% | |
| Human-computer interaction and pervasive computing
3 papers |
Usability and user experience research · 77% Haptics and multimodal interaction · 14% Interaction techniques and input · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › motion estimation
optical flow |
1.5 | 2 | 2025 | HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation Comparison · CVPR 2025 Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention Network · NeurIPS 2023 |
Computational photography and imaging › illumination analysis › computational illumination
light projection |
0.8 | 2 | 2019 | Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light Projection · IEEE Trans. Vis. Comput. Graph. 2019 Demonstration of Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light Projection · VR 2019 |
Computer animation and physical simulation
motion retargeting |
0.8 | 2 | 2019 | Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light Projection · IEEE Trans. Vis. Comput. Graph. 2019 Demonstration of Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light Projection · VR 2019 |
Computer vision › 3D vision
motion estimation |
0.7 | 1 | 2023 | Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention Network · NeurIPS 2023 |
Bioinformatics and computational biology
computational neuroscience |
0.7 | 1 | 2023 | Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention Network · NeurIPS 2023 |
Bioinformatics and computational biology › computational neuroscience › sensory processing
visual motion processing |
0.7 | 1 | 2023 | Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention Network · NeurIPS 2023 |
Computational photography and imaging
projection mapping |
0.7 | 1 | 2023 | Studying User Perceptible Misalignment in Simulated Dynamic Facial Projection Mapping · ISMAR 2023 |
Image and video processing › perceptual modeling
visual perception modeling |
0.4 | 1 | 2019 | Demonstration of Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light Projection · VR 2019 |
Virtual and augmented reality › 3d display
stereoscopic display |
0.3 | 1 | 2017 | Hiding of phase-based stereo disparity for ghost-free viewing without glasses · ACM Trans. Graph. 2017 |
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles |
0.3 | 1 | 2025 | HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation Comparison · CVPR 2025 |
Haptics and multimodal interaction
just noticeable difference |
0.2 | 1 | 2023 | Studying User Perceptible Misalignment in Simulated Dynamic Facial Projection Mapping · ISMAR 2023 |
Usability and user experience research › user perception
latency perception |
0.2 | 1 | 2023 | Studying User Perceptible Misalignment in Simulated Dynamic Facial Projection Mapping · ISMAR 2023 |
Interaction techniques and input
projector-based interaction |
0.1 | 1 | 2019 | Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light Projection · IEEE Trans. Vis. Comput. Graph. 2019 |
Virtual and augmented reality
depth perception |
0.1 | 1 | 2017 | Hiding of phase-based stereo disparity for ghost-free viewing without glasses · ACM Trans. Graph. 2017 |
Methods — techniques the papers use, named apart from their topics
psychophysical experiment · 1.7optical flow algorithms · 1.7weighted up-down two-alternative forced-choice · 1.3self-attention network · 1.3recurrent network · 1.3motion energy sensing · 1.3optimization · 1.1perceptual model · 0.8spatial subband phase shift · 0.3quadrature-phase pattern · 0.3psychophysical evaluation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation ComparisonabstractAs AI models are increasingly integrated into applications involving human interaction, understanding the alignment between human perception and machine vision has become essential. One example is the estimation of visual motion (optical flow) in dynamic applications such as driving assistance. While there are numerous optical flow datasets and benchmarks with ground truth information, human-perceived flow in natural scenes remains underexplored. We introduce HuPerFlow—a benchmark for human-perceived flow, measured at 2,400 locations across ten optical flow datasets, with ∼38,400 response vectors collected through online psychophysical experiments. Our data demonstrate that human-perceived flow aligns with ground truth in spatiotemporally smooth locations while also showing systematic errors influenced by various environmental properties. Additionally, we evaluated several optical flow algorithms against human-perceived flow, uncovering both similarities and unique aspects of human perception in complex natural scenes. HuPerFlow is the first large-scale human-perceived flow benchmark for alignment between computer vision models and human perception, as well as for scientific exploration of human motion perception in natural scenes. The HuPerFlow benchmark is publicly available on the HuPerFlow website. Yung-Hao Yang, Zitang Sun, Taiki Fukiage, Shin'ya Nishida |
CVPR | 4 |
| 2025 | Building Reasonable Inference for Vision-Language Models in Blind Image Quality Assessment
Zitang Sun, Yen-Ju Chen, Shin'ya Nishida |
ICONIP (2) | 4 |
| 2025 | HAPI: A Model for Learning Robot Facial Expressions from Human PreferencesabstractAutomatic robotic facial expression generation is crucial for human–robot interaction (HRI), as handcrafted methods based on fixed joint configurations often yield rigid and unnatural behaviors. Although recent automated techniques reduce the need for manual tuning, they tend to fall short by not adequately bridging the gap between human preferences and model predictions—resulting in a deficiency of nuanced and realistic expressions due to limited degrees of freedom and insufficient perceptual integration. In this work, we propose a novel learning-to-rank framework that leverages human feedback to address this discrepancy and enhanced the expressiveness of robotic faces. Specifically, we conduct pairwise comparison annotations to collect human preference data and develop the Human Affective Pairwise Impressions (HAPI) model, a Siamese RankNet-based approach that refines expression evaluation. Results obtained via Bayesian Optimization and online expression survey on a 35-DOF android platform demonstrate that our approach produces significantly more realistic and socially resonant expressions of Anger, Happiness, and Surprise than those generated by baseline and expert-designed methods. This confirms that our framework effectively bridges the gap between human preferences and model predictions while robustly aligning robotic expression generation with human affective responses. Dongsheng Yang 0009, Qianying Liu, Wataru Sato, Takashi Minato, Shin'ya Nishida |
IROS | 6 |
| 2025 | Human visual grouping based on within- and cross-area temporal correlationsabstractPerceptual organization in the human visual system involves neural mechanisms that spatially group and segment image areas based on local feature similarities, such as the temporal correlation of luminance changes. Successful segmentation models in computer vision, including graph-based algorithms and vision transformer, leverage similarity computations across all elements in an image, suggest that effective similarity-based grouping should rely on a global computational process. However, whether human vision employs a similarly global computation remains unclear due to the absence of appropriate methods for manipulating similarity matrices across multiple elements within a stimulus. To investigate how "temporal similarity structures" influence human visual segmentation, we developed a stimulus generation algorithm based on Vision Transformer. This algorithm independently controls within-area and cross-area similarities by adjusting the temporal correlation of luminance, color, and spatial phase attributes. To assess human segmentation performance with these generated texture stimuli, participants completed a temporal two-alternative forced-choice task, identifying which of two intervals contained a segmentable texture. The results showed that segmentation performance is significantly influenced by the configuration of both within- and cross-correlation across the elements, regardless of attribute type. Furthermore, human performance is closely aligned with predictions from a graph-based computational model, suggesting that human texture segmentation can be approximated by a global computational process that optimally integrates pairwise similarities across multiple elements. Yen-Ju Chen, Zitang Sun, Shin'ya Nishida |
PLoS Comput. Biol. | 3 |
| 2023 | Studying User Perceptible Misalignment in Simulated Dynamic Facial Projection MappingabstractHigh-speed dynamic facial projection mapping (DFPM) is an advanced technology that aims to create perceptual changes in facial appearance by overlapping images based on facial position and shape. Compared to traditional monitor-based augmented reality systems, DFPM offers a higher level of immersion because users can directly observe digital content on their faces. However, DFPM suffers from misalignment issues owing to a slight temporal delay from sensing to projection, which reduces the level of immersion. To the best of our knowledge, no previous study has established the necessary latency requirements to avoid perceptible misalignment and achieve an immersive experience. Furthermore, conventional DFPM works followed latency requirements that were not reported for the DFPM scenario. Therefore, this study measured the latency that provided a just-noticeable difference (JND) in DFPM under different facial motion conditions, using the weighted up-down two-alternative forced-choice method. The results showed that user-perceptible misalignment was influenced by facial motion types and their velocities. Additionally, it was found that an average latency of 3.87 ms was necessary to avoid perceptible misalignment in the DFPM system when the translation speed was 0.5 m/s, which contradicts the commonly held belief regarding the required latency threshold. Hao-Lun Peng, Shin'ya Nishida, Yoshihiro Watanabe |
ISMAR | 2 |
| 2023 | Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention NetworkabstractVisual motion processing is essential for humans to perceive and interact with dynamic environments. Despite extensive research in cognitive neuroscience, image-computable models that can extract informative motion flow from natural scenes in a manner consistent with human visual processing have yet to be established. Meanwhile, recent advancements in computer vision (CV), propelled by deep learning, have led to significant progress in optical flow estimation, a task closely related to motion perception. Here we propose an image-computable model of human motion perception by bridging the gap between biological and CV models. Specifically, we introduce a novel two-stages approach that combines trainable motion energy sensing with a recurrent self-attention network for adaptive motion integration and segregation. This model architecture aims to capture the computations in V1-MT, the core structure for motion perception in the biological visual system, while providing the ability to derive informative motion flow for a wide range of stimuli, including complex natural scenes. In silico neurophysiology reveals that our model's unit responses are similar to mammalian neural recordings regarding motion pooling and speed tuning. The proposed model can also replicate human responses to a range of stimuli examined in past psychophysical studies. The experimental results on the Sintel benchmark demonstrate that our model predicts human responses better than the ground truth, whereas the state-of-the-art CV models show the opposite. Our study provides a computational architecture consistent with human visual motion processing, although the physiological correspondence may not be exact. Zitang Sun, Yen-Ju Chen, Yung-Hao Yang, Shin'ya Nishida |
NeurIPS | 4 |
| 2023 | Decoupled spatiotemporal adaptive fusion network for self-supervised motion estimation
Zitang Sun, Zhengbo Luo, Shin'ya Nishida |
Neurocomputing | 3 |
| 2020 | Visual perception of liquids: Insights from deep neural networksabstractVisually inferring material properties is crucial for many tasks, yet poses significant computational challenges for biological vision. Liquids and gels are particularly challenging due to their extreme variability and complex behaviour. We reasoned that measuring and modelling viscosity perception is a useful case study for identifying general principles of complex visual inferences. In recent years, artificial Deep Neural Networks (DNNs) have yielded breakthroughs in challenging real-world vision tasks. However, to model human vision, the emphasis lies not on best possible performance, but on mimicking the specific pattern of successes and errors humans make. We trained a DNN to estimate the viscosity of liquids using 100.000 simulations depicting liquids with sixteen different viscosities interacting in ten different scenes (stirring, pouring, splashing, etc). We find that a shallow feedforward network trained for only 30 epochs predicts mean observer performance better than most individual observers. This is the first successful image-computable model of human viscosity perception. Further training improved accuracy, but predicted human perception less well. We analysed the network's features using representational similarity analysis (RSA) and a range of image descriptors (e.g. optic flow, colour saturation, GIST). This revealed clusters of units sensitive to specific classes of feature. We also find a distinct population of units that are poorly explained by hand-engineered features, but which are particularly important both for physical viscosity estimation, and for the specific pattern of human responses. The final layers represent many distinct stimulus characteristics-not just viscosity, which the network was trained on. Retraining the fully-connected layer with a reduced number of units achieves practically identical performance, but results in representations focused on viscosity, suggesting that network capacity is a crucial parameter determining whether artificial or biological neural networks use distributed vs. localized representations. Jan Jaap R. van Assen, Shin'ya Nishida, Roland W. Fleming |
PLoS Comput. Biol. | 2 |
| 2019 | Demonstration of Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light ProjectionabstractA recently developed light projection technique can add dynamic impressions to static real objects without changing their original visual attributes such as surface colors and textures. It produces illusory motion impressions in the projection target by projecting gray-scale motion-inducer patterns that selectively drive the motion detectors in the human visual system. However, with this technique, determining the best deformation sizes is often difficult: When users try to add a large deformation, the deviation in the projected patterns from the original surface pattern on the target object becomes apparent. Therefore, to obtain satisfactory results, they have to spend much time and effort to manually adjust the shift sizes. Here, to overcome this limitation, we propose an optimization framework that adaptively retargets the displacement vectors based on a perceptual model. The perceptual model predicts the subjective inconsistency between a projected pattern and an original one by simulating responses in the human visual system. The displacement vectors are adaptively optimized so that the projection effect is maximized within the tolerable range predicted by the model. In the research demonstration, we will present a demo tool that incorporates our optimization technique, where a user can interactively edit dynamic appearances of a real object without cumbersome manual adjustments of deformation sizes. Taiki Fukiage, Takahiro Kawabe, Shin'ya Nishida |
VR | 3 |
| 2019 | Perceptually Based Adaptive Motion Retargeting to Animate Real Objects by Light ProjectionabstractA recently developed light projection technique can add dynamic impressions to static real objects without changing their original visual attributes such as surface colors and textures. It produces illusory motion impressions in the projection target by projecting gray-scale motion-inducer patterns that selectively drive the motion detectors in the human visual system. Since a compelling illusory motion can be produced by an inducer pattern weaker than necessary to perfectly reproduce the shift of the original pattern on an object's surface, the technique works well under bright environmental light conditions. However, determining the best deformation sizes is often difficult: When users try to add a large deformation, the deviation in the projected patterns from the original surface pattern on the target object becomes apparent. Therefore, to obtain satisfactory results, they have to spend much time and effort to manually adjust the shift sizes. Here, to overcome this limitation, we propose an optimization framework that adaptively retargets the displacement vectors based on a perceptual model. The perceptual model predicts the subjective inconsistency between a projected pattern and an original one by simulating responses in the human visual system. The displacement vectors are adaptively optimized so that the projection effect is maximized within the tolerable range predicted by the model. We extensively evaluated the perceptual model and optimization method through a psychophysical experiment as well as user studies. Taiki Fukiage, Takahiro Kawabe, Shin'ya Nishida |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Material and shape perception based on two types of intensity gradient informationabstractVisual estimation of the material and shape of an object from a single image includes a hard ill-posed computational problem. However, in our daily life we feel we can estimate both reasonably well. The neural computation underlying this ability remains poorly understood. Here we propose that the human visual system uses different aspects of object images to separately estimate the contributions of the material and shape. Specifically, material perception relies mainly on the intensity gradient magnitude information, while shape perception relies mainly on the intensity gradient order information. A clue to this hypothesis was provided by the observation that luminance-histogram manipulation, which changes luminance gradient magnitudes but not the luminance-order map, effectively alters the material appearance but not the shape of an object. In agreement with this observation, we found that the simulated physical material changes do not significantly affect the intensity order information. A series of psychophysical experiments further indicate that human surface shape perception is robust against intensity manipulations provided they do not disturb the intensity order information. In addition, we show that the two types of gradient information can be utilized for the discrimination of albedo changes from highlights. These findings suggest that the visual system relies on these diagnostic image features to estimate physical properties in a distal world. Masataka Sawayama, Shin'ya Nishida |
PLoS Comput. Biol. | 2 |
| 2017 | Hiding of phase-based stereo disparity for ghost-free viewing without glassesabstractWhen a conventional stereoscopic display is viewed without stereo glasses, image blurs, or 'ghosts', are visible due to the fusion of stereo image pairs. This artifact severely degrades 2D image quality, making it difficult to simultaneously present clear 2D and 3D contents. To overcome this limitation (backward incompatibility), here we propose a novel method to synthesize ghost-free stereoscopic images. Our method gives binocular disparity to a 2D image, and drives human binocular disparity detectors, by the addition of a quadrature-phase pattern that induces spatial subband phase shifts. The disparity-inducer patterns added to the left and right images are identical except for the contrast polarity. Physical fusion of the two images cancels out the disparity-inducer components and makes only the original 2D pattern visible to viewers without glasses. Unlike previous solutions, our method perfectly excludes stereo ghosts without using special hardware. A simple algorithm can transform 3D contents from the conventional stereo format into ours. Furthermore, our method can alter the depth impression of a real object without its being noticed by naked-eye viewers by means of light projection of the disparity-inducer components onto the object's surface. Psychophysical evaluations have confirmed the practical utility of our method. Taiki Fukiage, Takahiro Kawabe, Shin'ya Nishida |
ACM Trans. Graph. | 3 |
| 2016 | Seeing jelly: judging elasticity of a transparent objectabstractTaking advantage of computer graphics technologies, recent psychophysical study on material perception has revealed how human vision estimates the mechanical property of objects, such as liquid viscosity, from image features. Here we consider how human perceive another important mechanical material property --- elasticity. We simulated scenes in which a transparent cube falling on the floor, while manipulating the elasticity of the cube. We asked observers to rate the elasticity using a 5-point scale. Human observers were quite sensitive to the change in the simulated elasticity of the cube. In comparison with the original condition, the elasticity was overestimated when only the cube contour deformation was visible, whereas underestimated when the cube contour deformation was hidden and only internal optical deformation was visible. The effects of contour and optical deformations on elasticity rating were almost the same when the observers viewed white noise fields that reproduced the optical flow fields of the cube movies. Increasing frame duration (which decreased image speed) also increased the apparent elasticity. These results suggest that human elasticity judgment is based on the pattern of image motion arising from contour and optical deformations. This scientific finding may provide a hint for computationally efficient rendering of perceptually realistic dynamic scenes. Takahiro Kawabe, Shin'ya Nishida |
SAP | 2 |
| 2016 | Deformation Lamps: A Projection Technique to Make Static Objects Perceptually DynamicabstractLight projection is a powerful technique that can be used to edit the appearance of objects in the real world. Based on pixel-wise modification of light transport, previous techniques have successfully modified static surface properties such as surface color, dynamic range, gloss, and shading. Here, we propose an alternative light projection technique that adds a variety of illusory yet realistic distortions to a wide range of static 2D and 3D projection targets. The key idea of our technique, referred to as (Deformation Lamps), is to project only dynamic luminance information, which effectively activates the motion (and shape) processing in the visual system while preserving the color and texture of the original object. Although the projected dynamic luminance information is spatially inconsistent with the color and texture of the target object, the observer's brain automatically combines these sensory signals in such a way as to correct the inconsistency across visual attributes. We conducted a psychophysical experiment to investigate the characteristics of the inconsistency correction and found that the correction was critically dependent on the retinal magnitude of the inconsistency. Another experiment showed that the perceived magnitude of image deformation produced by our techniques was underestimated. The results ruled out the possibility that the effect obtained by our technique stemmed simply from the physical change in an object's appearance by light projection. Finally, we discuss how our techniques can make the observers perceive a vivid and natural movement, deformation, or oscillation of a variety of static objects, including drawn pictures, printed photographs, sculptures with 3D shading, and objects with natural textures including human bodies. Takahiro Kawabe, Taiki Fukiage, Masataka Sawayama, Shin'ya Nishida |
ACM Trans. Appl. Percept. | 4 |
| 2014 | Rendering fine hair-like objects with Gaussian noiseabstractWhen synthesizing images of fine objects like hair, we usually adopt sub-pixel drawing techniques to improve the image quality. For this paper, we analyzed the statistical features of images of thin lines and found that the distributions of the pixel values tended to be Gaussian. A psychophysical experiment showed that images of stripes with the appropriate Gaussian noise added are perceived to be finer than the original ones. We applied this perceptional property to hair rendering and developed a fast fine hair drawing algorithm. Mikio Shinya, Shin'ya Nishida |
SAP | 2 |
| 2006 | Interactions and Integrations of Multiple Sensory Channels in Human BrainabstractThis paper describes a couple of new principles with regard to interactions and integrations of multiple sensory channels in the human brain. First, as opposed to the general belief that the perception of shape and that of color are relatively independent of motion processing, human visual system integrates shape and color signals along perceived motion trajectory in order to improve visibility of shape and color of moving objects. Second, when the human sensory system binds the outputs of different sensory channels, (including audio-visual signals) based on their temporal synchrony, it uses only sparse salient features rather than using the time courses of full sensory signals. We believe these principles are potentially useful for development of effective audiovisual processing and presentation devices Shin'ya Nishida |
ICME | 1 |
| 1994 | A computational model for shape estimation by integration of shading and edge information
Hideki Hayakawa, Shin'ya Nishida, Yasuhiro Wada, Mitsuo Kawato |
Neural Networks | 2 |