Takashi Matsuyama

dblp:70/3309 · DBLP profile ↗
← Back
74ranked-venue papers
11as first author
0since 2021 · last 2017
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 7 first-authorArtificial intelligence and machine learning · 48 · 5 first-authorHuman-computer interaction and ubiquitous computing · 6Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
17 papers
Geometric modeling and processing · 63% Computational photography and imaging · 12% Multimedia analysis and retrieval · 8%
Artificial intelligence
19 papers
3D vision · 95% Video understanding and tracking · 4% Information extraction and text analysis · 0%

Topics — the 30 heaviest of 54, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
camera calibration
0.632017
A Linear Extrinsic Calibration of Kaleidoscopic Imaging System from Single 3D Point · CVPR 2017
A Linear Generalized Camera Calibration from Three Intersecting Reference Planes · ICCV 2015
A new mirror-based extrinsic camera calibration using an orthogonality constraint · CVPR 2012
Geometric modeling and processing
shape descriptor
0.652014
Timing-Based Local Descriptor for Dynamic Surfaces · CVPR 2014
Intrinsic Characterization of Dynamic Surfaces · CVPR 2013
Topology dictionary with Markov model for 3D video content-based skimming and description · CVPR 2009
Computer vision › 3D vision › camera calibration
extrinsic calibration
0.422017
A Linear Extrinsic Calibration of Kaleidoscopic Imaging System from Single 3D Point · CVPR 2017
A new mirror-based extrinsic camera calibration using an orthogonality constraint · CVPR 2012
Computer vision › 3D vision › camera calibration
mirror-based calibration
0.422017
A Linear Extrinsic Calibration of Kaleidoscopic Imaging System from Single 3D Point · CVPR 2017
A new mirror-based extrinsic camera calibration using an orthogonality constraint · CVPR 2012
Geometric modeling and processing › discrete geometry › discrete differential geometry
geodesic mapping
0.322014
Geodesic Mapping for Dynamic Surface Alignment · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Dynamic surface matching by geodesic mapping for 3D animation transfer · CVPR 2010
Computational photography and imaging
camera calibration
0.212015
A Linear Generalized Camera Calibration from Three Intersecting Reference Planes · ICCV 2015
Geometric modeling and processing › topology › computational topology
reeb graph
0.232012
Topology dictionary with Markov model for 3D video content-based skimming and description · CVPR 2009
Topology matching for 3D video compression · CVPR 2007
Topology Dictionary for 3D Video Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.222009
Minimal 3D video · SIGGRAPH ASIA Sketches 2009
Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereo · ICCV 2009
Geometric modeling and processing
shape matching
0.212014
Geodesic Mapping for Dynamic Surface Alignment · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Geometric modeling and processing › shape registration
surface alignment
0.212014
Geodesic Mapping for Dynamic Surface Alignment · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer animation and physical simulation
motion transfer
0.112010
Dynamic surface matching by geodesic mapping for 3D animation transfer · CVPR 2010
Geometric modeling and processing › shape matching
surface matching
0.112010
Dynamic surface matching by geodesic mapping for 3D animation transfer · CVPR 2010
Computer vision › 3D vision
3d reconstruction
0.112009
Minimal 3D video · SIGGRAPH ASIA Sketches 2009
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction
0.112009
Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereo · ICCV 2009
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.112009
Minimal 3D video · SIGGRAPH ASIA Sketches 2009
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction
0.112009
Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereo · ICCV 2009
Computer vision › 3D vision › stereo vision › stereo matching
wide-baseline stereo
0.112009
Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereo · ICCV 2009
Geometric modeling and processing
3d reconstruction
0.112008
Simultaneous super-resolution and 3D video using graph-cuts · CVPR 2008
Image and video processing
super-resolution
0.112008
Simultaneous super-resolution and 3D video using graph-cuts · CVPR 2008
Image and video coding › video compression
3d video coding
0.112007
Topology matching for 3D video compression · CVPR 2007
Virtual and augmented reality › immersive display
immersive projection display
0.112007
Inter-Reflection Compensation for Immersive Projection Display · CVPR 2007
Virtual and augmented reality › immersive display
multi-projector display
0.112007
Inter-Reflection Compensation for Immersive Projection Display · CVPR 2007
Computational photography and imaging › camera calibration
photometric calibration
0.112007
Inter-Reflection Compensation for Immersive Projection Display · CVPR 2007
Virtual and augmented reality
3d video
0.122010
Dynamic surface matching by geodesic mapping for 3D animation transfer · CVPR 2010
Simultaneous super-resolution and 3D video using graph-cuts · CVPR 2008
Computer vision › 3D vision › 3d shape reconstruction
non-rigid surface reconstruction
0.112014
Geodesic Mapping for Dynamic Surface Alignment · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer vision › Video understanding and tracking › video analytics › behavior analysis
behavior recognition
0.022000
Multiobject Behavior Recognition by Event Driven Selective Attention Method · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Appearance Based Behavior Recognition by Event Driven Selective Attention · CVPR 1998
Computer vision › 3D vision
photometric stereo
0.012004
Difference Sphere: An Approach to Near Light Source Estimation · CVPR (1) 2004
Computer vision › Video understanding and tracking
multi-object tracking
0.012002
Real-time multitarget tracking by a cooperative distributed vision system · Proc. IEEE 2002
Computer vision › 3D vision
shape from shading
0.021997
Shape from Shading with Interreflections Under a Proximal Light Source: Distortion-Free Copying of an Unfolded Book · Int. J. Comput. Vis. 1997
Shape from Shading with Interreflections under Proximal Light Source: 3D Shape Reconstruction of Unfolded Book Surface from a Scanner Image · ICCV 1995
Computer vision › 3D vision
structure from motion
0.012009
Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereo · ICCV 2009

Methods — techniques the papers use, named apart from their topics

plane intersection · 0.4linear calibration · 0.4probabilistic optimization · 0.4global geodesic coordinates · 0.4geodesic diffeomorphism · 0.4coarse-to-fine correspondence · 0.4linear dynamical systems · 0.4linear estimation · 0.3kaleidoscopic imaging · 0.3coarse-to-fine strategy · 0.2bag-of-timings · 0.2bags of dynamical systems · 0.2p3p problem · 0.1orthogonality constraint · 0.1dynamic memory architecture · 0.0asynchronous module interaction · 0.0logical reasoning · 0.0algebraic reasoning · 0.0
YearPublicationVenuePosition
2017 A Linear Extrinsic Calibration of Kaleidoscopic Imaging System from Single 3D Point
abstract
This paper proposes a new extrinsic calibration of kaleidoscopic imaging system by estimating normals and distances of the mirrors. The problem to be solved in this paper is a simultaneous estimation of all mirror parameters consistent throughout multiple reflections. Unlike conventional methods utilizing a pair of direct and mirrored images of a reference 3D object to estimate the parameters on a per-mirror basis, our method renders the simultaneous estimation problem into solving a linear set of equations. The key contribution of this paper is to introduce a linear estimation of multiple mirror parameters from kaleidoscopic 2D projections of a single 3D point of unknown geometry. Evaluations with synthesized and real images demonstrate the performance of the proposed algorithm in comparison with conventional methods.
Kosuke Takahashi, Akihiro Miyata, Shohei Nobuhara, Takashi Matsuyama
CVPR4
2017 Power Flow Coloring System Over a Nanogrid With Fluctuating Power Sources and Loads
abstract
Due to the excessive increase in environmental costs of nonrenewable energy resources, there have been government legislation, policies, and new technologies to control carbon emissions. The increase in renewable resources requires new strategies for the operation and management of the electricity. This paper presents the power flow coloring system, which gives a unique ID to each power flow between a specific power source and a specific power load. It allows us to design multiple power flow patterns between distributed and fluctuating power sources and loads taking into account energy availability, cost, and carbon dioxide emission. For implementation, this paper proposes a cooperative distributed control method, where a master-slave role assignment scheme of power sources and a time-slot-based feedback control are introduced to cope with power fluctuations, while keeping the voltage stable. The experimental results clarify the practical feasibility of our proposed method in managing distributed fluctuating power sources and loads.
Saher Javaid, Takekazu Kato, Takashi Matsuyama
IEEE Trans. Ind. Informatics3
2016 A Single-Shot Multi-Path Interference Resolution for Mirror-Based Full 3D Shape Measurement with a Correlation-Based ToF Camera
abstract
This paper is aimed at presenting a new algorithm for multi-path interference resolutions under mirror-based full 3D capture using a single correlation-based ToF camera. Our algorithm does not require additional captures or device modifications, and resolves the interference using a single ToF sensing that is also used for the 3D reconstruction as well. Evaluations with real images prove the concept of the proposed algorithm qualitatively and quantitatively.
Shohei Nobuhara, Takashi Kashino, Takashi Matsuyama, Kouta Takeuchi, Kensaku Fujii
3DV3
2016 Voting-Based Backchannel Timing Prediction Using Audio-Visual Information
abstract
While many spoken dialog systems are recently developed, users need to summarize and convey what they want the system to do clearly. However, in a human dialog, a speaker often summarize what to say incrementally, provided that there is a good listener who responds to the speaker's utterances at appropriate timing. We consider that generating backchannel responses, where appropriate, overlapped with the user's utterances is crucial for an artificial listener system that can promote user's utterances since such overlaps are the norm in human dialogs. Toward the goal to realize such a listener system, in this paper, we propose a voting-based algorithm of predicting the end of utterances early (i.e., before the utterances end) using audio-visual information. In the evaluation, we demonstrate the effectiveness of using audio-visual information and the applicability of the voting-based prediction algorithm with some early results.
Tomoki Nishide, Kei Shimonishi, Hiroaki Kawashima, Takashi Matsuyama
HAI4
2016 Modulating Dynamic Models for Lip Motion Generation
abstract
Generation of natural human motion is one of key techniques for multimodal dialogue systems with a human-like avatar. In particular, natural and expressive lip motion synthesis is necessary to make conversation between a user and an avatar richer. However, such expressive lip motion is often difficult to be generated automatically because it can be changed depending on phonemic context and prosody. To address this difficulty, we introduce a novel motion generation method on the basis of the modulation of a set of dynamic models learned from neutral motion data. As a suitable model for lip motion generation, we adopt a hybrid dynamical system, which consists of linear dynamical systems for each motion unit and a symbolic automaton for switching between these units. We show that, from the viewpoint of control theory, it is possible to modulate linear dynamical systems for various types of motion. Early results demonstrate the applicability of the proposed method using lip motion synthesis for simple phoneme sequences.
Singo Sawa, Hiroaki Kawashima, Kei Shimonishi, Takashi Matsuyama
HAI4
2015 Interference-Free Epipole-Centered Structured Light Pattern for Mirror-Based Multi-view Active Stereo
abstract
This paper is aimed at proposing a new structured light pattern for mirror-based multi-view active stereo so that the patterns cast onto the object surface do not interfere even where the object is illuminated by the projector directly and indirectly via mirror. The key idea of our interference-free projection is to encode the projector pixel locations so that they do not collide with the code from other projector pixels by exploiting the epipolar geometry defined by the real and the virtual projectors. We prove that our new encoding does not generate code collisions between the direct and indirect patterns from the real and the virtual projectors respectively. Evaluations using real and synthesized datasets demonstrate that our approach can realize an interference-free projection without using specialized equipment such as orthographic projectors used in the state-of-the-art methods.
Tomu Tahara, Ryo Kawahara, Shohei Nobuhara, Takashi Matsuyama
3DV4
2015 A Linear Generalized Camera Calibration from Three Intersecting Reference Planes
abstract
This paper presents a new generalized (or ray-pixel, raxel) camera calibration algorithm for camera systems involving distortions by unknown refraction and reflection processes. The key idea is use of intersections of calibration planes, while conventional methods utilized collinearity constraints of points on the planes. We show that intersections of calibration planes can realize a simple linear algorithm, and that our method can be applied to any ray-distributions while conventional methods require knowing the ray-distribution class in advance. Evaluations using synthesized and real datasets demonstrate the performance of our method quantitatively and qualitatively.
Mai Nishimura, Shohei Nobuhara, Takashi Matsuyama, Shinya Shimizu, Kensaku Fujii
ICCV3
2015 Bayesian perspective-plane (BPP) with maximum likelihood searching for visual localization
Zhaozheng Hu, Takashi Matsuyama
Multim. Tools Appl.2
2015 Cell-based visual surveillance with active cameras for 3D human gaze computation
Zhaozheng Hu, Takashi Matsuyama, Shohei Nobuhara
Multim. Tools Appl.2
2015 Augmented Motion History Volume for Spatiotemporal Editing of 3-D Video in Multiparty Interaction Scenes
abstract
We present a novel method that performs spatio temporal editing of 3-D multiparty interaction scenes for free-viewpoint browsing, from separately captured 3-D video data. The main idea is to first propose the augmented motion history volume (aMHV) for individual motion representation. Then, by modeling the correlations between different aMHVs, we can define a multiparty interaction dictionary, describing the spatiotemporal constraints for different types of multi-party interaction events. Finally, a constraint satisfaction and a global optimization method synthesize natural and continuous 3-D multiparty interaction scenes. Evaluations with real data demonstrate the effectiveness of our method.
Qun Shi, Shohei Nobuhara, Takashi Matsuyama
IEEE Trans. Circuits Syst. Video Technol.3
2015 Invariant shape descriptor for 3D video encoding
Tony Tung, Takashi Matsuyama
Vis. Comput.2
2014 A 3D Shape Descriptor for Segmentation of Unstructured Meshes into Segment-Wise Coherent Mesh Series
abstract
This paper presents a novel shape descriptor for topology-based segmentation of 3D video sequence. 3D video is a series of 3D meshes without temporal correspondences which benefit for applications including compression, motion analysis, and kinematic editing. In 3D video, both 3D mesh connectivities and the global surface topology can change frame by frame. This characteristic prevents from making accurate temporal correspondences through the entire 3D mesh series. To overcome this difficulty, we propose a two-step strategy which decomposes the entire sequence into a series of topologically coherent segments using our new shape descriptor, and then estimates temporal correspondences on a per-segment basis. We demonstrate the robustness and accuracy of the shape descriptor on real data which consist of large non-rigid motion and reconstruction errors.
Tomoyuki Mukasa, Shohei Nobuhara, Tony Tung, Takashi Matsuyama
3DV4
2014 A Real-Time View-Dependent Shape Optimization for High Quality Free-Viewpoint Rendering of 3D Video
abstract
This paper is aimed at proposing a new high quality free-viewpoint rendering algorithm of 3D video. The main challenge on visualizing 3D video is how to utilize the original multi-view images used to estimate the 3D surface, and how to manage the mismatches between them due to calibration and reconstruction errors. The key idea to solve this problem is to optimize the 3D shape on a per-viewpoint basis on the fly. Given a virtual viewpoint for visualization, our algorithm optimizes the 3D shape so as to maximize the photo-consistency over the surface visible from the virtual viewpoint. An evaluation demonstrates that our method outperforms the state-of-the-art rendering qualitatively and quantitatively.
Shohei Nobuhara, Wei Ning, Takashi Matsuyama
3DV3
2014 Inlier Estimation for Moving Camera Motion Segmentation
Xuefeng Liang, Cuicui Zhang, Takashi Matsuyama
ACCV (4)3
2014 Timing-Based Local Descriptor for Dynamic Surfaces
abstract
In this paper, we present the first local descriptor designed for dynamic surfaces. A dynamic surface is a surface that can undergo non-rigid deformation (e.g., human body surface). Using state-of-the-art technology, details on dynamic surfaces such as cloth wrinkle or facial expression can be accurately reconstructed. Hence, various results (e.g., surface rigidity, or elasticity) could be derived by microscopic categorization of surface elements. We propose a timing-based descriptor to model local spatiotemporal variations of surface intrinsic properties. The low-level descriptor encodes gaps between local event dynamics of neighboring keypoints using timing structure of linear dynamical systems (LDS). We also introduce the bag-of-timings (BoT) paradigm for surface dynamics characterization. Experiments are performed on synthesized and real-world datasets. We show the proposed descriptor can be used for challenging dynamic surface classification and segmentation with respect to rigidity at surface keypoints.
Tony Tung, Takashi Matsuyama
CVPR2
2014 Geodesic Mapping for Dynamic Surface Alignment
abstract
This paper presents a novel approach that achieves dynamic surface alignment by geodesing mapping. The surfaces are 3D manifold meshes representing non-rigid objects in motion (e.g., humans) which can be obtained by multiview stereo reconstruction. The proposed framework consists of a geodesic mapping (i.e., geodesic diffeomorphism) between surfaces which carry a distance function (namely the global geodesic distance), and a geodesic-based coordinate system (namely the global geodesic coordinates) defined similarly to generalized barycentric coordinates. The coordinates are used to recursively choose correspondence points in non-ambiguous regions using a coarse-to-fine strategy to reliably locate all surface points and define a discrete mapping. Complete point-to-point surface alignment with smooth mapping is then derived by optimizing a piecewise objective function within a probabilistic framework. The proposed technique only relies on surface intrinsic geometrical properties, and does not require prior knowledge on surface appearance (e.g., color or texture), shape (e.g., topology) or parameterization (e.g., mesh connectivity or complexity). The method can be used for numerous applications, such as visual information (e.g., texture) transfer between surface models representing different objects, dense motion flow estimation of 3D dynamic surfaces, wide-timeframe matching, etc. Experiments show compelling results on challenging publicly available real-world datasets.
Tony Tung, Takashi Matsuyama
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Multiparty Interaction Understanding Using Smart Multimodal Digital Signage
abstract
This paper presents a novel multimodal system designed for multi-party human-human interaction analysis. The design of human-machine interfaces for multiple users is challenging because simultaneous processing of actions and reactions have to be consistent. The proposed system consists of a large display equipped with multiple sensing devices: microphone array, HD video cameras, and depth sensors. Multiple users positioned in front of the panel freely interact using voice or gesture while looking at the displayed content, without wearing any particular devices (such as motion capture sensors or head mounted devices). Acoustic and visual information is captured and processed jointly using established and state-of-the-art techniques to obtain individual speech and gaze direction. Furthermore, a new framework is proposed to model A/V multimodal interaction between verbal and nonverbal communication events. Dynamics of audio signals obtained from speaker diarization and head poses extracted from video images are modeled using hybrid dynamical systems (HDS). We show that HDS temporal structure characteristics can be used for multimodal interaction level estimation, which is useful feedback that can help to improve multi-party communication experience. Experimental results using synthetic and real-world datasets of group communication such as poster presentations show the feasibility of the proposed multimodal system.
Tony Tung, Randy Gomez, Tatsuya Kawahara, Takashi Matsuyama
IEEE Trans. Hum. Mach. Syst.4
2013 Augmented Motion History Volume for Spatiotemporal Editing of 3D Video in Multi-party Interaction Scenes
abstract
In this paper we present a novel method that performs spatiotemporal editing of 3D multi-party interaction scenes for free-viewpoint browsing, from separately captured data. The main idea is to first propose the augmented Motion History Volume (aMHV) for motion representation. Then by modeling the correlations between different aMHVs we can define a multi-party interaction dictionary, describing the spatiotemporal constraints for different types of multiparty interaction events. Finally, constraint satisfaction and global optimization methods are proposed to synthesize natural and continuous multi-party interaction scenes. Experiments with real data illustrate the effectiveness of our method.
Qun Shi, Shohei Nobuhara, Takashi Matsuyama
3DV3
2013 Intrinsic Characterization of Dynamic Surfaces
abstract
This paper presents a novel approach to characterize deformable surface using intrinsic property dynamics. 3D dynamic surfaces representing humans in motion can be obtained using multiple view stereo reconstruction methods or depth cameras. Nowadays these technologies have become capable to capture surface variations in real-time, and give details such as clothing wrinkles and deformations. Assuming repetitive patterns in the deformations, we propose to model complex surface variations using sets of linear dynamical systems (LDS) where observations across time are given by surface intrinsic properties such as local curvatures. We introduce an approach based on bags of dynamical systems, where each surface feature to be represented in the codebook is modeled by a set of LDS equipped with timing structure. Experiments are performed on datasets of real-world dynamical surfaces and show compelling results for description, classification and segmentation.
Tony Tung, Takashi Matsuyama
CVPR2
2013 Predicting where we look from spatiotemporal gaps
abstract
When we are watching videos, there exist spatiotemporal gaps between where we look and what we focus on, which result from temporally delayed responses and anticipation in eye movements. We focus on the underlying structures of those gaps and propose a novel method to predict points of gaze from video data. In the proposed methods, we model the spatiotemporal patterns of salient regions that tend to be focused on and statistically learn which types of the patterns strongly appear around the points of gaze with respect to each type of eye movements. It allows us to exploit the structures of gaps affected by eye movements and salient motions for the gaze-point prediction. The effectiveness of the proposed method is confirmed with several public datasets.
Ryo Yonetani, Hiroaki Kawashima, Takashi Matsuyama
ICMI3
2012 Invariant Surface-Based Shape Descriptor for Dynamic Surface Encoding
Tony Tung, Takashi Matsuyama
ACCV (1)2
2012 A new mirror-based extrinsic camera calibration using an orthogonality constraint
abstract
This paper is aimed at calibrating the relative posture and position, i.e. extrinsic parameters, of a stationary camera against a 3D reference object which is not directly visible from the camera. We capture the reference object via a mirror under three different unknown poses, and then calibrate the extrinsic parameters from 2D appearances of reflections of the reference object in the mirrors. The key contribution of this paper is to present a new algorithm which returns a unique solution of three P3P problems from three mirrored images. While each P3P problem has up to four solutions and therefore a set of three P3P problems has up to 64 solutions, our method can select a solution based on an orthogonality constraint which should be satisfied by all families of reflections of a single reference object. In addition we propose a new scheme to compute the extrinsic parameters by solving a large system of linear equations. These two points enable us to provide a unique and robust solution. We demonstrate the advantages of the proposed method against a state-of-the-art by qualitative and quantitative evaluations using synthesized and real data.
Kosuke Takahashi, Shohei Nobuhara, Takashi Matsuyama
CVPR3
2012 Multi-mode saliency dynamics model for analyzing gaze and attention
abstract
We present a method to analyze a relationship between eye movements and saliency dynamics in videos for estimating attentive states of users while they watch the videos. The multi-mode saliency-dynamics model (MMSDM) is introduced to segment spatio-temporal patterns of the saliency dynamics into multiple sequences of primitive modes underlying the saliency patterns. The MMSDM enables us to describe the relationship by the local saliency dynamics around gaze points, which is modeled by a set of distances between gaze points and salient regions characterized by the extracted modes. Experimental results show the effectiveness of the proposed model to classify the attentive states of users by learning the statistical difference of the local saliency dynamics on gaze-paths at each level of attentiveness.
Ryo Yonetani, Hiroaki Kawashima, Takashi Matsuyama
ETRA3
2012 Multi-subregion face recognition using coarse-to-fine Quad-tree decomposition
Cuicui Zhang, Xuefeng Liang, Takashi Matsuyama
ICPR3
2012 Topology Dictionary for 3D Video Understanding
abstract
This paper presents a novel approach that achieves 3D video understanding. 3D video consists of a stream of 3D models of subjects in motion. The acquisition of long sequences requires large storage space (2 GB for 1 min). Moreover, it is tedious to browse data sets and extract meaningful information. We propose the topology dictionary to encode and describe 3D video content. The model consists of a topology-based shape descriptor dictionary which can be generated from either extracted patterns or training sequences. The model relies on 1) topology description and classification using Reeb graphs, and 2) a Markov motion graph to represent topology change states. We show that the use of Reeb graphs as the high-level topology descriptor is relevant. It allows the dictionary to automatically model complex sequences, whereas other strategies would require prior knowledge on the shape and topology of the captured subjects. Our approach serves to encode 3D video sequences, and can be applied for content-based description and summarization of 3D video sequences. Furthermore, topology class labeling during a learning process enables the system to perform content-based event recognition. Experiments were carried out on various 3D videos. We showcase an application for 3D video progressive summarization using the topology dictionary.
Tony Tung, Takashi Matsuyama
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Dynamic surface matching by geodesic mapping for 3D animation transfer
abstract
This paper presents a novel approach that achieves complete matching of 3D dynamic surfaces. Surfaces are captured from multi-view video data and represented by sequences of 3D manifold meshes in motion (3D videos). We propose to perform dense surface matching between 3D video frames using geodesic diffeomorphisms. Our algorithm uses a coarse-to-fine strategy to derive a robust correspondence map, then a probabilistic formulation is coupled with a voting scheme in order to obtain local unicity of matching candidates and a smooth mapping. The significant advantage of the proposed technique compared to existing approaches is that it does not rely on a color-based feature extraction process. Hence, our method does not lose accuracy in poorly textured regions and is not bounded to be used on video sequences of a unique subject. Therefore our complete surface mapping can be applied to: (1) texture transfer between surface models extracted from different sequences, (2) dense motion flow estimation in 3D video, and (3) motion transfer from a 3D video to an unanimated 3D model. Experiments are performed on challenging publicly available real-world datasets and show compelling results.
Tony Tung, Takashi Matsuyama
CVPR2
2010 3D video performance segmentation
abstract
We present a novel approach that achieves segmentation of subject body parts in 3D videos. 3D video consists in a free-viewpoint video of real-world subjects in motion immersed in a virtual world. Each 3D video frame is composed of one or several 3D models. A topology dictionary is used to cluster 3D video sequences with respect to the model topology and shape. The topology is characterized using Reeb graph-based descriptors and no prior explicit model on the subject shape is necessary to perform the clustering process. In this framework, the dictionary consists in a set of training input poses with a priori segmentation and labels. As a consequence, all identified frames of 3D video sequences can be automatically segmented. Finally, motion flows computed between consecutive frames are used to transfer segmented region labels to unidentified frames. Our method allows us to perform robust body part segmentation and tracking in 3D cinema sequences.
Tony Tung, Takashi Matsuyama
ICIP2
2010 Gaze Probing: Event-Based Estimation of Objects Being Focused On
abstract
We propose a novel method to estimate the object that a user is focusing on by using the synchronization between the movements of objects and a user's eyes as a cue. We first design an event as a characteristic motion pattern, and we then embed it within the movement of each object. Since the user's ocular reactions to these events are easily detected using a passive camera-based eye tracker, we can successfully estimate the object that the user is focusing on as the one whose movement is most synchronized with the user's eye reaction. Experimental results obtained from the application of this system to dynamic content (consisting of scrolling images) demonstrate the effectiveness of the proposed method over existing methods.
Ryo Yonetani, Hiroaki Kawashima, Takatsugu Hirayama, Takashi Matsuyama
ICPR4
2010 Speech estimation in non-stationary noise environments using timing structures between mouth movements and sound signals
abstract
A variety of methods for audio-visual integration, which inte-grate audio and visual information at the level of either features, states, or classifier outputs, have been proposed for the purpose of robust speech recognition. However, these methods do not al-ways fully utilize auditory information when the signal-to-noise ratio becomes low. In this paper, we propose a novel approach to estimate speech signal in noise environments. The key idea behind this approach is to exploit clean speech candidates gen-erated by using timing structures between mouth movements and sound signals. We first extract a pair of feature sequences of media signals and segment each sequence into temporal inter-vals. Then, we construct a cross-media timing-structure model of human speech by learning the temporal relations of overlap-ping intervals. Based on the learned model, we generate clean speech candidates from the observed mouth movements. Index Terms: multimodal, non-stationary noise, timing, linear dynamical system, particle filtering
Hiroaki Kawashima, Yu Horii, Takashi Matsuyama
INTERSPEECH3
2009 Topology dictionary with Markov model for 3D video content-based skimming and description
abstract
This paper presents a novel approach to skim and describe 3D videos. 3D video is an imaging technology which consists in a stream of 3D models in motion captured by a synchronized set of video cameras. Each frame is composed of one or several 3D models, and therefore the acquisition of long sequences at video rate requires massive storage devices. In order to reduce the storage cost while keeping relevant information, we propose to encode 3D video sequences using a topology-based shape descriptor dictionary. This dictionary is either generated from a set of extracted patterns or learned from training input sequences with semantic annotations. It relies on an unsupervised 3D shape-based clustering of the dataset by Reeb graphs, and features a Markov network to characterize topological changes. The approach allows content-based compression and skimming with accurate recovery of sequences and can handle complex topological changes. Redundancies are detected and skipped based on a probabilistic discrimination process. Semantic description of video sequences is then automatically performed. In addition, forthcoming frame encoding is achieved using a multiresolution matching scheme and allows action recognition in 3D. Our experiments were performed on complex 3D video sequences. We demonstrate the robustness and accuracy of the 3D video skimming with dramatic low bitrate coding and high compression ratio.
Tony Tung, Takashi Matsuyama
CVPR2
2009 Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereo
abstract
This paper presents a novel approach to achieve accurate and complete multi-view reconstruction of dynamic scenes (or 3D videos). 3D videos consist in sequences of 3D models in motion captured by a surrounding set of video cameras. To date 3D videos are reconstructed using multiview wide baseline stereo (MVS) reconstruction techniques. However it is still tedious to solve stereo correspondence problems: reconstruction accuracy falls when stereo photo-consistency is weak, and completeness is limited by self-occlusions. Most MVS techniques were indeed designed to deal with static objects in a controlled environment and therefore cannot solve these issues. Hence we propose to take advantage of the image content stability provided by each single-view video to recover any surface regions visible by at least one camera. In particular we present an original probabilistic framework to derive and predict the true surface of models. We propose to fuse multi-view structure-from-motion with robust 3D features obtained by MVS in order to significantly improve reconstruction completeness and accuracy. A min-cut problem where all exact features serve as priors is solved in a final step to reconstruct the 3D models. In addition, experimental results were conducted on synthetic and challenging real world datasets to illustrate the robustness and accuracy of our method.
Tony Tung, Shohei Nobuhara, Takashi Matsuyama
ICCV3
2009 Minimal 3D video
abstract
We present a new concept that achieves the 3D reconstruction of dynamic scenes from multi-view video cameras (or 3D videos) using a minimal number of cameras, as opposed to the present state of the art approaches which require either several tens of cameras or high definition devices. A 3D video consists of a sequence of 3D models in motion captured by a surrounding set of video cameras. The result is a video where observers can choose freely their viewpoints. It is a markerless motion capture system where subjects do not need to wear special equipment. Hence, this system suits to a very wide range of applications (e.g. entertainment, medicine, sports, and so on). The 3D models are obtained using image-based multi-view stereo reconstruction techniques (or MVS). The performance of MVS relies on the quality and quantity of images taken from different viewpoints. As stereo correspondences have to be found between the images, the reconstruction fails in the case of weak stereo photo-consistency due to lack of camera views or lighting variations: consistent information is necessary.
Tony Tung, Takashi Matsuyama
SIGGRAPH ASIA Sketches2
2009 Difference sphere: An approach to near light source estimation
Takeshi Takai, Atsuto Maki, Koichiro Niinuma, Takashi Matsuyama
Comput. Vis. Image Underst.4
2009 The Multiple-Camera 3-D Production Studio
abstract
Multiple-camera systems are currently widely used in research and development as a means of capturing and synthesizing realistic 3-D video content. Studio systems for 3-D production of human performance are reviewed from the literature, and the practical experience gained in developing prototype studios is reported across two research laboratories. System design should consider the studio backdrop for foreground matting, lighting for ambient illumination, camera acquisition hardware, the camera configuration for scene capture, and accurate geometric and photometric camera calibration. A ground-truth evaluation is performed to quantify the effect of different constraints on the multiple-camera system in terms of geometric accuracy and the requirement for high-quality view synthesis. As changing camera height has only a limited influence on surface visibility, multiple-camera sets or an active vision system may be required for wide area capture, and accurate reconstruction requires a camera baseline of 25deg, and the achievable accuracy is 5-10-mm at current camera resolutions. Accuracy is inherently limited, and view-dependent rendering is required for view synthesis with sub-pixel accuracy where display resolutions match camera resolutions. The two prototype studios are contrasted and state-of-the-art techniques for 3-D content production demonstrated.
Jonathan Starck, Atsuto Maki, Shohei Nobuhara, Adrian Hilton 0001, Takashi Matsuyama
IEEE Trans. Circuits Syst. Video Technol.5
2008 Simultaneous super-resolution and 3D video using graph-cuts
abstract
This paper presents a new method to increase the quality of 3D video, a new media developed to represent 3D objects in motion. This representation is obtained from multi-view reconstruction techniques that require images recorded simultaneously by several video cameras. All cameras are calibrated and placed around a dedicated studio to fully surround the models. The limited quality and quantity of cameras may produce inaccurate 3D model reconstruction with low quality texture. To overcome this issue, first we propose super-resolution (SR) techniques for 3D video: SR on multi-view images and SR on single-view video frames. Second, we propose to combine both super-resolution and dynamic 3D shape reconstruction problems into a unique Markov Random Field (MRF) energy formulation. The MRF minimization is performed using graph-cuts. Thus, we jointly compute the optimal solution for super-resolved texture and 3D shape model reconstruction. Moreover, we propose a coarse-to-fine strategy to iteratively produce 3D video with increasing quality. Our experiments show the accuracy and robustness of the proposed technique on challenging 3D video sequences.
Tony Tung, Shohei Nobuhara, Takashi Matsuyama
CVPR3
2008 Person-independent face tracking based on dynamic AAM selection
abstract
We have developed a high-precision method that selects an appropriate model of a video image in order to track an unknown face in front of a large display. Currently, Active Appearance Models (AAMs) are used to track non-rigid objects, such as a faces, because the models efficiently learn the correlation between shape and texture. The problem with an AAM is that when it tracks an unknown face, excessive training data increases tracking errors because there is an intermediate model size beyond which the reduction in fitting performance outweighs the gains from any improved representational power of the model. To increases the accuracy with which an unknown face is tracked, we built clustered models from training datasets and select a cluster that includes a face which is similar to the unknown face. Our method of clustering and cluster selecting is based on the Mutual Subspace Method (MSM). We demonstrated the effectiveness of our method by using the leave-one-out cross-validation.
Akihiro Kobayashi, Junji Satake, Takatsugu Hirayama, Hiroaki Kawashima, Takashi Matsuyama
FG5
2008 Complex human motion estimation using visibility
abstract
This paper presents a novel algorithm for estimating complex human motion from 3D video. We base our algorithm on a model-based approach which uses a complete surface mesh of a 3D human model to be matched with 3D video data. This type of method usually works well against partly incomplete input data. However, it fails to estimate what we call ldquocomplex motionrdquo: where some parts of the body touch each other for a long period. It is because the touching deteriorates the visibility of the neighbouring surface, which causes matching failures. In order to solve this problem, we introduce a ldquovisibilityrdquo measure for each mesh vertex that represents how it is occluded or missed on the observed surface. Using the ldquovisibilityrdquo we selectively suppress the outliers caused by low observability while traditional surface matching algorithms try to find corresponding area for the entire surface and cannot converge to the real posture by definition. Our algorithm shows improvements over naive surface matching algorithm on both synthesized and real 3D video.
Tomoyuki Mukasa, Arata Miyamoto, Shohei Nobuhara, Atsuto Maki, Takashi Matsuyama
FG5
2008 Tracking features on a moving object using local image bases
abstract
This paper presents a new method for tracking feature points on the textureless surface of a moving object. We employ a local image basis as a descriptor of each point for dealing with intensity variances due to the relative motion of the object to the light source. In particular, we propose to adaptively update the basis for enhancing its capability as the tracking proceeds. We show that the performance of the method further improves when it is coupled with Harris feature detector at a very large integration scale.
Atsuto Maki, Yosuke Hatanaka, Takashi Matsuyama
ICPR3
2007 Inter-Reflection Compensation for Immersive Projection Display
abstract
This paper proposes an effective method for compensating inter-reflection in immersive projection displays (IPDs). Because IPDs project images onto a screen, which surrounds a viewer, we have perform out both geometric and photometric corrections. Our method compensates inter-reflection on the screen. It requires no special device, and approximates both diffuse and specular reflections on the screen using block-based photometric calibration.
Hitoshi Habe, Nobuo Saeki, Takashi Matsuyama
CVPR3
2007 Topology matching for 3D video compression
abstract
This paper presents a new technique to reduce the storage cost of high quality 3D video. In 3D video, a sequence of 3D objects represents scenes in motion. Every frame is composed by one or several accurate 3D meshes with attached high fidelity properties such as color and texture. Each frame is acquired at video rate. The entire video sequence requires a huge amount of free disk space. To overcome this issue, we propose an original approach using Reeb graphs, which are well-known topology based shape descriptors. In particular, we take advantage of the augmented multiresolution Reeb graph properties to store the relevant information of the 3D model of each frame. This graph structure has shown its efficiency as a motion descriptor, being able to track similar nodes all along the 3D video sequence. Therefore we can describe and reconstruct the 3D models of all frames with a very low-cost data size. The algorithm has been implemented as a fully automatic 3D video compression system. Our experiments show the robustness and accuracy of the proposed technique by comparing reconstructed sequences against challenging real ones.
Tony Tung, Francis J. M. Schmitt, Takashi Matsuyama
CVPR3
2007 Phase-Based Feature Matching Under Illumination Variances
Masaaki Nishino, Atsuto Maki, Takashi Matsuyama
IEA/AIE3
2006 Parallel Pipeline Volume Intersection for Real-Time 3D Shape Reconstruction on a PC Cluster
abstract
The human activity monitoring is one of the major tasks in the field of computer vision. Recently, not only the 2D images but also 3D shapes of a moving person are desired in kinds of cases, such as motion analysis, security monitoring, 3D video creation and so on. In this paper, we propose a parallel pipeline system on a PC cluster for reconstructing the 3D shape of a moving person in real-time. For the 3D shape reconstruction, we have extended the volume intersection method to the 3-base-plane volume intersection. By thus extension, the computation is accelerated greatly for arbitrary camera layouts. We also parallelized the 3-base-plane method and implemented it on a PC cluster. On each node, the pipeline processing is adopted to improve the throughput. To decrease the CPU idle time caused by I/O processing, image capturing, communications over nodes and so on, we implement the pipeline using multiple threads. So that, all stages can be executed concurrently. However, there exists resource conflicts between stages in a real system. To avoid the conflicts while keeping high percentage of CPU running time, we propose a tree structured thread control model. As a result, We achieve the performance as obtaining the full 3D volumes of a moving person at about 12 frames per second, where the voxel size is 5×5×5 [mm^3]. The effectiveness of the thread tree model in such real-time computation is also proved by the experimental results.
Osamu Takizawa, Takashi Matsuyama
ICVS3
2005 Measurement of Human Concentration with Multiple Cameras
Kazuhiko Sumi, Koichi Tanaka, Takashi Matsuyama
KES (4)3
2005 Real-time cooperative multi-target tracking by communicating active vision agents
Norimichi Ukita, Takashi Matsuyama
Comput. Vis. Image Underst.2
2004 Difference Sphere: An Approach to Near Light Source Estimation
Takeshi Takai, Koichiro Niinuma, Atsuto Maki, Takashi Matsuyama
CVPR (1)4
2004 Real-time 3D shape reconstruction, dynamic 3D mesh deformation, and high fidelity visualization for 3D video
Takashi Matsuyama, Takeshi Takai, Shohei Nobuhara
Comput. Vis. Image Underst.1
2004 Real-time dynamic 3-D object shape reconstruction and high-fidelity texture mapping for 3-D video
abstract
Three-dimensional (3-D) video is a real 3-D movie recording the object's full 3-D shape, motion, and precise surface texture. This paper first proposes a parallel pipeline processing method for reconstructing a dynamic 3-D object shape from multiview video images, by which a temporal series of full 3-D voxel representations of the object behavior can be obtained in real time. To realize the real-time processing, we first introduce a plane-based volume intersection algorithm: first represent an observable 3-D space by a group of parallel plane slices, then back-project observed multiview object silhouettes onto each slice, and finally apply two-dimensional silhouette intersection on each slice. Then, we propose a method to parallelize this algorithm using a PC cluster, where we employ five-stage pipeline processing in each PC as well as slice-by-slice parallel silhouette intersection. Several results of the quantitative performance evaluation are given to demonstrate the effectiveness of the proposed methods. In the latter half of the paper, we present an algorithm of generating video texture on the reconstructed dynamic 3-D object surface. We first describe a naive view-independent rendering method and show its problems. Then, we improve the method by introducing image-based rendering techniques. Experimental results demonstrate the effectiveness of the improved method in generating high fidelity object images from arbitrary viewpoints.
Takashi Matsuyama, Takeshi Takai, Toshikazu Wada
IEEE Trans. Circuits Syst. Video Technol.1
2002 Real-time multitarget tracking by a cooperative distributed vision system
abstract
Target detection and tracking is one of the most important and fundamental technologies to develop real-world computer vision systems such as security and traffic monitoring systems. This paper first categorizes target tracking systems based on characteristics of scenes, tasks, and system architectures. Then we present a real-time cooperative multitarget tracking system. The system consists of a group of active vision agents (AVAs), where an AVA is a logical model of a network-connected computer with an active camera. All AVAs cooperatively track their target objects by dynamically exchanging object information with each other With this cooperative tracking capability, the system as a whole can track multiple moving objects persistently even under complicated dynamic environments in the real world. In this paper we address the technologies employed in the system and demonstrate their effectiveness.
Takashi Matsuyama, Norimichi Ukita
Proc. IEEE1
2000 Dynamic Memory: Architecture for Real Time Integration of Visual Perception, Camera Action, and Network Communication
abstract
In a Cooperative Distributed Vision system a group of communicating Active Vision Agents (AVA, in short, i.e. real time image processor with an active video camera and high speed network interface) cooperate to fulfil a meaningful task such as moving object tracking and dynamic scene visualization. A key issue to design and implement an AVA rests in the dynamic integration of Visual Perception, Camera Action, and Network Communication. This paper proposes a novel dynamic system architecture named Dynamic Memory Architecture, where perception, action, and communication modules share what we call the Dynamic Memory. It maintains not only temporal histories of state variables such as pan-tilt angles of the camera and the target object location but also their predicted values in the future. Perception, action, and communication modules are implemented as parallel processes which dynamically read from and write into the memory according to their own individual dynamics. The dynamic memory supports such asynchronous dynamic interactions (i.e. data exchanges between the modules) without wasting time for synchronization. This no-wait asynchronous module interaction capability greatly facilitates the implementation of real time reactive systems such as moving object tracking. Moreover, the dynamic memory supports the virtual synchronization between multiple AVAs, which facilitates the cooperative object tracking by communicating AVAs. A prototype system for real time moving object tracking demonstrated the effectiveness of the proposed idea.
Takashi Matsuyama, Shinsaku Hiura, Toshikazu Wada, Kazuyuki Murase, A. Yoshioka
CVPR1
2000 Human Head Tracking Using Adaptive Appearance Models with a Fixed-Viewpoint Pan-Tilt-Zoom Camera
abstract
We propose a method for detecting and tracking a human head in real time from an image sequence. The proposed method has three advantages: (1) we employ a fixed-viewpoint pan-tilt-zoom camera to acquire image sequences; with the camera, we eliminate the variations in the head appearance due to camera rotations with respect to the viewpoint; (2) we prepare a variety of contour models of the head appearances and relate them to the camera parameters; this allows us to adaptively select the model to deal with the variations in the head appearance due to human activities; (3) we use the model parameters obtained by detecting the head in the previous image to estimate those to be fitted in the current image; this estimation facilitates computational time for the head detection. Accordingly, the accuracy of the detection and required computational time are both improved and, at the same time, the robust head detection and tracking are realized in almost real time. Experimental results in the real situation show the effectiveness of our method.
Kiyotake Yachi, Toshikazu Wada, Takashi Matsuyama
FG3
2000 Multilinear Relationships between the Coordinates of Corresponding Image Conics
abstract
This paper presents a study, based on conic correspondences, on the relationship between multiple images acquired by uncalibrated cameras. Representing image conics as points in the five-dimensional projective space allows one to handle image conics in the same way as image points. We show that the coordinates of corresponding image conics satisfy the multilinear constraints, as shown in the case for points and lines. To be more specific, the coordinates of two corresponding image conics satisfy bilinear constraints. When a third image comes in, the coordinates of three corresponding image conics satisfy trilinear constraints. Moreover, these constraints are naturally extended to the case where more images are available.
Akihiro Sugimoto, Takashi Matsuyama
ICPR2
2000 Incremental Observable-Area Modeling for Cooperative Tracking
abstract
We propose an observable-area model of the scene for real-time cooperative object tracking by multiple cameras. The knowledge of partners' abilities is necessary for cooperative action whatever task is defined. In particular, for the tracking a moving object in the scene, every active vision agent (AVA), a rational model of the network-connected computer with an active camera, should therefore know the area in the scene that is observable by each AVA. Each AVA should then decide its target object and gazing direction taking into account other AVAs' actions. To realize such a cooperative gazing, the system gathers all the observable-area information to incrementally generate the observable-area model at each frame during the tracking. Hence, the system cooperatively tracks the object by utilizing both the observable-area model and the object's motion estimated at each frame. Experimental results demonstrate the effectiveness of the cooperation among the AVAs with the help of the proposed observable-area model.
Norimichi Ukita, Takashi Matsuyama
ICPR2
2000 Multiobject Behavior Recognition by Event Driven Selective Attention Method
abstract
This paper presents a multiobject behaviour recognition approach based on assumption generation and verification, i.e., feasible assumptions about the present behaviors consistent with the input image and behavior models are dynamically generated and verified by finding their supporting evidence in input images. This can be realized by an architecture called the selective attention model, which consists of a state-dependent event detector and an event sequence analyzer. The former detects image variation (event) in a limited image region (focusing region), which is not affected by occlusions and outliers. The latter analyzes sequences of detected events and activates all feasible states representing assumptions about multiobject behaviors. We further extend the system by introducing colored-token propagation to discriminate different objects in state space, and integration of multiviewpoint image sequences to disambiguate the single-view recognition results. Extensive experiments of human behavior recognition in real world environments demonstrate the soundness and robustness of our architecture.
Toshikazu Wada, Takashi Matsuyama
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Depth Measurement by the Multi-Focus Camera
abstract
In this paper, we first introduce the multi-focus camera, a new image sensor used for depth from defocus (DFD) range measurement. It can capture three images with different focus values simultaneously. We then propose two different depth measurement methods using the camera. The first method, an augmented version of the one proposed by N. Asada et al. (1998), employs a noniterative optimization process to compute depth values on edge points. The second one incorporates a coded aperture with the camera; and applies model-based pattern matching to estimate depth values of textured surfaces. Here we propose two types of coded apertures and corresponding analysis algorithms: 1D Fourier analysis to acquire a depth map and a blur-free image from three defocused images taken with a pair of pinholes, and 2D convolution based model matching for the fast and precise depth measurement using a coded aperture with four pinholes. Experimental results showed that the multi-focus camera works well as a practical DFD range sensor and that the coded apertures much improve its range estimation capability for real world scenes.
Shinsaku Hiura, Takashi Matsuyama
CVPR2
1998 Appearance Based Behavior Recognition by Event Driven Selective Attention
abstract
Most of behavior recognition methods proposed so far share the limitations of bottom-up analysis, and single-object assumption; the bottom-up analysis can be confused by erroneous and missing image features and the single-object assumption prevents us from analyzing image sequences including multiple moving objects. This paper presents a robust behavior recognition method free from these limitations. Our method is best characterized by 1) top-down image feature extraction by selective attention mechanism, 2) object discrimination by colored-token propagation, and 3) integration of multi-viewpoint images. Extensive experiments of human behavior recognition in real world environments demonstrate the soundness and robustness of our method.
Toshikazu Wada, Takashi Matsuyama
CVPR2
1998 Robust color segmentation using the dichromatic reflection model
abstract
This paper proposes a robust color segmentation method for real-world scenes. The robustness comes from a physics-based color analysis and the use of robust statistics. We analyze data distributions in the RGB color space to identify (non-Gaussian) object color clusters which conform to the dichromatic reflection model. Such a physics-based approach enables the detection of diffusion and specular interface reflections as well as body reflection, whose mixtures often confuse traditional statistics-based color segmentation algorithms. Experimental results show that accurate image segmentation can be realized and object colors can be correctly estimated without bias.
Chun-Kiat Ong, Takashi Matsuyama
ICPR2
1998 Edge and Depth from Focus
Naoki Asada, Hisanaga Fujiwara, Takashi Matsuyama
Int. J. Comput. Vis.3
1998 Seeing Behind the Scene: Analysis of Photometric Properties of Occluding Edges by the Reversed Projection Blurring Model
abstract
This paper analyzes photometric properties of occluding edges and proves that an object surface behind a nearer object is partially observable beyond the occluding edges. We first discuss a limitation of the image blurring model using the convolution, and then present an optical flux based blurring model named the reversed projection blurring (RPB) model. Unlike the multicomponent blurring model proposed by Nguyen et al., the RPB model enables us to explore the optical phenomena caused by a shift-variant point spread function that appears at a depth discontinuity. Using the RPB model, theoretical analysis of occluding edge properties are given and two characteristic phenomena are shown: (1) a blurred occluding edge produces the same brightness profiles as would be predicted for a surface edge on the occluding object when the occluded surface radiance is uniform and (2) a nonmonotonic brightness transition would be observed in blurred occluding edge profiles when the occluded object has a surface edge. Experimental results using real images have demonstrated the validity of the RPB model as well as the observability of the characteristic phenomena of blurred occluding edges.
Naoki Asada, Hisanaga Fujiwara, Takashi Matsuyama
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 Shape from Shading with Interreflections Under a Proximal Light Source: Distortion-Free Copying of an Unfolded Book
Toshikazu Wada, Hiroyuki Ukida, Takashi Matsuyama
Int. J. Comput. Vis.3
1997 Cooperative Spatial Reasoning for Image Understanding
abstract
Spatial Reasoning, reasoning about spatial information (i.e. shape and spatial relations), is a crucial function of image understanding and computer vision systems. This paper proposes a novel spatial reasoning scheme for image understanding and demonstrates its utility and effectiveness in two different systems: region segmentation and aerial image understanding systems. The scheme is designed based on a so-called Multi-Agent/Cooperative Distributed Problem Solving Paradigm, where a group of intelligent agents cooperate with each other to fulfill a complicated task. The first part of the paper describes a cooperative distributed region segmentation system, where each region in an image is regarded as an agent. Starting from seed regions given at the initial stage, region agents deform their shapes dynamically so that the image is partitioned into mutually disjoint regions. The deformation of each individual region agent is realized by the snake algorithm14 and neighboring region agents cooperate with each other to find common region boundaries between them. In the latter part of the paper, we first give a brief description of the cooperative spatial reasoning method used in our aerial image understanding system SIGMA. In SIGMA, each recognized object such as a house and a road is regarded as an agent. Each agent generates hypotheses about its neighboring objects to establish spatial relations and to detect missing objects. Then, we compare its reasoning method with that used in the region segmentation system. We conclude the paper by showing further utilities of the Multi-gent/Cooperative Distributed Problem Solving Paradigm for image understanding.
Takashi Matsuyama, Toshikazu Wada
Int. J. Pattern Recognit. Artif. Intell.1
1996 Appearance sphere: background model for pan-tilt-zoom camera
abstract
Background subtraction is a simple and effective method to detect anomalous regions in images. In spite of the effectiveness, it cannot be used with an active (moving) camera head because the background image varies with camera-parameter control. This paper presents a background subtraction method with pan-tilt-zoom control. The proposed method consists of an omnidirectional background model called appearance sphere and parallax free sensing. Based on this model, precise background images can be generated and background subtraction can be performed for any combination of pan-tilt-zoom parameters without restoring 3D scene information.
Toshikazu Wada, Takashi Matsuyama
ICPR2
1995 Seeing Behind the Scene: Analysis of Photometric Properties of Occluding Edges by the Reversed Projection Blurring Model
abstract
This paper analyzes photometric properties of occluding edges and proves that (1) we can observe surface edges on the farther object located close to the occluding edge even if they are occluded by the nearer object, (2) the image of an occluding edge coincides with that of a surface edge on the nearer object if the brightness of the farther object is uniform around the occluding edge. First, we propose a blurring model named the reversed projection blurring model to analyze photometric properties of blurring phenomena of an occluding edge. Using this model, the theoretical proof of the two properties mentioned above is given. Finally, experimental results in real world environments demonstrate the validity of our blurring model as well as the observability of the photometric properties of occluding edges.>
Naoki Asada, Hisanaga Fujiwara, Takashi Matsuyama
ICCV3
1995 Shape from Shading with Interreflections under Proximal Light Source: 3D Shape Reconstruction of Unfolded Book Surface from a Scanner Image
abstract
We address the problem to recover the 3D shape of an unfolded book surface from the shading information in a scanner image. From a technical point of view, this shape from shading problem in real world environments is characterized by (1) proximal light source, (2) interreflections, (3) moving light source, (4) specular reflection, and (5) nonuniform albedo distribution. Taking all these factors into account, we first formulate the problem based on an iterative nonlinear optimization scheme. Then we introduce piecewise polynomial models of the 3D shape. Image restoration experiments for a real book surface demonstrated that geometric and photometric distortions are almost completely removed by the proposed method.>
Toshikazu Wada, Hiroyuki Ukida, Takashi Matsuyama
ICCV3
1995 Geometric Theorem Proving by Integrated Logical and Algebraic Reasoning
Takashi Matsuyama, Tomoaki Nitta
Artif. Intell.1
1992 Color image analysis by varying camera aperture
abstract
A set of images taken by systematically varying camera parameters conveys useful information which cannot or can hardly be extracted from a single image. This paper proposes a method to obtain the reliable color information, i.e. chromaticity and brightness from multiple color images taken with different aperture sizes, i.e. multi-iris color images. Based on the fundamental characteristics of optical systems, the authors first define an evaluation function to measure the reliability of color information at each pixel. They then synthesize a color image which consists of those pixels with the most reliable color information. In the experiments, they quantitatively examine the stability and reliability of color information in the synthesized image. They also demonstrate the usefulness of multi-iris color images: estimation of light source chromaticity and recovery of depth information from blurred edges.>
Naoki Asada, Takashi Matsuyama
ICPR (1)2
1992 γ-ω Hough transform-elimination of quantization noise and linearization of voting curves in the ρ-θ parameter space
abstract
It is known that the rho - theta parameter space has inherent bias, and it has been treated as the appearance of the white noise in the image space. In this paper, the authors first show that the bias is caused by the uniform quantization of the parameter space. To eliminate the bias, a new parameter gamma representing a nonuniform quantization along the rho -axis is introduced and the gamma - theta parameter space is constructed. In this space, the uniform quantization does not introduce any bias. Then, by a nonlinear transformation of theta , the gamma - omega parameter space is derived, in which a voting curve becomes a pair of straight lines preserving the unbiasedness.>
Toshikazu Wada, Takashi Matsuyama
ICPR (3)2
1991 Random Closed Sets: a Unified Approach to the Representation of Imprecision and Uncertainty
Philippe Quinio, Takashi Matsuyama
ECSQARU2
1989 Expert systems for image processing: Knowledge-based composition of image analysis processes
Takashi Matsuyama
Comput. Vis. Graph. Image Process.1
1986 Hypothesis integration in image understanding systems
Vincent Shang-Shouq Hwang, Larry Davis 0001, Takashi Matsuyama
Comput. Vis. Graph. Image Process.3
1985 SIGMA: A Framework for Image Understanding - Integration of Bottom-Up and Top-Down Analysis
Takashi Matsuyama, Vincent Shang-Shouq Hwang
IJCAI1
1984 A file organization for geographic information systems based on spatial proximity
Takashi Matsuyama, Le Viet Hao, Makoto Nagao
Comput. Vis. Graph. Image Process.1
1983 Structural analysis of natural textures by Fourier transformation
Takashi Matsuyama, Shu-Ichi Miura, Makoto Nagao
Comput. Vis. Graph. Image Process.1
1982 A structural analyzer for regularly arranged textures
Takashi Matsuyama, Kinjiro Saburi, Makoto Nagao
Comput. Graph. Image Process.1
1979 Structural Analysis of Complex Aerial Photographs
Makoto Nagao, Takashi Matsuyama, Hisayuki Mori
IJCAI2