Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Steven Zhiying Zhou

dblp:66/8417 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
3since 2021 · last 2022
0000-0001-8455-1915ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 since 2021Artificial intelligence and machine learning · 6Human-computer interaction and ubiquitous computing · 6

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
3D vision · 54% Video understanding and tracking · 43% Learning paradigms · 3%
Computer graphics and multimedia
8 papers
Virtual and augmented reality · 37% Image and video processing · 28% Multimedia analysis and retrieval · 14%
Human-computer interaction and pervasive computing
2 papers
Interaction techniques and input · 50% Collaborative and social computing · 25% User interface design and tools · 25%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › stereo vision
stereo matching
1.122022
ToF and Stereo Data Fusion Using Dynamic Search Range Stereo Matching · IEEE Trans. Multim. 2022
Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical Flow · IEEE Trans. Image Process. 2021
Computer vision › 3D vision
depth estimation
0.612022
ToF and Stereo Data Fusion Using Dynamic Search Range Stereo Matching · IEEE Trans. Multim. 2022
Computer vision › Video understanding and tracking
video anomaly detection
0.612022
Contrastive Attention for Video Anomaly Detection · IEEE Trans. Multim. 2022
Computer vision › Video understanding and tracking › video anomaly detection
weakly supervised video anomaly detection
0.612022
Contrastive Attention for Video Anomaly Detection · IEEE Trans. Multim. 2022
Computer vision › 3D vision › feature matching › hierarchical matching
coarse-to-fine matching
0.512021
Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical Flow · IEEE Trans. Image Process. 2021
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction
0.212015
Simultaneous video defogging and stereo reconstruction · CVPR 2015
Image and video processing
image restoration
0.212015
Simultaneous video defogging and stereo reconstruction · CVPR 2015
Data mining
clustering
0.212014
SCAMS: Simultaneous Clustering and Model Selection · CVPR 2014
Mathematical optimization › continuous optimization › convex optimization › proximal methods
alternating direction method of multipliers
0.212014
SCAMS: Simultaneous Clustering and Model Selection · CVPR 2014
Machine learning › Learning paradigms
multiple instance learning
0.212022
Contrastive Attention for Video Anomaly Detection · IEEE Trans. Multim. 2022
Computer vision › Video understanding and tracking
motion segmentation
0.212013
Perspective Motion Segmentation via Collaborative Clustering · ICCV 2013
Computer vision › Video understanding and tracking › motion segmentation
multi-frame motion segmentation
0.212013
Perspective Motion Segmentation via Collaborative Clustering · ICCV 2013
Computer vision › 3D vision
structure from motion
0.212013
Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013
Image and video processing › image matting
alpha matting
0.212013
Image Matting with Local and Nonlocal Smooth Priors · CVPR 2013
Virtual and augmented reality › mixed reality
diminished reality
0.212013
Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013
Image and video processing
image matting
0.212013
Image Matting with Local and Nonlocal Smooth Priors · CVPR 2013
Visual content generation and editing › image completion
object removal
0.212013
Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013
Computer vision › 3D vision › motion estimation
optical flow
0.112021
Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical Flow · IEEE Trans. Image Process. 2021
Geometric modeling and processing
3d reconstruction
0.112010
Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010
Virtual and augmented reality › tracking
model-based tracking
0.112010
Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010
Multimedia analysis and retrieval
object tracking
0.112010
Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010
Virtual and augmented reality
tracking and registration
0.112010
Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010
Collaborative and social computing
collaborative design
0.112010
MTMR: A conceptual interior design framework integrating Mixed Reality with the Multi-Touch tabletop interface · ISMAR 2010
Interaction techniques and input › input sensing › tracking
hand tracking
0.112010
A real-time multi-cue hand tracking algorithm based on computer vision · VR 2010
Interaction techniques and input › input sensing › tracking › hand tracking
vision-based hand tracking
0.112010
A real-time multi-cue hand tracking algorithm based on computer vision · VR 2010
Virtual and augmented reality
augmented reality
0.112009
Consistent real-time lighting for virtual objects in augmented reality · ISMAR 2009
Virtual and augmented reality › visual realism
illumination consistency
0.112009
Consistent real-time lighting for virtual objects in augmented reality · ISMAR 2009
Rendering
shadow rendering
0.112009
Consistent real-time lighting for virtual objects in augmented reality · ISMAR 2009
Computer vision › 3D vision
object pose estimation
0.012013
Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013
Computer vision › 3D vision
visual localization
0.012013
Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013

Methods — techniques the papers use, named apart from their topics

markov random field · 0.8matting laplacian · 0.6upsampling module · 0.6end-to-end neural network · 0.6contrastive attention · 0.6coarse-to-fine matching · 0.6attention consistency loss · 0.6neural network · 0.5differentiable upsampling · 0.5photo-consistency · 0.4sparsity regularization · 0.4rank minimization · 0.4frobenius inner product · 0.4ADMM · 0.4trajectory co-saliency · 0.2texture synthesis · 0.2plane warping · 0.2graph model · 0.2
YearPublicationVenuePosition
2022 Contrastive Attention for Video Anomaly Detection
abstract
We consider weakly-supervised video anomaly detection in this work. This task aims to learn to localize video frames containing anomaly events with only binary video-level annotation,i.e., anomaly vs. normal. Traditional approaches usually formulate it as a multiple instance learning problem, which ignore the intrinsic data imbalance issue that positive samples are very scarce compared to negative ones. In this paper, we focus on addressing this issue to boost detection performance further. We develop a new light-weight anomaly detection model that fully utilizes enough normal videos to train a classifier with a good discriminative ability for normal videos, and we employ it to improve the selectivity for anomalous segments and filter out normal segments. Specifically, in addition to boosting anomalous prediction, a novel contrastive attention module additionally produces a converted normal feature from anomalous video to refined anomalous predictions by maximizing the classifier making a mistake. Moreover, to remove the stubborn normal segments selected by the attention module, we also design an attention consistency loss to employ the classifier with high confidence for normal features to guide the attention module. Extensive experiments on two large-scale datasets, UCF-Crime, ShanghaiTech and XD-Violence, clearly demonstrate that our model largely improves frame-level AUC over the state-of-the-art. Code is released athttps://github.com/changsn/Contrastive-Attention-for-Video-Anomaly-Detection.
Shuning Chang, Shengmei Shen, Jiashi Feng, Steven Zhiying Zhou
IEEE Trans. Multim.5
2022 ToF and Stereo Data Fusion Using Dynamic Search Range Stereo Matching
abstract
Time-of-Flight (ToF) sensors and stereo vision systems are both widely used for capturing depth data. They have some complementary strengths and limitations, which have been exploited in prior research to produce more accurate depth maps by fusing data from the two sources. However, among these diverse data fusion approaches, none of them provides an end-to-end neural network solution. In this work, we propose the first end-to-end ToF and stereo data fusion network using the coarse-to-fine matching framework, where the prior of ToF depth is integrated into the stereo matching process by constraining the search range of stereo matching within an interval around the ToF camera depth measurement. We adopt a dynamic search range for each pixel according to an estimated ToF error map, which is more efficient and effective than a constant one when handling various errors. The ToF error map is estimated by the ToF error estimator branching out from the stereo matching network. Both ToF error estimation and stereo matching are performed in a joint framework, with the two tasks assisting each other mutually. We also propose an upsampling module to replace the naive bilinear upsampling in the coarse-to-fine stereo matching network, which reduces the error caused by the upsampling. The proposed deep network is trained end-to-end on synthetic datasets and generalizable to real-world datasets without further fine-tuning. Experimental results show that our fusion method achieves higher accuracy than either ToF or stereo alone, and outperforms state-of-the-art fusion methods on both synthetic and real data.
Yong Deng 0006, Jimin Xiao, Steven Zhiying Zhou
IEEE Trans. Multim.3
2021 Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical Flow
abstract
The Coarse-To-Fine (CTF) matching scheme has been widely applied to reduce computational complexity and matching ambiguity in stereo matching and optical flow tasks by converting image pairs into multi-scale representations and performing matching from coarse to fine levels. Despite its efficiency, it suffers from several weaknesses, such as tending to blur the edges and miss small structures like thin bars and holes. We find that the pixels of small structures and edges are often assigned with wrong disparity/flow in the upsampling process of the CTF framework, introducing errors to the fine levels and leading to such weaknesses. We observe that these wrong disparity/flow values can be avoided if we select the best-matched value among their neighborhood, which inspires us to propose a novel differentiable Neighbor-Search Upsampling (NSU) module. The NSU module first estimates the matching scores and then selects the best-matched disparity/flow for each pixel from its neighbors. It effectively preserves finer structure details by exploiting the information from the finer level while upsampling the disparity/flow. The proposed module can be a drop-in replacement of the naive upsampling in the CTF matching framework and allows the neural networks to be trained end-to-end. By integrating the proposed NSU module into a baseline CTF matching network, we design our Detail Preserving Coarse-To-Fine (DPCTF) matching network. Comprehensive experiments demonstrate that our DPCTF can boost performances for both stereo matching and optical flow tasks. Notably, our DPCTF achieves new state-of-the-art performances for both tasks - it outperforms the competitive baseline (Bi3D) by 28.8% (from 0.73 to 0.52) on EPE of the FlyingThings3D stereo dataset, and ranks first in KITTI flow 2012 benchmark. The code is available at https://github.com/Deng-Y/DPCTF.
Yong Deng 0006, Jimin Xiao, Steven Zhiying Zhou, Jiashi Feng
IEEE Trans. Image Process.3
2015 Simultaneous video defogging and stereo reconstruction
abstract
We present a method to jointly estimate scene depth and recover the clear latent image from a foggy video sequence. In our formulation, the depth cues from stereo matching and fog information reinforce each other, and produce superior results than conventional stereo or defogging algorithms. We first improve the photo-consistency term to explicitly model the appearance change due to the scattering effects. The prior matting Laplacian constraint on fog transmission imposes a detail-preserving smoothness constraint on the scene depth. We further enforce the ordering consistency between scene depth and fog transmission at neighboring points. These novel constraints are formulated together in an MRF framework, which is optimized iteratively by introducing auxiliary variables. The experiment results on real videos demonstrate the strength of our method.
Zhuwen Li, Robby T. Tan, Danping Zou, Steven Zhiying Zhou, Loong Fah Cheong
CVPR5
2014 Consistent Foreground Co-segmentation
Jiaming Guo, Loong Fah Cheong, Robby T. Tan, Steven Zhiying Zhou
ACCV (4)4
2014 SCAMS: Simultaneous Clustering and Model Selection
abstract
While clustering has been well studied in the past decade, model selection has drawn less attention. This paper addresses both problems in a joint manner with an indicator matrix formulation, in which the clustering cost is penalized by a Frobenius inner product term and the group number estimation is achieved by a rank minimization. As affinity graphs generally contain positive edge values, a sparsity term is further added to avoid the trivial solution. Rather than adopting the conventional convex relaxation approach wholesale, we represent the original problem more faithfully by taking full advantage of the particular structure present in the optimization problem and solving it efficiently using the Alternating Direction Method of Multipliers. The highly constrained nature of the optimization provides our algorithm with the robustness to deal with the varying and often imperfect input affinity matrices arising from different applications and different group numbers. Evaluations on the synthetic data as well as two real world problems show the superiority of the method across a large variety of settings.
Zhuwen Li, Loong Fah Cheong, Steven Zhiying Zhou
CVPR3
2013 Image Matting with Local and Nonlocal Smooth Priors
abstract
In this paper we propose a novel alpha matting method with local and nonlocal smooth priors. We observe that the manifold preserving editing propagation [4] essentially introduced a nonlocal smooth prior on the alpha matte. This nonlocal smooth prior and the well known local smooth prior from matting Laplacian complement each other. So we combine them with a simple data term from color sampling in a graph model for nature image matting. Our method has a closed-form solution and can be solved efficiently. Compared with the state-of-the-art methods, our method produces more accurate results according to the evaluation on standard benchmark datasets.
Xiaowu Chen 0001, Dongqing Zou, Steven Zhiying Zhou, Qinping Zhao
CVPR3
2013 Video Co-segmentation for Meaningful Action Extraction
abstract
Given a pair of videos having a common action, our goal is to simultaneously segment this pair of videos to extract this common action. As a preprocessing step, we first remove background trajectories by a motion-based figure ground segmentation. To remove the remaining background and those extraneous actions, we propose the trajectory co saliency measure, which captures the notion that trajectories recurring in all the videos should have their mutual saliency boosted. This requires a trajectory matching process which can compare trajectories with different lengths and not necessarily spatiotemporally aligned, and yet be discriminative enough despite significant intra-class variation in the common action. We further leverage the graph matching to enforce geometric coherence between regions so as to reduce feature ambiguity and matching errors. Finally, to classify the trajectories into common action and action outliers, we formulate the problem as a binary labeling of a Markov Random Field, in which the data term is measured by the trajectory co-saliency and the smoothness term is measured by the spatiotemporal consistency between trajectories. To evaluate the performance of our framework, we introduce a dataset containing clips that have animal actions as well as human actions. Experimental results show that the proposed method performs well in common action extraction.
Jiaming Guo, Zhuwen Li, Loong Fah Cheong, Steven Zhiying Zhou
ICCV4
2013 Perspective Motion Segmentation via Collaborative Clustering
abstract
This paper addresses real-world challenges in the motion segmentation problem, including perspective effects, missing data, and unknown number of motions. It first formulates the 3-D motion segmentation from two perspective views as a subspace clustering problem, utilizing the epipolar constraint of an image pair. It then combines the point correspondence information across multiple image frames via a collaborative clustering step, in which tight integration is achieved via a mixed norm optimization scheme. For model selection, we propose an over-segment and merge approach, where the merging step is based on the property of the ell_1-norm of the mutual sparse representation of two over-segmented groups. The resulting algorithm can deal with incomplete trajectories and perspective effects substantially better than state-of-the-art two-frame and multi-frame methods. Experiments on a 62-clip dataset show the significant superiority of the proposed idea in both segmentation accuracy and model selection.
Zhuwen Li, Jiaming Guo, Loong Fah Cheong, Steven Zhiying Zhou
ICCV4
2013 Diminished reality using appearance and 3D geometry of internet photo collections
abstract
This paper presents a new system level framework for Diminished Reality, leveraging for the first time both the appearance and 3D information provided by large photo collections on the Internet. Recent computer vision techniques have made it possible to automatically reconstruct 3-D structure-from-motion points from large and unordered photo collections. Using these point clouds and a prior provided by GPS, reasonably accurate 6 degree of freedom camera poses can be obtained, thus allowing localization. Once the camera (and hence the user) is correctly localized, photos depicting scenes visible from the user's viewpoint can be used to remove unwanted objects indicated by the user in the video sequences. Existing methods based on texture synthesis bring undesirable artifacts and video inconsistency when the background is heterogeneous; the task is rendered even harder for these methods when the background contains complex structures. On the other hand, methods based on plane warping fail when the background has arbitrary shape. Unlike these methods, our algorithm copes with these problems by making use of internet photos, registering them in 3D space and obtaining the 3D scene structure in an offline process. We carefully design the various components during the online phase so as to meet both speed and quality requirements of the task. Experiments on real data collected demonstrate the superiority of our system.
Zhuwen Li, Yuxi Wang 0002, Jiaming Guo, Loong Fah Cheong, Steven Zhiying Zhou
ISMAR5
2010 Model-based localization and drift-free user tracking for outdoor augmented reality
abstract
In this paper we present a novel model-based hybrid technique for user localization and drift-free tracking in urban environments. In outdoor augmented reality, instantaneous 6-DoF user localization is achieved with position and orientation sensors such as GPS and gyroscopes. Initial pose obtained with these sensors is dominated by large positional errors due to coarse granularity of GPS data. Subsequent tracking is also erroneous as gyroscopes are prone to drifts and often need recalibration. We propose to use model-to-image registration technique to refine initial rough estimate for accurate user localization. Large positional errors in user localization are mitigated by aligning silhouettes of the model with that of the camera image using shape context descriptors as they are invariant to translation, scale and rotational errors. Once initialized, drift-free tracking is achieved by combining frame-to-frame and model-to-frame feature tracking. Frame-to-frame tracking is done by matching corners whereas edges are used for model-to-frame silhouette tracking. Final camera pose is obtained with M-estimators.
Jayashree Karlekar, Steven Zhiying Zhou, Yuta Nakayama, Weiquan Lu, Loh Zhi Chang, Daniel Hii
ICME2
2010 Positioning, tracking and mapping for outdoor augmentation
abstract
This paper presents a novel approach for user positioning, robust tracking and online 3D mapping for outdoor augmented reality applications. As coarse user pose obtained from GPS and orientation sensors is not sufficient for augmented reality applications, sub-meter accurate user pose is then estimated by a one-step silhouette matching approach. Silhouette matching of the rendered 3D model and camera data is carried out with shape context descriptors as they are invariant to translation, scale and rotational errors, giving rise to a non-iterative registration approach. Once the user is correctly positioned, further tracking is carried out with camera data alone. Drifts associated with vision based approaches are minimized by combining different feature modalities. Robust visual tracking is maintained by fusing frame-to-frame and model-to-frame feature matches. Frame-to-frame tracking is accomplished with corner matching while edges are used for model-to-frame registration. Results from individual feature tracker are fused using a pose estimate obtained from an extended Kalman filter (EKF) and a weighted M-estimator. In scenarios where dense 3D models of the environment are not available, online 3D incremental mapping and tracking is proposed to track the user in unprepared environments. Incremental mapping prepares the 3D point cloud of the outdoor environment for tracking.
Jayashree Karlekar, Steven Zhiying Zhou, Weiquan Lu, Loh Zhi Chang, Yuta Nakayama, Daniel Hii
ISMAR2
2010 MTMR: A conceptual interior design framework integrating Mixed Reality with the Multi-Touch tabletop interface
abstract
This paper introduces a conceptual interior design framework - Multi-Touch Mixed Reality (MTMR), which integrates mixed reality with the multi-touch tabletop interface, to provide an intuitive and efficient interface for collaborative design and an augmented 3D view to users at the same time. Under this framework, multiple designers can carry out design work simultaneously on the top view displayed on the tabletop, while live video of the ongoing design work is captured and augmented by overlaying virtual 3D furniture models to their 2D virtual counterparts, and shown on a vertical screen in front of the tabletop. Meanwhile, the remote client's camera view of the physical room is augmented with the interior design layout in real time, that is, as the designers place, move, and modify the virtual furniture models on the tabletop, the client sees the corresponding life-size 3D virtual furniture models residing, moving, and changing in the physical room through the camera view on his/her screen. By adopting MTMR, which we argue may also apply to other kinds of collaborative work, the designers can expect a good working experience in terms of naturalness and intuitiveness, while the client can be involved in the design process and view the design result without moving around heavy furniture. By presenting MTMR, we hope to provide reliable and precise freehand interactions to mixed reality systems, with multi-touch inputs on tabletop interfaces.
Dong Wei 0004, Steven Zhiying Zhou, Du Xie
ISMAR2
2010 A real-time multi-cue hand tracking algorithm based on computer vision
abstract
Although hand tracking algorithm has been widely used in virtual reality and HCI system, it is still a challenging problem in vision-based research area. Due to the robustness and real-time requirements in VR applications, most hand tracking algorithms require special device to achieve satisfactory results. In this paper, we propose an easy-to-use and inexpensive approach to track the hands accurately with a single normal webcam. Outstretched hand is detected by contour & curvature based detection techniques to initialize the tracking region. Robust multi-cue hand tracking is then achieved by velocity-weighted features and color cue. Experiments show that the proposed multi-cue hand tracking approach achieves continuous real-time results even for the situation of cluttered background. The approach fulfills the speed and accuracy requirements of frontal-view vision-based human computer interactions.
Kangde Guo, Xing Tang 0002, Steven Zhiying Zhou
VR7
2010 Immersive Multiplayer Games With Tangible and Physical Interaction
abstract
In this paper, we present a new immersive multiplayer game system developed for two different environments, namely, virtual reality (VR) and augmented reality (AR). To evaluate our system, we developed three game applications-a first-person-shooter game (for VR and AR environments, respectively) and a sword game (for the AR environment). Our immersive system provides an intuitive way for users to interact with the VR or AR world by physically moving around the real world and aiming freely with tangible objects. This encourages physical interaction between players as they compete or collaborate with other players. Evaluation of our system consists of users' subjective opinions and their objective performances. Our design principles and evaluation results can be applied to similar immersive game applications based on AR/VR.
Jefry Tedjokusumo, Steven Zhiying Zhou, Stefan Winkler 0001
IEEE Trans. Syst. Man Cybern. Part A2
2009 Consistent real-time lighting for virtual objects in augmented reality
abstract
We present a technique for rendering realistic shadows of virtual objects in a mixed reality environment by recovering the light source distribution of a scene in real-time, through the segmentation and analysis of a known occluding object's shadows. A fiducial marker provides information about the position of the occluding object and the plane of the surface on which shadows are cast, and serves as the origin of a marker coordinate system. A new shadow segmentation approach is carried out on the shadow image and is able to recover geometrical information on multiple faint shadows. Using normalised iterative reinforcement, noise and artifacts can be suppressed in the final shadow map. The scene's light source distribution is then extrapolated using geometrical data from both the occluding object and its cast shadows. Virtual light sources in a game engine are used to mimic real light sources and achieve consistent illumination and increase the realism of the augmented reality scene.
Ryan Christopher Yeoh, Steven Zhiying Zhou
ISMAR2