EDBT 2026 Demo / reviewers in the wild / expert
Steven Zhiying Zhou
dblp:66/8417
· DBLP profile ↗
16ranked-venue papers
0as first author
3since 2021 · last 2022
0000-0001-8455-1915ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 since 2021Artificial intelligence and machine learning · 6Human-computer interaction and ubiquitous computing · 6
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 54% Video understanding and tracking · 43% Learning paradigms · 3% | |
| Computer graphics and multimedia
8 papers |
Virtual and augmented reality · 37% Image and video processing · 28% Multimedia analysis and retrieval · 14% | |
| Human-computer interaction and pervasive computing
2 papers |
Interaction techniques and input · 50% Collaborative and social computing · 25% User interface design and tools · 25% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › stereo vision
stereo matching |
1.1 | 2 | 2022 | ToF and Stereo Data Fusion Using Dynamic Search Range Stereo Matching · IEEE Trans. Multim. 2022 Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical Flow · IEEE Trans. Image Process. 2021 |
Computer vision › 3D vision
depth estimation |
0.6 | 1 | 2022 | ToF and Stereo Data Fusion Using Dynamic Search Range Stereo Matching · IEEE Trans. Multim. 2022 |
Computer vision › Video understanding and tracking
video anomaly detection |
0.6 | 1 | 2022 | Contrastive Attention for Video Anomaly Detection · IEEE Trans. Multim. 2022 |
Computer vision › Video understanding and tracking › video anomaly detection
weakly supervised video anomaly detection |
0.6 | 1 | 2022 | Contrastive Attention for Video Anomaly Detection · IEEE Trans. Multim. 2022 |
Computer vision › 3D vision › feature matching › hierarchical matching
coarse-to-fine matching |
0.5 | 1 | 2021 | Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical Flow · IEEE Trans. Image Process. 2021 |
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction |
0.2 | 1 | 2015 | Simultaneous video defogging and stereo reconstruction · CVPR 2015 |
Image and video processing
image restoration |
0.2 | 1 | 2015 | Simultaneous video defogging and stereo reconstruction · CVPR 2015 |
Data mining
clustering |
0.2 | 1 | 2014 | SCAMS: Simultaneous Clustering and Model Selection · CVPR 2014 |
Mathematical optimization › continuous optimization › convex optimization › proximal methods
alternating direction method of multipliers |
0.2 | 1 | 2014 | SCAMS: Simultaneous Clustering and Model Selection · CVPR 2014 |
Machine learning › Learning paradigms
multiple instance learning |
0.2 | 1 | 2022 | Contrastive Attention for Video Anomaly Detection · IEEE Trans. Multim. 2022 |
Computer vision › Video understanding and tracking
motion segmentation |
0.2 | 1 | 2013 | Perspective Motion Segmentation via Collaborative Clustering · ICCV 2013 |
Computer vision › Video understanding and tracking › motion segmentation
multi-frame motion segmentation |
0.2 | 1 | 2013 | Perspective Motion Segmentation via Collaborative Clustering · ICCV 2013 |
Computer vision › 3D vision
structure from motion |
0.2 | 1 | 2013 | Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013 |
Image and video processing › image matting
alpha matting |
0.2 | 1 | 2013 | Image Matting with Local and Nonlocal Smooth Priors · CVPR 2013 |
Virtual and augmented reality › mixed reality
diminished reality |
0.2 | 1 | 2013 | Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013 |
Image and video processing
image matting |
0.2 | 1 | 2013 | Image Matting with Local and Nonlocal Smooth Priors · CVPR 2013 |
Visual content generation and editing › image completion
object removal |
0.2 | 1 | 2013 | Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013 |
Computer vision › 3D vision › motion estimation
optical flow |
0.1 | 1 | 2021 | Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical Flow · IEEE Trans. Image Process. 2021 |
Geometric modeling and processing
3d reconstruction |
0.1 | 1 | 2010 | Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010 |
Virtual and augmented reality › tracking
model-based tracking |
0.1 | 1 | 2010 | Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010 |
Multimedia analysis and retrieval
object tracking |
0.1 | 1 | 2010 | Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010 |
Virtual and augmented reality
tracking and registration |
0.1 | 1 | 2010 | Positioning, tracking and mapping for outdoor augmentation · ISMAR 2010 |
Collaborative and social computing
collaborative design |
0.1 | 1 | 2010 | MTMR: A conceptual interior design framework integrating Mixed Reality with the Multi-Touch tabletop interface · ISMAR 2010 |
Interaction techniques and input › input sensing › tracking
hand tracking |
0.1 | 1 | 2010 | A real-time multi-cue hand tracking algorithm based on computer vision · VR 2010 |
Interaction techniques and input › input sensing › tracking › hand tracking
vision-based hand tracking |
0.1 | 1 | 2010 | A real-time multi-cue hand tracking algorithm based on computer vision · VR 2010 |
Virtual and augmented reality
augmented reality |
0.1 | 1 | 2009 | Consistent real-time lighting for virtual objects in augmented reality · ISMAR 2009 |
Virtual and augmented reality › visual realism
illumination consistency |
0.1 | 1 | 2009 | Consistent real-time lighting for virtual objects in augmented reality · ISMAR 2009 |
Rendering
shadow rendering |
0.1 | 1 | 2009 | Consistent real-time lighting for virtual objects in augmented reality · ISMAR 2009 |
Computer vision › 3D vision
object pose estimation |
0.0 | 1 | 2013 | Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013 |
Computer vision › 3D vision
visual localization |
0.0 | 1 | 2013 | Diminished reality using appearance and 3D geometry of internet photo collections · ISMAR 2013 |
Methods — techniques the papers use, named apart from their topics
markov random field · 0.8matting laplacian · 0.6upsampling module · 0.6end-to-end neural network · 0.6contrastive attention · 0.6coarse-to-fine matching · 0.6attention consistency loss · 0.6neural network · 0.5differentiable upsampling · 0.5photo-consistency · 0.4sparsity regularization · 0.4rank minimization · 0.4frobenius inner product · 0.4ADMM · 0.4trajectory co-saliency · 0.2texture synthesis · 0.2plane warping · 0.2graph model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Contrastive Attention for Video Anomaly DetectionabstractWe consider weakly-supervised video anomaly detection in this work. This task aims to learn to localize video frames containing anomaly events with only binary video-level annotation,i.e., anomaly vs. normal. Traditional approaches usually formulate it as a multiple instance learning problem, which ignore the intrinsic data imbalance issue that positive samples are very scarce compared to negative ones. In this paper, we focus on addressing this issue to boost detection performance further. We develop a new light-weight anomaly detection model that fully utilizes enough normal videos to train a classifier with a good discriminative ability for normal videos, and we employ it to improve the selectivity for anomalous segments and filter out normal segments. Specifically, in addition to boosting anomalous prediction, a novel contrastive attention module additionally produces a converted normal feature from anomalous video to refined anomalous predictions by maximizing the classifier making a mistake. Moreover, to remove the stubborn normal segments selected by the attention module, we also design an attention consistency loss to employ the classifier with high confidence for normal features to guide the attention module. Extensive experiments on two large-scale datasets, UCF-Crime, ShanghaiTech and XD-Violence, clearly demonstrate that our model largely improves frame-level AUC over the state-of-the-art. Code is released athttps://github.com/changsn/Contrastive-Attention-for-Video-Anomaly-Detection. Shuning Chang, Shengmei Shen, Jiashi Feng, Steven Zhiying Zhou |
IEEE Trans. Multim. | 5 |
| 2022 | ToF and Stereo Data Fusion Using Dynamic Search Range Stereo MatchingabstractTime-of-Flight (ToF) sensors and stereo vision systems are both widely used for capturing depth data. They have some complementary strengths and limitations, which have been exploited in prior research to produce more accurate depth maps by fusing data from the two sources. However, among these diverse data fusion approaches, none of them provides an end-to-end neural network solution. In this work, we propose the first end-to-end ToF and stereo data fusion network using the coarse-to-fine matching framework, where the prior of ToF depth is integrated into the stereo matching process by constraining the search range of stereo matching within an interval around the ToF camera depth measurement. We adopt a dynamic search range for each pixel according to an estimated ToF error map, which is more efficient and effective than a constant one when handling various errors. The ToF error map is estimated by the ToF error estimator branching out from the stereo matching network. Both ToF error estimation and stereo matching are performed in a joint framework, with the two tasks assisting each other mutually. We also propose an upsampling module to replace the naive bilinear upsampling in the coarse-to-fine stereo matching network, which reduces the error caused by the upsampling. The proposed deep network is trained end-to-end on synthetic datasets and generalizable to real-world datasets without further fine-tuning. Experimental results show that our fusion method achieves higher accuracy than either ToF or stereo alone, and outperforms state-of-the-art fusion methods on both synthetic and real data. Yong Deng 0006, Jimin Xiao, Steven Zhiying Zhou |
IEEE Trans. Multim. | 3 |
| 2021 | Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical FlowabstractThe Coarse-To-Fine (CTF) matching scheme has been widely applied to reduce computational complexity and matching ambiguity in stereo matching and optical flow tasks by converting image pairs into multi-scale representations and performing matching from coarse to fine levels. Despite its efficiency, it suffers from several weaknesses, such as tending to blur the edges and miss small structures like thin bars and holes. We find that the pixels of small structures and edges are often assigned with wrong disparity/flow in the upsampling process of the CTF framework, introducing errors to the fine levels and leading to such weaknesses. We observe that these wrong disparity/flow values can be avoided if we select the best-matched value among their neighborhood, which inspires us to propose a novel differentiable Neighbor-Search Upsampling (NSU) module. The NSU module first estimates the matching scores and then selects the best-matched disparity/flow for each pixel from its neighbors. It effectively preserves finer structure details by exploiting the information from the finer level while upsampling the disparity/flow. The proposed module can be a drop-in replacement of the naive upsampling in the CTF matching framework and allows the neural networks to be trained end-to-end. By integrating the proposed NSU module into a baseline CTF matching network, we design our Detail Preserving Coarse-To-Fine (DPCTF) matching network. Comprehensive experiments demonstrate that our DPCTF can boost performances for both stereo matching and optical flow tasks. Notably, our DPCTF achieves new state-of-the-art performances for both tasks - it outperforms the competitive baseline (Bi3D) by 28.8% (from 0.73 to 0.52) on EPE of the FlyingThings3D stereo dataset, and ranks first in KITTI flow 2012 benchmark. The code is available at https://github.com/Deng-Y/DPCTF. Yong Deng 0006, Jimin Xiao, Steven Zhiying Zhou, Jiashi Feng |
IEEE Trans. Image Process. | 3 |
| 2015 | Simultaneous video defogging and stereo reconstructionabstractWe present a method to jointly estimate scene depth and recover the clear latent image from a foggy video sequence. In our formulation, the depth cues from stereo matching and fog information reinforce each other, and produce superior results than conventional stereo or defogging algorithms. We first improve the photo-consistency term to explicitly model the appearance change due to the scattering effects. The prior matting Laplacian constraint on fog transmission imposes a detail-preserving smoothness constraint on the scene depth. We further enforce the ordering consistency between scene depth and fog transmission at neighboring points. These novel constraints are formulated together in an MRF framework, which is optimized iteratively by introducing auxiliary variables. The experiment results on real videos demonstrate the strength of our method. Zhuwen Li, Robby T. Tan, Danping Zou, Steven Zhiying Zhou, Loong Fah Cheong |
CVPR | 5 |
| 2014 | Consistent Foreground Co-segmentation
Jiaming Guo, Loong Fah Cheong, Robby T. Tan, Steven Zhiying Zhou |
ACCV (4) | 4 |
| 2014 | SCAMS: Simultaneous Clustering and Model SelectionabstractWhile clustering has been well studied in the past decade, model selection has drawn less attention. This paper addresses both problems in a joint manner with an indicator matrix formulation, in which the clustering cost is penalized by a Frobenius inner product term and the group number estimation is achieved by a rank minimization. As affinity graphs generally contain positive edge values, a sparsity term is further added to avoid the trivial solution. Rather than adopting the conventional convex relaxation approach wholesale, we represent the original problem more faithfully by taking full advantage of the particular structure present in the optimization problem and solving it efficiently using the Alternating Direction Method of Multipliers. The highly constrained nature of the optimization provides our algorithm with the robustness to deal with the varying and often imperfect input affinity matrices arising from different applications and different group numbers. Evaluations on the synthetic data as well as two real world problems show the superiority of the method across a large variety of settings. Zhuwen Li, Loong Fah Cheong, Steven Zhiying Zhou |
CVPR | 3 |
| 2013 | Image Matting with Local and Nonlocal Smooth PriorsabstractIn this paper we propose a novel alpha matting method with local and nonlocal smooth priors. We observe that the manifold preserving editing propagation [4] essentially introduced a nonlocal smooth prior on the alpha matte. This nonlocal smooth prior and the well known local smooth prior from matting Laplacian complement each other. So we combine them with a simple data term from color sampling in a graph model for nature image matting. Our method has a closed-form solution and can be solved efficiently. Compared with the state-of-the-art methods, our method produces more accurate results according to the evaluation on standard benchmark datasets. Xiaowu Chen 0001, Dongqing Zou, Steven Zhiying Zhou, Qinping Zhao |
CVPR | 3 |
| 2013 | Video Co-segmentation for Meaningful Action ExtractionabstractGiven a pair of videos having a common action, our goal is to simultaneously segment this pair of videos to extract this common action. As a preprocessing step, we first remove background trajectories by a motion-based figure ground segmentation. To remove the remaining background and those extraneous actions, we propose the trajectory co saliency measure, which captures the notion that trajectories recurring in all the videos should have their mutual saliency boosted. This requires a trajectory matching process which can compare trajectories with different lengths and not necessarily spatiotemporally aligned, and yet be discriminative enough despite significant intra-class variation in the common action. We further leverage the graph matching to enforce geometric coherence between regions so as to reduce feature ambiguity and matching errors. Finally, to classify the trajectories into common action and action outliers, we formulate the problem as a binary labeling of a Markov Random Field, in which the data term is measured by the trajectory co-saliency and the smoothness term is measured by the spatiotemporal consistency between trajectories. To evaluate the performance of our framework, we introduce a dataset containing clips that have animal actions as well as human actions. Experimental results show that the proposed method performs well in common action extraction. Jiaming Guo, Zhuwen Li, Loong Fah Cheong, Steven Zhiying Zhou |
ICCV | 4 |
| 2013 | Perspective Motion Segmentation via Collaborative ClusteringabstractThis paper addresses real-world challenges in the motion segmentation problem, including perspective effects, missing data, and unknown number of motions. It first formulates the 3-D motion segmentation from two perspective views as a subspace clustering problem, utilizing the epipolar constraint of an image pair. It then combines the point correspondence information across multiple image frames via a collaborative clustering step, in which tight integration is achieved via a mixed norm optimization scheme. For model selection, we propose an over-segment and merge approach, where the merging step is based on the property of the ell_1-norm of the mutual sparse representation of two over-segmented groups. The resulting algorithm can deal with incomplete trajectories and perspective effects substantially better than state-of-the-art two-frame and multi-frame methods. Experiments on a 62-clip dataset show the significant superiority of the proposed idea in both segmentation accuracy and model selection. Zhuwen Li, Jiaming Guo, Loong Fah Cheong, Steven Zhiying Zhou |
ICCV | 4 |
| 2013 | Diminished reality using appearance and 3D geometry of internet photo collectionsabstractThis paper presents a new system level framework for Diminished Reality, leveraging for the first time both the appearance and 3D information provided by large photo collections on the Internet. Recent computer vision techniques have made it possible to automatically reconstruct 3-D structure-from-motion points from large and unordered photo collections. Using these point clouds and a prior provided by GPS, reasonably accurate 6 degree of freedom camera poses can be obtained, thus allowing localization. Once the camera (and hence the user) is correctly localized, photos depicting scenes visible from the user's viewpoint can be used to remove unwanted objects indicated by the user in the video sequences. Existing methods based on texture synthesis bring undesirable artifacts and video inconsistency when the background is heterogeneous; the task is rendered even harder for these methods when the background contains complex structures. On the other hand, methods based on plane warping fail when the background has arbitrary shape. Unlike these methods, our algorithm copes with these problems by making use of internet photos, registering them in 3D space and obtaining the 3D scene structure in an offline process. We carefully design the various components during the online phase so as to meet both speed and quality requirements of the task. Experiments on real data collected demonstrate the superiority of our system. Zhuwen Li, Yuxi Wang 0002, Jiaming Guo, Loong Fah Cheong, Steven Zhiying Zhou |
ISMAR | 5 |
| 2010 | Model-based localization and drift-free user tracking for outdoor augmented realityabstractIn this paper we present a novel model-based hybrid technique for user localization and drift-free tracking in urban environments. In outdoor augmented reality, instantaneous 6-DoF user localization is achieved with position and orientation sensors such as GPS and gyroscopes. Initial pose obtained with these sensors is dominated by large positional errors due to coarse granularity of GPS data. Subsequent tracking is also erroneous as gyroscopes are prone to drifts and often need recalibration. We propose to use model-to-image registration technique to refine initial rough estimate for accurate user localization. Large positional errors in user localization are mitigated by aligning silhouettes of the model with that of the camera image using shape context descriptors as they are invariant to translation, scale and rotational errors. Once initialized, drift-free tracking is achieved by combining frame-to-frame and model-to-frame feature tracking. Frame-to-frame tracking is done by matching corners whereas edges are used for model-to-frame silhouette tracking. Final camera pose is obtained with M-estimators. Jayashree Karlekar, Steven Zhiying Zhou, Yuta Nakayama, Weiquan Lu, Loh Zhi Chang, Daniel Hii |
ICME | 2 |
| 2010 | Positioning, tracking and mapping for outdoor augmentationabstractThis paper presents a novel approach for user positioning, robust tracking and online 3D mapping for outdoor augmented reality applications. As coarse user pose obtained from GPS and orientation sensors is not sufficient for augmented reality applications, sub-meter accurate user pose is then estimated by a one-step silhouette matching approach. Silhouette matching of the rendered 3D model and camera data is carried out with shape context descriptors as they are invariant to translation, scale and rotational errors, giving rise to a non-iterative registration approach. Once the user is correctly positioned, further tracking is carried out with camera data alone. Drifts associated with vision based approaches are minimized by combining different feature modalities. Robust visual tracking is maintained by fusing frame-to-frame and model-to-frame feature matches. Frame-to-frame tracking is accomplished with corner matching while edges are used for model-to-frame registration. Results from individual feature tracker are fused using a pose estimate obtained from an extended Kalman filter (EKF) and a weighted M-estimator. In scenarios where dense 3D models of the environment are not available, online 3D incremental mapping and tracking is proposed to track the user in unprepared environments. Incremental mapping prepares the 3D point cloud of the outdoor environment for tracking. Jayashree Karlekar, Steven Zhiying Zhou, Weiquan Lu, Loh Zhi Chang, Yuta Nakayama, Daniel Hii |
ISMAR | 2 |
| 2010 | MTMR: A conceptual interior design framework integrating Mixed Reality with the Multi-Touch tabletop interfaceabstractThis paper introduces a conceptual interior design framework - Multi-Touch Mixed Reality (MTMR), which integrates mixed reality with the multi-touch tabletop interface, to provide an intuitive and efficient interface for collaborative design and an augmented 3D view to users at the same time. Under this framework, multiple designers can carry out design work simultaneously on the top view displayed on the tabletop, while live video of the ongoing design work is captured and augmented by overlaying virtual 3D furniture models to their 2D virtual counterparts, and shown on a vertical screen in front of the tabletop. Meanwhile, the remote client's camera view of the physical room is augmented with the interior design layout in real time, that is, as the designers place, move, and modify the virtual furniture models on the tabletop, the client sees the corresponding life-size 3D virtual furniture models residing, moving, and changing in the physical room through the camera view on his/her screen. By adopting MTMR, which we argue may also apply to other kinds of collaborative work, the designers can expect a good working experience in terms of naturalness and intuitiveness, while the client can be involved in the design process and view the design result without moving around heavy furniture. By presenting MTMR, we hope to provide reliable and precise freehand interactions to mixed reality systems, with multi-touch inputs on tabletop interfaces. Dong Wei 0004, Steven Zhiying Zhou, Du Xie |
ISMAR | 2 |
| 2010 | A real-time multi-cue hand tracking algorithm based on computer visionabstractAlthough hand tracking algorithm has been widely used in virtual reality and HCI system, it is still a challenging problem in vision-based research area. Due to the robustness and real-time requirements in VR applications, most hand tracking algorithms require special device to achieve satisfactory results. In this paper, we propose an easy-to-use and inexpensive approach to track the hands accurately with a single normal webcam. Outstretched hand is detected by contour & curvature based detection techniques to initialize the tracking region. Robust multi-cue hand tracking is then achieved by velocity-weighted features and color cue. Experiments show that the proposed multi-cue hand tracking approach achieves continuous real-time results even for the situation of cluttered background. The approach fulfills the speed and accuracy requirements of frontal-view vision-based human computer interactions. Kangde Guo, Xing Tang 0002, Steven Zhiying Zhou |
VR | 7 |
| 2010 | Immersive Multiplayer Games With Tangible and Physical InteractionabstractIn this paper, we present a new immersive multiplayer game system developed for two different environments, namely, virtual reality (VR) and augmented reality (AR). To evaluate our system, we developed three game applications-a first-person-shooter game (for VR and AR environments, respectively) and a sword game (for the AR environment). Our immersive system provides an intuitive way for users to interact with the VR or AR world by physically moving around the real world and aiming freely with tangible objects. This encourages physical interaction between players as they compete or collaborate with other players. Evaluation of our system consists of users' subjective opinions and their objective performances. Our design principles and evaluation results can be applied to similar immersive game applications based on AR/VR. Jefry Tedjokusumo, Steven Zhiying Zhou, Stefan Winkler 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2009 | Consistent real-time lighting for virtual objects in augmented realityabstractWe present a technique for rendering realistic shadows of virtual objects in a mixed reality environment by recovering the light source distribution of a scene in real-time, through the segmentation and analysis of a known occluding object's shadows. A fiducial marker provides information about the position of the occluding object and the plane of the surface on which shadows are cast, and serves as the origin of a marker coordinate system. A new shadow segmentation approach is carried out on the shadow image and is able to recover geometrical information on multiple faint shadows. Using normalised iterative reinforcement, noise and artifacts can be suppressed in the final shadow map. The scene's light source distribution is then extrapolated using geometrical data from both the occluding object and its cast shadows. Virtual light sources in a game engine are used to mimic real light sources and achieve consistent illumination and increase the realism of the augmented reality scene. Ryan Christopher Yeoh, Steven Zhiying Zhou |
ISMAR | 2 |