Zhan Song

dblp:53/1204 · DBLP profile ↗
← Back
47ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Explicit Analytical Reconstruction and Global Geometric Constraints for Micron-Level Telecentric 3D Metrology
Yuping Ye, Jixin Liang, Feifei Gu, Zhan Song
ICPR (12)5
2026 ReVEAL: GNN-Guided Reverse Engineering for Formal Verification of Optimized Multipliers
Chen Chen 0172, Daniela Kaufmann, Chenhui Deng, Zhan Song, Hongce Zhang, Cunxi Yu
TACAS (2)4
2026 CF-GAT: Curvature-Fused Graph Attention Network for High-Precision Unordered Facial Point Cloud Landmark Detection
abstract
Due to the lack of large-scale, accurately annotated 3D facial datasets, most current 3D facial landmark detection algorithms rely on 2D texture assistance or non-real digital 3D faces. The performance of these algorithms is limited by the accuracy of texture mapping and the difference between digital faces and actual faces. To tackle these challenges, we built a large-scale, high-precision, and accurately annotated 3D facial database using a structured light system. Building upon this foundation, we proposed a novel point cloud sampling method and 3D facial landmark detection algorithm. This method utilizes a curvature-fused graph attention network to directly predict landmark coordinates from 3D point clouds. Firstly, we extracted a simplifying point set carrying curvature information from the original 3D facial point cloud via geometric point sampling. Then, we incorporated curvature-encoded positional information as the learning component of the attention module and employed it as a feature extractor to construct the network. We evaluated the performance of the method on BU-3DFE, FaceScape and our custom dataset. Compared to existing facial landmark detection algorithms, our model achieved higher accuracy. To facilitate future research on face related applications, we have made the database available on Github1.
Juncheng Han, Yuping Ye, Xintong Yang, Zhan Song
IEEE Trans. Circuits Syst. Video Technol.5
2025 BoolE: Exact Symbolic Reasoning via Boolean Equality Saturation
abstract
Boolean symbolic reasoning for gate-level netlists is a critical step in verification, logic and datapath synthesis, and hardware security. Specifically, reasoning datapath and adder tree in bit-blasted Boolean networks is particularly crucial for verification and synthesis, and challenging. Conventional approaches either fail to accurately (exactly) identify the function blocks of the designs in gate-level netlist with structural hashing and symbolic propagation, or their reasoning performance is highly sensitive to structure modifications caused by technology mapping or logic optimization. This paper introduces BoolE, an exact symbolic reasoning framework for Boolean netlists using equality saturation. BoolE optimizes scalability and performance by integrating domain-specific Boolean ruleset for term rewriting. We incorporate a novel extraction algorithm into BoolE to enhance its structural insight and computational efficiency, which adeptly identifies and captures multi-input, multi-output high-level structures (e.g., full adder) in the reconstructed e-graph. Our experiments show that BoolE surpasses state-of-the-art symbolic reasoning baselines, including the conventional functional approach (ABC) and machine learning-based method (Gamora). Specifically, we evaluated its performance on various multiplier architecture with different configurations. Our results show that BoolE identifies $3.53 \times$ and $3.01 \times$ more exact full adders than ABC in carry-save array and Booth-encoded multipliers, respectively. Additionally, we integrated BoolE into multiplier formal verification tasks, where it significantly accelerates the performance of traditional formal verification tools using computer algebra, demonstrated over four orders of magnitude runtime improvements.
Zhan Song, Qihao Hu, Cunxi Yu
DAC2
2025 Accurate 3D Facial Paralysis Analysis Using Multi-View Infrared Structured Light System
abstract
Facial paralysis is a prevalent disorder affecting the facial nerve. In clinical settings, the severity of facial paralysis is typically assessed by physicians based on their subjective experience, evaluating the range of facial muscle movements and facial symmetry. To address these limitations, this paper proposes a method for the quantifiable evaluation of facial paralysis. We have developed a multi-view real-time facial acquisition system utilizing three infrared structured light units, enabling high-precision dynamic capture. Through non-rigid registration of the collected point cloud sequences, we generated a unified topological mesh sequence. Subsequently, we employed a novel facial asymmetry operator to quantitatively assess facial paralysis. Extensive experimental results demonstrate that the proposed method is both effective and accurate.
Yuping Ye, Jixin Liang, Shiyang Long, Zhan Song
ICASSP5
2025 e-boost: Boosted E-Graph Extraction with Adaptive Heuristics and Exact Solving
abstract
E-graphs have attracted growing interest in many fields, particularly in logic synthesis and formal verification. E-graph extraction is a challenging NP-hard combinatorial optimization problem. It requires identifying optimal terms from exponentially many equivalent expressions, serving as the primary performance bottleneck in e-graph based optimization tasks. However, traditional extraction methods face a critical trade-off: heuristic approaches offer speed but sacrifice optimality, while exact methods provide optimal solutions but face prohibitive computational costs on practical problems. We present e-boost, a novel framework that bridges this gap through three key innovations: (1) parallelized heuristic extraction that leverages weak data dependence to compute DAG costs concurrently, enabling efficient multi-threaded performance without sacrificing extraction quality; (2) adaptive search space pruning that employs a parameterized threshold mechanism to retain only promising candidates, dramatically reducing the solution space while preserving near-optimal solutions; and (3) initialized exact solving that formulates the reduced problem as an Integer Linear Program with warm-start capabilities, guiding solvers toward high-quality solutions faster.Across the diverse benchmarks in formal verification and logic synthesis fields, e-boost demonstrates 558× runtime speedup over traditional exact approaches (ILP) and 19.04% performance improvement over the state-of-the-art extraction framework (SmoothE). In realistic logic synthesis tasks, e-boost produces 7.6% and 8.1% area improvements compared to conventional synthesis tools with two different technology mapping libraries. e-boost is available at https://github.com/Yu-Maryland/e-boost.
Zhan Song, Yaohui Cai, Zhiru Zhang, Cunxi Yu
ICCAD2
2025 HEC: Equivalence Verification Checking for Code Transformation via Equality Saturation
Zhan Song, Nicolas Bohm Agostini, Antonino Tumeo, Cunxi Yu
USENIX ATC2
2025 Illumination Map Estimation via Sparse Bright Channel for Enhancing Under-Exposed Images
abstract
This paper presents a novel image enhancement approach to avoid common artifacts such as over-exposure, color cast, and unnatural results. The key innovation lies in estimating the illumination map of an underexposed image using a sparse bright channel. Our approach includes an algorithm that enforces the sparsity of the inverted bright channel, enabling the indirect estimation of a coarse but suitable initial illumination map. This initial map is refined using an updated weight-constrained regularization with joint local exposure and detail feedback constraints, producing a piece-wise smooth, structure-preserving illumination map. Computer simulations show that the proposed method is competitive with or even outperforms several state-of-the-art enhancement methods in terms of both subjective and objective evaluations.
Shiqian Wu, Dianwei Wang, Sos S. Agaian, Zhan Song
IEEE Trans. Intell. Transp. Syst.5
2025 Weakly Aligned Feature Fusion for Multimodal Object Detection
abstract
To achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data often suffer from the position shift problem, i.e., the image pair is not strictly aligned, making one object has different positions in different modalities. For the deep learning method, this problem makes it difficult to fuse multimodal features and puzzles the convolutional neural network (CNN) training. In this article, we propose a general multimodal detector named aligned region CNN (AR-CNN) to tackle the position shift problem. First, a region feature (RF) alignment module with adjacent similarity constraint is designed to consistently predict the position shift between two modalities and adaptively align the cross-modal RFs. Second, we propose a novel region of interest (RoI) jitter strategy to improve the robustness to unexpected shift patterns. Third, we present a new multimodal feature fusion method that selects the more reliable feature and suppresses the less useful one via feature reweighting. In addition, by locating bounding boxes in both modalities and building their relationships, we provide novel multimodal labeling named KAIST-Paired. Extensive experiments on 2-D and 3-D object detection, RGB-T, and RGB-D datasets demonstrate the effectiveness and robustness of our method.
Lu Zhang 0054, Zhiyong Liu 0001, Xiangyu Zhu 0001, Zhan Song, Xu Yang 0004, Zhen Lei 0001, Hong Qiao
IEEE Trans. Neural Networks Learn. Syst.4
2024 High-precision 3D Facial Landmark Detection with Curvature-fused Graph Attention Network
abstract
With the rapid advancement of deep learning, 2D facial landmark detection algorithms have achieved satisfying results. However, due to the absence of large-scale accurately annotated 3D facial datasets, most current 3D facial landmark detection algorithms rely on 2D texture assistance or non-real digital 3D faces. The performance of these algorithms is limited by the accuracy of 2D texture mapping onto 3D faces and the difference between digital faces and actual faces. To tackle these challenges, we have built a large-scale, high-precision 3D facial database using a structured light system. Facial landmarks within the database are marked multiple times to ensure accuracy. Building upon this foundation, we proposed a novel point cloud sampling method and 3D facial landmark detection algorithm. This method utilizes a curvature-fused graph attention network (CGAT) to directly predict landmark coordinates from 3D point clouds. Initially, we extracted a subsampled point set carrying curvature information from the original 3D facial point cloud via geometric point sampling (GPS). Then, we used curvature-encoded positional information as the learning component of the attention module and employed it as a feature extractor to construct the CGAT. We assessed the performance of CGAT on two datasets, BU-3DFE and CIE-H3DF (Ours). Compared to the existing facial landmark detection algorithms, CGAT achieves higher accuracy.
Juncheng Han, Yuping Ye, Shiyang Long, Zhan Song
BIBM5
2024 Adaptive Head Pose Estimation with Real-Time Structured Light
abstract
Head pose estimation (HPE) is a crucial task in pose recognition, but the existing HPE methods suffer from low robustness, low accuracy and inconvenience of contact measurement. In this paper, we build up an infrared structured light system for head pose estimation and propose an adaptive head pose estimation method to improve robustness and accuracy without the need for training. Firstly, we utilize a dynamic structured light system to acquire both the standard model and real-time dynamic data. Then, an iterative registration algorithm is proposed to adaptively segment the facial region which remains stable excluding distractions such as expressions and estimate the head pose. We experimentally evaluate the effectiveness of our method under conventional and large-angular head motion, and different expressions. The results demonstrate our method achieves highly accurate and robust real-time head pose estimation in a contactless manner.
Yuping Ye, Feifei Gu, Zhan Song, Xiaodong Bai
ICASSP4
2024 Retargeting of facial model for unordered dense point cloud
Yuping Ye, Juncheng Han, Jixin Liang, Zhan Song
Comput. Graph.5
2024 High-efficiency automated triaxial robot grasping system for motor rotors using 3D structured light sensor
Jixin Liang, Yuping Ye, Zhan Song
Mach. Vis. Appl.5
2022 CSIE: Coded Strip-Patterns Image Enhancement Embedded in Structured Light-Based Methods
Yuping Ye, Chu Shi, Zhan Song
ACCV (3)4
2022 High-fidelity 3D real-time facial animation using infrared structured light sensing system
Yuping Ye, Zhan Song
Comput. Graph.2
2021 Precise grabbing of overlapping objects system based on end-to-end deep neural network
Xining Cui, Zhan Song, Feifei Gu
Comput. Commun.3
2020 Normal-Based Bas-Relief Modelling via Near-Lighting Photometric Stereo
abstract
Abstract We present a near‐lighting photometric stereo (NL‐PS) system to produce digital bas‐reliefs from a physical object (set) directly. Unlike both the 2D image and 3D model‐based modelling methods that require complicated interactions and transformations, the technique using NL‐PS is easy to use with cost‐effective hardware, providing users with a trade‐off between abstract and representation when creating bas‐reliefs. Our algorithm consists of two steps: normal map acquisition and constrained 3D reconstruction. First, we introduce a lighting model, named the quasi‐point lighting model (QPLM), and provide a two‐step calibration solution in our NL‐PS system to generate a dense normal map. Second, we filter the normal map into a detail layer and a structure layer, and formulate detail‐ or structure‐preserving bas‐relief modelling as a constrained surface reconstruction problem of solving a sparse linear system. The main contribution is a WYSIWYG (i.e. what you see is what you get) way of building new solvers that produces multi‐style bas‐reliefs with their geometric structures and/or details preserved. The performance of our approach is experimentally validated via comparisons with the state‐of‐the‐art methods.
Mingqiang Wei, Zhan Song, Ying Nie 0006, Jianhuang Wu, Zhongping Ji, Yanwen Guo 0001, Haoran Xie 0001, Jun Wang 0039, Fu Lee Wang
Comput. Graph. Forum2
2020 Bigflow: A General Optimization Layer for Distributed Computing Frameworks
Yuncong Zhang, Xiaoyang Wang 0006, Guangyu Sun 0003, Gong-Lin Zheng, Shan-Hui Yin, Xian-Jin Ye, Zhan Song, Dong-Dong Miao
J. Comput. Sci. Technol.12
2019 Vehicle Positioning and Ranging with Static Traffic Camera based on 2D-3D Tracking and Re-Projection
abstract
Vehicle positioning and ranging are the current research hotspots. To attain the competition goal of the MMSP Witcomm Challenge 2019 and promote the development of the autonomous driving technology, a novel framework is proposed in this paper, which combines the 2D object tracking and 3D reprojection methodologies. Firstly, a correlation filter using the deep convolutional features is designed to detect the bounding box of the moving objects which is achieved via finding the maximum response of the initial object in the convolutional feature maps. Next the homography matrix is calculated based on the image coordinate points and the corresponding world coordinate points to project the vehicle 2D boundary into the real world. The validity of the proposed framework is verified using the dataset provided by the MMSP Witcomm Challenge 2019 competition, where our method was awarded the second-place prize.
Zhan Song, Yipeng Liu 0003, Yiling Xu, Le Yang 0001
MMSP1
2019 Parametric 3D modeling of a symmetric human body
Yin Chen 0003, Zhan Song, Weiwei Xu 0003, Ralph R. Martin, Zhi-Quan Cheng
Comput. Graph.2
2019 A closed-form single-pose calibration method for the camera-projector system
Zexi Feng, Zhi-Quan Cheng, Zhan Song
Mach. Vis. Appl.3
2018 Non-iterative multiple data registration method based on the motion screw theory and trackable features
abstract
Registration of 3D point clouds is an important issue in the field of 3D reconstruction. In this work, we proposed a non-iterative registration method based on the motion screw theory and trackable features. The screw theory is derived from the theory of rigid body mechanics, which holds the idea that the motion of rigid body can be regarded as a kind of spiral motion and can be effectively represented by an angular velocity vector and a linear velocity vector. It has not been utilized in 3D data registration before as far as we know. In this paper, 3D data registration based on the motion screw theory is specifically introduced, and a searching strategy based on the trackable features in image sequences is presented to improve the accuracy of 3D registration. The proposed method has been successfully tested on real multi-view data. Experimental results showed that it could simplify the computational process, accelerate the speed of registration, and achieve higher precision than other methods.
Feifei Gu, Zhan Song
ICPR3
2018 Capture of hair geometry using white structured light
Yin Chen 0003, Zhan Song, Ralph R. Martin, Zhi-Quan Cheng
Comput. Aided Des.2
2018 Parametric modeling of 3D human body shape - A survey
Zhi-Quan Cheng, Yin Chen 0003, Ralph R. Martin, Zhan Song
Comput. Graph.5
2018 Learning a Multiple Kernel Similarity Metric for kinship verification
Yanguo Zhao, Zhan Song, Feng Zheng 0001, Ling Shao 0001
Inf. Sci.2
2017 An Automatic 3D Textured Model Building Method Using Stripe Structured Light System
Hualie Jiang, Yuping Ye, Zhan Song, Suming Tang
ICVS3
2017 Calibration of a Structured Light Measurement System Using Binary Shape Coding
Hai Zeng, Suming Tang, Zhan Song, Feifei Gu
ICVS3
2016 A practical means for the optimization of structured light system calibration parameters
abstract
This paper presents a novel approach for the optimization of calibration parameters in structured light system (SLS). Different with conventional calibration algorithms, the proposed optimization algorithm is implemented in 3D space instead of 2D image space. The object used for parameter optimization can be a simple plane with some markers. A global optimal function is constructed to contain all the intrinsic and extrinsic parameters of the SLS. Using the primary calibration parameters by conventional methods as initial values, the optimal function can be solved by minimizing the 3D measurement errors like distance, angle between markers, and the planarity of the reference plane. Experimental results show that, 3D reconstruction accuracy can be greatly improved by the proposed approach in comparison with traditional SLS calibration methods.
Yuping Ye, Zhan Song
ICIP2
2016 A novel photometric stereo method with nonisotropic point light sources
abstract
This paper presents a photometric stereo method with nonisotropic point light sources. Subject to the non-uniform lighting conditions produced by the nonisotropic point sources, each incident light ray should be precisely determined so as to realize an accurate calculation of surface normal. In the proposed method, radiance model of the light source is firstly introduced to the classical photometric stereo framework. By considering the distance and angular attenuations of incident light rays, a precise description for the lighting field can be established. Based on the initial 3D reconstruction result, an iterative process is introduced to optimize the primary light model parameters with respect to the unknown distance factor. The experimental setup is quite simple, which only consists of some LEDs and one camera. And the experimental results show that, with the proposed method, accuracy of the reconstructed surface normal can be greatly improved in comparison with some conventional light models.
Ying Nie 0006, Zhan Song
ICPR2
2016 Multi-class kernel margin maximization for kernel learning
Yanguo Zhao, Ronald Chung, Zhan Song
Neurocomputing4
2016 A single-shot structured light means by encoding both color and geometrical features
Zhan Song
Pattern Recognit.3
2015 Laplacian Auto-Encoders: An explicit learning of nonlinear data manifold
Kui Jia, Lin Sun 0004, Shenghua Gao, Zhan Song, Bertram E. Shi
Neurocomputing4
2014 Hand posture recognition using approximate vanishing ideal generators
abstract
This paper represents a hand posture recognition method that combines both skin and shape cues. In the algorithm, implicit skin image is firstly computed to suppress the background disturbances as well as to enhance the hand region; and vanishing component analysis (VCA) algorithm is applied to each posture category to learn a group of Approximate Vanishing Ideal Generators (AVIGs) which are used for the follow-up feature extraction. Each generator is essentially an effective characterization for geometrical structure of corresponding hand posture. In recognition phase, features acquired from implicit skin image and AVIGs are inputted to a softmax model for classification. A dataset comprising 5 hand postures are constructed for its evaluation. The proposed algorithm is demonstrated to be robust to complex environment and challenging illuminations.
Yanguo Zhao, Zhan Song
ICIP2
2014 Mean shift-based single image dehazing with re-refined transmission map
abstract
Bad weather (eg., fog and haze) significantly degrades the quality of outdoor images taken by camera, leading to the fact that most automatic systems, which strongly depends on the definition of the input images, fail to work normally. Thus, the improvement of the dehazing technology is highly desired. To overcome the disadvantages of traditional dark channel priorbased algorithm, we propose a more efficient dehazing algorithm combining dark channel prior and mean shift segmentation. Firstly, we take the operation of white balance on the input haze image to reduce the negative influence of color cast. Secondly, we use the mean shift segmentation algorithm to separate the sky regions from the foreground in the transmission map, which is obtained with the dark channel prior-based approach. Thirdly, we enhance the brightness of the sky regions in the transmission map independently and use guided image filtering to smooth the map. Finally, we restore the image with the re-refined transmission map. The experimental results demonstrate that the proposed approach is better able to handle the sky regions and solve the problem of color cast compared with the typical dehazing algorithms.
Qingsong Zhu 0001, Jiaming Mai, Zhan Song, Lei Wang 0029
SMC3
2014 An effective quad-dominant meshing method for unorganized point clouds
Xufang Pang, Zhan Song, Rynson W. H. Lau
Graph. Model.2
2013 SUPERCUT: An accurate and effective interactive image segmentation algorithm
abstract
The task of interactive image segmentation has attracted a significant attention in recent years. The ultimate goal is to extract an object with as few user interactions as possible. In this paper, we present SUPERCUT, a novel interactive algorithm for foreground object extraction and segmentation in images. In the algorithm, the mean shift algorithm with a boundary confidence prior is introduced to efficiently pre-segment the original image into super-pixels with precise boundary. Secondly, a Bayes decision theory is introduced to model and cluster the super-pixels so as to obtain an initial effective classification of super-pixels. To achieve a more accurate object segmentation result, a boundary refinement using Interactive rectangle box with GMM learning is adopted. Experimental results on a benchmark data set show that the proposed framework is highly effective and can accurately segment a wide variety of natural images with ease.
Qingsong Zhu 0001, Ling Shao 0001, Zhan Song, Yaoqin Xie
ICIP3
2013 Whitening central projection descriptor for affine-invariant shape description
abstract
A novel descriptor, referred to as the whitening central projection predictor (WCPD), is developed for affine‐invariant shape description. The proposed descriptor is based on central projection transform (CPT) and whitening transform (WT). Dislike contour‐based or region‐based approaches, an object is first converted to a closed curve by CPT, which is called the general curve (GC). The derived GC not only keeps the affine transform information, but also is very robust to noise. Then WT is performed to the GC with the purpose that the affine transformation is simplified to a rotation only. Finally, Fourier descriptors are employed to remove the rotation, and WCPD is obtained. One advantage of using WCPD for affine‐invariant description lies in that it is applicable to objects consisting of several components. Furthermore, the approach used on the GC is contour‐based, and is of small computational complexity. Several experiments have been conducted to evaluate the performance of the proposed method. Experimental results show that the proposed method has a powerful discrimination ability, and is more robust to noise.
Rushi Lan, Colin Fyfe, Zhan Song
IET Image Process.5
2013 A semi-supervised approach for dimensionality reduction with distributional similarity
Feng Zheng 0001, Zhan Song, Ling Shao 0001, Ronald Chung, Kui Jia
Neurocomputing2
2013 A Novel Nonlinear Regression Approach for Efficient and Accurate Image Matting
abstract
Current image matting approaches are often implemented based upon color samples under various local assumptions. In this letter, a novel image matting algorithm is investigated by treating the alpha matting as a regression problem. Specifically, we learn spatially-varying relations between pixel features and alpha values using support vector regression. Via the learning-based approach, limitations caused by local image assumptions can be greatly relieved. In addition, the computed confidence terms in learning phase can be conveniently integrated with other matting approaches for the matting accuracy improvement. Qualitative and quantitative evaluations are implemented with a public matting benchmark. And the results are compared with some recent matting algorithms to show its advantages in both efficiency and accuracy.
Qingsong Zhu 0001, Zhan Song, Yaoqin Xie, Lei Wang 0029
IEEE Signal Process. Lett.3
2012 An efficient r-KDE model for the segmentation of dynamic scenes
Qingsong Zhu 0001, Zhan Song, Yaoqin Xie
ICPR2
2012 An affine invariant discriminate analysis with canonical correlation analysis
Rushi Lan, Zhan Song, Yuan Yan Tang
Neurocomputing4
2012 A Novel Recursive Bayesian Learning-Based Method for the Efficient and Accurate Segmentation of Video With Dynamic Background
abstract
Segmentation of video with dynamic background is an important research topic in image analysis and computer vision domains. In this paper, we present a novel recursive Bayesian learning-based method for the efficient and accurate segmentation of video with dynamic background. In the algorithm, each frame pixel is represented as the layered normal distributions which correspond to different background contents in the scene. The layers are associated with a confident term and only the layers satisfy the given confidence which will be updated via the recursive Bayesian estimation. This makes learning of background motion trajectories more accurate and efficient. To improve the segmentation quality, the coarse foreground is obtained via simple background subtraction first. Then, a local texture correlation operator is introduced to fill the vacancies and remove the fractional false foreground regions. Extensive experiments on a variety of public video datasets and comparisons with some classical and recent algorithms are used to demonstrate its improvements in both segmentation accuracy and efficiency.
Qingsong Zhu 0001, Zhan Song, Yaoqin Xie, Lei Wang 0029
IEEE Trans. Image Process.2
2011 Recent advances and trends in visual tracking: A review
Hanxuan Yang 0001, Ling Shao 0001, Feng Zheng 0001, Liang Wang 0001, Zhan Song
Neurocomputing5
2010 Dynamic video segmentation via a novel recursive Bayesian learning method
abstract
Segmentation of an interesting target from a dynamic video has been an important research topic in computer vision. In this work, we present a novel recursive Bayesian learning method for dynamic video segmentation. In the algorithm, each frame pixel is represented as layered normal distributions and the recursive Bayesian estimation is used to update the background parameters so as to obtain a robust background model. In the segmentation, foreground is separated by simple background subtraction method firstly. And then, a local texture correlation operator is proposed to remove vacancies in the separated foreground to refine the segmentation result. Experiments with two typical video clips are used to demonstrate that the proposed method can outperform traditional methods in both segmentation result and converging speed.
Zhan Song
ICIP2
2010 Determining Both Surface Position and Orientation in Structured-Light-Based Sensing
abstract
Position and orientation profiles are two principal descriptions of shape in space. We describe how a structured light system, coupled with the illumination of a pseudorandom pattern and a suitable choice of feature points, can allow not only the position but also the orientation of individual surface elements to be determined independently. Unlike traditional designs which use the centroids of the illuminated pattern elements as the feature points, the proposed design uses the grid points between the pattern elements instead. The grid points have the essences that their positions in the image data are inert to the effect of perspective distortion, their individual extractions are not directly dependent on one another, and the grid points possess strong symmetry that can be exploited for their precise localization in the image data. Most importantly, the grid lines of the illuminated pattern that form the grid points can aid in determining surface normals. In this paper, we describe how each of the grid points can be labeled with a unique color code, what symmetry they possess and how the symmetry can be exploited for their precise localization at subpixel accuracy in the image data, and how 3D orientation in addition to 3D position can be determined at each of them. Both the position and orientation profiles can be determined with only a single pattern illumination and a single image capture.
Zhan Song, Ronald Chung
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 Nonstructured light-based sensing for 3D reconstruction
Zhan Song, Ronald Chung
Pattern Recognit.1
2008 Grid point extraction exploiting point symmetry in a pseudo-random color pattern
abstract
Structured light system together with the use of pseudo-randomly coded pattern in the projection is an effective solution for 3D reconstruction; it requires only a single image capture to operate. In this article, a 2D pseudo-random pattern consisting of rhombic color elements is proposed, and the grid-points between the pattern elements as opposed to the centroids of the elements are adopted as the feature points. Two possible types of grid-point are described, and a scheme that allows each grid-point to be uniquely distinguished by a codeword and the grid-point type it belongs to is described. We also present a grid-point detector that is based upon a certain symmetry the grid-points of rhombic pattern own — the 2-fold rotation symmetry — which is largely preserved under pattern projection, reflection, perspective distortion, image noise, and image blur.
Zhan Song, Ronald Chung
ICIP1