Zhebin Zhang

dblp:59/2896 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Details Enhancement in Unsigned Distance Field Learning for High-fidelity 3D Surface Reconstruction
abstract
While Signed Distance Fields (SDF) are well-established for modeling watertight surfaces, Unsigned Distance Fields (UDF) broaden the scope to include open surfaces and models with complex inner structures. Despite their flexibility, UDFs encounter significant challenges in high-fidelity 3D reconstruction, such as non-differentiability at the zero level set, difficulty in achieving the exact zero value, numerous local minima, vanishing gradients, and oscillating gradient directions near the zero level set. To address these challenges, we propose Details Enhanced UDF (DEUDF) learning that integrates normal alignment and the SIREN network for capturing fine geometric details, adaptively weighted Eikonal constraints to address vanishing gradients near the target surface, unconditioned MLP-based UDF representation to relax non-negativity constraints, and DCUDF for extracting the local minimal average distance surface. These strategies collectively stabilize the learning process from unoriented point clouds and enhance the accuracy of UDFs. Our computational results demonstrate that DEUDF outperforms existing UDF learning methods in both accuracy and the quality of reconstructed surfaces.
Fei Hou 0001, Wencheng Wang 0001, Hong Qin 0001, Zhebin Zhang, Ying He 0001
AAAI5
2025 A Lightweight UDF Learning Framework for 3D Reconstruction Based on Local Shape Functions
abstract
Unsigned distance fields (UDFs) provide a versatile framework for representing a diverse array of 3D shapes, encompassing both watertight and non-watertight geometries. Traditional UDF learning methods typically require extensive training on large 3D shape datasets, which is costly and necessitates re-training for new datasets. This paper presents a novel neural framework, LoSF-UDF, for reconstructing surfaces from 3D point clouds by leveraging local shape functions to learn UDFs. We observe that 3D shapes manifest simple patterns in localized regions, prompting us to develop a training dataset of point cloud patches characterized by mathematical functions that represent a continuum from smooth surfaces to sharp edges and corners. Our approach learns features within a specific radius around each query point and utilizes an attention mechanism to focus on the crucial features for UDF estimation. Despite being highly lightweight, with only 653 KB of trainable parameters and a modest-sized training dataset with 0.5 GB storage, our method enables efficient and robust surface reconstruction from point clouds without requiring for shape-specific training. Furthermore, our method exhibits enhanced resilience to noise and outliers in point clouds compared to existing methods. We conduct comprehensive experiments and comparisons across various datasets, including synthetic and real-scanned point clouds, to validate our method’s efficacy. Notably, our lightweight framework offers rapid and reliable initialization for other unsupervised iterative approaches, improving both the efficiency and accuracy of their reconstructions. Our project and code are available at https://jbhu67.github.io/LoSF-UDF.github.io/.
Jiangbei Hu, Yanggeng Li, Fei Hou 0001, Junhui Hou, Zhebin Zhang, Shengfa Wang, Na Lei, Ying He 0001
CVPR5
2025 mmTAA: A Contact-Less Thoracoabdominal Asynchrony Measurement System Based on mmWave Sensing
abstract
Thoracoabdominal Asynchrony (TAA) is a key metric in respiration monitoring, which characterizes the non-parallel periodical motion of human's rib cage (RC) and abdomen (AB) during each breath. Long-term measurement of TAA plays a significant role in respiration health tracking. Existing TAA measurement methods including Respiratory Inductive Plethysmography (RIP) and Optoelectronic Plethysmography (OEP) all intrusive to subjects and have certain requirements on operation conditions, which limit their usage to hospital scenario. To address this gap, we proposemmTAA, the first mmWave-based, non-intrusive TAA measurement system ready for ubiquitous usage in daily-life. InmmTAA, we design a Two-stage RC-AB centroid finding module, aiming to identify the most probable location of RC-AB centroid, which can best represent RC and AB in mmWave sensing scenario. Subsequently, we design TAANet, a novel Convolutional Neural Network (CNN)-based architecture with residual modules, tailored for TAA measurement. Meanwhile, in order to address the imbalance of continuous data, we add imbalance information equalizer including feature and label equalizer during network training. We implementmmTAAon a commonly used multi-antenna mmWave radar. We prototype, deploy and evaluatemmTAAon 25 subjects and 25.7h data in total.mmTAAachieves 4.01$^{\circ }$MAE and 1.56$^{\circ }$average error, close to OEP method.
Fenglin Zhang, Zhebin Zhang, Anfu Zhou, Huadong Ma
IEEE Trans. Mob. Comput.2
2025 DCUDF2: Improving Efficiency and Accuracy in Extracting Zero Level Sets From Unsigned Distance Fields
abstract
Unsigned distance fields (UDFs) provide a flexible representation for models with complex topologies, but accurately extracting their zero level sets remains challenging, particularly in preserving topological correctness and fine geometric details. We present DCUDF2, an enhanced method that builds upon DCUDF to address these limitations. Our approach introduces an accuracy-aware loss function with self-adaptive weights, enabling precise geometric fitting while avoiding over-smoothing. To improve robustness, we propose a topology correction strategy that reduces the sensitivity to hyper-parameter settings. Furthermore, we develop new operations leveraging self-adaptive weights to accelerate convergence and improve runtime efficiency. Extensive experiments on diverse datasets demonstrate that DCUDF2 consistently outperforms DCUDF and existing methods in both geometric fidelity and topological accuracy.
Fugang Yu, Fei Hou 0001, Wencheng Wang 0001, Zhebin Zhang, Ying He 0001
IEEE Trans. Vis. Comput. Graph.5
2024 Practical Measurements of Translucent Materials with Inter-Pixel Translucency Prior
abstract
Material appearance is a key component of photorealism, with a pronounced impact on human perception. Although there are many prior works targeting at measuring opaque materials using light-weight setups (e.g., consumer-level cameras), little attention is paid on acquiring the optical properties of translucent materials which are also quite common in nature. In this paper, we present a practical method for acquiring scattering properties of translucent materials, based solely on ordinary images captured with unknown lighting and camera parameters. The key to our method is an inter-pixel translucency prior which states that image pixels of a given homogeneous translucent material typically form curves (dubbed translucent curves) in the RGB space, of which the shapes are determined by the parameters of the material. We leverage this prior in a specially-designed convolutional neural network comprising multiple encoders, a translucency-aware feature fusion module and a cascaded decoder. We demonstrate, through both visual comparisons and quantitative evaluations, that high accuracy can be achieved on a wide range of real-world translucent materials.
Zhenyu Chen 0001, Jie Guo 0001, Shuichang Lai, Ruoyu Fu, Mengxun Kong, Chen Wang 0149, Hongyu Sun 0001, Zhebin Zhang, Chen Li 0062, Yanwen Guo 0001
CVPR8
2024 Multi-View Attentive Contextualization for Multi-View 3D Object Detection
abstract
We present Multi-View Attentive Contextualization (MvACon), a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed in the field of query-based MV3D object detection, prior art often suffers from either the lack of exploiting high-resolution 2D features in dense attention-based lifting, due to high computational costs, or from insufficiently dense grounding of 3D queries to multi-scale 2D features in sparse attention-based lifting. Our proposed MvACon hits the two birds with one stone using a representationally dense yet computationally sparse attentive feature contextualization scheme that is agnostic to specific 2D-to-3D feature lifting approaches. In experiments, the proposed MvA-Con is thoroughly tested on the nuScenes benchmark, using both the BEVFormer and its recent 3D deformable attention (DFA3D) variant, as well as the PETR, showing consistent detection performance improvement, especially in enhancing performance in location, orientation, and velocity prediction. It is also tested on the Waymo-mini benchmark using BEVFormer with similar improvement. We qualitatively and quantitatively show that global cluster-based contexts effectively encode dense scene-level contexts for MV3D object detection. The promising results of our proposed MvA-Con reinforces the adage in computer vision - “(contextualized) feature matters”.
Xianpeng Liu, Ming Qian, Nan Xue 0001, Chen Chen 0001, Zhebin Zhang, Tianfu Wu 0001
CVPR6
2024 Real-time volume rendering with octree-based implicit surface representation
Luo Zhang 0002, Jiangbei Hu, Zhebin Zhang, Gaochao Song, Ying He 0001
Comput. Aided Geom. Des.4
2024 GS-Octree: Octree-based 3D Gaussian Splatting for Robust Object-level 3D Reconstruction Under Strong Lighting
abstract
Abstract The 3D Gaussian Splatting technique has significantly advanced the construction of radiance fields from multi‐view images, enabling real‐time rendering. While point‐based rasterization effectively reduces computational demands for rendering, it often struggles to accurately reconstruct the geometry of the target object, especially under strong lighting conditions. Strong lighting can cause significant color variations on the object's surface when viewed from different directions, complicating the reconstruction process. To address this challenge, we introduce an approach that combines octree‐based implicit surface representations with Gaussian Splatting. Initially, it reconstructs a signed distance field (SDF) and a radiance field through volume rendering, encoding them in a low‐resolution octree. This initial SDF represents the coarse geometry of the target object. Subsequently, it introduces 3D Gaussians as additional degrees of freedom, which are guided by the initial SDF. In the third stage, the optimized Gaussians enhance the accuracy of the SDF, enabling the recovery of finer geometric details compared to the initial SDF. Finally, the refined SDF is used to further optimize the 3D Gaussians via splatting, eliminating those that contribute little to the visual appearance. Experimental results show that our method, which leverages the distribution of 3D Gaussians with SDFs, reconstructs more accurate geometry, particularly in images with specular highlights caused by strong lighting. The source code can be downloaded from https://github.com/LaoChui999/GS-Octree .
Zhengyu Wen, Luo Zhang 0002, Jiangbei Hu, Fei Hou 0001, Zhebin Zhang, Ying He 0001
Comput. Graph. Forum6
2024 Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs Prior
abstract
Existing neural rendering-based text-to-3D-portrait generation methods typically make use of human geometry prior and diffusion models to obtain guidance. However, relying solely on geometry information introduces issues such as the Janus problem, over-saturation, and over-smoothing. We present Portrait3D , a novel neural rendering-based framework with a novel joint geometry-appearance prior to achieve text-to-3D-portrait generation that overcomes the aforementioned issues. To accomplish this, we train a 3D portrait generator, 3DPortraitGAN, as a robust prior. This generator is capable of producing 360° canonical 3D portraits, serving as a starting point for the subsequent diffusion-based generation process. To mitigate the "grid-like" artifact caused by the high-frequency information in the feature-map-based 3D representation commonly used by most 3D-aware GANs, we integrate a novel pyramid tri-grid 3D representation into 3DPortraitGAN. To generate 3D portraits from text, we first project a randomly generated image aligned with the given prompt into the pre-trained 3DPortraitGAN's latent space. The resulting latent code is then used to synthesize a pyramid tri-grid. Beginning with the obtained pyramid tri-grid , we use score distillation sampling to distill the diffusion model's knowledge into the pyramid tri-grid. Following that, we utilize the diffusion model to refine the rendered images of the 3D portrait and then use these refined images as training data to further optimize the pyramid tri-grid , effectively eliminating issues with unrealistic color and unnatural artifacts. Our experimental results show that Portrait3D can produce realistic, high-quality, and canonical 3D portraits that align with the prompt.
Hao Xu 0049, Xiangjun Tang, Xien Chen, Siyu Tang 0001, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001
ACM Trans. Graph.6
2024 FusionDeformer: text-guided mesh deformation using diffusion models
Hao Xu 0049, Xiangjun Tang, Jing Zhang 0038, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001
Vis. Comput.6
2024 Publisher Correction: FusionDeformer: text-guided mesh deformation using diffusion models
Hao Xu 0049, Xiangjun Tang, Jing Zhang 0038, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001
Vis. Comput.6
2023 IAG: Induction-Augmented Generation Framework for Answering Reasoning Questions
abstract
Retrieval-Augmented Generation (RAG), by incorporating external knowledge with parametric memory of language models, has become the state-of-the-art architecture for opendomain QA tasks.However, common knowledge bases are inherently constrained by limited coverage and noisy information, making retrieval-based approaches inadequate to answer implicit reasoning questions.In this paper, we propose an Induction-Augmented Generation (IAG) framework that utilizes inductive knowledge along with the retrieved documents for implicit reasoning.We leverage large language models (LLMs) for deriving such knowledge via a novel prompting method based on inductive reasoning patterns.On top of this, we implement two versions of IAG named IAG-GPT and IAG-Student, respectively.IAG-GPT directly utilizes the knowledge generated by GPT-3 for answer prediction, while IAG-Student gets rid of dependencies on GPT service at inference time by incorporating a student inductor model.The inductor is firstly trained via knowledge distillation and further optimized by back-propagating the generator feedback via differentiable beam scores.Experimental results show that IAG outperforms RAG baselines as well as ChatGPT on two Open-Domain QA tasks.Notably, our best models have won the first place in the official leaderboards of CSQA2.0 (since Nov 1, 2022) and StrategyQA (since Jan 8, 2023).
Zhebin Zhang, Yuanhang Ren, Saijiang Shi, Yongkang Wu, Ruofei Lai, Zhao Cao
EMNLP1
2023 Lightweight Fisher Vector Transfer Learning for Video Deduplication
abstract
Video deduplication in cloud and on devices is a key challenge for storage and communication efficiency. The lifetime of video content creation, communication/sharing, and consumption can generate multiple versions of the same content with variations in coding and editing effects. In this work, we develop a lightweight and robust deduplication feature based on the fisher vector aggregation of Scale-Invariant Feature Transform (SIFT) keypoints. The fisher vector representation is used for a deduplication transfer learning process that utilizes a lightweight Multilayer Perceptron (MLP) network with center loss to learn a compact and distinctive feature. Simulation on the CC_WEB_VIDEO dataset demonstrated that the proposed feature is extremely robust in deduplication with respect to typical editing effects and coding/transcoding degenerations while being computationally very lightweight compared to other solutions.
Chris Henry, Rijun Liao, Ruiyuan Lin, Zhebin Zhang, Hongyu Sun 0001, Zhu Li 0001
ICASSP4
2023 Interpretability for reliable, efficient, and self-cognitive DNNs: From theories to applications
Xu Kang 0002, Jie Guo 0008, Bin Song 0001, Binghuang Cai, Hongyu Sun 0001, Zhebin Zhang
Neurocomputing6
2021 BERT-JAM: Maximizing the utilization of BERT for neural machine translation
Zhebin Zhang, Sai Wu, Dawei Jiang, Gang Chen 0001
Neurocomputing1
2021 Refiner: A Reliable Incentive-Driven Federated Learning System Powered by Blockchain
abstract
Modern mobile applications often produce decentralized data, i.e., a huge amount of privacy-sensitive data distributed over a large number of mobile devices. Techniques for learning models from decentralized data must properly handle two natures of such data, namely privacy and massive engagement. Federated learning (FL) is a promising approach for such a learning task since the technique learns models from data without exposing privacy. However, traditional FL methods assume that the participating mobile devices are honest volunteers. This assumption makes traditional FL methods unsuitable for applications where two kinds of participants are engaged: 1) self-interested participants who, without economical stimulus, are reluctant to contribute their computing resources unconditionally, and 2) malicious participants who send corrupt updates to disrupt the learning process. This paper proposes Refiner, a reliable federated learning system for tackling the challenges introduced by massive engagements of self-interested and malicious participants. Refiner is built upon Ethereum, a public blockchain platform. To engage self-interested participants, we introduce an incentive mechanism which rewards each participant in terms of the amount of its training data and the performance of its local updates. To handle malicious participants, we propose an audit scheme which employs a committee of randomly chosen validators for punishing them with no reward and preclude corrupt updates from the global model. The proposed incentive and audit scheme is implemented with cryptocurrency and smart contract, two primitives offered by Ethereum. This paper demonstrates the main features of Refiner by training a digit classification model on the MNIST dataset.
Zhebin Zhang, Dajie Dong, Yilong Ying, Dawei Jiang, Ke Chen 0005, Lidan Shou, Gang Chen 0001
Proc. VLDB Endow.1
2015 Image guided label map propagation in video sequences
abstract
In this paper, we propose a novel method to transmit the label maps by propagating from a key frame to non-key frames. The label map of a non-key frame is initialized by warping the label map of its corresponding key frame according to the motion estimation between them. Subsequently, the initialized label map is optimized with the guidance of its texture image. The optimization process minimizes an energy function which takes two constraints into consideration: (i) the data term measuring the similarity between an estimated label map and its initialized one, (ii) the regularization term enforcing the local smoothness in the label map and the consistency of region boundaries between the estimated label map and its corresponding texture image. Graph cuts based computation process is finally performed to generate the optimized label map. Experimental results show that our method achieves higher accuracy and better visual quality comparing with the state-of-the-art method.
Shuolin Di, Zhebin Zhang, Shiqi Wang 0001, Nan Zhang 0015, Siwei Ma 0001
ISCAS2
2014 Depth map propagation with the texture image guidance
abstract
We propose a novel method to propagate the depth maps from key-frames to the non-key frames in a video shot. The depths of non-key frames are initialized by warping the depth maps of their corresponding key frames according to the motion estimation between them. Such initial depth maps are then refined with the guidance of their texture frames in a convex optimization process. The energy function in the optimization involves three kinds of constraints, (i) the similarity between the initial depths and the estimated ones, (ii) the spatial constraint to maintain the smoothness of homogeneous regions and region boundaries, and (iii) the temporal constraint to guarantee the temporal smoothness between adjacent frames. In the experiments, we evaluate our method on six video sequences with ground truth depth maps. The experiment results show that our method is robust to many complex scenes and obtain lower error rates of the depth propagation than the baseline method.
Zhebin Zhang
ICIP2
2014 A Compact Representation for Compressing Converted Stereo Videos
abstract
We propose a novel representation for stereo videos namely 2D-plus-depth-cue. This representation is able to encode stereo videos compactly by leveraging the by-product of a stereo video conversion process. Specifically, the depth cues are derived from an interactive labeling process during 2D-to-stereo video conversion—they are contour points of image regions and their corresponding depth models, and so forth. Using such cues and the image features of 2D video frames, the scene depth can be reliably recovered. Experimental results demonstrate that the bit rate can be saved about 10%–50% in coding a stereo video compared with multiview video coding and the 2D-plus-depth methods. In addition, since the objects are segmented in the conversion process, it is convenient to adopt the region-of-interest (ROI) coding in the proposed stereo video coding system. Experimental results show that using ROI coding, the bit rate is reduced by 30%–40% or the video quality is increased by 1.5–4 dB with the fixed bit rate.
Zhebin Zhang, Ronggang Wang, Yizhou Wang 0001, Wen Gao 0001
IEEE Trans. Image Process.1
2013 Interactive Stereoscopic Video Conversion
abstract
This paper presents a system of converting conventional monocular videos to stereoscopic ones. In the system, an input monocular video is firstly segmented into shots so as to reduce operations on similar frames. An automatic depth estimation method is proposed to compute the depth maps of the video frames utilizing three monocular depth cues - depth-from-defocus, aerial perspective, and motion. Foreground/background objects can be interactively segmented on selected key frames and their depth values can be adjusted by users. Such results are propagated from key frames to nonkey frames within each video shot. Equipped with a depth-to-disparity conversion module, the system synthesizes the counterpart (either left or right) view for stereoscopic display by warping the original frames according to their disparity maps. The quality of converted videos is evaluated by human mean opinion scores, and experiment results demonstrate that the proposed conversion method achieves encouraging performance.
Zhebin Zhang, Yizhou Wang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2012 A Compact Stereoscopic Video Representation for 3D Video Generation and Coding
abstract
We propose a novel compact representation for stereoscopic videos - a 2D video and its depth cues. Depth cues are derived from an interactive labeling process during 2D-to-3D video conversion, they are contour points of foreground objects and a background geometric model. By using such cues and image features of 2D video frames, depth maps of the frames can be recovered. Compared with traditional 3D video representation, the proposed one is more compact. We also design algorithms to encode and decode the depth cues. The representation benefits both 3D video generation and coding. Experimental results demonstrate that the bit rate can be saved about 10%-50% in coding 3D videos compared with multi-view video coding and 2D+depth methods. A system coupling 2D-to-3D video conversion and coding (CVCC) is proposed to verify advantages of the representation.
Zhebin Zhang, Ronggang Wang, Yizhou Wang 0001, Wen Gao 0001
DCC1
2012 An interactive system of stereoscopic video conversion
abstract
With the recent booming of 3DTV industry, more and more stereoscopic videos are demanded by the market. This paper presents a system of converting conventional monocular videos to stereoscopic ones. In this system, an input video is firstly segmented into shots to reduce operations on similar frames. Then, automatic depth estimation and interactive image segmentation are integrated to obtain depth maps and foreground/background segments on selected key frames. Within each video shot, such results are propagated from key frames to non-key frames. Combined with a depth-to-disparity conversion method, the system synthesizes the counterpart (either left or right) view for stereoscopic display by warping the original frame according to disparity maps. For evaluation, we use human labeled depth map as the reference and compute both the mean opinion score (MOS) and Peak signal-to-noise ratio (PSNR) to valuate the converted video quality. Experiment results demonstrate that the proposed conversion system and methods achieves encouraging performance.
Zhebin Zhang, Bo Xin, Yizhou Wang 0001, Wen Gao 0001
ACM Multimedia1
2011 Visual pertinent 2D-to-3D video conversion by multi-cue fusion
abstract
We describe an approach to2D-to-3D video conversion for the stereoscopic display. Targeting the problem of synthesizing the frames of a virtual 'right view' from the original monocular 2D video, we generate the stereoscopic video in steps as following. (1) A 2.5D depth map is first estimated in a multi-cue fusion manner by leveraging motion cues and photometric cues in video frames with a depth prior of spatial and temporal smoothness. (2) The depth map is converted to a disparity map with considering both the displaying device size and human's stereoscopic visual perception constraints. (3) We fix the original 2D frames as the 'left view' ones, and warp them to "virtually viewed" right ones according to the predicted disparity value. The main contribution of this method is to combine motion and photometric cues together to estimate depth map. In the experiments, we apply our method to converting several movie clips of well-known films into stereoscopic 3D video and get good results1.
Zhebin Zhang, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001
ICIP1
2011 Stereoscopic learning for disparity estimation
abstract
In this paper, we propose a learning based approach to estimating pixel disparities from the motion information extracted out of input monoscopic video sequences. We represent each video frame with superpixels, and extract the motion features from the superpixels and the frame boundary. These motion features account for the motion pattern of the superpixel as well as camera motion. In the learning phase, given a pair of stereoscopic video sequences, we employ a state-of-the-art stereo matching method to compute the disparity map of each frame as ground truth. Then a multi-label SVM is trained from the estimated disparities and the corresponding motion features. In the testing phase, we use the learned SVM to predict the disparity for each superpixel in a monoscopic video sequence. Experiment results show that the proposed method achieves low error rate in disparity estimation.
Zhebin Zhang, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001
ISCAS1
2010 An interactive method for curve extraction
abstract
We introduce a curve process framework to solve the challenging problem of curve extraction from “non-traceable” curve groups. We propose a comprehensive curve model, which consists of the geometric, photometric and topological sub-models. Two typical categories of the non-traceable curve groups are considered. First, for the interlaced curves with complex structures, we show how to use the proposed curve model especially the topological sub-model to extract curves from the group. Second, for the non-interlaced but over-dense or faint curves we leverage the curve group pattern priors in addition, and extract the whole pattern in a global optimization. Applications and experiments demonstrate the competence of our models and methods.
Ge Guo 0002, Luoqi Liu, Zhebin Zhang, Yizhou Wang 0001, Wen Gao 0001
ICIP3
2005 Video2Cartoon: generating 3D cartoon from broadcast soccer video
abstract
In this demonstration, a prototype system for generating 3D cartoon from broadcast soccer video is proposed. This system takes advantage of computer vision (CV) and computer graphics (CG) techniques to provide users new experience that can not be obtained from original video. Firstly, it uses CV techniques to obtain 3D positions of the players and ball. Then, CG techniques are applied to model the playfield, players, and ball. Finally, 3D cartoon is generated. Our system allows users to watch the game at any point of view using a 3D viewer based on OpenGL.
Dawei Liang, Yang Liu 0006, Qingming Huang, Guangyu Zhu 0002, Shuqiang Jiang, Zhebin Zhang, Wen Gao 0001
ACM Multimedia6