EDBT 2026 Demo / reviewers in the wild / expert
Dong Tian
dblp:51/2308
· DBLP profile ↗
66ranked-venue papers
11as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | T2VTree: User-Centered Visual Analytics for Agent-Assisted Thought-to-Video Authoring
Zhuoyun Zheng, Yu Dong 0001, Gaorong Liang, Guan Li 0002, Guihua Shan, Dong Tian, Jianlong Zhou, Christy Jie Liang |
PacificVis | 7 |
| 2026 | Gaussian Mixture Model-Based Splatting for Rapid Rendering and Time Series Analysis of Large-Scale Particle DataabstractDriven by advances in supercomputing, the scale of scientific simulation data has grown dramatically. In fields such as cosmology, particle data have become a common representation, with state-of-the-art simulations now exceeding the trillion-particle mark. Consequently, the challenge of visually analyzing such massive datasets has become increasingly urgent. The traditional visual analysis workflow typically follows a "compression $\rightarrow$→ storage $\rightarrow$→ reconstruction $\rightarrow$→ visualization" pipeline. However, this process is hampered by an extremely time-consuming reconstruction stage, which severely impedes real-time interactive visualization. Moreover, in multi-time-step analyses, the enormous volume of reconstructed data creates significant I/O bottlenecks. In this work, we draw inspiration from 3D Gaussian splatting and compress the simulation data using Gaussian Mixture Models (GMMs), treating the resulting Gaussian kernels as fundamental rendering primitives. Our method renders billion-scale particles for each timestep in approximately 32 ms, requiring only 645 MB of GPU memory per timestep - nearly 20× smaller than the original 12 GB raw data. This eliminates costly reconstruction, accelerates the visual analysis pipeline, and overcomes I/O bottlenecks in multi-time-step analysis. Extensive experiments and comparisons across multiple datasets validate the effectiveness of our method. Ruixiao Peng, Guan Li 0002, Zhe Wang 0059, Yu Dong 0001, Tianchi Zhang 0003, Xuyi Lu, Yifei Jia, Guihua Shan, Dong Tian |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2025 | Enhancing INR-Based Super-Resolution Performance in Scientific Visualization via a Priori and a Posteriori Constraints
Yang Liu 0469, Guan Li 0002, Weiqun Cao, Guihua Shan, Dong Tian, Zhe Wang 0059 |
CGI (3) | 5 |
| 2025 | Geometry Regularized Point Cloud AutoencoderabstractPoint cloud is a prevalent format in representing 3D geometry. Regardless of the recent advances, unsupervised learning for 3D point clouds remains arduous for various tasks due to its unorganized and sparsely distributed nature. To address this challenge, we propose a geometry regularized point cloud autoencoder, aiming to preserve local geometry structure. In particular, based on the Mahalanobis distance, we propose a point cloud geometry metric counting the local statistics. It endeavors to maximize the posterior probability of the reconstruction conditioned on the input point cloud. Our eigenspace analysis reveals the adaptivity of the developed metric—it behaves differently given different local structures. Moreover, a coarse-to-fine training strategy by varying the metric granularity is applied, leading to our proposed geometry regularized point cloud autoencoder. By applying our proposal to several off-the-shelf point cloud autoencoders, we show an improved point cloud reconstruction quality. In addition, the superior representability of our learned features is also demonstrated via an object classification task. Ritwik Sadhu, Jiahao Pang, Dong Tian |
ICIP | 3 |
| 2025 | TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement LearningabstractThis work introduces Transformer-based Off-Policy Episodic Reinforcement Learning (TOP-ERL), a novel algorithm that enables off-policy updates in the ERL framework. In ERL, policies predict entire action trajectories over multiple time steps instead of single actions at every time step. These trajectories are typically parameterized by trajectory generators such as Movement Primitives (MP), allowing for smooth and efficient exploration over long horizons while capturing high-level temporal correlations. However, ERL methods are often constrained to on-policy frameworks due to the difficulty of evaluating state-action values for entire action sequences, limiting their sample efficiency and preventing the use of more efficient off-policy architectures. TOP-ERL addresses this shortcoming by segmenting long action sequences and estimating the state-action values for each segment using a transformer-based critic architecture alongside an n-step return estimation. These contributions result in efficient and stable training that is reflected in the empirical results conducted on sophisticated robot learning environments. TOP-ERL significantly outperforms state-of-the-art RL methods. Thorough ablation studies additionally show the impact of key design choices on the model performance. Dong Tian, Hongyi Zhou, Xinkai Jiang, Rudolf Lioutikov, Gerhard Neumann |
ICLR | 2 |
| 2025 | A Visual Analysis Approach for Deep Learning-based Precipitation ForecastingabstractMeteorological experts achieved promising results in improving the effectiveness of neural weather networks in precipitation forecasting by incorporating precipitation-optimized objectives into the loss function. However, since neural networks are not based on explicit physical processes, meteorological experts may lack understanding and trust in the model’s predictions, and they cannot perform bias correction by adjusting physical parameters. Additionally, precipitation events are highly imbalanced, with heavy rainfall events being relatively rare but of greater importance. Therefore, traditional metrics for evaluating deep models are insufficient to fully assess precipitation forecasting performance. In this paper, we present a neural weather network visual analysis system designed to help domain experts understand and comprehensively compare the impact of different loss functions on neural networks. We customize the model evaluation process based on the characteristics of precipitation forecasting tasks and provide a reference for bias correction using historically similar data. To validate our approach, we perform two case studies using real-world reanalysis datasets, with feedback from domain experts further confirming its effectiveness. Xuyi Lu, Guan Li 0002, Yu Dong 0001, Dong Tian, Guihua Shan |
PacificVis | 5 |
| 2025 | UH-PCC: Unified Octree and Feature Coding for Hierarchical Point Cloud Geometry Compression
Muhammad Asad Lodhi, Jiahao Pang, Junghyun Ahn, Yuning Huang, Dong Tian |
PCS | 5 |
| 2025 | ClayVolume: A progressive refinement interaction system for immersive visualizationabstractImmersive visualization has become an important tool for discovering hidden patterns and obtaining insights from data. Target acquisition in immersive visualization is a fundamental step in visual analysis. However, limited visual encoding attributes and the presence of stacking and occlusion in immersive environments pose challenges in discovering valuable targets and making unambiguous selections. In this paper, we present ClayVolume, an interactive system designed for immersive visualization. It comprises metaphorical tools for customizing regions of interest (ROIs) and multiple views that serve as interactive and analytical mediums. ClayVolume empowers analysts to efficiently acquire valuable targets through a progressive refinement of interactive methods, enabling further extraction of insights. We evaluate ClayVolume in the scenario of immersive visualization of network data and perform a comparative analysis of its performance against other techniques in target selection tasks. The results indicate that ClayVolume enables flexible target selection in immersive visualization and provides fast target discovery and localization capabilities. • A selection tool for defining ROIs in immersive visualization with depth awareness. • A navigation tool using WiM technology, featuring destination previews for accuracy. • A multi-view tool for data awareness, optimizing target acquisition with spatial cues. Zhenyuan Wang, Guihua Shan, Dong Tian |
Vis. Informatics | 7 |
| 2024 | PIVOT-Net: Heterogeneous Point-Voxel-Tree-based Framework for Point Cloud CompressionabstractThe universality of the point cloud format enables many 3D applications, making the compression of point clouds a critical phase in practice. Sampled as discrete 3D points, a point cloud approximates 2D surface(s) embedded in 3D with a finite bit-depth. However, the point distribution of a practical point cloud changes drastically as its bit-depth increases, requiring different methodologies for effective consumption/analysis. In this regard, a heterogeneous point cloud compression (PCC) framework is proposed. We unify typical point cloud representations-pointbased, voxel-based, and tree-based representations-and their associated backbones under a learning-based framework to compress an input point cloud at different bit-depth levels. Having recognized the importance of voxel-domain processing, we augment the framework with a proposed context-aware upsampling for decoding and an enhanced voxel transformer for feature aggregation. Extensive experimentation demonstrates the state-of-the-art performance of our proposal on a wide range of point clouds. Jiahao Pang, Kevin Bui, Dong Tian |
3DV | 3 |
| 2024 | WrappingNet: Mesh Autoencoder Via Deep Sphere DeformationabstractThere have been recent efforts to learn more meaningful representations via fixed length codewords from mesh data, since a mesh serves as a complete model of underlying 3D shape compared to a point cloud. However, the mesh connectivity presents new difficulties when constructing a deep learning pipeline for meshes. Previous mesh unsupervised learning approaches typically assume category-specific templates, e.g., human face/body templates. It restricts the learned latent codes to only be meaningful for objects in a specific category, so the learned latent spaces are unable to be used across different types of objects. In this work, we present WrappingNet, the first mesh autoencoder enabling general mesh unsupervised learning over heterogeneous objects. It introduces a novel base graph in the bottleneck dedicated to representing mesh connectivity, which is shown to facilitate learning a shared latent space representing object shape. The superiority of WrappingNet mesh learning is further demonstrated via improved reconstruction quality and competitive classification compared to point cloud learning, as well as latent interpolation between meshes of different categories. The code is available at https://github.com/InterDigitalInc/WrappingNet. Eric Lei, Muhammad Asad Lodhi, Jiahao Pang, Junghyun Ahn, Dong Tian |
ICIP | 5 |
| 2024 | Towards Reproducible Learning-Based CompressionabstractA deep learning system typically suffers from a lack of reproducibility that is partially rooted in hardware or software implementation details. The irreproducibility leads to skepticism in deep learning technologies and it can hinder them from being deployed in many applications. In this work, the irreproducibility issue is analyzed where deep learning is employed in compression systems while the encoding and decoding may be run on devices from different manufacturers. The decoding process can even crash due to a single bit difference, e.g., in a learning-based entropy coder. For a given deep learning-based module with limited resources for protection, we first suggest that reproducibility can only be assured when the mismatches are bounded. Then a safeguarding mechanism is proposed to tackle the challenges. The proposed method may be applied for different levels of protection either at the reconstruction level or at a selected decoding level. Furthermore, the overhead introduced for the protection can be scaled down accordingly when the error bound is being suppressed. Experiments demonstrate the effectiveness of the proposed approach for learning-based compression systems, e.g., in image compression and point cloud compression. Jiahao Pang, Muhammad Asad Lodhi, Junghyun Ahn, Yuning Huang, Dong Tian |
MMSP | 5 |
| 2024 | An Improved Genetic-XGBoost Classifier for Customer Consumption Behavior PredictionabstractAbstract In an increasingly competitive market, predicting the customer’s consumption behavior has a vital role in customer relationship management. In this study, a new classifier for customer consumption behavior prediction is proposed. The proposed methods are as follows: (i) A feature selection method based on least absolute shrinkage and selection operator (Lasso) and Principal Component Analysis (PCA), to achieve efficient feature selection and eliminate correlations between variables. (ii) An improved genetic-eXtreme Gradient Boosting (XGBoost) for customer consumption behavior prediction, to improve the accuracy of prediction. Furthermore, the global search ability and flexibility of the genetic mechanism are used to optimize the XGBoost parameters, which avoids inaccurate parameter settings by manual experience. The adaptive crossover and mutation probabilities are designed to prevent the population from falling into the local extremum. Moreover, the grape-customer consumption behavior dataset is employed to compare the six Lasso-based models from the original, normalized and standardized data sources with the Isometric Mapping, Locally Linear Embedding, Multidimensional Scaling, PCA and Kernel Principal Component Analysis methods. The improved genetic-XGBoost is compared with several well-known parameter optimization algorithms and state-of-the-art classification approaches. Furthermore, experiments are conducted on the University of California Irvine datasets to verify the improved genetic-XGBoost algorithm. All results show that the proposed methods outperform the existing ones. The prediction results provide the decision-making basis for enterprises to formulate better marketing strategies. Yue Li 0054, Jianfang Qi, Haibin Jin, Dong Tian, Weisong Mu, Jianying Feng |
Comput. J. | 4 |
| 2024 | Guest Editorial Special Section on Recent Standardization Efforts for Learning-Based Visual Data CodingabstractVisual data coding is an enabling technology for various applications and is now ubiquitously adopted in modern image processing, communications, and computer vision systems. To enable interoperability between devices manufactured and services provided by different enterprises, a series of standards targeting visual data coding have been crafted in the past three decades. Several standardization organizations, such as ISO/IEC JTC 1/SC 29 consisting of Joint Picture Experts Group (JPEG) and Moving Picture Experts Group (MPEG),1ITU-T SG 16 Video Coding Experts Group (VCEG),2IEEE Data Compression Standards Committee Audio Video Coding Working Group (1857 WG),3MPAI Community,4have been creating these standards from many contributions of academia and industry. While most of these visual coding standards have been successfully deployed in many applications, there are more challenges nowadays, especially to accommodate the large volume of visual data in limited storage and limited bandwidth transmission links. Compression efficiency improvements are still needed, especially considering emerging data representation formats ranging from 8K/HDR image/video to rich plenoptic data. Dong Liu 0002, Shan Liu 0001, João Ascenso, Dong Tian, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Sparse Convolution Based Octree Feature Propagation for Lidar Point Cloud CompressionabstractWith the advent of new 3D scanning technologies, point clouds have become a crucial way to depict real and virtual objects/scenes. Point clouds represent the continuous sur-faces of underlying object/scene through a collection (usually millions) of discrete, irregular, and often sparsely distributed 3D samples on the surface of the objects, e.g. LiDAR scans. This nature of point cloud data presents a considerable challenge to not only store but also understand and extract the topology of object(s) from the point cloud data. In this regard, our work presents a point cloud compression procedure that leverages sparse 3D convolutions to extract features at various octree scales for lossless compression of octree representation of point clouds. For hierarchical flow of information between octree levels, our proposed method named SparseContextNet (SCN) also propagates features from a lower resolution scale to higher resolution scale via 3D upsampling convolutions. Our experiments with LiDAR datasets reveal competitive performance of our proposal compared to the state-of-the-art. Muhammad Asad Lodhi, Jiahao Pang, Dong Tian |
ICASSP | 3 |
| 2023 | RecMon: A Deep Learning-based Data Recovery System for Network MonitoringabstractNetwork monitoring systems struggle with the issue that the measurement data is incomplete, with only a subset of origin-destination (OD) pairs or time slots observed, due to the high deployment and measurement cost. Recent studies show that the missing data can be inferred from partial measurements using neural network models and tensor methods. However, these recovery approaches fail to achieve accuracy, adaptability and high speed, simultaneously. In this paper, we propose RecMon, a deep learning-based data recovery system that satisfies the above three criteria. A global spatio-temporal attention mechanism and a data augmentation algorithm are proposed to improve the recovery accuracy. A semi-supervised learning-based scheme is devised for fast and effective model updates. We conduct extensive experiments on three real-world datasets to compare RecMon with four state-of-the-art methods in terms of online recovery performance. The experimental results show that RecMon can adapt to the latest state of the network and accurately recover network measurement data in less than 100 milliseconds. When 90% of the data is missing, the recovery accuracy of RecMon improves over the strongest baseline method by 22.7%, 16.0%, and 8.2% in the three datasets, respectively. Huaiyi Zhao, Xinyi Zhang 0004, Kun Xie 0001, Dong Tian, Gaogang Xie |
INFOCOM | 4 |
| 2023 | DDA-Net: Deep Distribution-Aware Network for Point Cloud CompressionabstractDeep neural networks have been recently applied to point cloud compression (PCC). The features extracted via deep neural networks are essential for compression performance. Different from high level tasks such as point cloud classification or segmentation which homogenizes descriptors within same classes, PCC requires low level features discriminative for point-level 3D reconstructions. With this motivation, we first adopt Gaussian distribution to model the shape of feature elements. Then, we propose a deep distribution-aware network (DDA-Net) which manipulates distributions of feature elements on-the-fly to favor the point cloud reconstruction with high fidelity. Moreover, a residual network is integrated to enhance the modification of the Gaussian models. The proposed DDA-Net is incorporated into an end-to-end PCC system. Experimental results show that our DDA-Net significantly improves the compression performance across a wide range of point clouds. Junghyun Ahn, Jiahao Pang, Muhammad Asad Lodhi, Dong Tian |
ISCAS | 4 |
| 2023 | A survey of immersive visualization: Focus on perception and interactionabstractImmersive visualization utilizes virtual reality, mixed reality devices, and other interactive devices to create a novel visual environment that integrates multimodal perception and interaction. This technology has been maturing in recent years and has found broad applications in various fields. Based on the latest research advancements in visualization, this paper summarizes the state-of-the-art work in immersive visualization from the perspectives of multimodal perception and interaction in immersive environments, additionally discusses the current hardware foundations of immersive setups.By examining the design patterns and research approaches of previous immersive methods, the paper reveals the design factors for multimodal perception and interaction in current immersive environments. Furthermore, the challenges and development trends of immersive multimodal perception and interaction techniques are discussed, and potential areas of growth in immersive visualization design directions are explored. Zhenyuan Wang, Guihua Shan, Dong Tian |
Vis. Informatics | 5 |
| 2022 | Graph Signal Processing for Geometric Data and Beyond: Theory and ApplicationsabstractGeometric data acquired from real-world scenes,e.g., 2D depth images, 3D point clouds, and 4D dynamic point clouds, have found a wide range of applications including immersive telepresence, autonomous driving, surveillance,etc. Due to irregular sampling patterns of most geometric data, traditional image/video processing methodologies are limited, while Graph Signal Processing (GSP)—a fast-developing field in the signal processing community—enables processing signals that reside on irregular domains and plays a critical role in numerous applications of geometric data from low-level processing to high-level analysis. To further advance the research in this field, we provide the first timely and comprehensive overview of GSP methodologies for geometric data in a unified manner by bridging the connections between geometric data and graphs, among the various geometric data modalities, and with spectral/nodal graph filtering techniques. We also discuss the recently developed Graph Neural Networks (GNNs) and interpret the operation of these networks from the perspective of GSP. We conclude with a brief discussion of open problems and challenges. Wei Hu 0003, Jiahao Pang, Xianming Liu 0005, Dong Tian, Chia-Wen Lin, Anthony Vetro |
IEEE Trans. Multim. | 4 |
| 2021 | TearingNet: Point Cloud Autoencoder To Learn Topology-Friendly RepresentationsabstractTopology matters. Despite the recent success of point cloud processing with geometric deep learning, it remains arduous to capture the complex topologies of point cloud data with a learning model. Given a point cloud dataset containing objects with various genera, or scenes with multiple objects, we propose an autoencoder, TearingNet, which tackles the challenging task of representing the point clouds using a fixed-length descriptor. Unlike existing works directly deforming predefined primitives of genus zero (e.g., a 2D square patch) to an object-level point cloud, our TearingNet is characterized by a proposed Tearing network module and a Folding network module interacting with each other iteratively. Particularly, the Tearing network module learns the point cloud topology explicitly. By breaking the edges of a primitive graph, it tears the graph into patches or with holes to emulate the topology of a target point cloud, leading to faithful reconstructions. Experimentation shows the superiority of our proposal in terms of reconstructing point clouds as well as generating more topology-friendly representations than benchmarks. Jiahao Pang, Duanshun Li, Dong Tian |
CVPR | 3 |
| 2021 | FESTA: Flow Estimation via Spatial-Temporal Attention for Scene Point CloudsabstractScene flow depicts the dynamics of a 3D scene, which is critical for various applications such as autonomous driving, robot navigation, AR/VR, etc. Conventionally, scene flow is estimated from dense/regular RGB video frames. With the development of depth-sensing technologies, precise 3D measurements are available via point clouds which have sparked new research in 3D scene flow. Nevertheless, it remains challenging to extract scene flow from point clouds due to the sparsity and irregularity in typical point cloud sampling patterns. One major issue related to irregular sampling is identified as the randomness during point set abstraction/feature extraction—an elementary process in many flow estimation scenarios. A novel Spatial Abstraction with Attention (SA2) layer is accordingly proposed to alleviate the unstable abstraction problem. Moreover, a Temporal Abstraction with Attention (TA2) layer is proposed to rectify attention in temporal domain, leading to benefits with motions scaled in a larger range. Extensive analysis and experiments verified the motivation and significant performance gains of our method, dubbed as Flow Estimation via Spatial-Temporal Attention (FESTA), when compared to several state-of-the-art benchmarks of scene flow estimation. Haiyan Wang 0019, Jiahao Pang, Muhammad Asad Lodhi, Yingli Tian, Dong Tian |
CVPR | 5 |
| 2020 | Automatic Generation of Electromyogram Diagnosis ReportabstractElectrophysiological tests, especially, electromyogram (EMG) and nerve conduction velocity (NCV) test are commonly used in clinical practice for diagnosis of muscle and nerve diseases. Report-writing of these tests can be problematic for under-experienced physicians and time-consuming for experienced physicians. In this paper, we apply several neural based natural language generation (NLG) methods to automatically generate diagnosis reports, a first attempt in this domain. Specifically, we use tabular diagnostic records of electrophysiological tests to generate Findings & Impression, which together constitute the diagnostic report. We further use gram-based metrics to evaluate our models and conduct a case study for the result. Qizheng Gu, Cong Nie, Ruixiang Zou, Wei Chen 0088, Chaojun Zheng, Dongqing Zhu, Xiaojun Mao, Zhongyu Wei, Dong Tian |
BIBM | 9 |
| 2020 | 3D Point Cloud Enhancement Using Graph-Modelled Multiview Depth MeasurementsabstractA 3D point cloud is often synthesized from depth measurements collected by sensors at different viewpoints. The acquired measurements are typically both coarse in precision and corrupted by noise. To improve quality, previous works denoise a synthesized 3D point cloud a posteriori, after projecting the imperfect depth data onto the 3D space. Instead, we enhance depth measurements on the sensed images a priori, exploiting inherent 3D geometric correlation across views, before synthesizing a 3D point cloud from the improved measurements. By enhancing closer to the actual sensing process, we benefit from optimization targeting specifically the depth image formation model, before subsequent processing steps that can further obscure measurement errors. Mathematically, for each pixel row in a pair of rectified viewpoint depth images, we first construct a graph reflecting inter-pixel similarities via metric learning using data in previous enhanced rows. To optimize left and right viewpoint images simultaneously, we write a non-linear mapping function from left pixel row to the right based on 3D geometry relations. We formulate a MAP optimization problem, which, after suitable linear approximations, results in an unconstrained convex and differentiable objective, solvable using fast gradient method (FGM). Experimental results show that our method noticeably outperforms recent denoising algorithms that enhance after 3D point clouds are synthesized. Xue Zhang 0008, Gene Cheung, Jiahao Pang, Dong Tian |
ICIP | 4 |
| 2020 | Acceleration of multi-task cascaded convolutional networksabstractMulti‐task cascaded convolutional neural network (MTCNN) is a human face detection architecture which uses a cascaded structure with three stages (P‐Net, R‐Net and O‐Net). The authors intend to reduce the computation time of the whole process of the MTCNN. They find that the non‐maximum suppression (NMS) processes after the P‐Net occupy over half of the computation time. Therefore, the authors propose a self‐fine‐tuning method which makes the control of computation time for the NMS process easier. Self‐fine‐tuning is a training trick which uses hard samples generated by P‐Net to retrain P‐Net. After self‐fine‐tuning, the distribution of human face probabilities generated by P‐Net is changed, and the tail of distribution becomes thinner. The control of the number of NMS input boxes can be made easier when the distribution has a thinner tail, and choosing a suitable threshold to filter the face boxes will generate less boxes. So the computation time can be reduced. In order to keep the performance of MTCNN, the authors still propose a landmark data set augmentation, which can enhance the performance of the self‐fine‐tuned MTCNN. From the experiments, it is found that the proposed scheme can significantly reduce the computation time of MTCNN. Longhua Ma, Hang-Yu Fan, Zheming Lu 0001, Dong Tian |
IET Image Process. | 4 |
| 2020 | An attentional spatial temporal graph convolutional network with co-occurrence feature learning for action recognition
Dong Tian, Zheming Lu 0001, Longhua Ma |
Multim. Tools Appl. | 1 |
| 2020 | Deep Unsupervised Learning of 3D Point Clouds via Graph Topology Inference and FilteringabstractWe propose a deep autoencoder with graph topology inference and filtering to achieve compact representations of unorganized 3D point clouds in an unsupervised manner. Many previous works discretize 3D points to voxels and then use lattice-based methods to process and learn 3D spatial information; however, this leads to inevitable discretization errors. In this work, we try to handle raw 3D points without such compromise. The proposed networks follow the autoencoder framework with a focus on designing the decoder. The encoder of the proposed networks adopts similar architectures as in PointNet, which is a well-acknowledged method for supervised learning of 3D point clouds. The decoder of the proposed networks involves three novel modules: the folding module, the graph-topology-inference module, and the graph-filtering module. The folding module folds a canonical 2D lattice to the underlying surface of a 3D point cloud, achieving coarse reconstruction; the graph-topology-inference module learns a graph topology to represent pairwise relationships between 3D points, pushing the latent code to preserve both coordinates and pairwise relationships of points in 3D point clouds; and the graph-filtering module couples the above two modules, refining the coarse reconstruction through a learnt graph topology to obtain the final reconstruction. The proposed decoder leverages a learnable graph topology to push the codeword to preserve representative features and further improve the unsupervised-learning performance. We further provide theoretical analyses of the proposed architecture. We provide an upper bound for the reconstruction loss and further show the superiority of graph smoothness over spatial smoothness as a prior to model 3D point clouds. In the experiments, we validate the proposed networks in three tasks, including 3D point cloud reconstruction, visualization, and transfer classification. The experimental results show that (1) the proposed networks outperform the state-of-the-art methods in various tasks, including reconstruction and transfer classification; (2) a graph topology can be inferred as auxiliary information without specific supervision on graph topology inference; (3) graph filtering refines the reconstruction, leading to better performances; and (4) designing a powerful decoder could improve the unsupervised-learning performance, just like a powerful encoder. Siheng Chen, Chaojing Duan, Yaoqing Yang 0002, Duanshun Li, Chen Feng 0002, Dong Tian |
IEEE Trans. Image Process. | 6 |
| 2019 | Graph Based Skeleton Modeling for Human Activity AnalysisabstractUnderstanding human activity based on sensor information is required in many applications and has been an active research area. With the advancement of depth sensors and tracking algorithms, systems for human motion activity analysis can be built by combining off-the-shelf motion tracking systems with application-dependent learning tools to extract higher semantic level information. Many of these motion tracking systems provide raw motion data registered to the skeletal joints in the human body. In this paper, we propose novel representations for human motion data using the skeleton-based graph structure along with techniques in graph signal processing. Methods for graph construction and their corresponding basis functions are discussed. The proposed representations can achieve comparable classification performance in action recognition tasks while additionally being more robust to noise and missing data. Jiun-Yu Kao, Antonio Ortega, Dong Tian, Hassan Mansour, Anthony Vetro |
ICIP | 3 |
| 2018 | Mining Point Cloud Local Structures by Kernel Correlation and Graph PoolingabstractUnlike on images, semantic learning on 3D point clouds using a deep network is challenging due to the naturally unordered data structure. Among existing works, PointNet has achieved promising results by directly learning on point sets. However, it does not take full advantage of a point's local neighborhood that contains fine-grained structural information which turns out to be helpful towards better semantic learning. In this regard, we present two new operations to improve PointNet with a more efficient exploitation of local structures. The first one focuses on local 3D geometric structures. In analogy to a convolution kernel for images, we define a point-set kernel as a set of learnable 3D points that jointly respond to a set of neighboring data points according to their geometric affinities measured by kernel correlation, adapted from a similar technique for point cloud registration. The second one exploits local high-dimensional feature structures by recursive feature aggregation on a nearest-neighbor-graph computed from 3D positions. Experiments show that our network can efficiently capture local information and robustly achieve better performances on major datasets. Our code is available at http://www.merl.com/research/license#KCNet. Yiru Shen, Chen Feng 0002, Yaoqing Yang 0002, Dong Tian |
CVPR | 4 |
| 2018 | FoldingNet: Point Cloud Auto-Encoder via Deep Grid DeformationabstractRecent deep networks that directly handle points in a point set, e.g., PointNet, have been state-of-the-art for supervised learning tasks on point clouds such as classification and segmentation. In this work, a novel end-to-end deep auto-encoder is proposed to address unsupervised learning challenges on point clouds. On the encoder side, a graph-based enhancement is enforced to promote local structures on top of PointNet. Then, a novel folding-based decoder deforms a canonical 2D grid onto the underlying 3D object surface of a point cloud, achieving low reconstruction errors even for objects with delicate structures. The proposed decoder only uses about 7% parameters of a decoder with fully-connected neural networks, yet leads to a more discriminative representation that achieves higher linear SVM classification accuracy than the benchmark. In addition, the proposed decoder structure is shown, in theory, to be a generic architecture that is able to reconstruct an arbitrary point cloud from a 2D grid. Our code is available at http://www.merl.com/research/license#FoldingNet. Yaoqing Yang 0002, Chen Feng 0002, Yiru Shen, Dong Tian |
CVPR | 4 |
| 2017 | Contour-enhanced resampling of 3D point clouds via graphsabstractTo reduce storage and computational cost for processing and visualizing large-scale 3D point clouds, an efficient resampling strategy is needed to select a representative subset of 3D points that can preserve contours in the original 3D point cloud. We tackle this problem by using graph-based techniques as graphs can represent underlying surfaces and lend themselves well to efficient computation. We first construct a general graph for a 3D point cloud and then propose a graph-based metric to quantify the contour information via high-pass graph filtering. Finally, we obtain an optimal resampling distribution that preserves the contour information by solving an optimization problem. When browsing, the proposed graph-based resampling performs better than uniform resampling both for toy point clouds as well as real large-scale point clouds. Furthermore, as neither mesh construction nor surface normal calculation is involved, the proposed graph-based method is computationally more efficient than the mesh-based methods. Siheng Chen, Dong Tian, Chen Feng 0002, Anthony Vetro, Jelena Kovacevic |
ICASSP | 2 |
| 2017 | Disc-GLasso: Discriminative graph learning with sparsity regularizationabstractLearning graph topology from data is challenging. Previous work leads to learning graphs on which the graph signals used for training are smooth. In this paper, we propose an optimization framework for learning multiple graphs, each associated to a class of signals, such that representation of signals within a class and discrimination of signals in different classes are both taken into consideration. A Fisher-LDA-like term is included in the optimization objective function in addition to the conventional Gaussian ML objective. A block coordinate descent algorithm is then developed to estimate optimal graphs for different categories of signals, which are then used to efficiently classify the different signals. Experiments on synthetic data demonstrate that our proposed method can achieve better discrimination between the learned graphs, leading to improvements in subsequent classification tasks. Jiun-Yu Kao, Dong Tian, Hassan Mansour, Antonio Ortega, Anthony Vetro |
ICASSP | 2 |
| 2017 | Compression of 3-D point clouds using hierarchical patch fittingabstractFor applications such as virtual reality and mobile mapping, point clouds are an effective means for representing 3-D environments. The need for compressing such data is rapidly increasing, given the widespread use and precision of these systems. This paper presents a method for compressing organized point clouds. 3-D point cloud data is mapped to a 2-D organizational grid, where each element on the grid is associated with a point in 3-D space and its corresponding attributes. The data on the 2-D grid is hierarchically partitioned, and a Bezier patch is fit to the 3-D coordinates associated with each partition. Residual values are quantized and signaled along with data necessary to reconstruct the patch hierarchy in the decoder. We show how this method can be used to process point clouds captured by a mobile-mapping system, in which laser-scanned point locations are organized and compressed. The performance of the patch-fitting codec exceeds or is comparable to that of an octree-based codec. Robert A. Cohen, Maja Krivokuca, Chen Feng 0002, Yuichi Taguchi, Hideaki Ochimizu, Dong Tian, Anthony Vetro |
ICIP | 6 |
| 2017 | Geometric distortion metrics for point cloud compressionabstractIt is challenging to measure the geometry distortion of point cloud introduced by point cloud compression. Conventionally, the errors between point clouds are measured in terms of point-to-point or point-to-surface distances, that either ignores the surface structures or heavily tends to rely on specific surface reconstructions. To overcome these drawbacks, we propose using point-to-plane distances as a measure of geometric distortions on point cloud compression. The intrinsic resolution of the point clouds is proposed as a normalizer to convert the mean square errors to PSNR numbers. In addition, the perceived local planes are investigated at different scales of the point cloud. Finally, the proposed metric is independent of the size of the point cloud and rather reveals the geometric fidelity of the point cloud. From experiments, we demonstrate that our method could better track the perceived quality than the point-to-point approach while requires limited computations. Dong Tian, Hideaki Ochimizu, Chen Feng 0002, Robert A. Cohen, Anthony Vetro |
ICIP | 1 |
| 2016 | Point Cloud Attribute Compression Using 3-D Intra Prediction and Shape-Adaptive TransformsabstractWith the increased proliferation of applications using 3-D capture technologies for applications such as virtual reality, mobile mapping, scanning of historical artifacts, and 3-D printing, representing these kinds of data as 3-Dpoint clouds has become a popular method for storing and conveying the data independently of how it was captured. A point cloud consists of a set of coordinates indicating the location of each point, along with one or more attributes such as color associated with each point. Because the size of point cloud data can be quite large, compression is needed to efficiently store or transmit this data. This paper, motivated by techniques currently being used for image and video coding, proposes methods using 3-D block-based prediction and transform coding to compress point cloud attributes. Experimental results using a modified shape-adaptive DCT tailored for use in 3-D point clouds and a benchmark using 3-D graph transforms are shown. Robert A. Cohen, Dong Tian, Anthony Vetro |
DCC | 2 |
| 2016 | Geometric-guided label propagation for moving object detectionabstractMoving object segmentation in video has uses in many applications and is a particularly challenging task when the video is acquired by a moving camera. Typical approaches that rely on principal component analysis (PCA) tend to extract scattered sparse components of the moving objects and generally fail in extracting dense object segmentations. In this paper, a novel label propagation framework based on motion vanishing point (MVP) analysis is proposed to address the challenges. A weighted graph is constructed with image pixels as nodes and the MVP-guided approach is used to define the graph weights. Label propagation is then performed by incorporating the graph Laplacian. In addition, a PCA result is used to initialize the foreground/background labels. Experiments on the Hopkins data set of outdoor sequences captured by a hand-held moving camera demonstrate that the proposed label propagation method outperforms state-of-the-art PCA and spectral clustering methods for a dense segmentation task. Moreover, the framework is capable of correcting mislabeled foreground pixels and thus does not require accurate initial label assignment. Jiun-Yu Kao, Dong Tian, Hassan Mansour, Anthony Vetro, Antonio Ortega |
ICASSP | 2 |
| 2016 | Attribute compression for sparse point clouds using graph transformsabstractWith the recent improvements in 3-D capture technologies for applications such as virtual reality, preserving cultural artifacts, and mobile mapping systems, new methods for compressing 3-D point cloud representations are needed to reduce the amount of bandwidth or storage consumed. For point clouds having attributes such as color associated with each point, several existing methods perform attribute compression by partitioning the point cloud into blocks and reducing redundancies among adjacent points. If, however, many blocks are sparsely populated, few or no points may be adjacent, thus limiting the compression efficiency of the system. In this paper, we present two new methods using block-based prediction and graph transforms to compress point clouds that contain sparsely-populated blocks. One method compacts the data to guarantee one DC coefficient for each graph-transformed block, and the other method uses a K-nearest-neighbor extension to generate more efficient graphs. Robert A. Cohen, Dong Tian, Anthony Vetro |
ICIP | 2 |
| 2016 | Moving object segmentation using depth and optical flow in car driving sequencesabstractSegmentation of moving objects in a scene is difficult for non-stationary cameras, and especially challenging in the presence of fast and unstable egomotion, e.g., as encountered with car-mounted cameras or wearable devices. Based on an analysis of motion vanishing points of the scene and estimated depth, a geometric model that relates extracted 2D motion to a 3D motion field relative to the camera is derived. Observing that the 3D motion field is piece-wise smooth, a constrained optimization problem that considers group sparsity is formulated to recover the 3D motion field from the 2D motion. The recovered 3D motion field is then clustered to provide the segmentation of moving objects. Experiments are performed using the KITTI Vision Benchmark Suite and demonstrate that the proposed framework provides a dense segmentation of moving objects that is robust to the challenging conditions inherent with car driving sequences. Jiun-Yu Kao, Dong Tian, Hassan Mansour, Anthony Vetro, Antonio Ortega |
ICIP | 2 |
| 2016 | Keypoint trajectory coding on compact descriptor for video analysisabstractIn contrast to still image analysis, motion information offers a powerful means to analyze video. In particular, motion trajectories determined from keypoints have become very popular in recent years for a variety of video analysis tasks, including search, retrieval and classification. Additionally, cloud-based analysis of media content has been gaining momentum, so efficient communication of salient video information to perform the necessary analysis of video at the cloud server is needed. This paper describes a novel framework to efficiently represent the keypoint trajectories. In particular, an interframe prediction is designed with the option to operate in a low-delay mode. Additionally, a scalable coding method is proposed that allows for a subset of the coded trajectories in a video segment to be easily accessed. Experimental results on several popular datasets including Stanford MAR and Hopkin155 demonstrate a significant rate saving of up to 25% with our proposed trajectory coding approaches relative to a state-of-the-art reference approach. Dong Tian, Huifang Sun, Anthony Vetro |
ICIP | 1 |
| 2016 | Robust low rank dynamic mode decomposition for compressed domain crowd and traffic flow analysisabstractIn this paper, we develop a dynamic mode decomposition algorithm that is robust to both inlier and outlier noise in the data. One application of our algorithm is the identification of multiple crowd or traffic flows from compressed video streams. Our method uses motion vectors that are readily available in the compressed bitstream, and do not require computationally expensive optical flow. These motion vectors are known to be very noisy, however, our algorithm is able to extract the underlying dynamical systems that define the flows. We formulate a rank regularized dynamic mode decomposition problem with total least squares constraints to estimate the Koopman modes of the motion dynamics. The estimated Koopman modes are then used to analyze the stability of the system and extract steady state and transient flows. We demonstrate the improved performance of our approach compared to state of the art schemes and illustrate it applicability in identifying transient and steady-state flows in real video sequences. Caglayan Dicle, Hassan Mansour, Dong Tian, Mouhacine Benosman, Anthony Vetro |
ICME | 3 |
| 2016 | An accurate eye pupil localization approach based on adaptive gradient boosting decision treeabstractEye pupil localization is an important part in computer vision applications such as face recognition, gaze estimation and so on. In this paper, we propose an improved method for precise and fast eye pupil localization. Based on gradient boosting decision tree(GBDT) algorithm, a more accurate localization is achieved by increasing the weight of the training samples with larger errors in a moderate rate. Furthermore, a pruning strategy is utilized to avoid overfitting and reduce the localization time without accuracy loss. Experimental results show that the improved method achieves an accuracy of 92.39% at a speed as fast as 1.7ms to locate in the range of eye pupil on BioID database. The proposed method outperforms most state-of-the-art methods in terms of localization accuracy and consumed time. Dong Tian, Jiaxiang Wu 0002, Hongtao Chen |
VCIP | 1 |
| 2015 | Depth-weighted group-wise principal component analysis for video foreground/background separationabstractWe propose a depth-weighted group-wise PCA (DG-PCA) approach to separate moving foreground pixels from the background of a video acquired by a moving camera. Our approach utilizes a corresponding depth signal in addition to the video signal. The problem is formulated as a weighted l2,1-norm PCA problem with depth-based group sparsity being introduced. In particularly, dynamic groups are first generated solely based on depth, and then an iterative solution using depth to define the weights in l2,1-norm is developed. In addition, we propose a depth-enhanced homography model for global motion compensation before the DG-PCA method is executed. We demonstrate through experiments on an RGB-D dataset the superiority of the proposed DG-PCA approach over conventional robust PCA methods. Dong Tian, Hassan Mansour, Anthony Vetro |
ICIP | 1 |
| 2015 | Graph spectral motion segmentation based on motion vanishing point analysisabstractMotion segmentation relies on identifying coherent relationships between image pixels that are associated with motion vectors. However, perspective differences can often deteriorate the performance of conventional techniques. In this paper, we develop a motion segmentation scheme that utilizes the motion map of a single frame to identify motion representations based on motion vanishing points. Segmentation is achieved using graph spectral clustering where a novel graph is constructed using the motion representation distances in the motion vanishing point image associated with the image pixels. Experimental results show that the proposed graph spectral motion segmentation algorithm outperforms state-of-the-art methods for dense segmentation on image sequences with strong perspective effects using motion vectors between only two images. Dong Tian, Jiun-Yu Kao, Hassan Mansour, Anthony Vetro |
MMSP | 1 |
| 2015 | Depth Map Coding Optimization Using Rendered View Distortion for 3D Video CodingabstractIn order to improve 3D video coding efficiency, we propose methods to estimate rendered view distortion in synthesized views as a function of the depth map quantization error. Our approach starts by calculating the geometric error caused by the depth map error based on the camera parameters. Then, we estimate the rendered view distortion based on the local video characteristics. The estimated rendered view distortion is used in the rate-distortion optimized mode selection for depth map coding. A Lagrange multiplier is derived using the proposed distortion metric, which is estimated based on an autoregressive model. Experimental results show the efficiency of the proposed methods, with average savings of 43% in depth map bitrate as compared with encoding the depth maps using the same coding tools but with the rate-distortion optimization based on the conventional distortion metric. Woo-Shik Kim, Antonio Ortega, PoLin Lai, Dong Tian |
IEEE Trans. Image Process. | 4 |
| 2014 | A graph-based joint bilateral approach for depth enhancementabstractDepth images are often presented at a lower spatial resolution, either due to limitations in the acquisition of the depth or to increase compression efficiency. As a result, upsampling low-resolution depth images to a higher spatial resolution is typically required prior to depth image based rendering. In this paper, depth enhancement and up-sampling techniques are proposed using a graph-based formulation. In one scheme, the depth is first upsampled using a conventional method, then followed by a graph-based joint bilateral filtering to enhance edges and reduce noise. A second scheme avoids the two-step processing and upsamples the depth directly using the proposed graph-based joint bilateral upsampling. Both filtering and interpolation problems are formulated as regularization problems and the solutions are different from conventional approaches. Further, we also studied operations on different graph structures such as star graph and 8-connected graph. Experimental results show that the proposed methods produce slightly more accurate depth at the full resolution with improved rendering quality of intermediate views. Yongzhe Wang, Antonio Ortega, Dong Tian, Anthony Vetro |
ICASSP | 3 |
| 2014 | Depth-assisted stereo video enhancement using graph-based approachesabstractIn stereo video applications, the quality of the two views may vary based on different camera capturing conditions and setup, compression/transmission, and sensor noise. Although some studies show that the perceived video quality may not be significantly affected by the lower quality view, maintaining a similar video quality is still desired in order to prevent eye strain during extended viewing sessions. In this paper, we study a graph-based approach to enhance the lower quality views by referring to the high quality view in addition to an accompanying depth map. We construct a graphical signal model with joint bilateral edge weights and show that graph-based joint bilateral filtering can better suppress several types of noises, e.g., Gaussian, motion as well as quantization noise. Dong Tian, Hassan Mansour, Anthony Vetro, Yongzhe Wang, Antonio Ortega |
ICIP | 1 |
| 2014 | View Synthesis Prediction in the 3-D Video Coding Extensions of AVC and HEVCabstractAdvanced multiview video systems are able to generate intermediate viewpoints of a 3-D scene. To enable low-complexity free view generation, texture and its associated depth are used as input data for each viewpoint. To improve the coding efficiency of such content, view synthesis prediction (VSP) is proposed to further reduce interview redundancy in addition to traditional disparity compensated prediction. This paper describes and analyzes rate-distortion optimized VSP designs, which were adopted in the 3-D extensions of both Advanced Video Coding (AVC) and High Efficiency Video Coding (HEVC). In particular, we propose a novel backward-VSP scheme using a derived disparity vector, as well as efficient signalling methods in the context of AVC and HEVC. In addition, we put forward a novel depth-assisted motion vector prediction method to optimize the coding efficiency. A thorough analysis of coding performance is provided using different VSP schemes and configurations. Experimental results demonstrate average bit rate reductions of 2.5% and 1.2% in AVC and HEVC coding frameworks, respectively, with up to 23.1% bit rate reduction for dependent views. Feng Zou 0006, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au, Shinya Shimizu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | An Analytical Model for Synthesis Distortion Estimation in 3D VideoabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. The model relates errors in the depth images to the synthesis quality, taking into account texture image characteristics, texture image quality, and the rendering process. Especially, we decompose the synthesis distortion into texture-error induced distortion and depth-error induced distortion. We analyze the depth-error induced distortion using an approach combining frequency and spatial domain techniques. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Thus, the model can be used to estimate the rendering quality for different system designs. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au |
IEEE Trans. Image Process. | 3 |
| 2013 | Disparity estimation of misaligned images in a scanline optimization frameworkabstractModern, state-of-the-art disparity estimation techniques are able to very accurately estimate the disparity for a wide variety of scene types. However all of these methods assume that the input images are epipolar rectified. When an image pair is not rectified, it must be pre-processed before any estimation can be done. In this paper we propose a disparity estimation scheme that is able to handle non-rectified images without requiring a rectification step. We show how a minor modification to an existing estimation framework can allow for any disparity estimation framework to produce disparity maps for non-rectified images. Richard Rzeszutek, Dong Tian, Anthony Vetro |
ICASSP | 2 |
| 2013 | Synthesis distortion estimation in 3D video using frequency and spatial analysisabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. Specifically, we estimate the depth-error induced distortion using an approach that combines frequency and spatial domain analysis. We also propose to decompose the spatial-variant video signals into gradient-based representations to capture the interaction between image gradients, depth errors and synthesis distortion. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Lu Yu 0003 |
ICIP | 3 |
| 2013 | Novel distortion metric for depth coding of 3D videoabstractIn state-of-the-art HEVC-based 3D video codec, multiview video plus associated depth maps are used. In order to achieve better coding performance, instead of the conventional sum of squared errors (SSE), view synthesis optimization (VSO) is proposed and included in the anchor encoder software to calculate view synthesis distortion in rate-distortion optimization (RDO) of depth coding. The anchor VSO achieves high rate-distortion (RD) performance. However, it requires partial rendering and is quite complex and time-consuming. On the other hand, simple SSE metric is fast but RD performance is low. In this paper, we propose a new distortion metric to be used in RDO for depth coding. The complexity of the proposed method is slightly higher than SSE, while its RD performance remains competitive. With a good trade-off between complexity and performance, the proposed method can replace the conventional SSE metric in RDO for depth coding, and can be used as a low-complexity alternative to the anchor VSO. Ngai-Man Cheung, Oscar C. Au, Dong Tian |
ICIP | 4 |
| 2013 | Backward view synthesis prediction for 3D-HEVCabstractView synthesis prediction provides an effective way to reduce inter-view redundancy of multiview video in addition to conventional disparity compensated prediction. Traditional forward warping techniques incur high complexity since an entire picture is typically warped from one viewpoint to another. To reduce this complexity, block-based backward warping is considered as an alternative solution. One difficulty with this approach is that it requires depth information of the current block prior to its encoding. To solve this problem, a novel method is proposed to derive the depth information from neighboring blocks with high accuracy. With the proposed approach, backward warping is enabled using depth information derived from either neighboring spatial (inter-view) compensated blocks or temporal compensated blocks. As a generalization, the proposed backward warping scheme is not only applied at the pixel level, but at the sub-block level as well. Simulation results demonstrate that the proposed scheme achieves an average bitrate savings of 1.2% for coded video vs coded video bitrate, 1.1% for coded video vs total bitrate, and 1.0% for synthesized video vs total bitrate under common test conditions for 3D video coding using HEVC, with maximum gains of greater than 10% for dependent views. Dong Tian, Feng Zou 0006, Anthony Vetro |
ICIP | 1 |
| 2013 | View synthesis prediction using adaptive depth quantization for 3D video codingabstractAdvanced multiview video systems are able to generate intermediate viewpoints of a 3D scene. In addition to the texture content, corresponding depth is associated with each viewpoint. To improve the coding efficiency of such content, view synthesis prediction can be used to further reduce inter-view redundancy in addition to traditional disparity compensated prediction. However, the predictor generated from the view synthesis process is affected by several factors, including signal properties of the texture, the accuracy of the depth and complexity of the scene, as well as coding errors in both the texture and depth. This paper presents an analysis of view synthesis prediction performance considering these factors. Based on this analysis, an adaptive depth quantization scheme is proposed to improve the depth coding, leading to better view synthesis prediction and overall coding efficiency gains. The proposed scheme is able to achieve an average bit rate savings of 0.9% on the coded and synthesized video with a maximum gain of up to 11.7% on the dependent views in the context of an HEVC-based codec. Feng Zou 0006, Dong Tian, Anthony Vetro, Antonio Ortega |
ICIP | 2 |
| 2013 | View synthesis prediction using skip and merge candidates for HEVC-based 3D video codingabstractTraditional multi-view coding (MVC) systems compress the texture content captured from different view points, where temporal and inter-view redundancy are exploited to improve MVC coding efficiency. The advanced 3D video coding systems compress both the texture content and its corresponding depth captured from different view points, known as multiview video plus depth (MVD), to support low complexity free view point applications. However, MVD systems consist of a large amount of data including both texture and depth to be compressed and transmitted. To improve the coding efficiency of MVD systems, view synthesis prediction (VSP) can be used to further reduce inter-view redundancy using synthetic views as predictors. In this paper, an in-loop view synthesis framework is proposed, where the synthesized predictor is encoded as a special motion compensated predictor and the motion information is encoded as one of the motion predictors in skip/merge candidate list for HEVC-based 3D video coding. The proposed scheme is applicable to both texture coding and depth coding. The experimental results show that the proposed framework improved the coding performance up to 12.1% for dependent views. Feng Zou 0006, Dong Tian, Anthony Vetro |
ISCAS | 2 |
| 2012 | On modeling the rendering error in 3D videoabstractWe propose an analytical model to estimate the rendering quality in 3D video. The model relates errors in the depth images to the rendering quality, taking into account texture image characteristics, texture image quality, the camera configuration and the rendering process. Specifically, we derive position (disparity) errors from the depth errors, and the probability distribution of the position errors is used to calculate the power spectral density of the rendering errors. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that the model can accurately estimate the synthesis noise up to a constant offset. Thus, the model can be used to estimate the change in rendering quality for different system designs. Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun |
ICIP | 2 |
| 2012 | Local depth image enhancement scheme for view synthesisabstractThe quality of the depth map is crucial for depth image based rendering (DIBR) which enables a variety of advanced 3D video related applications such as perceived depth adjustment for stereoscopic video and intermediate view generation for multiview auto-stereoscopic displays. However, the input depth map for DIBR may suffer from errors and noise, which could seriously impact the rendered view quality. In order to reduce the errors and suppress the noise in the depth map, a local depth image enhancement technique is proposed that leverages trellis-based optimization techniques. A cost function is used to evaluate candidate depth values based on the stereo cost as well as color and depth consistency. Sparse depth features are also used in the enhancement process. The experimental results show notable subjective improvements in terms of rendering quality. Yongzhe Wang, Dong Tian, Anthony Vetro |
ICIP | 2 |
| 2011 | A trellis-based approach for robust view synthesisabstractView synthesis is an essential function for a number of 3D video applications including free-viewpoint navigation and view generation for auto-stereoscopic displays. Depth Image Based Rendering (DIBR) techniques are typically applied for this purpose. However, the quality of the rendered views is very sensitive to the quality of the depth image. In this paper, a novel trellis-based view synthesis framework is proposed to overcome the above limitations in depth images and reduce artifacts in the rendered picture. Our results demonstrate that the proposed approach offers visible improvements in rendering quality compared to existing view synthesis techniques. Dong Tian, Anthony Vetro, Matthew Brand |
ICIP | 1 |
| 2010 | NN-SA Based Dynamic Failure Detector for Services Composition in Distributed Environment
Changze Wu, Kaigui Wu, Dong Tian |
ADMA (2) | 4 |
| 2010 | An Identity-Based Authentication Protocol for Clustered ZigBee Network
Xiaoshuan Zhang, Dong Tian, Zetian Fu |
ICIC (2) | 3 |
| 2010 | Sparse dyadic mode for depth map compressionabstractIn order to enable new video applications such as 3DTV and free-viewpoint video, new data formats including both 2D video sequences and corresponding depth map sequences have been proposed. One major characteristic making the depth maps different from video frames is that they typically consist of homogeneous areas separated by sharp edges representing depth discontinuities. Another characteristic of depth map sequences is that the edges exhibit quite similar boundary behaviors as the edges in the corresponding video frames. In this paper, we propose a novel sparse dyadic mode in the design of an efficient depth map compression algorithm through appropriately exploiting these characteristics. With sparse representations of depth blocks and effective reference of edge information from the corresponding video frames, sparse dyadic mode can achieve up to 1.5 dB gain on rendering quality as compared to depth sequences coded using MVC at the same bitrate. Shujie Liu 0001, PoLin Lai, Dong Tian, Cristina Gomila, Chang Wen Chen |
ICIP | 3 |
| 2010 | Depth map processing with iterative joint multilateral filteringabstractDepth maps estimated using stereo matching between frames from different video views typically exhibit false contours and noisy artifacts around object boundaries. In this paper, iterative joint multilateral filtering is proposed to deal with these artifacts. The proposed filter consists of multiple filter kernels. Knowing that the estimated depth maps are erroneous, besides the kernels which measure the proximity of depth samples and the similarity between depth sample values, we further develop kernels which measure similarity between the corresponding video pixel values. To increase reliability, these novel kernels operate on the color (RGB) domain instead of only on the luminance domain. Furthermore, the filter shapes are designed to adapt brightness variations. Finally, to tackle large misalignment between boundaries in depth maps and in the corresponding video frames, iterative approach is utilized. Our results demonstrate that the proposed method can significantly improve the boundaries in depth maps and can reduce false contours. With the processed depth maps, it is observed that the quality of object boundaries in synthesized views can be improved. PoLin Lai, Dong Tian, Patrick Lopez |
PCS | 2 |
| 2010 | Suppressing texture-depth misalignment for boundary noise removal in view synthesisabstractDuring view synthesis based on depth maps, also known as Depth-Image-Based Rendering (DIBR), annoying artifacts are often generated around foreground objects, yielding the visual effects that slim silhouettes of foreground objects are scattered into the background. The artifacts are referred as the boundary noises. We investigate the cause of boundary noises, and find out that they result from the misalignment between texture and depth information along object boundaries. Accordingly, we propose a novel solution to remove such boundary noises by applying restrictions during forward warping on the pixels within the texture-depth misalignment regions. Experiments show this algorithm can effectively eliminate most boundary noises and it is also robust for view synthesis with compressed depth and texture information. Yin Zhao, Dong Tian, Ce Zhu, Lu Yu 0003 |
PCS | 3 |
| 2010 | Joint trilateral filtering for depth map compressionabstractNew data formats including 2D video and the corresponding depth maps enable new video applications in which virtual views can be rendered, such as 3DTV and free-viewpoint video (FVV). Different from video frames, depth maps typically consist of homogeneous areas (with no textures) separated by sharp edges representing depth value changes such as between foreground and background. Conventional video coding techniques with transforms followed by quantization typically result in large artifacts along such sharp edges. To suppress these coding artifacts while preserving edges, we propose in this paper a novel filtering method for depth coding, joint trilateral filter. The main contribution in the proposed filter design is the utilization of edge information in the collocated video frame as well as in the depth map. The filtering weights are determined by the following three factors: a domain (spatial) filter which measures the proximity of pixel positions, and two range filters. One range filter takes into account the similarity among depth samples and the other one considers the similarity among the collocated pixels in the video frame. By replacing the deblocking filter in H.264/AVC with the proposed trilateral filter, simulation results demonstrate up to 0.8 dB gain in rendering quality at given bitrate for depth signal. Shujie Liu 0001, PoLin Lai, Dong Tian, Cristina Gomila, Chang Wen Chen |
VCIP | 3 |
| 2009 | Depth map distortion analysis for view rendering and depth codingabstractVideo representations that support view synthesis based on depth maps, such as multiview plus depth (MVD), have been recently proposed, raising interest in efficient tools for depth map coding. In this paper, we derive a new distortion metric that takes into consideration camera parameters and global video characteristics in order to quantify the effect of lossy coding of depth maps on synthesized view quality. In addition, a new skip mode selection method is proposed based on local video characteristics. Experimental results with the proposed mode selection scheme show coding gains of up to 2 dB for the synthesized views, as well as better subjective quality. Woo-Shik Kim, Antonio Ortega, PoLin Lai, Dong Tian, Cristina Gomila |
ICIP | 4 |
| 2009 | Improving the quality of depth image based rendering for 3D Video systemsabstractIn 3D video (3DV) applications, a reduced number of views plus depth maps are transmitted or stored. When there is a need to render virtual views in between the actual views, the technique of depth image based rendering (DIBR) can be used to generate the intermediate views. To address the problem of noisy depth information in 3DV systems, we propose novel methods that can be easily incorporated into DIBR to improve synthesized image quality. These include: (1) a heuristic scheme with adaptive spatting that blends multiple warped reference pixels based on their depth, warped pixel positions and camera parameters; (2) an approximation of the first scheme with up-sampling for fast processing; (3) boundary only splatting; and (4) view weighting based on hole distribution. Experiment results show that the proposed methods can improve synthesis quality significantly. Zefeng Ni, Dong Tian, Sitaram Bhagavathy, Joan Llach, B. S. Manjunath |
ICIP | 2 |
| 2009 | Applying evolutionary prototyping model in developing FIDSS: An intelligent decision support system for fish disease/health management
Xiaoshuan Zhang, Zetian Fu, Wengui Cai, Dong Tian |
Expert Syst. Appl. | 4 |
| 2007 | Fuzzy-Grey Prediction Based Dynamic Failure Detector for Distributed Systems
Dong Tian, Taiping Mao |
ICA3PP | 1 |
| 2002 | Coding of faded scene transitionsabstractCoding of a scene transition is often a challenging problem, from the compression efficiency point of view, because motion compensation may not be a powerful enough method to represent changes between pictures in the transition. This paper proposes a overlay coding technique for coding faded scene transitions. As shown by extensive simulations, over 50% bit-rate savings in both cross-fades and through-black fades compared to earlier techniques can be achieved. Overlay coding suits situations where video is edited manually or automatically. Dong Tian, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj |
ICIP (2) | 1 |