Tao Yang 0006

dblp:67/1120-6 · DBLP profile ↗
← Back
49ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-5180-2316ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 2 since 2021Computer networks · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 LD3DGS-SLAM: Long-Distance Monocular SLAM With 3-D Gaussian Splatting and GNSS-Aided Localization for UAVs
abstract
Integrating neural rendering with simultaneous localization and mapping (SLAM) has shown great promise for achieving high-precision localization and photorealistic scene reconstruction. However, the dynamic perspectives of uncrewed aerial vehicles (UAVs), compounded by Global Navigation Satellite System (GNSS) signal interference and the accumulation of localization errors, pose significant challenges to existing monocular simultaneous localization and mapping (SLAM) methods in complex environments. To address these limitations, we propose LD3DGS-SLAM, a novel framework that integrates monocular SLAM, GNSS, and 3D Gaussian splatting (3DGS) [1] to enhance UAV-based localization and mapping. The system integrates traditional geometric feature-based localization with the multiview synthesis capability of 3D Gaussian rendering, achieving a breakthrough in accuracy while enhancing robustness. First, we construct a multisensor fusion graph optimization model that tightly integrates GNSS and monocular vision data, effectively mitigating the cumulative drift typical of traditional SLAM pipelines. Afterward, we introduce a 3D Gaussian-based mapping strategy that incrementally refines a dense scene representation using SLAM-generated point clouds, fusing geometric and texture information to reconstruct highly accurate and detailed maps. To further enhance localization performance in GNSS-denied scenarios, we develop a dynamic frame interpolation and tracking approach based on 3D Gaussian rendering. By synthesizing novel viewpoints, this method improves observation matching and strengthens loop closure detection, enabling robust relocalization in large-scale environments. To validate the effectiveness of the proposed framework, we implement a UAV test system and evaluate its performance on both the publicly available TUM [2] dataset and a self-collected aerial dataset spanning several kilometers. Experimental results show that LD3DGS-SLAM achieves a state-of-the-art average localization error of only 0.8 meters using just a monocular camera, even during long-range flight missions exceeding tens of kilometers and altitudes up to 500 meters. Overall, LD3DGS-SLAM effectively addresses the limitations of monocular SLAM in aerial scenarios, providing a robust, accurate, and cost-effective localization solution for urban air mobility and aerial Internet of Things (IoT) applications.
Dongdong Li 0008, Tao Yang 0006, Haidong Qin, Shixiong Fan, Shuanghan Zhang, Jing Li 0010
IEEE Internet Things J.2
2026 GaussHead: Real-Time 3-D Head Avatar Driving for Mobile With Compact Gaussian Representation
Xiaoshi Zhou, Tao Yang 0006, Baogang Song, Chengwei Cao, Hongwei Bao, Jing Li 0010
IEEE Internet Things J.2
2025 Beyond Implicit Representations: Exploring Gaussian Splatting for Next-Generation SLAM, Introduction, and Review
abstract
Simultaneous localization and mapping (SLAM) has attracted tremendous interest from members of the research community in recent years because of its ability to make the robot truly independent in navigation. The integration of neural rendering and the SLAM system showed promising results in joint localization and photorealistic view reconstruction. However, currently available methods that fully rely on implicit representations are resource hungry. Recent work has shown that 3D Gaussians enable high-quality reconstruction and real-time rendering of scenes using multiple posed cameras. Owing to the increasing popularity of Gaussian splatting (GS) and its expanding research area, we present a comprehensive survey of GS based SLAM papers. This survey addresses the challenges of neural based systems by exploring the application of GS within SLAM frameworks. We present a comprehensive analysis of Gaussian Splatting for scene representation in SLAM, focusing on the past three years of research. Our survey delves into the advantages of GS based SLAM, including efficiency, unified representation for various tasks, suitability for resource-constrained devices, and high accuracy and robustness for dense scene reconstruction. Additionally, the survey includes a broader range of research, including semantic GS techniques, to provide a holistic understanding of this approach. By offering an in-depth exploration of Gaussian Splatting in SLAM, a benchmark comparison of performance, and a focus on 3D Gaussian Splatting (3DGS) based SLAM advantages, this survey aims to be a valuable resource for researchers and developers seeking to leverage Gaussian Splatting for efficient and robust SLAM.
Besufekad T. Hadero, Abdulmoiz Ahsan, Dongdong Li 0008, Tao Yang 0006
IEEE Internet Things J.4
2025 RT3DHVC: A Real-Time Human Holographic Video Conferencing System With a Consumer RGB-D Camera Array
abstract
In this paper, we present an end-to-end holographic video conferencing system that enables real-time high-quality free-viewpoint rendering of participants in different spatial regions, placing them in a unified virtual space for a more immersive display. Our system offers a cost-effective, complete holographic conferencing process, including multiview 3D data capture, RGB-D stream compression and transmission, high-quality rendering, and immersive display. It employs a sparse set of commodity RGB-D cameras that capture 3D geometric and textural information. We then remotely transmit color and depth maps via standard video encoding and transmission protocols. We propose a GPU-parallelized rendering pipeline based on an image-based virtual view synthesis algorithm to achieve real-time and high-quality scene rendering. This algorithm uses an on-the-fly Truncated Signed Distance Function (TSDF) approach, which marches along virtual rays within a computed precise search interval to determine surface intersections. We then design a multiweight projective texture mapping method to fuse color information from multiple views. Furthermore, we introduce a method that uses a depth confidence map to weight the rendering results from different views, which mitigates the impact of sensor noise and inaccurate measurements on the rendering results. Finally, our system places conference participants from different spaces into a virtual conference environment with a global coordinate system through coordinate transformation, which simulates a real conference scene in physical space, providing an immersive remote conferencing experience. Experimental evaluations confirm our system’s real-time, low-latency, high-quality, and immersive capabilities.
Jing Li 0010, Yanran Dai, Haidong Qin, Xiaoshi Zhou, Kefan Yan, Tao Yang 0006
IEEE Trans. Circuits Syst. Video Technol.9
2025 ECC-NeRF: Anti-Aliasing Neural Radiance Fields With Elliptic Cone-Casting for Diverse Camera Models
abstract
Anti-aliasing is a crucial research topic in computer graphics, which can significantly enhance the rendering quality of neural radiance fields (NeRF). Recent studies have introduced effective anti-aliasing NeRF methods, utilizing cone-casting to replace ray-casting and modeling the 3D observation area of pixels as circular cones. The cone-casting strategy has successfully reduced blurring and aliasing in novel view rendering. However, we have observed that the light cones are not standard circular cones because the camera projection model distorts them into elliptic cones of diverse sizes and shapes. This finding motivates us to model pixel light cones as anisotropic elliptic cones and propose an elliptic cone-casting-based anti-aliasing NeRF method called “ECC-NeRF". Specifically, we first derive the elliptic cone models for common pinhole, fisheye, and panoramic cameras based on their camera projection models. Then, we integrate the proposed elliptic cone-casting into two representative cone-casting-based anti-aliasing NeRF methods: Mip-NeRF and Zip-NeRF. Our experimental evaluations on multiple datasets demonstrate that our method can achieve more accurate multi-scale anisotropic representation and better novel view rendering quality with negligible additional computation cost.
Haidong Qin, Tao Yang 0006, Xiaoshi Zhou, Dongdong Li 0008, Yanran Dai, Jing Li 0010
IEEE Trans. Circuits Syst. Video Technol.2
2024 Real-time distance field acceleration based free-viewpoint video synthesis for large sports fields
abstract
Free-viewpoint video allows the user to view objects from any virtual perspective, creating an immersive visual experience. This technology enhances the interactivity and freedom of multimedia performances. However, many free-viewpoint video synthesis methods hardly satisfy the requirement to work in real time with high precision, particularly for sports fields having large areas and numerous moving objects. To address these issues, we propose a free-viewpoint video synthesis method based on distance field acceleration. The central idea is to fuse multi-view distance field information and use it to adjust the search step size adaptively. Adaptive step size search is used in two ways: for fast estimation of multi-object three-dimensional surfaces, and synthetic view rendering based on global occlusion judgement. We have implemented our ideas using parallel computing for interactive display, using CUDA and OpenGL frameworks, and have used real-world and simulated experimental datasets for evaluation. The results show that the proposed method can render free-viewpoint videos with multiple objects on large sports fields at 25 fps. Furthermore, the visual quality of our synthetic novel viewpoint images exceeds that of state-of-the-art neural-rendering-based methods.
Yanran Dai, Jing Li 0010, Haidong Qin, Bang Liang, Shikuan Hong, Haozhe Pan, Tao Yang 0006
Comput. Vis. Media8
2024 GS-SFS: Joint Gaussian Splatting and Shape-From-Silhouette for Multiple Human Reconstruction in Large-Scale Sports Scenes
abstract
We introduce GS-SFS, a method that utilizes a camera array with wide baselines for high-quality multiple human mesh reconstruction in large-scale sports scenes. Traditional human reconstruction methods in sports scenes, such as Shape-from-Silhouette (SFS), struggle with sparse camera setups and small human targets, making it challenging to obtain complete and accurate human representations. Despite advances in differentiable rendering, including 3D Gaussian Splatting (3DGS), which can produce photorealistic novel-view renderings with dense inputs, accurate depiction of surfaces and generation of detailed meshes is still challenging. Our approach uniquely combines 3DGS's view synthesis with an optimized SFS method, thereby significantly enhancing the quality of multiperson mesh reconstruction in large-scale sports scenes. Specifically, we introduce body shape priors, including the human surface point clouds extracted through SFS and human silhouettes, to constrain 3DGS to a more accurate representation of the human body only. Then, we develop an improved mesh reconstruction method based on SFS, mainly by adding additional viewpoints through 3DGS and obtaining a more accurate surface to achieve higher-quality reconstruction models. We implement a high-density scene resampling strategy based on spherical sampling of human bounding boxes and render new perspectives using 3D Gaussian Splatting to create precise and dense multi-view human silhouettes. During mesh reconstruction, we integrate the human body's 2D Signed Distance Function (SDF) into the computation of the SFS's implicit surface field, resulting in smoother and more accurate surfaces. Moreover, we enhance mesh texture mapping by blending original and rendered images with different weights, preserving high-quality textures while compensating for missing details. The experimental results from real basketball game scenarios demonstrate the significant improvements of our approach for multiple human body model reconstruction in complex sports settings.
Jing Li 0010, Haidong Qin, Yanran Dai, Jing Liu 0006, Canbin Zhang, Tao Yang 0006
IEEE Trans. Multim.8
2023 Calibration-Free Cross-Camera Target Association Using Interaction Spatiotemporal Consistency
abstract
In this paper, we propose a novel calibration-free cross-camera target association algorithm that aims to relate local visual data of the same object across cameras with overlapping FOVs. Unlike other methods using object's own characteristics, our approach makes full use of the interactions between objects and explores their spatiotemporal consistency in projection transformation to associate cameras. It has wider applicability in deployed overlapping multi-camera systems with unknown or rarely available calibration data, especially if there is a large perspective gap between cameras. Specifically, we first extract trajectory intersection which is one of the typical object-object interactive behaviors from each camera for feature vector construction. Then, based on the consistency of object-object interactions, we propose a multi-camera spatiotemporal alignment method via wide-domain cross-correlation analysis. It realizes time synchronization and spatial calibration of the multi-camera system simultaneously. After that, we introduce a cross-camera target association approach using aligned object-object interactions. The local data of the same target are successfully associated across cameras without any additional calibration. Extensive experimental evaluations on different databases verify the effectiveness and robustness of our proposed method.
Jing Li 0010, Yuguang Xie, Jiayang Nie, Tao Yang 0006, Zhaoyang Lu
IEEE Trans. Multim.5
2023 Bullet-Time Video Synthesis Based on Virtual Dynamic Target Axis
abstract
Bullet-time videos have been widely used in movies, TV advertisements, and computer games, and can produce an immersive and smooth orbital free-viewpoint of frozen action. However, existing bullet-time video synthesis methods remain challenging in practical applications, especially in complex situations with poor camera calibration and a variety of camera array structures. This paper proposes a novel bullet-time video synthesis method based on a virtual dynamic target axis. We adopt an image similarity transformation strategy to eliminate image distortion in the bullet-time video. We use a high-order polynomial curve fitting strategy to reserve more bullet-time video frame content. The proposed dynamic target axis strategy can support various camera array structures, including camera arrays with and without a common field of view. In addition, this strategy can also tolerate poor camera calibration situations with unevenly distributed reprojection errors to some extent and synthesize smooth bullet-time videos without high-precision camera calibration. Qualitative and quantitative experiments in real environments and on simulation platforms demonstrate the high performance of our bullet-time video synthesis method. Compared with the state-of-the-art methods, the proposed method shows superiority.
Haidong Qin, Jing Li 0010, Yanran Dai, Shikuan Hong, Tao Yang 0006
IEEE Trans. Multim.8
2022 Multi-camera joint spatial self-organization for intelligent interconnection surveillance
Jing Li 0010, Yuguang Xie, Jiayang Nie, Tao Yang 0006, Zhaoyang Lu
Eng. Appl. Artif. Intell.5
2022 Accurate localization of moving objects in dynamic environment for small unmanned aerial vehicle platform using global averaging
abstract
Abstract Small unmanned aerial vehicles (UAVs) have developed rapidly and are widely used for disaster relief, traffic monitoring and military surveillance. To perform these tasks better, it is necessary to improve the environmental perception ability of UAVs in a dynamic environment, including their static and dynamic perception ability. Specifically, both three‐dimensional reconstruction for a static scene and localization for moving objects are required. Simultaneous Localization And Mapping technology has made great progress in static scene structure reconstruction and UAV self‐motion estimation. However, accurate real‐time localization of moving objects is still challenging. In this article, a global averaging based localization method is proposed to locate moving objects for a small UAV platform. Inspired by global structure from motion, this idea is applied to the localization of moving objects. To solve moving object localization, the relative motion estimation and global position optimisation methods are proposed. The proposed method was tested in various scenarios with a several trajectories. The extensive experimental results demonstrate the robustness and effectiveness of the proposed method.
Xiuchuan Xie, Tao Yang 0006, Yanning Zhang 0001, Bang Liang, Linfeng Liu 0002
IET Comput. Vis.2
2022 Online Ground Multitarget Geolocation Based on 3-D Map Construction Using a UAV Platform
abstract
Geolocating multiple targets of interest on the ground from an aerial platform is an important activity in many applications, such as visual surveillance. However, due to the limited measurement accuracy of commonly used airborne sensors (including altimeters, accelerometer, gyroscopes, and so on) and the small size, complex motion, and a large number of ground targets in aerial images, most of the current unmanned aerial vehicle (UAV)-based ground target geolocation algorithms have difficulty in obtaining accurate geographic location coordinates online, especially at middle and high altitudes. To solve these problems, in this article, a novel online ground multitarget geolocation framework using a UAV platform is proposed, which minimizes the introduction of sensor error sources and uses only monocular aerial image sequences and global positioning system (GPS) data to perform parallel processing of target detection and rapid 3-D sparse geographic map construction and target geographic location estimations, thereby improving the accuracy and speed of ground multitarget online geolocation. In this framework, a detection algorithm based on deep learning is first adopted to improve the accuracy and robustness of small target detection in aerial images by constructing an aerial image dataset. Then, we propose a novel target geolocation algorithm based on 3-D map construction, which combines continuous images and GPS data collected online using a UAV platform to generate a 3-D geographic map and accurately estimates the GPS location of the target center pixel through projection and triangulation of the map points on the image. Finally, we design a data transmission architecture that selects multiple processes to perform image acquisition, target detection, and target geolocation tasks in parallel and utilizes database communication between the processes to achieve accurate online geolocation of ground targets in aerial images, regardless of whether the target is in a static or moving state. To evaluate the effectiveness of the proposed framework, we build an online ground multitarget geolocation system using a quad-rotor UAV and carry out a large number of experiments in simulations and real environments. Qualitative and quantitative experimental results proved that the framework can accurately locate ground targets in various complex environments, such as parks, highways, schools, cities, different flight altitudes (50–2000 m), and different attitude angles, and the average positioning error is approximately 1 m at 2000 m for cities with rich 3-D structures.
Fangbing Zhang, Tao Yang 0006, Yajia Ning, Jinghui Fan, Dongdong Li 0008
IEEE Trans. Geosci. Remote. Sens.2
2021 Image-Only Real-Time Incremental UAV Image Mosaic for Multi-Strip Flight
abstract
Limited by aircraft flight altitude and camera parameters, it is necessary to obtain wide-angle panoramas quickly by stitching aerial images, which is helpful in rapid disaster investigation, recovery after earthquakes, and aerial reconnaissance. However, most existing stitching algorithms do not simultaneously meet practical real-time, robustness, and accuracy requirements, especially in the case of a long-distance multistrip flight. In this paper, we propose a novel image-only real-time UAV image mosaic framework for long-distance multistrip flights that does not require any auxiliary information, such as GPS or GCPs. The framework has a complete structure, mainly consisting of the three tasks of automatic initialization, current frame tracking, and real-time mosaic generation. The stitching plane is determined in the initialization process, the homography transformation of the current image is estimated in the tracking task, and the image is mapped to the stitching plane to generate and update the panorama in the real-time mosaic process. The core idea is that, in the tracking task, we introduce and develop a keyframe insertion strategy to generate a keyframe list and, on this basis, design a homography matrix estimation based on a local optimization strategy to reduce the accumulated error when continuously stitching image sequences collected online by UAVs and to realize real-time, effective UAV image mosaic construction. In addition, this framework has good scalability, which is not limited to a specific algorithm. To evaluate the effectiveness of the proposed framework, we carry out a large number of experiments on the AirSim simulation platform and present an exhaustive evaluation in some sequences from a popular dataset. Qualitative and quantitative experimental results in simulation and real environments demonstrate that our algorithm can obtain an effective and robust mosaic image in real-time. Through strategy comparison experiments, it is proven that the keyframe insertion strategy and the local optimization strategy both improve the stitching performance. Compared with five state-of-art image stitching approaches, the mosaic effect of the proposed method is comparable or better. In terms of algorithm speed, its performance is superior to them. Additionally, experiments of illumination change and feature replacement in the framework verify the good adaptability and scalability of the algorithm.
Fangbing Zhang, Tao Yang 0006, Linfeng Liu 0002, Bang Liang, Jing Li 0010
IEEE Trans. Multim.2
2020 Deep Image-to-Video Adaptation and Fusion Networks for Action Recognition
abstract
Existing deep learning methods for action recognition in videos require a large number of labeled videos for training, which is labor-intensive and time-consuming. For the same action, the knowledge learned from different media types, e.g., videos and images, may be related and complementary. However, due to the domain shifts and heterogeneous feature representations between videos and images, the performance of classifiers trained on images may be dramatically degraded when directly deployed to videos. In this paper, we propose a novel method, named Deep Image-to-Video Adaptation and Fusion Networks (DIVAFN), to enhance action recognition in videos by transferring knowledge from images using video keyframes as a bridge. The DIVAFN is a unified deep learning model, which integrates domain-invariant representations learning and cross-modal feature fusion into a unified optimization framework. Specifically, we design an efficient cross-modal similarities metric to reduce the modality shift among images, keyframes and videos. Then, we adopt an autoencoder architecture, whose hidden layer is constrained to be the semantic representations of the action class names. In this way, when the autoencoder is adopted to project the learned features from different domains to the same space, more compact, informative and discriminative representations can be obtained. Finally, the concatenation of the learned semantic feature representations from these three autoencoders are used to train the classifier for action recognition in videos. Comprehensive experiments on four real-world datasets show that our method outperforms some state-of-the-art domain adaptation and action recognition methods.
Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006
IEEE Trans. Image Process.4
2019 Hierarchically Learned View-Invariant Representations for Cross-View Action Recognition
abstract
Recognizing human actions from varied views is challenging due to huge appearance variations in different views. The key to this problem is to learn discriminant view-invariant representations generalizing well across views. In this paper, we address this problem by learning view-invariant representations hierarchically using a novel method, referred to as joint sparse representation and distribution adaptation. To obtain robust and informative feature representations, we first incorporate a sample-affinity matrix into the marginalized Stacked Denoising Autoencoder to obtain shared features that are then combined with the private features. In order to make the feature representations of videos across views transferable, we then learn a transferable dictionary pair simultaneously from pairs of videos taken at different views to encourage each action video across views to have the same sparse representation. However, the distribution difference across views still exists because a unified subspace, where the sparse representations of one action across views are the same, may not exist when the view difference is large. Therefore, we propose a novel unsupervised distribution adaptation method that learns a set of projections that project the source and target views data into respective low-dimensional subspaces, where the marginal and conditional distribution differences are reduced simultaneously. Therefore, the finally learned feature representation is view-invariant and robust for substantial distribution difference across views even though the view difference is large. Experimental results on four multi-view datasets show that our approach outperforms the state-of-the-art approaches.
Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006
IEEE Trans. Circuits Syst. Video Technol.4
2019 Joint Deep and Depth for Object-Level Segmentation and Stereo Tracking in Crowds
abstract
Tracking multiple people in crowds is a fundamental and essential task in the multimedia field. It is often hindered by difficulties, such as dynamic occlusion between objects, cluttered background, and abrupt illumination changes. To respond to this need, in this paper, we combine deep and depth to build a stereo tracking system for crowds. The core of the system is the fusion of the advantages of deep learning and depth information, which is exploited to achieve object segmentation and improve the multiobject tracking performance in severe occlusion. More specifically, first, to obtain more accurate detection observations in the tracking system, we present a novel object-level segmentation method. This method combines the effective detection results of deep learning with depth information to obtain precise object segmentation results. Then, we integrate the segmentation results and three-dimensional (3-D) information to extract 2-D and 3-D characteristics to represent the target, and design three similarity models to realize a stereo tracking method through data association in crowds. Finally, we build a diverse stereo dataset including various challenging indoor and outdoor scenes. The comprehensive experiments verify the effective and robust tracking performance of our system in various scenarios, and the system has rich output results including segmentation results, target distance, and tracking results. Moreover, the qualitative and quantitative comparison results show that the proposed algorithm not only has good object segmentation performance but also improves the tracking performance of completely and partially occluded objects, which is superior to the tested state-of-the-art tracking approaches.
Jing Li 0010, Lisong Wei, Fangbing Zhang, Tao Yang 0006, Zhaoyang Lu
IEEE Trans. Multim.4
2018 Global Temporal Representation Based CNNs for Infrared Action Recognition
abstract
Infrared human action recognition has many advantages, i.e., it is insensitive to illumination change, appearance variability, and shadows. Existing methods for infrared action recognition are either based on spatial or local temporal information, however, the global temporal information, which can better describe the movements of body parts across the whole video, is not considered. In this letter, we propose a novel global temporal representation named optical-flow stacked difference image (OFSDI) and extract robust and discriminative feature from the infrared action data by considering the local, global, and spatial temporal information together. Due to the small size of the infrared action dataset, we first apply convolutional neural networks on local, spatial, and global temporal stream respectively to obtain efficient convolutional feature maps from the raw data rather than train a classifier directly. Then these convolutional feature maps are aggregated into effective descriptors named three-stream trajectory-pooled deep-convolutional descriptors by trajectory-constrained pooling. Furthermore, we improve the robustness of these features by using the locality-constrained linear coding (LLC) method. With these features, a linear support vector machine (SVM) is adopted to classify the action data in our scheme. We conduct the experiments on infrared action recognition datasets InfAR and NTU RGB+D. The experimental results show that the proposed approach outperforms the representative state-of-the-art handcrafted features and deep learning features based methods for the infrared action recognition.
Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006
IEEE Signal Process. Lett.4
2017 Research on big data management and analysis method of multi-platform avionics system
abstract
The avionics system is an important part of modern fighters, and the analysis and processing of large data generated by the avionics system can provide some guidance for the pilot's decision-making. After analyzing the existing distributed framework, a multi-platform avionics data system is designed and implemented to solve the problem that heterogeneous real-time data generated by a large number of sensors is difficult to be managed efficiently. With the help of the cloud platform, the system supports functions of data collection, data classification management, data storage and data analysis. It can establish the related model based on the historical data and it can make real-time prediction with real-time data and correlation model. The test results show that the system functional testing requirements coverage rate is 100%, and performance and stability are both in line with the requirements.
Miao Wang 0008, Zheng Dai, Hangyu Guo, Tao Yang 0006
ICIS4
2017 Multi-platform fire control strike track planning method based on deep enhance learning
abstract
The modem war has been transforming from center-based to cyber-based. The cyber war is an information network system, which is composed by detective system, communication system, command and control system, and weapon system. In such system, commander can see all battlefield situations, change combat information, design and implement combat plan. Cloud-based platform would be the development trend of next generation avionics system in cyber combat. In order to improve system combat efficiency, in this paper, we propose a multi-platform fire control strike track planning method based on deep enhance learning. It uses deep enhance learning model to generate higher hit rate fire control strike track. The experiment shows our proposed method is more efficient.
Miao Wang 0008, Hangyu Guo, Tao Yang 0006
ICIS4
2017 Cube surface modeling for human detection in crowd
abstract
Human detection in dense crowds poses to be a demanding task owing to complex background and serious occlusion. In this paper, we propose a novel real-time and reliable human detection system. We solve the human detection problem by presenting a novel cube surface model captured by a binocular stereo vision camera. We first propose a cube surface model to estimate the 3D background cubes in the surveillance area. We then develop a shadow-free strategy for cube surface model updating. Thereafter, we present a shadow weighted clustering method to efficiently search for human as well as remove false alarms. Ultimately, we have developed a highly robust human detection system, and we carefully evaluate our system in many real challenge indoor and outdoor scenes. Expensive experiments demonstrate our system achieves real-time performance, higher detection rate and lower face alarms in comparison with state-of-the-art human detection methods.
Jing Li 0010, Fangbing Zhang, Lisong Wei, Tao Yang 0006, Zhongzhen Li
ICME4
2017 From online to offline: Charactering user's online music listening behavior for efficacious offline radio program arrangement
abstract
Efficacious radio music program arrangement with appropriate music played in the right time can greatly improve users' listening experience and attract more listeners for traditional radio stations. In this paper, we propose a method for efficacious offline radio program arrangement by mining users' online listening behavior characteristics in popular music websites. Firstly, we collect user's listening behavior trace of specific songs from two popular music sites in China. Secondly, the characteristics of basic listening behavior are analyzed and the concept of soaring degree is employed to quantify the popularity of a specific song. Different forecasting methods are also proposed for popularity prediction. Finally, based on the measurement and analysis results, we propose a simple and feasible strategy for offline program arrangement and the feedbacks from Shaanxi music station (FM 98.8) verify the efficiency and correctness of the proposed methods.
Tao Qin 0002, Chenxu Wang 0001, Tao Yang 0006
ISCC4
2017 Synthetic aperture photography using a moving camera-IMU system
Xiaoqiang Zhang 0002, Yanning Zhang 0001, Tao Yang 0006, Yee-Hong Yang
Pattern Recognit.3
2016 A new model for nickname detection based on network structure and similarity propagation
abstract
Users can participate in variety topics and express their opinions using different kinds of online applications, and the IDs (nickname) they used are usually virtual and difficult for finding the physical person. Which pose great challenges for network security management and user's online behavior supervision. Focus on this problem, we proposed methods for nickname detection based on the user's online connection structure and similarity propagation model. Firstly, we collected user's profile information from two popular online applications, sina microblog and RenRen network. Then we mark several matched pairs which belong to the same person from the applications with hardly manually effort, and those IDs are selected as seed set for nickname detection. Secondly, we proposed an nickname detection model based on the connection structure and similarity propagation. We selected one matched pairs from the seed set and obtain all their neighbors. Then we calculated the similarity of each pairs from the neighbor set and calculate their neighbors' connection similarity and neighbors' location similarity. If the similarity is bigger than a selected threshold, we claim they are matched pairs and insert them into the seed sets. On one hand, the correlation results can propagated based on the updated seed set. On the other hand, the computational complexity are greatly reduced as we only employ the neighbors' profiles to calculate the similarity. Experimental results verify the efficiency of the proposed method, which can lay a solid foundation for the online network management and user's behavior supervision.
Zhaoli Liu, Tao Qin 0002, Xiaohong Guan, Tao Yang 0006
ISCC5
2016 Compressive Tracking based on Superpixel Segmentation
Ting Chen 0004, Hichem Sahli, Yanning Zhang 0001, Tao Yang 0006, Lingyan Ran
MoMM4
2016 Autonomous Near Ground Quadrone Navigation with Uncalibrated Spherical Images Using Convolutional Neural Networks
Lingyan Ran, Yanning Zhang 0001, Tao Yang 0006, Ting Chen 0004
MoMM3
2016 Modeling heterogeneous and correlated human dynamics of online activities with double Pareto distributions
Chenxu Wang 0001, Xiaohong Guan, Tao Qin 0002, Tao Yang 0006
Inf. Sci.4
2016 Tracking with dynamic weighted compressive model
Ting Chen 0004, Yanning Zhang 0001, Tao Yang 0006, Hichem Sahli
J. Vis. Commun. Image Represent.3
2016 Kinect based real-time synthetic aperture imaging through occlusion
Tao Yang 0006, Wenguang Ma, Sibing Wang, Jing Li 0010, Jingyi Yu 0001, Yanning Zhang 0001
Multim. Tools Appl.1
2015 Multi-Object Tracking in Airborne Video Imagery based on Compressive Tracking Detection Responses
abstract
Multi-object tracking (MOT) in airborne video is a challenging problem due to the uncertain airborne vehicle motion as well as mounted camera vibrations. Most approaches addressing tracking in such type of scenario, use data association based on motion detection responses. Such approaches fail tracking objects with low speed or static ones. To alleviate the motion detection failures, in this paper we propose a multi-object tracking system based on combining motion-detection and Compressive Tracking detection responses.
Ting Chen 0004, Hichem Sahli, Yanning Zhang 0001, Tao Yang 0006
MoMM4
2015 An easy-to-implement Benchmarking Tool for Mobile Tablet-PC Visual Pose Estimation
abstract
Recently, researchers are interested in mobile device based computer vision applications. An accurate visual pose estimation is a common and important subtask. To quantitatively evaluate the accuracy, ground truth visual pose datasets are needed. However, the lack of inexpensive and easy benchmarking tool for mobile device based visual poses estimation makes it difficult, if not impossible, to quantitatively evaluate the estimated visual poses. In this paper, a novel and easy-to-implement experimental setup is proposed to generate ground truth visual pose data for handheld tablet-PC. The tablet-PC screen is leveraged to display a calibration pattern every time the on-board camera captures an image. The tablet-PC screen image is captured by another camera and is used to estimate the visual pose of the tablet-PC. An experimental environment is setup for parameter calibration and pose accuracy verification. Extensive experimental results with quantitative analysis demonstrate the accuracy and the generality of our tool.
Xiaoqiang Zhang 0002, Yanning Zhang 0001, Tao Yang 0006, Ting Chen 0004, Yee-Hong Yang
MoMM3
2014 All-In-Focus Synthetic Aperture Imaging
Tao Yang 0006, Yanning Zhang 0001, Jingyi Yu 0001, Jing Li 0010, Wenguang Ma, Xiaomin Tong, Lingyan Ran
ECCV (6)1
2014 Simultaneous active camera array focus plane estimation and occluded moving object imaging
Tao Yang 0006, Yanning Zhang 0001, Xiaoqiang Zhang 0002, Ting Chen 0004, Lingyan Ran, Zhengxi Song, Wenguang Ma
Image Vis. Comput.1
2013 Unstructured Synthetic Aperture Photograph Based Occluded Object Imaging
abstract
Recently the camera array synthetic aperture imaging has been proved to be a powerful technology for occluded object imaging. However, synthetic aperture imaging with a camera mounted on a moving platform is still a significant challenging task for many computer vision applications. This paper presents a novel unstructured synthetic aperture imaging approach to solve the above problem. The main characteristics of our approach include: (1) To the best of our knowledge, our proposed method is the first one for synthetic aperture imaging through occlusion on a moving platform. (2) In contrast to "averaging data", we creatively use the clustering method to estimate the pixel color with high visible probability in the virtual focal plane. A new moving synthetic aperture imaging system has been built to capture unstructured light field in complex outdoor scene, and experimental results demonstrate that our approach successfully generates clear image of hidden object even under severe occlusion, which outperform the traditional synthetic aperture imaging method.
Wenguang Ma, Tao Yang 0006, Yanning Zhang 0001, Xiaomin Tong
ICIG2
2013 Occluded object imaging via optimal camera selection
abstract
High performance occluded object imaging in cluttered scenes is a significant challenging task for many computer vision applications. Recently the camera array synthetic aperture imaging is proved to be an effective way to seeing object through occlusion. However, the imaging quality of occluded object is often significantly decreased by the shadows of the foreground occluder. Although some works have been presented to label the foreground occluder via object segmentation or 3D reconstruction, these methods will fail in the case of complicated occluder and severe occlusion. In this paper, we present a novel optimal camera selection algorithm to solve the above problem. The main characteristics of this algorithm include: (1) Instead of synthetic aperture imaging, we formulate the occluded object imaging problem as an optimal camera selection and mosaicking problem. To the best of our knowledge, our proposed method is the first one for occluded object mosaicing. (2) A greedy optimization framework is presented to propagate the visibility information among various depth focus planes. (3) A multiple label energy minimization formulation is designed in each plane to select the optimal camera. The energy is estimated in the synthetic aperture image volume and integrates the multi-view intensity consistency, previous visibility property and camera view smoothness, which is minimized via Graph cuts. We compare our method with the state-of-the-art synthetic aperture imaging algorithms, and extensive experimental results with qualitative and quantitative analysis demonstrate the effectiveness and superiority of our approach.
Tao Yang 0006, Yanning Zhang 0001, Xiaomin Tong, Wenguang Ma
ICMV1
2013 Dynamic Compressive Tracking
abstract
Real-Time Compressive Tracking utilizes a very spare measurement matrix to extract the features for the appearance model. Such model performs well when the tracked objects are well defined. However, when the objects are low-grain, low-resolution, or small, a fixed size sparse measurement matrix is not sufficient enough to preserve the image structure of the object. In this work, we propose a Dynamic Compressive Tracking algorithm that employs adaptive random projections that preserve the image structure of the objects during tracking. The proposed tracker uses a dynamic importance ranking weight to evaluate the classification results obtained by each of the sparse measurement matrices and complete the tracking with the optimal sparse matrix. Extensive experimental results, on challenging publicly available data sets, shows that the proposed dynamic compressible tracking algorithm outperforms conventional compressive tracker.
Ting Chen 0004, Yanning Zhang 0001, Tao Yang 0006, Hichem Sahli
MoMM3
2013 A New Hybrid Synthetic Aperture Imaging Model for Tracking and Seeing People Through Occlusion
abstract
Robust detection and tracking of multiple people in cluttered and crowded scenes with severe occlusion is a significant challenge for many computer vision applications. In this paper, we present a novel hybrid synthetic aperture imaging model to solve this problem. The main characteristics of this approach are as follows. 1) To the best of our knowledge, this is the first attempt to solve the occluded people imaging and tracking problem in a joint multiple camera synthetic aperture imaging domain. 2) A multiple model framework is designed to achieve seamless interaction among the detection, imaging and tracking modules. 3) In the object detection module, a multiple constraints-based approach is presented for people localization and ghost objects removal in a 3-D foreground silhouette synthetic aperture imaging volume. 4) In the synthetic imaging module, a novel occluder removal-based synthetic imaging approach is proposed to significantly improve the imaging quality of objects even under severe occlusion. 5) In the object tracking module, a camera array is used for robust people tracking in color synthetic aperture images. A network-camera-based hybrid synthetic aperture imaging system has been set up, and experimental results with qualitative and quantitative analyses demonstrate that the method can reliably locate and see people in challenging scenes.
Tao Yang 0006, Yanning Zhang 0001, Xiaomin Tong, Xiaoqiang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2012 A novel multi-object detection method in complex scene using synthetic aperture imaging
Zhao Pei, Yanning Zhang 0001, Tao Yang 0006, Xiuwei Zhang 0001, Yee-Hong Yang
Pattern Recognit.3
2011 Continuously tracking and see-through occlusion based on a new hybrid synthetic aperture imaging model
abstract
Robust detection and tracking of multiple people in cluttered and crowded scenes with severe occlusion is a significant challenging task for many computer vision applications. In this paper, we present a novel hybrid synthetic aperture imaging model to solve this problem. The main characteristics of this approach include: (1) To the best of our knowledge, this algorithm is the first time to solve the occluded people imaging and tracking problem in a joint multiple camera synthetic aperture imaging domain. (2) A multiple model framework is designed to achieve seamless interaction among the detection, imaging and tracking modules. (3)In the object detection module, a multiple constraints based approach is presented for people localizing and ghost objects removal in a 3D foreground silhouette synthetic aperture imaging volume. (4) In the synthetic imaging module, a novel occluder removal based synthetic imaging approach is proposed to continuously obtain object clear image even under severe occlusion. (5) In the object tracking module, a camera array is used for robust people tracking in color synthetic aperture images. A network camera based hybrid synthetic aperture imaging system has been set up, and experimental results with qualitative and quantitative analysis demonstrate that the method can reliably locate and see people in challenge scene.
Tao Yang 0006, Yanning Zhang 0001, Xiaomin Tong, Xiaoqiang Zhang 0002
CVPR1
2009 Silhouette-Based 2D Human Pose Estimation
abstract
In this paper we present a novel silhouette-based method to estimate 2D human pose. It takes a pre-defined human skeleton model as the prior information and a video sequence as the data source, and estimates human pose in each frame by the following steps: Firstly, the Gaussian Mixture Background Model (GMM) is adopted to extract silhouette from an image and this silhouette will be the human body data set after treatment. Then the Distance Transform (DT) and Principal Component Analysis (PCA) are introduced to locate the base point of the human skeleton model, with which the human skeleton model is initialized automatically. Afterwards an iterative process, based on the Expectation Maximization (EM), is constructed out to cluster the human body data set and estimate the parameters of the skeleton iteratively. Finally the pose of human body is figured out in the form of a corresponding skeleton model. Extensive experiments show that this method is robust and precise, and is feasible to apply in the real-time system.
Tao Yang 0006, Runping Xi, Zenggang Lin
ICIG2
2009 A Novel Multi-planar Homography Constraint Algorithm for Robust Multi-people Location with Severe Occlusion
abstract
Multi-view approach has been proposed to solve occlusion and lack of visibility in crowded scenes. However, the problem is that too much redundancy information might bring about false alarm. Although researchers have done many efforts on how to use the multi-view information to track people accurately, it is particularly hard to wipe off the false alarm. Our approach is to use multiple views cooperatively to detect objects and use objects silhouette on planes of different height to remove false alarm. To achieve this we adopt a novel multi-planar homography constraint to resolve occlusions and false alarm. Experimental results show that our algorithm is able to accurately locate people in crowded scene maintaining correct correspondences across views. Moreover, the false alarm rate is obviously reduced.
Xiaomin Tong, Tao Yang 0006, Runping Xi, Dapei Shao, Xiuwei Zhang 0001
ICIG2
2009 A Multi-camera Network System for Markerless 3D Human Body Voxel Reconstruction
abstract
This paper presents a fully automated system for realtime 3D human visual hull reconstruction and skeleton voxels extraction. The main contributions include: (1) A novel network based system is presented, which uses AXIS network cameras as video capture device, and performs a parallel processing among data capture, 3D voxel reconstruction and display. (2) A new human visual hull reconstruction algorithm is given. This approach firstly segments the foreground accurately by an efficient Gaussian Mixture Model (GMM) and a shadow model in HSV color space, then extends the standard Shape-From-Silhouette (SFS) lgorithm with online Region-of-Interest (ROI) estimation and binary searching, and finally construct skeleton probability visual hull with distance transform. Experiments with real video sequences show that the system can process eleven 640 × 480 video sequences at a frame rate of 15 fps, and construct human body voxels reliably in complex scenarios with cast shadows, various body configurations and multiple persons.
Tao Yang 0006, Yanning Zhang 0001, Dapei Shao, Xingong Zhang
ICIG1
2009 Real-Time Camera Pose Estimation Based on Multiple Planar Markers
abstract
Vision-based registration techniques for augmented reality(AR) systems have been the subject of intensive research recently due to their potential to accurately align virtual objects with the real world. The downfall of these vision-based approaches, however, is their high computational cost and lack of robustness. To address these shortcomings, a robust pose estimation algorithm based on artificial planar markers is adopted. This algorithm solves the problem of camera pose ambiguities and is able to draw a unique and robust solution. Experiments show the robustness and effectiveness of this method in the context of real-time AR tracking.
Tao Yang 0006, Jiangbin Zheng 0001, Xingong Zhang
ICIG2
2009 Image Registration Based on Rectangle Pattern
abstract
To against the complexity in finding feature-pair in image registration caused by traditional features: corners, lines and image edge, a novel method for image registration based on rectangle pattern is proposed. Unlike traditional features, a rectangle pattern can be described as its four vertexes and center, which can afford five pair-wise points for any kind of image transformation, and it holds stable in different weather condition, time and imaging way, which formed by building angular, widely found in aero-image. Firstly, image edge is detected by canny, and then distance transform is applied on the result of canny, by follows, a thresholding and mask convoluting are used on the distance transform result to avoid the interfering complex lines and edge. The center of the rectangles is obtained by clustering on the prior result, and then the four vertexes is calculated by geometric restrict on rectangle pattern. Consequently, we calculate the centers pair of the rectangles by their slope difference which aims to get the correct pair-wise points set between two images. Finally, we use the four vertexes pair and center pair of the rectangle pattern as the input of RANSAC algorithm, to solve an affine transform. The proposed algorithm is proved to be effective and accurate on the translation, rotation and scale between electro-optic(EO) images pair and SAR-EO images pair. 1.3 pixels of registration accuracy result is obtained in the experiment.
Xingong Zhang, Runping Xi, Xiuwei Zhang 0001, Tao Yang 0006
ICIG5
2009 A Convenient Multi-camera Self-Calibration Method Based on Human Body Motion Analysis
abstract
A novel and convenient multi-camera self-calibration method is proposed in this paper. Different from other calibration methods, our method is done by analyzing human body motion. The only constraint is that several people of different heights are needed to walk around the experimental environment one by one in the calibration period. By this way, two kinds of corresponding points are extracted from synchronous video sequences. One is the centroid of the moving human body. The other is points on the floor, which is extracted by matching floor planes in video sequences. The floor planes registration is based on shadow detection and co-motion feature. Based on these corresponding points, camera parameters and 3D points observed are estimated. The proposed method is tested in our own experimental environment. Experimental results show the accuracy of our calibration method. Our method can satisfy many applications of multi-view computer vision.
Xiuwei Zhang 0001, Yanning Zhang 0001, Xingong Zhang, Tao Yang 0006, Xiaomin Tong, Haichao Zhang 0001
ICIG4
2007 Trailblazing: Video Playback Control by Direct Object Manipulation
abstract
We describe a new interaction technique that allows users to control nonlinear video playback by directly manipulating objects seen in the video. This interaction technique is similar to video "scrubbing" where the user adjusts the playback time by moving the mouse along a slider. Our approach is superior to variable-scale scrubbing in that the user can concentrate on interesting objects and does not have to guess how long the objects will stay in view. Our method relies on a video tracking system that tracks objects in fixed cameras, maps them into 3D space, and handles hand-offs between cameras. In addition to dragging objects visible in video windows, users may also drag iconic object representations on a floor plan. In that case, the best video views are selected for the dragged objects.
Don Kimber, Anthony Dunnigan, Andreas Girgensohn, Frank M. Shipman III, Thea Turner, Tao Yang 0006
ICME6
2007 Robust People Detection and Tracking in a Multi-Camera Indoor Visual Surveillance System
abstract
In this paper we describe the analysis component of an indoor, real-time, multi-camera surveillance system. The analysis includes: (1) a novel feature-level foreground segmentation method which achieves efficient and reliable segmentation results even under complex conditions, (2) an efficient greedy search based approach for tracking multiple people through occlusion, and (3) a method for multi-camera handoff that associates individual trajectories in adjacent cameras. The analysis is used for an 18 camera surveillance system that has been running continuously in an indoor business over the past several months. Our experiments demonstrate that the processing method for people detection and tracking across multiple cameras is fast and robust.
Tao Yang 0006, Francine Chen 0001, Don Kimber, Jim Vaughan
ICME1
2007 DOTS: support for effective video surveillance
abstract
DOTS (Dynamic Object Tracking System) is an indoor, real-time, multi-camera surveillance system, deployed in a real office setting. DOTS combines video analysis and user interface components to enable security personnel to effectively monitor views of interest and to perform tasks such as tracking a person. The video analysis component performs feature-level foreground segmentation with reliable results even under complex conditions. It incorporates an efficient greedy-search approach for tracking multiple people through occlusion and combines results from individual cameras into multi-camera trajectories. The user interface draws the users. attention to important events that are indexed for easy reference at a later time. Different views within the user interface provide spatial information for easier navigation. Our system, with over twenty video cameras installed in hallways and other public spaces in our office building, has been in constant use for almost a year.
Andreas Girgensohn, Don Kimber, Jim Vaughan, Tao Yang 0006, Frank M. Shipman III, Thea Turner, Eleanor Gilbert Rieffel, Lynn Wilcox, Francine Chen 0001, Anthony Dunnigan
ACM Multimedia4
2005 Real-Time Multiple Objects Tracking with Occlusion Handling in Dynamic Scenes
abstract
This work presents a real-time system for multiple objects tracking in dynamic scenes. A unique characteristic of the system is its ability to cope with long-duration and complete occlusion without a prior knowledge about the shape or motion of objects. The system produces good segment and tracking results at a frame rate of 15-20 fps for image size of 320 /spl times/ 240, as demonstrated by extensive experiments performed using video sequences under different conditions indoor and outdoor with long-duration and complete occlusions in changing background.
Tao Yang 0006, Stan Z. Li, Quan Pan 0001, Jing Li 0010
CVPR (1)1
2004 Multiple layer based background maintenance in complex environment
abstract
A fast and efficient multiple layer background maintenance model is built to conserve the original and the current background separately. Fusing the properties of object motion in image pixels and the changes between the input video and the multiple background layers, this method could handle various sources of scene changes, including ghosts, abandon objects and illumination changes. An intelligent video surveillance system is developed to test the performance of the algorithm. Experiments are performed using long video sequences under different conditions indoor and outdoor. The results show that the proposed algorithm is effective and efficient in real-time and accurate background maintenance in complex environment.
Tao Yang 0006, Quan Pan 0001, Stan Z. Li, Jing Li 0010
ICIG1