Pin Tao

dblp:87/119 · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
18since 2021 · last 2025
0000-0003-2687-7997ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2025 UV-Mamba: A DCN-Enhanced State Space Model for Urban Village Boundary Identification in High-Resolution Remote Sensing Images
abstract
Due to the diverse geographical environments, intricate landscapes, and high-density settlements, the automatic identification of urban village boundaries using remote sensing images remains a highly challenging task. This paper proposes a novel and efficient neural network model called UV-Mamba for accurate boundary detection in high-resolution remote sensing images. UV-Mamba mitigates the memory loss problem in lengthy sequence modeling, which arises in state space models (SSM) with increasing image size, by incorporating deformable convolutions (DCN). Its architecture utilizes an encoder-decoder framework and includes an encoder with four deformable state space augmentation (DSSA) blocks for efficient multi-level semantic extraction and a decoder to integrate the extracted semantic information. We conducted experiments on two large datasets showing that UV-Mamba achieves state-of-the-art performance. Specifically, our model achieves 73.3% and 78.1% IoU on the Beijing and Xi’an datasets, respectively, representing improvements of 1.2% and 3.4% IoU over the previous best model while also being 6× faster in inference speed and 40× smaller in parameter count. Source code and pre-trained models are available at https://github.com/Devin-Egber/UV-Mamba.
Lulin Li, Xuechao Zou, Junliang Xing, Pin Tao
ICASSP5
2025 Dynamic Dictionary Learning for Remote Sensing Image Segmentation
Xuechao Zou, Kai Li 0023, Pin Tao, Junliang Xing, Congyan Lang
ICCV6
2025 Knowledge Transfer and Domain Adaptation for Fine-Grained Remote Sensing Image Segmentation
abstract
Fine-Grained remote sensing image segmentation is essential for accurately identifying detailed objects in remote sensing images. Recently, vision transformer models (VTMs) pretrained on large-scale datasets have demonstrated strong zero-shot generalization. However, directly applying them to specific tasks may lead to domain shift. We introduce a novel end-to-end learning paradigm combining knowledge guidance with domain refinement to enhance performance. We present two key components: the Feature Alignment Module (FAM) and the Feature Modulation Module (FMM). FAM aligns features from a CNN-based backbone with those from the pretrained VTM’s encoder using channel transformation and spatial interpolation, and transfers knowledge via KL divergence and L2 normalization constraint. FMM further adapts the knowledge to the specific domain to address domain shift. We also introduce a fine-grained grass segmentation dataset and demonstrate, through experiments on two datasets, that our method achieves a significant improvement of 2.57 mIoU on the grass dataset and 3.73 mIoU on the cloud dataset. The results highlight the potential of combining knowledge transfer and domain adaptation to overcome domain-related challenges and data limitations. The project page is available at https://xavierjiezou.github.io/KTDA/.
Xuechao Zou, Kai Li 0023, Congyan Lang, Pin Tao
ICME6
2025 Asymmetric interaction preference induces cooperation in human-agent hybrid game
Danyang Jia, Xiangfeng Dai, Junliang Xing, Pin Tao, Yuanchun Shi, Zhen Wang 0004
Sci. China Inf. Sci.4
2025 KDA: Knowledge Distillation Adversarial Framework With Vision Foundation Models for Landslide Segmentation
abstract
Landslides pose severe threats to infrastructure and safety, and their segmentation in remote sensing imagery remains challenging due to irregular boundaries, scale variation, and complex terrain. Traditional lightweight models often struggle to capture rich semantic features under these conditions. To address this, we leverage vision foundation models (VFMs) as teachers and propose a knowledge distillation adversarial (KDA) framework to transfer high-capacity knowledge into compact student models. Additionally, we introduce a dynamic cross-layer fusion (DCF) decoder to enhance global–local feature interaction. The experimental results demonstrate that, compared to the previous best-performing model SegNeXt [89.92% precision and 84.78% mean intersection over union (mIoU)], our method achieves a precision of 91.93% and mIoU of 86.53%, yielding improvements of 2.01% and 1.75%, respectively. Source code is available athttps://github.com/PreWisdom/KDA
Lulin Li, Xuan Dong 0001, Lei Shi 0002, Pin Tao
IEEE Geosci. Remote. Sens. Lett.5
2025 A Novel Two-Stream Algorithm for Spatio-Temporal Fusion in Remote Sensing
abstract
The spatio-temporal fusion technology can effectively solve the problem of missing time series data. However, existing algorithms often struggle to accurately capture surface feature changes and fine details. To overcome these limitations, this letter introduces FemSTF, a flexible two-stream algorithm for remote sensing spatio-temporal fusion, which consists of three main stages: image sharpening, image fusion, and edge optimization. In the first stage, techniques such as interpolation, classification, and blending are employed to address missing pixel values caused by minimal input data. In the second stage, a two-stream strategy is introduced to enhance fusion accuracy, enabling the algorithm to capture regional detail changes with high sensitivity to surface variations. The third stage focuses on edge optimization, generating high-resolution predictive images under minimal input conditions (three images) while retaining rich texture and edge details. Notably, FemSTF achieves a performance comparable to that of methods requiring five input images. Experimental results demonstrate that FemSTF outperforms mainstream algorithms across multiple datasets. Ablation experiments confirm the effectiveness of the two-stream strategy, highlighting its role in enhancing remote sensing data processing accuracy. This study offers an efficient solution for spatio-temporal fusion, demonstrating great potential in the field.
Xuechao Zou, Pin Tao
IEEE Geosci. Remote. Sens. Lett.7
2025 Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images
abstract
Cloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated powerful generalization capabilities across various visual tasks. In this paper, we present a parameter-efficient adaptive approach, termed Cloud-Adapter, designed to enhance the accuracy and robustness of cloud segmentation. Our method leverages a VFM pretrained on general domain data, which remains frozen, eliminating the need for additional training. Cloud-Adapter incorporates a lightweight spatial perception module that initially utilizes a convolutional neural network (ConvNet) to extract dense spatial representations. These multi-scale features are then aggregated and serve as contextual inputs to an adapting module, which modulates the frozen transformer layers within the VFM. Experimental results demonstrate that the Cloud-Adapter approach, utilizing only 0.6% of the trainable parameters of the frozen backbone, achieves substantial performance gains. Cloud-Adapter consistently achieves state-of-the-art performance across various cloud segmentation datasets from multiple satellite sources, sensor series, data processing levels, land cover scenarios, and annotation granularities. Code and model checkpoints are available at https://xavierjiezou.github.io/Cloud-Adapter/.
Xuechao Zou, Kai Li 0023, Junliang Xing, Lei Jin 0003, Congyan Lang, Pin Tao
IEEE Trans. Geosci. Remote. Sens.8
2024 Advancing Video Synchronization with Fractional Frame Analysis: Introducing a Novel Dataset and Model
abstract
Multiple views play a vital role in 3D pose estimation tasks. Ideally, multi-view 3D pose estimation tasks should directly utilize naturally collected videos for pose estimation. However, due to the constraints of video synchronization, existing methods often use expensive hardware devices to synchronize the initiation of cameras, which restricts most 3D pose collection scenarios to indoor settings. Some recent works learn deep neural networks to align desynchronized datasets derived from synchronized cameras and can only produce frame-level accuracy. For fractional frame video synchronization, this work proposes an Inter-Frame and Intra-Frame Desynchronized Dataset (IFID), which labels fractional time intervals between two video clips. IFID is the first dataset that annotates inter-frame and intra-frame intervals, with a total of 382,500 video clips annotated, making it the largest dataset to date. We also develop a novel model based on the Transformer architecture, named InSynFormer, for synchronizing inter-frame and intra-frame. Extensive experimental evaluations demonstrate its promising performance. The dataset and source code of the model are available at https://github.com/yuxuan-cser/InSynFormer.
Haizhou Ai, Junliang Xing, Xuri Li, Pin Tao
AAAI6
2024 LEFormer: A Hybrid CNN-Transformer Architecture for Accurate Lake Extraction from Remote Sensing Imagery
abstract
Lake extraction from remote sensing images is challenging due to the complex lake shapes and inherent data noises. Existing methods suffer from blurred segmentation boundaries and poor foreground modeling. This paper proposes a hybrid CNN-Transformer architecture, called LEFormer, for accurate lake extraction. LEFormer contains three main modules: CNN encoder, Transformer encoder, and cross-encoder fusion. The CNN encoder effectively recovers local spatial information and improves fine-scale details. Simultaneously, the Transformer encoder captures long-range dependencies between sequences of any length, allowing them to obtain global features and context information. The cross-encoder fusion module integrates the local and global features to improve mask prediction. Experimental results show that LEFormer consistently achieves state-of-the-art performance and efficiency on the Surface Water and the Qinghai-Tibet Plateau Lake datasets. Specifically, LEFormer achieves 90.86% and 97.42% mIoU on two datasets with a parameter count of 3.61M, respectively, while being 20× minor than the previous best lake extraction method. The source code is available at https://github.com/BastianChen/LEFormer.
Xuechao Zou, Yu Zhang 0165, Jiayu Li 0008, Kai Li 0023, Junliang Xing, Pin Tao
ICASSP7
2024 High-Fidelity Lake Extraction Via Two-Stage Prompt Enhancement: Establishing A Novel Baseline and Benchmark
abstract
Lake extraction from remote sensing imagery is a complex challenge due to the varied lake shapes and data noise. Current methods rely on multispectral image datasets, making it challenging to learn lake features accurately from pixel arrangements. This, in turn, affects model learning and the creation of accurate segmentation masks. This paper introduces a prompt-based dataset construction approach that provides approximate lake locations using point, box, and mask prompts. We also propose a two-stage prompt enhancement framework, LEPrompter, with prompt-based and prompt-free stages during training. The prompt-based stage employs a prompt encoder to extract prior information, integrating prompt tokens and image embedding through self- and cross-attention in the prompt decoder. Prompts are deactivated to ensure independence during inference, enabling automated lake extraction without introducing additional parameters and GFlops. Extensive experiments showcase performance improvements of our proposed approach compared to the previous state- of-the-art method. The source code is available at https://github.com/BastianChen/LEPrompter.
Xuechao Zou, Yu Zhang 0165, Junliang Xing, Pin Tao
ICME6
2024 A Parallel Attention Network For Cattle Face Recognition
abstract
Cattle face recognition holds paramount significance in domains such as animal husbandry and behavioral research. Despite significant progress in confined environments, applying these accomplishments in wild settings remains challenging. Thus, we create the first large-scale cattle face recognition dataset, ICRWE, for wild environments. It encompasses 483 cattle and 9,816 high-resolution image samples. Each sample undergoes annotation for face features, light conditions, and face orientation. Furthermore, we introduce a novel parallel attention network, PANet. Comprising several cascaded Transformer modules, each module incorporates two parallel Position Attention Modules (PAM) and Feature Mapping Modules (FMM). PAM focuses on local and global features at each image position through parallel channel attention, and FMM captures intricate feature patterns through non-linear mappings. Experimental results indicate that PANet achieves a recognition accuracy of 88.03% on the ICRWE dataset, establishing itself as the current state-of-the-art approach. The source code is available at https://github.com/1jy-0124/PANet
Jiayu Li 0008, Xuechao Zou, Junliang Xing, Pin Tao
ICME6
2024 Absorb What You Need: Accelerating Exploration via Valuable Knowledge Extraction
abstract
Leveraging external knowledge and extracting valuable insights are efficient human practices when handling various tasks. In contrast, current artificial intelligence lacks this capability. Recent research aims to teach Reinforcement Learning (RL) agents to incorporate external knowledge in the form of natural language to accelerate exploration. A common assumption in many of these approaches is that all introduced external knowledge is inherently valuable. To eliminate this assumption, we introduce the Knowledge Extraction Exploration Framework (KEEF). KEEF comprises two key components: 1) a knowledge extractor, designed to filter useful external knowledge based on three dimensions, task relevance, environment relevance, and achievement difficulty, through a prediction network and a policy network; and 2) a policy executor, which is a knowledge-conditioned network facilitating joint reasoning between the useful knowledge extracted by the knowledge extractor and the current state of the environment. In eight challenging sparse reward BabyAI environments, KEEF has consistently demonstrated superior sampling efficiency compared to knowledge-based and traditional RL methods.
Renye Yan, Pin Tao, Junliang Xing
IJCNN4
2024 Freedom of choice disrupts cyclic dominance but maintains cooperation in voluntary prisoner's dilemma game
Danyang Jia, Chen Shen 0006, Xiangfeng Dai, Xinyu Wang 0022, Junliang Xing, Pin Tao, Yuanchun Shi, Zhen Wang 0004
Knowl. Based Syst.6
2024 DiffCR: A Fast Conditional Diffusion Framework for Cloud Removal From Optical Satellite Images
abstract
Optical satellite images are a critical data source; however, cloud cover often compromises their quality, hindering image applications and analysis. Consequently, effectively removing clouds from optical satellite images has emerged as a prominent research direction. Recent advances in deep learning-based cloud removal methods have been significant, but image generation quality still needs improvement. Diffusion models have demonstrated remarkable success in diverse image-generation tasks, showcasing their potential in addressing this challenge. This paper presents a novel framework called DiffCR, which leverages conditional guided diffusion with deep convolutional networks for high-performance cloud removal for optical satellite imagery. Specifically, we introduce a decoupled encoder for conditional image feature extraction, providing a robust color representation to ensure the close similarity of appearance information between the conditional input and the synthesized output. Moreover, we propose a novel and efficient time and condition fusion block within the cloud removal model to accurately simulate the correspondence between the appearance in the conditional image and the target image at a low computational cost. Extensive experimental evaluations on three commonly used benchmark datasets demonstrate that DiffCR consistently achieves state-of-the-art performance on all metrics, with parameter and computational complexities amounting to only 5.1% and 5.4%, respectively, of those previous best methods. The source code, pre-trained models, and all the experimental results will be publicly available at https://github.com/XavierJiezou/DiffCR upon the paper’s acceptance of this work.
Xuechao Zou, Kai Li 0023, Junliang Xing, Yu Zhang 0165, Lei Jin 0003, Pin Tao
IEEE Trans. Geosci. Remote. Sens.7
2023 PMAA: A Progressive Multi-Scale Attention Autoencoder Model for High-Performance Cloud Removal from Multi-Temporal Satellite Imagery
abstract
Satellite imagery analysis plays a pivotal role in remote sensing; however, information loss due to cloud cover significantly impedes its application. Although existing deep cloud removal models have achieved notable outcomes, they scarcely consider contextual information. This study introduces a high-performance cloud removal architecture, termed Progressive Multi-scale Attention Autoencoder (PMAA), which concurrently harnesses global and local information to construct robust contextual dependencies using a novel Multi-scale Attention Module (MAM) and a novel Local Interaction Module (LIM). PMAA establishes long-range dependencies of multi-scale features using MAM and modulates the reconstruction of fine-grained details utilizing LIM, enabling simultaneous representation of fine- and coarse-grained features at the same level. With the help of diverse and multi-scale features, PMAA consistently outperforms the previous state-of-the-art model CTGAN on two benchmark datasets. Moreover, PMAA boasts considerable efficiency advantages, with only 0.5% and 14.6% of the parameters and computational complexity of CTGAN, respectively. These comprehensive results underscore PMAA’s potential as a lightweight cloud removal network suitable for deployment on edge devices to accomplish large-scale cloud removal tasks. Our source code and pre-trained models are available at https://github.com/XavierJiezou/PMAA.
Xuechao Zou, Kai Li 0023, Junliang Xing, Pin Tao, Yachao Cui
ECAI4
2023 Mnemonic Dictionary Learning for Intrinsic Motivation in Reinforcement Learning
abstract
Reinforcement learning for hard-exploration tasks remains challenging due to the long-term dependence and sparse-and-delay rewards in complex environments. In these challenging tasks, intrinsic motivation has become a dominant paradigm to enable the agent to explore the environment when no external reward feedback is available. In this work, inspired by studies from the human memory mechanism, we present a mnemonic dictionary learning (MDL) model for intrinsic motivation in reinforcement learning. The MDL model leverages sparse dictionary learning to incremental abstract the exploration histories into a compact memory-like dictionary, providing an excellent intrinsic motivation model. This mnemonic dictionary model not only drives the agent to explore novel stats in the environments indicated by the memory reconstruction error but also helps the agent to remember the key states and structure of the environments using its learned bases and reconstruction coefficients. The proposed MDL model can serve as a generative module for existing exploration methods. Extensive experimental results on typical sparse-reward tasks demonstrate its effectiveness and applicability over several competing algorithms. We will release the source code and trained models to facilitate further studies in this research direction.
Renye Yan, Yuan Zhan, Pin Tao, Zongwei Wang 0001, Yimao Cai, Junliang Xing
IJCNN4
2023 TLMIX: Twin Leader Mixing Network for Cooperative Multi-Agent Reinforcement Learning
abstract
Recent methods of cooperative multiagent rein-forcement learning built upon the individual global max value decomposition principle show promising results using variants of deep mixing networks, and credit assignment plays a crucial role in it. However, each agent in a multiagent system requires not only credit assignment but also credit feedback which tells each agent how many rewards it should obtain to maximize expected cumulative global rewards. In this work, we propose TLMIX, a novel Twin Leader Mixing Network for multiagent cooperation reinforcement learning while maintaining the centralized training and decentralized execution paradigm. TLMIX introduces a leader network to address the credit feedback issue by utilizing global information to provide reasonable objectives for agent networks. TLMIX also introduces a twin mixing network to find a more accurate target function from the Q-value functions, which avoids the rapid increase in parameter scale caused by introducing individual agents' twin networks and effectively mitigates the accumulation of high overestimation errors caused by temporal difference updates. Extensive results on SMAC experimental scenarios and the Predator-Prey environment demonstrate that TLMIX significantly outperforms comparable benchmark algorithms on convergence speed and performance.
Yu Zhang 0165, Pengyu Gao, Yusheng Jiang, Junliang Xing, Pin Tao
IJCNN6
2023 Pseudo Value Network Distillation for High-Performance Exploration
abstract
Solving hard exploration tasks with sparse rewards is notoriously challenging in reinforcement learning (RL), which needs to address two key issues simultaneously: exploiting past successful experiences and exploring the unknown environment. Many prior works take expert demonstrations as successful experiences and learn to imitate them directly. However, these demonstrations are often not available in practice. Recently, curiosity-driven RL methods provide intrinsic rewards, encouraging the agent to explore states with high novelty. Nonetheless, they lack a mechanism for leveraging past good experiences effectively. This work presents a Pseudo Value Network Distillation (PVND) framework to balance the RL agent's exploitative and exploratory behaviors effectively and automatically. In particular, PVND learns to set high exploitation bonuses to the critical states in rewarded trajectories from past experiences and high exploration bonuses to the novel states that agents rarely visit during exploration. We theoretically demonstrate that PVND gives larger positive intrinsic rewards to more critical states. Furthermore, PVND automatically finds meaningful and critical hierarchical sub-tasks for agents to accomplish the final goal progressively. Competitive results in several hard exploration sparse reward problems have verified its effectiveness and efficiency.
Enmin Zhao, Junliang Xing, Kai Li 0022, Yongxin Kang, Pin Tao
IJCNN5
2019 Livestock detection in aerial images using a fully convolutional network
abstract
In order to accurately count the number of animals grazing on grassland, we present a livestock detection algorithm using modified versions of U-net and Google Inception-v4 net. This method works well to detect dense and touching instances. We also introduce a dataset for livestock detection in aerial images, consisting of 89 aerial images collected by quadcopter. Each image has resolution of about 3000×4000 pixels, and contains livestock with varying shapes, scales, and orientations. We evaluate our method by comparison against Faster RCNN and Yolo-v3 algorithms using our aerial livestock dataset. The average precision of our method is better than Yolo-v3 and is comparable to Faster RCNN.
Pin Tao, Ralph R. Martin
Comput. Vis. Media2
2015 Research on system level innovative practice in embedded laboratory education
abstract
Embedded System Curriculum at Tsinghua University is target at advanced undergraduate senior students in CS department. At the end of this course, an overall understanding about System is expected to be obtained by students who have gotten perfect software knowledge from software series curriculums and hardware knowledge from hardware series curriculums. Embedded System is unique for its close association with actual application, thus an effective laboratory education is of great importance for embedded system curriculum. In this paper, we are dedicated to reform existing experimental teaching models, methods and laboratory platform to improve system level innovation practice for CS undergraduate students. Firstly, we mainly analyzed the characteristics of embedded laboratory education and current generation students. Based on this analysis, some principles as well as a concept model that guiding the designing of embedded laboratory platform are proposed. And then we present the whole innovative embedded laboratory system in CS department of Tsinghua University. Supported by the system, some typical projects will be presented, include intelligent Greenhouse, Sunflower, intelligence handcart, intelligence spider and rotorcraft.
Ninghan Zheng, Pin Tao
FIE3
2015 Improvement of re-sample template matching for lossless screen content video
abstract
Screen Content (SC) video coding becomes more important for screen sharing and screen broadcasting applications. There are many easy to see different characters between screen content video and camera-captured video. We proposed a template matching prediction method for lossless SC intra picture coding with higher compression ratio. The pixels are re-sampled to form the Virtual Largest Coding Unit (VLCU) firstly. About 80% pixels in VLCU can be predicted exactly by template matching with zero error. Then, pixels with non-zero prediction error should be coded with three information, index, position and value. Among these three, position will consume the most bits than the other two. In order to handle this challenge, we propose to apply the similarity of non-zero prediction error pixel positions of neighbor VLCU which can greatly help to improve the compression performance. The VLCU can be divided into sub CU as the same as the division in the standard High Efficiency Video Coding(HEVC) intra coding, and RDO is applied to find the best coding efficiency.
Pin Tao, Lixin Feng, Sichao Song 0002, Jiangtao Wen, Shiqiang Yang
ICME1
2015 Rearrangement pixel granularity template matching for lossy screen content picture intra coding
abstract
Screen Content (SC) video coding is becoming more important in screen sharing and screen broadcasting applications. There are many distinctly different characteristics between the SC video and the camera-captured video. In this paper, we propose a rearrangement pixel granularity template matching method for lossy SC intra picture coding with higher compression ratio. Firstly, the original picture will be rearranged for lossy compression and parallel processing purpose. Then every pixel is predicted with template matching method. The template matching method can achieve high prediction accuracy on SC pictures. Our proposed method improves about 4.5dB and saves 11.7% bitrate in comparison with the HEVC range extension software HM12.0+RExt4.1.
Pin Tao, Lixin Feng
VCIP2
2015 Efficient Software H.264/AVC to HEVC Transcoding on Distributed Multicore Processors
abstract
The latest High Efficiency Video Coding (HEVC) standard achieves a significant compression efficiency improvement over the H.264/Advanced Video Coding (AVC) standard, but with a much higher computational complexity. In this paper, we propose a novel framework for software-based H.264/AVC to HEVC transcoding, integrated with tools such as wavefront parallel processing that are useful for achieving higher levels of parallelism on multicore processors and distributed systems. By utilizing information extracted from the input H.264/AVC bitstream, the transcoding process can be greatly accelerated with a visual quality loss that is modest for many applications. Based on the HEVC HM 14.0 reference software and using standard HEVC test bitstreams, the proposed transcoder can achieve up to 60× speedup on a Quad Core 8-thread server over decoding-re-encoding based on FFMPEG and the HM software with a BD-rate loss of 15%-20%. By implementing a group of picture-level task distribution on a distributed system with nine processing units, the proposed software transcoder can achieve a speed for transcoding 720 p at 30 Hz in real time.
Yucong Chen, Ziyu Wen, Jiangtao Wen, Minhao Tang, Pin Tao
IEEE Trans. Circuits Syst. Video Technol.5
2013 Cross Segment Decoding for Improved Quality of Experience for Video Applications
abstract
In this paper, we present an improved algorithm for decoding live streamed or pre-encoded video bit streams with time-varying qualities. The algorithm extracts information available to the decoder from a high visual quality segment of the clip that has already been received and decoded, but was encoded independently from the current segment. The proposed decoder is capable of significantly improve the Quality of Experience of the user without incurring significant overhead to the storage and computational complexities of both the encoder and the decoder. We present simulation results using the HEVC reference encoder and standard test clips, and discuss areas of improvements to the algorithm and potential ways of incorporating the technique to a video streaming system or standards.
Jiangtao Wen, Shunyao Li, Yao Lu 0006, Meiyuan Fang, Xuan Dong 0001, Huiwen Chang, Pin Tao
DCC7
2011 Harmonicare: a novel wind instrument easy to learn and play
abstract
In this paper, we present a novel design of the wind instrument, Harmonicare, which makes everybody playing the wind instrument easily, having a lot of fun and doing the respiratory training simultaneously. The device contains many valves for all reed chambers of harmonica, they control when and which note should be played according to the music notation. What the player needs to do is exhaling and inhaling rhythmically according to the hints on the device screen. The design combines the technology, musical arts and health care together. A bundle of Volunteers' experiences show that our novel design is easy to master and full of fun.
Pin Tao, Yinqiao Wang
UbiComp1
2010 Horizontal Spatial Prediction for High Dimension Intra Coding
abstract
Macroblock level Horizontal Spatial Prediction(HSP) based intra frame coding scheme for High Dimension(HD) video sequences was proposed in this paper. According to the correlation experiment on HD sequences, most HD pictures have the stronger horizontal spatial correlation than the vertical spatial correlation, about 2dB stronger. This phenomena drop a valuable hint to us that the horizontal spatial prediction can be used in HD video intra coding without considering the vertical spatial prediction. An adaptive divide and predict intra frame coding scheme has been proposed by Piao which has the similar idea. But this method divided the whole picture into several parts which is not conform to the conventional macroblock based video coding framework and it has the high computation complexity in motion estimation procedure.
Pin Tao, Wenting Wu, Chao Wang 0063, Mou Xiao, Jiangtao Wen
DCC1
2010 Fast Rate Distortion Optimized Quantization for H.264/AVC
abstract
In this paper, a fast RDO (rate-distortion optimization) quantization algorithm for H.264/AVC is proposed. In this algorithm, the searching space of level adjustments is reduced by filtering the input quantized coefficients in a hierarchical way. The well quantized coefficients is first filtered out, and then the RD tradeoff of each level adjustment to each of the rest coefficients is examined to select some good candidates with their associated level adjustments. Finally these good candidates are combined to find the best combination of level adjustments which gives the minimal rate-distortion cost. Furthermore, a fast rate estimation technique is adopted to save the rate-distortion estimation time. Experimental results show that about 44% quantization time on average can be saved at the cost of negligible PSNR loss compared with RDO quantization algorithm implemented in JM.
Jiangtao Wen, Mou Xiao, Pin Tao, Chao Wang 0063
DCC4
2010 HFAG: Hierarchical Frame Affinity Group for video retrieval on very large video dataset
abstract
Content-based video retrieval systems are desired to fast and accurately find the nearest-neighbors of user input examples from very large video datasets. This poses a great challenge since exhaustive and redundant computation of similarities is required. Cluster based index approaches can be used to address this problem, but the similarity computation and clustering methods for videos are very time-consuming, thus preventing it from indexing very large video datasets. In this paper, we propose the Hierarchical Frame Affinity Group (HFAG), which is a hierarchy of frame clusters built using affinity propagation (AP) method, to represent video clusters. Our proposed video similarity metric and AP method guarantee the high performance of forming HFAG. We then build the cluster-based index structure to support retrieval of the nearest-neighbors of video sequences. The experiments on real large video datasets prove the effectiveness and efficiency of our approach.
Yin-Jun Miao, Chao Wang 0063, Peng Cui 0001, Lifeng Sun, Pin Tao, Shiqiang Yang
ICIP5
2010 Improved intra prediction for high definition video using localized horizontal spatial prediction
abstract
In this paper, a localized horizontal spatial prediction (HSP) based algorithm was proposed for Intra coding of high definition (HD) inputs. In the algorithm, a block of size 32×16 is divided into two 16×16 MBs, consisting of the pixels from the even-numbered columns and the odd-numbered columns (termed the even MB and the odd MB) respectively. The even MB is encoded using conventional Intra coding techniques and then its reconstruction is used for the prediction of the odd MB. Experimental results show that up to 0.79 dB and on average 0.39 dB gain can be achieved for HD sequences with the proposed framework at lower complexity than H.264 Intra coding.
Wenting Wu, Pin Tao, Mou Xiao, Jiangtao Wen, Ruiping Li
ICIP2
2010 AN efficient algorithm for joint QP and quantization optimization for H.264/AVC
abstract
We describe an efficient algorithm for jointly optimizing the quantization parameter (QP) and the quantization decisions in H.264/AVC. Based on bitrate estimation for quantized and entropy coded coefficients in H.264/AVC, a technique was introduced to select only a small subset of transform coefficients which have the most significant impact on the overall quantization performance. Only the quantization decisions for the coefficients in this subset need to be explicitly examined along with the QP. Simulation results showed that the proposed algorithm achieves similar quantization performances as existing state-of-the-art optimized algorithms at roughly half the complexity.
Mou Xiao, Jiangtao Wen, Pin Tao
ICIP4
2010 Macroblock level hybrid temporal-spatial prediction for H.264/AVC
abstract
In this paper, a novel macroblock level hybrid temporal-spatial video coding framework is proposed. In this framework, a new Hybrid Temporal Spatial Prediction (HTSP) coding mode is adopted for each macroblock. Macroblock is divided into two partitions. The first partition is temporally predicted using motion compensation and encoded, while the second partition is spatially predicted using the reconstruction of the first partition. Experimental results show that up to 0.4 dB coding gain and 0.2 dB on average can be achieved for HD sequences.
Mou Xiao, Pin Tao, Wenting Wu, Jiangtao Wen
ISCAS2
2010 Aware service base on the assembly of multiple weak location sensors
abstract
Indoor location information is important for the awareness computing environment. One single location sensor always is not accurate and applicable enough for providing the right service to the user. In this paper, an OSGi based simulation system is introduced firstly, and then we propose a framework to support the combination of several weak location sensors. With this combination model, multiple weak location sensors can work together to provide the better location aware service to the user. The experiment platform which consists of RFID array location sensor, RF sensor, pressure sensor and other sensors can work together in the OSGi based simulation system, experiment results show that the assembly of multiple weak sensors do better than the solo sensor. It can help the designers to achieve the better cost-performance for the awareness applications between various possible combinations.
Pin Tao, Lei Zhang 0060, Yu Chen 0004
SMC1
2008 An adaptive frame interpolation algorithm using statistic analysis of motions and residual energy
abstract
In this paper, a new motion-compensated frame interpolation (MCFI) algorithm using adaptive criterions to correct the motion vector field is proposed. First, an effective pre-processing scheme is done to the transmitted motion vectors. Then, unlike the conventional MCFI algorithms using fixed criterion to analyse the reliability of motion vectors, we proposed to select the parameters and thresholds by analysing the statistical characterization of motion vectors and residual energy, thus thresholds can be changed adaptively during the decoding process. Meanwhile, our new criterions consider both reliability of motion vectors and the smoothness of the region, which avoid the unnecessary motion estimation and reduce the complexity. Experimental results show that the proposed algorithm has 1.4~5.5 dB increase comparing with vector median filter algorithm in average PSNR, and greatly improves the subjective visual quality.
Chunbo Yang, Pin Tao, Shiqiang Yang
MMSP2
2006 The Analysis of Offloading H.264 Video Encoder on Mobile Devices for Energy Saving
abstract
As more and more people want to communicate with each other vividly, video encoder application will be very popular on mobile devices. But it is also a resource-killer application. Meanwhile the mobile devices are powered by battery, and energy resource is very important for mobile devices. So when using this application on mobile devices, we must consider energy conserving. The computation offloading method is to offload some computations of application from mobile device to powerful server around mobile device via wireless network. So it can relieve the resource constraint and save energy. In this paper, we use the computation offloading method to H.264 video encoder on mobile device and propose three offloading schemes. Then we apply them to encode different video sequences. Finally, we analyze the efficiency of offloading schemes with different encoder parameters and different video sequences.
Pin Tao, Shiqiang Yang
SMC2