VLDB 2026 Research / reviewers in the wild / expert
Qian Huang 0008
dblp:07/4378-8
· DBLP profile ↗
55ranked-venue papers
10as first author
43since 2021 · last 2026
0000-0001-5625-0402ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 10 first-author · 31 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparsely guided adaptive pseudo-supervised learning for nucleus segmentation
Qian Huang 0008, Zhijian Wang 0002, Meng Geng |
Expert Syst. Appl. | 2 |
| 2026 | Spatial-Temporal Self-Compensating Graph Convolutional Network for Skeleton-Based Action Recognition Under Data ConstraintsabstractSkeleton-based human action recognition has emerged as a prominent research focus in computer vision, with significant progress achieved in recent years. However, existing methods often suffer substantial performance degradation under real-world data constraints, such as body occlusion, missing frames, and noise. These limitations critically undermine the robustness of related techniques in practical applications. To address these challenges, we propose a Spatial Temporal Self-compensating Graph Convolutional Network (STSc-GCN), which skillfully utilizes the systematic and regular nature of human movement to mitigate performance degradation caused by data constraints through a data self-compensation mechanism. Specifically, STSc-GCN comprises two key modules: 1) collaborative motion spatial compensation (CMSC). This module designs multiple distinct topological relationships, primarily including Walk-probability Generality Topology and Self-organizing Particularity Topology, respectively, to deeply explore the universal and personalized collaborative relationships between human joints. These relationships help compensate for the lack of information caused by spatial data constraints and 2) meta-action sharpening temporal Compensation (MSTC). This module introduces a novel motion sharpening mechanism that enhances key dynamic information within the meta-action sequences through cross-attention technology, thereby improving model adaptability to missing-frame scenarios. STSc-GCN achieves state-of-the-art performance on four constrained datasets and shows superior results on three widely used standard datasets, confirming its effectiveness in both constrained and general scenarios. Code will be available at https://github.com/XingLi1012/STSc-GCN.git. Xing Li 0005, Qian Huang 0008, Xin Li 0090, Jinhui Tang 0001, Qiaolin Ye |
IEEE Trans. Image Process. | 3 |
| 2026 | MD-PCSN: Meta-Motion Decoupling Point Cloud Sequence Network for Privacy-Preserving Human Action Recognition in AI MachinesabstractIn next-generation communication networks and Industry 5.0 based applications, ensuring robust security and reliability in human-computer interaction (HCI) constitutes a fundamental prerequisite for safety-critical AI machine systems. Point cloud sequence-based human action recognition demonstrates intrinsic advantages in privacy-preserving HCI, leveraging its non-intrusive sensing modality to mitigate data vulnerability while maintaining high-precision action interpretation in industrial environments. Existing spatio-temporal encoding methods for point cloud sequence-based action recognition suffer from two fundamental limitations: (1) rigid neighborhood constraints impair multi-scale feature extraction for heterogeneous body parts, and (2) independent spatial-temporal decomposition introduces motion representation distortion. We propose a Meta-motion Decoupling Point Cloud Sequence Network (MD-PCSN) that addresses these challenges through: (1) logarithmic spatio-temporal point convolution for hierarchical meta-motion construction at variable granularities, and (2) a novel Gated-KANsformer architecture with differential motion encoding to explicitly model both short-term displacements and long-term spatio-temporal dependencies. The proposed meta-motion decoupling mechanism significantly enhances robustness against sensor perturbations, making the framework particularly suitable for security-critical applications. Extensive experiments on three benchmark datasets demonstrate MD-PCSN’s superior performance. It outperforms classic PST-Transformer by 1.5% on MSR Action3D and 4.14% on UTD-MHAD. Under the NTU RGB+D 60, it achieves 2.9% cross-view gain over the latest PointActionCLIP. Xing Li 0005, Xin Li 0090, Qian Huang 0008 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2026 | Dual-Scale Transformer with Variable Bitrate Synchronization for Neural Video CompressionabstractNeural video compression (NVC) has emerged as a promising paradigm for improving rate-distortion performance. However, existing neural video codecs predominantly rely on convolutional neural networks (CNNs) with limited local receptive fields to generate the latent representations, often neglecting global–local spatial correlations. This leads to suboptimal feature modeling and redundancy in the latent space. To address this limitation, we propose a novel Dual-Scale Transformer (DST) block specifically tailored for NVC, which effectively enhances coding efficiency. The DST block incorporates a Global–Local (Shifted) Window-based Self-Attention (GL(S)WSA) mechanism to jointly capture global structure information and local texture details. Moreover, we design a Cross-Gated Feed-Forward Network (CGFFN) to adaptively modulate complementary components, producing more compact and expressive latent representations. Furthermore, to overcome the drawbacks of traditional asynchronous training and further boost rate-distortion performance, we introduce a Variable Bitrate Synchronization (VBRS) strategy that leverages multi-GPU parallel training, with each GPU dedicated to a specific bitrate and synchronized via gradient backpropagation for joint optimization. Experimental results demonstrate that our proposed method achieves the higher coding performance compared to the previous state-of-the-art (SOTA) methods and significantly outperforms H.266/VVC (VTM-13.2) under various low delay B (LDB) coding configurations. Yiming Wang 0008, Yaojun Wu 0001, Zhaobin Zhang, Qian Huang 0008, Bin Tang 0002, Zhangjing Yang, Kai Zhang 0007, Li Zhang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Learned Video Compression With Refined Adaptive Flow Pyramid And Coordinate-Aware AttentionabstractIn video compression, motion estimation and motion compensation are critical for achieving efficient encoding. Although the commonly used SpyNet and bilinear interpolation have contributed in improving the compression efficiency, they still have limitations. SpyNet often loses details and fails to fully utilize the feature extraction capabilities of deep networks. Furthermore, bilinear interpolation inherently attenuates high-frequency information, leading to frame blurring and distortion. In this paper, we propose a novel video compression algorithm. To overcome the limitations of SpyNet, we propose a refined adaptive flow pyramid network. This network uses a multi-scale feature pyramid to capture more details. Moreover, we use an iterative cost volume refinement engine that improves the feature representation of the network and iteratively improve the accuracy of motion estimation. In addition, to overcome the limitation of bilinear interpolation, we propose a coordinate-aware attention module, which captures more high-frequency information to improve the accuracy of motion compensation. Experimental results show that our method outperforms VTM-13.2 (LDP) in terms of PSNR. Qian Huang 0008, Xin Li 0090, Yiming Wang 0008 |
ICASSP | 1 |
| 2025 | Learned Video Compression with Spatial Correlation Priors and Hierarchical Temporal AttentionabstractAccurately predicting the probability distribution of quantized latent representations is a critical challenge for entropy models in learned video compression (LVC). Existing mainstream LVC methods typically adopt ready-made entropy models based on image compression, which fail to fully exploit the information of spatial-temporal correlation. To address this issue, we propose a spatial correlation priors and hierarchical temporal attention (SCP-HTA) model, which exploits the spatial correlation information from the current video frames and refine the temporal information from the context. First, we extract the spatial correlation of the current frame to guide the generation of masks, enabling the frame to leverage more information during encoding and decoding process. Additionally, to obtain more accurate temporal information, we introduce a hierarchical temporal attention module at channel level when we generate the context. Experimental results demonstrate that the proposed SCP-HTA model achieve 15.76% bitrate saving in PSNR and 61.82% in MS-SSIM on average across all test datasets when compared with VTM-13.2 (LDP). Qian Huang 0008, Wenchao Shan, Zaipeng Xie, Yiming Wang 0008 |
ICIP | 1 |
| 2025 | Improved Cervical Cell Detection Model Based on Hybrid-Domain Feature Pyramid NetworkabstractThe automated detection of cervical cells plays a vital role in cervical cancer screening. Cancer cells typically exhibit more prominent edge features compared to normal cells. However, during the feature fusion process, blurred edges often lack precise high-frequency information, which hinders the model’s ability to accurately identify cancer cells. To address this challenge, we propose the Hybrid-Domain Feature Pyramid Network (HD-FPN), which enhances the model’s sensitivity to cellular edge information. Specifically, we propose the Frequency-Aware Sampling that preserves critical texture features by leveraging high-frequency decomposition using wavelet transform. Furthermore, we design a multi-domain feature fusion approach to progressively refine edge features and minimize background noise. Our method achieves state-of-the-art performance on the CDetector and CRIC datasets in terms of both detection accuracy and efficiency. Qian Huang 0008, Ziyang Yin, Hao Lu 0013, Shaoling Qin |
ICIP | 1 |
| 2025 | Neighbor-Aware Feature-Driven Motion Compensation for Learned Video CompressionabstractLearned video compression (LVC) methods typically align spatial-temporal transformation features with optical flow to perform motion compensation. However, existing LVC methods typically rely on a single reference feature to provide local detail features. This ignores global structural information, resulting in limited capabilities when dealing with fast motion or occlusion scenarios. In addition, with the encoding of P-frames, errors accumulate, leading to a degradation in reconstruction quality. To address these issues, we propose the Neighbor-Aware Feature-Driven Motion Compensation (NAFD-MC) that utilizes the spatial-temporal correlations of the neighboring features to explore the global structural information and the local detail information. Furthermore, we introduce the Synergy Filtering Module (SFM) to enhance inter-frame consistency and alleviate the error accumulation. Experimental results demonstrate that our method outperforms the H.266/VVC reference software VTM-13.2 in public benchmark datasets. Hao Lu 0013, Qian Huang 0008, Ziyang Yin, Zaipeng Xie, Yiming Wang 0008 |
ICIP | 2 |
| 2025 | Efficient Local-Global Collaboration Transcoding for JPEG AIabstractIn the past decade, learning-based image compression has made significant advancements, with the Joint Photographic Experts Group (JPEG) working towards the launch of the first neural network-based image coding standard, JPEG AI. However, traditional codecs like JPEG and HEVC intra coding remain the most widely used. A key challenge arises when attempting to compress images pre-encoded with these traditional codecs using JPEG AI, as such images often carry artifacts that JPEG AI is not fully optimized to handle, leading to significant loss in performance. This paper addresses this issue by proposing a novel transcoding framework designed to optimize pre-encoded images for JPEG AI, bridging the gap between traditional and learning-based compression methods. The framework employs two branches for processing luminance and chrominance components, with a local-global modulation module (LGMM) for luminance and a dynamic fusion module (DFM) for chrominance. Experimental results demonstrate that the proposed scheme achieves up to 22.36% rate savings on images pre-encoded with HEVC intra and up to 22.29% rate savings on images pre-encoded with JPEG using JPEG AI reference software, highlighting its effectiveness in enhancing the performance of JPEG AI for real-world applications. Yiming Wang 0008, Zhaobin Zhang, Yaojun Wu 0001, Qian Huang 0008, Bin Tang 0002, Kai Zhang 0007, Li Zhang 0136 |
ICME | 4 |
| 2025 | Geo-CF2Net: Geometry-Prior Cross-Frequency Interactive Fusion Network for 3D Human Action RecognitionabstractDynamic point cloud-based human action recognition has garnered increasing attention due to its inherent advantages in privacy preservation and structural completeness. Current methods typically rely on nested point spatio-temporal convolutions to understand motion semantics in a bottom-up manner, which is intractable for capturing high-fidelity human dynamics disentangled from spatio-temporal interference. Motivated by this, designing a practical spatio-temporal factorization backbone is essential. However, the repeated coarsening of aggregated features along the spatial dimension often leads to the degradation of intrinsic geometric texture relations within point cloud data. Moreover, discretizing continuous visual data into isolated temporal hyperpoints significantly diminishes temporal continuity, resulting in the fragmentation of human action. To circumvent above limitations, we propose a novel Geometry-Prior Cross-Frequency Interactive Fusion Network (Geo-CF2Net). Specifically, we investigate a Spatial-Geometry Pose Prior (SGPP) module, which compensates for pose information loss during spatial downsampling by explicitly modeling geometric constraints among neighboring points. In addition, we elaborate on a Temporal Motion Unit Interactive Coordination (TMIC) module to track the interactive composite semantics of low-frequency steady-state venations and high-frequency transient-state details within a high-dimensional pose evolution flow. Extensive experiments on three public benchmarks substantiate the superiority of Geo-CF2Net over state-of-the-art methods. Qian Huang 0008, Xing Li 0005, Shihao Han, Yirui Wu, Xin Li 0090, Ziyang Yin |
ACM Multimedia | 2 |
| 2025 | STFE-VC: Spatio-temporal feature enhancement for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Expert Syst. Appl. | 2 |
| 2025 | A Systematic Review on Cell Nucleus Instance SegmentationabstractABSTRACT Cell nucleus instance segmentation plays a pivotal role in medical research and clinical diagnosis by providing insights into cell morphology, disease diagnosis, and treatment evaluation. Despite significant efforts from researchers in this field, there remains a lack of a comprehensive and systematic review that consolidates the latest advancements and challenges in this area. In this survey, we offer a thorough overview of existing approaches to nucleus instance segmentation, exploring both traditional and deep learning‐based methods. Traditional methods include watershed, thresholding, active contour model, and clustering algorithms, while deep learning methods include one‐stage methods and two‐stage methods. For these methods, we examine their principles, procedural steps, strengths, and limitations, offering guidance on selecting appropriate techniques for different types of data. Furthermore, we comprehensively investigate the formidable challenges encountered in the field, including ethical implications, robustness under varying imaging conditions, computational constraints, and the scarcity of annotated data. Finally, we outline promising future directions for research, such as privacy‐preserving and fair AI systems, domain generalization and adaptation, efficient and lightweight model design, learning from limited annotations, as well as advancing multimodal segmentation models. Qian Huang 0008, Meng Geng, Zhijian Wang 0002 |
IET Image Process. | 2 |
| 2025 | Multiscale motion-aware and spatial-temporal-channel contextual coding network for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Knowl. Based Syst. | 2 |
| 2025 | Review of cervical cell segmentation
Qian Huang 0008, Junzhou Chen 0002 |
Multim. Tools Appl. | 1 |
| 2024 | FedDGL: Federated Dynamic Graph Learning for Temporal Evolution and Data Heterogeneity
Zaipeng Xie, Likun Li, Xiangbin Chen, Qian Huang 0008 |
ACML | 5 |
| 2024 | Locality-Enhanced Transformer for Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractTransformers have emerged as a transformative tool in various computer vision tasks, excelling at capturing long-range dependencies. Their potential applicability and scalability in the interpretation of high-resolution remote sensing images (HRRSIs) have thus garnered substantial interest. However, unlike natural images, HRRSIs present intricate scenes characterized by scale variations and diverse appearances. These challenges underscore the importance of enabling networks to effectively assimilate both local intricacies and global context. In this letter, we introduce LETFormer, a semantic segmentation transformer. LETFormer balances capturing longrange dependencies with preserving local details through its unique LETFormer block, featuring an anchor token. This token aggregates localized contextual information within a designated window and promotes meaningful interactions among anchor tokens. With a mask transformer decoder, LETFormer gains ample contextual cues for precise semantic mask prediction. Empirical findings based on evaluations using the ISPRS Potsdam and LoveDA benchmarks unequivocally establish LETFormer’s superiority over state-of-the-art models. Additionally, we analyze the parameter size and floating-point operations per second (FLOPs) of LETFormer. Xin Li 0090, Feng Xu 0008, Runliang Xia, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Qian Huang 0008, Xin Lyu 0001 |
ICASSP | 7 |
| 2024 | Learned Video Compression with Spatial-Temporal OptimizationabstractPrevious optical flow based video compression is gradually replaced by unsupervised deformable convolution (DCN) based method. This is mainly due to the fact that the motion vector (MV) estimated by the existing optical flow network is not accurate and may introduce extra artifacts. However, DCN based method is difficult for training owing to the lack of explicit guidance in the feature space. In this work, we propose a learned video compression with spatial-temporal optimization. Specifically, we first propose the spatial-temporal motion refinement module to improve the accuracy of MV estimated by the optical flow network for prediction. Then, we propose the In-loop filter module to remove compression artifacts and improve the reconstructed frame quality. Finally, comprehensive experimental results demonstrate our proposed method outperforms the recent learned methods on three benchmark datasets. Moreover, our method also beats the H.266/VVC in terms of MS-SSIM metrics. Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Wenchao Shan |
ICASSP | 2 |
| 2024 | EchoGCN: An Echo Graph Convolutional Network for Skeleton-Based Action Recognition
Weiwen Qian, Qian Huang 0008, Zhongqi Chen, Yingchi Mao |
ICPR (15) | 2 |
| 2024 | A Purified Stacking Ensemble Framework for Cytology Classification
Linyi Qian, Qian Huang 0008, Junzhou Chen 0002 |
MMM (2) | 2 |
| 2024 | Multi-granular spatial-temporal synchronous graph convolutional network for robust action recognition
Qian Huang 0008, Yingchi Mao, Xing Li 0005, Jie Wu 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Temporal context video compression with flow-guided feature prediction
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Huashan Sun, Zhuang Miao |
Expert Syst. Appl. | 2 |
| 2024 | A survey of feature matching methodsabstractAbstract Feature matching plays a crucial role in computer vision, with applications in visual localization, simultaneous localization and mapping (SLAM), image stitching, and more. It establishes correspondences between sets of feature points from multiple images, enabling various tasks. Over the years, feature matching has witnessed significant development, with an increasing number of methods being applied. However, different methods exhibit different degrees of applicability in different scenarios and requirements due to their different rationales. To cope with these issues, a comprehensive analysis and comparison of matching methods are essential. Existing reviews often lack coverage of deep learning models and focus more on feature detection and description, neglecting the matching process. This survey investigates feature detection, description, and matching techniques within the feature‐based image‐matching pipeline. Representative methods, their mechanisms, and application scenarios are also briefly introduced. In addition, comprehensive evaluations of classical and state‐of‐the‐art methods are conducted through extensive experiments on representative datasets. Particularly, matching‐based applications are compared to fully demonstrate the advantages of the methods. Lastly, this survey highlights current problems and development directions in matching methods, serving as a reference for researchers in the field. Qian Huang 0008, Yiming Wang 0008, Huashan Sun |
IET Image Process. | 1 |
| 2024 | DMCVS: Decomposed motion compensation-based video stabilizationabstractAbstract With the popularity of handheld devices, video stabilization is becoming increasingly important. In previous studies, many methods have been proposed to stabilize shaky videos. However, these methods fail to balance between image content integrity and stability. Some methods sacrifice image content for better stability. Other methods ignore the subtle jitters, which leads to poor stability. This work innovatively proposes a video stabilization method based on decomposed motion compensation. First, a grid‐based motion statistics method is adopted for motion estimation, which obtains more accurate motion vectors according to matched likelihood estimates. Then, the motion compensation is inherently decomposed into two parts: linear motion compensation and auxiliary motion compensation. Linear motion compensation removes complex jitter by constructing linear path constraints to obtain a more stable camera path. Auxiliary motion compensation uses a moving average filter to remove the high‐frequency jitter as a supplement and preserve more image content. The two components are combined with individual weights to derive the final transform matrix and warp the original frames. Experimental results show that our method outperforms the previous methods on NUS and DeepStab datasets qualitatively and quantitatively. Qian Huang 0008, Jiwen Liu, Chuanxu Jiang, Yiming Wang 0008 |
IET Image Process. | 1 |
| 2024 | Fusing angular features for skeleton-based action recognition using multi-stream graph convolution networkabstractAbstract Distinguishing similar actions has been a challenging challenge in skeleton‐based action recognition. Since the joint coordinates in these actions are similar, it is difficult to accomplish the recognition task using traditional joint features. To address this issue, the use of angle features to capture subtle nuances in various body parts, along with a critical angle enhancement module that assigns weights to different angle feature representations for a given action are proposed, highlighting the critical angle feature representation. The approach is evaluated using a three‐stream ensemble method on three large action recognition datasets, NTU‐RGB+D, NTU‐RGB+D 120, and Kinetics‐400. The experimental results demonstrate that incorporating angular information can effectively complement joint and skeletal features, leading to improved recognition of similar actions and enhanced model performance and robustness. Qian Huang 0008, Mingzhou Shang, Yiming Wang 0008 |
IET Image Process. | 1 |
| 2024 | Ship detection based on YOLO algorithm for visible imagesabstractAbstract Ship detection is a crucial task for waterway surveillance and channel optimization, especially in close proximity to the shore. However, detecting ship in visible image‐based detection remains a challenge due to the limited nature of visible image datasets. To address this issue, the Inland Ships Data Set (ISDS) is constructed to facilitate research on ship identification. On the other hand, most detection methods struggle to accurately identify ships that are small in size. Therefore, a visible image‐based ship detection model is proposed that employs a multi‐scale weighted feature fusion structure with the YOLOv4 detection model to improve the efficacy of small ship detection. Specifically, the YOLOv4 model is improved through fusing multi‐scale feature, redesigning priori frame, and enhancing loss function. The model, named YOLOv4‐MSW (i.e. YOLOv4 based on Multi‐Scale Weighted feature fusion), exhibits improved performance on ship detection in experiments conducted on the ISDS dataset, outperforming the original YOLOv4 model by improving the average precision (AP) by 4.87% and the recall rate by 10.03%. Meanwhile, the model achieve better detection accuracy and improve the average precision rate by at least 0.86% compared to existing learned object detection methods. The code related to this work are released at https://github.com/Sunhuashan/YOLOv4‐MSW . The whole dataset is available at https://drive.google.com/drive/folders/1fzJ2fcqiko6lFwqIEGghMceoQgv‐8jBy . Qian Huang 0008, Huashan Sun, Yiming Wang 0008 |
IET Image Process. | 1 |
| 2024 | A Real-Time Emotion-Aware System Based on Wireless Body Area Network for IoMT ApplicationsabstractThe Internet of Medical Things (IoMT) stimulates the development of intelligent medical applications. As mental disorders become a global problem, emotion recognition has received widespread attention, as it can contribute to more comprehensive mental health monitoring and psychological assessment. Physiological signal-based emotion-aware monitoring is a particularly promising application due to its noninvasive and objective data collection. Recently, multimodal emotion recognition has been enhanced with wireless body area network (WBAN) access to IoMT, where wireless medical sensors are interconnected and abundant signals are acquired conveniently. However, how to synthesize these multisource physiological signals to facilitate emotion recognition is a challenging problem due to their heterogeneity and interference. To solve this problem, we propose a real-time differential multimodal transformer (Diff-MT), where the main components are the differential hyperinformation extraction (DHE) module, the multimodal global cross-attention encoder (MGCE), and the difference-augmented feature fusion (DFF). Ultimately, we endow the system with emotional awareness and distribute the state to IoMT devices. Extensive experiments demonstrate that the proposed Diff-MT exhibits superior performance compared to existing methods on the WESAD and DEAP datasets and is appropriate for IoMT-based healthcare. Yingchi Mao, Qian Huang 0008, Weiliang Xie, Xiaoming He 0004, Jie Wu 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Adaptive video stabilization based on feature point detection and full-reference stability assessment
Yiming Wang 0008, Qian Huang 0008, Jiwen Liu, Chuanxu Jiang, Mingzhou Shang |
Multim. Tools Appl. | 2 |
| 2024 | Scale-Aware Graph Convolutional Network With Part-Level Refinement for Skeleton-Based Human Action RecognitionabstractGraph Convolutional Networks (GCNs) have been widely used in skeleton-based human action recognition and have achieved promising results. However, current GCN-based methods are limited by their inability to refine semantic-guided joint relations and perform adaptive multi-scale analysis. These limitations impair their performance, particularly for analogical actions involving the interaction of the same body parts (e.g., drinking water and eating) as well as deficient actions with limited spatial-temporal information (e.g., subtle action writing and transient action sneezing). To solve these problems, we propose Part-level Refined Spatial Graph Convolution (PR-SGC) and Scale-aware Temporal Graph Convolution (Sa-TGC) for optimal action representation. The PR-SGC divides the skeleton into body parts and embeds this high-level semantics to refine the physical adjacency matrix. The Sa-TGC leverages the dynamic scale-aware mechanism to extract context-dependent multi-scale features. On this basis, we develop a novel Scale-aware Graph Convolutional Network with Part-level Refinement (SaPR-GCN), which is on par with state-of-the-art benchmarks on NTU RGB+D 60, NTU RGB+D 120, and NW-UCLA datasets. Yingchi Mao, Qian Huang 0008, Jie Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | FGC-VC: Flow-Guided Context Video CompressionabstractDeep video compression has attracted more and more attention in recent years. Previous works rely on feature space operations, which may cause the offset maps overflow degrading reconstructed frame quality. In this work, we propose a flow- guided module to guide the offset maps learning explicitly and alleviate offset maps overflow. Moreover, we introduce a context scheme to explore the temporal prior and fuse the hyper prior model to improve the compression ratio. For coding speed, we drop the time-consuming auto regressive module. Experimental results demonstrate that our method out-performs the previous learning-based schemes and traditional codecs. Compared to x265 with medium preset, our approach brings average 38.53% and 54.67% bit rate savings in PSNR and MS-SSIM metrics, respectively. Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Huashan Sun |
ICIP | 2 |
| 2023 | DD-GCN: Directed Diffusion Graph Convolutional Network for Skeleton-based Human Action RecognitionabstractGraph Convolutional Networks (GCNs) have been widely used in skeleton-based human action recognition. In GCN-based methods, the spatio-temporal graph is fundamental for capturing motion patterns. However, existing approaches ignore the physical dependency and synchronized spatio-temporal correlations between joints, which limits the representation capability of GCNs. To solve these problems, we construct the directed diffusion graph for action modeling and introduce the activity partition strategy to optimize the weight sharing mechanism of graph convolution kernels. In addition, we present the spatio-temporal synchronization encoder to embed synchronized spatio-temporal semantics. Finally, we propose Directed Diffusion Graph Convolutional Network (DD-GCN) for action recognition, and the experiments on three public datasets: NTU-RGB+D, NTU-RGB+D 120, and NW-UCLA, demonstrate the state-of-the-art performance of our method. Qian Huang 0008, Yingchi Mao |
ICME | 2 |
| 2023 | MA-Net: Multi-Attention Network for Skeleton-Based Action RecognitionabstractGraph Convolution Networks (GCNs) have become the main-stream framework for skeleton-based action recognition tasks. Aiming at the problem of redundant spatial-temporal feature information and neighborhood constraints obtained in GCNs, we propose a novel method called Multi-Attention Network (MA-Net) to explore crucial skeleton information, including two main modules: Combined Attention Graph Convolution (CAGC) and Multi-layer Transposed Attention Encoding (MTAE). The CAGC utilizes multi-dimensional combination attention to capture more valuable information and enhance feature performance. The MTAE adopts self-attention to encode feature maps, effectively establishing long-range dependency and capturing global information. Centre on the attention mechanism, these two modules combine the complementary advantages of GCN (i.e., local topology and temporal dynamics) and Transformer (i.e., global context and dynamic attention). Extensive experiments on the challenging NTU-RGB+D 60 and Kinetics-Skeleton datasets demonstrate that our model performs excellently. Jingwen Cui, Qian Huang 0008 |
MMAsia | 2 |
| 2023 | End-to-End Variable-Rate Image Compression with Bi-Resolution Spatial-Channel Context AggregationabstractRecently, neural network-based image compression techniques have demonstrated remarkable compression performance. The use of context-adaptive entropy models greatly enhances the rate-distortion (R-D) performance by effectively capturing spatial redundancy in latent representations. However, latent representations still contain some spatial correlations(e.g. same spatial structure), it needs to be eliminated by further processing. And many compression models are single-rate model, which is difficult to cover a big range of bitrate. In order to address this issue, we propose a novel variable-rate image compression algorithm that efficiently leverages bi-resolution spatial-channel information through learned mechanisms. In this paper, we first proposed a BRP network to divide our latent representations and side information into HR and LR components, eliminating the spatial redundancy in same location. Combining the spatial-channel context, we proposed a BSC context model, including a decreasing-granularity checkerboard pattern and channel grouping based on cosine slicing strategy. To cover a wide range of bitrate, we take a weight map as input to control bit allocation, achieving multiple compression rates. Our experimental results show that our method provides a better rate-distortion trade-off than BPG, JPEG and other recent image compression methods based on deep learning. Qian Huang 0008, Yiming Wang 0008, Huashan Sun |
MMAsia | 2 |
| 2023 | Optical Flow based Feature Prediction and Decomposed Context for Video CompressionabstractIn recent years, there have been a growing interest in developing end-to-end neural video codecs. Previous works generally use a past decoded frame as reference directly, utilizing the motion information between it and the input frame to reduce temporal redundancy. However, this approach may lead to high bit rate consumption of the motion and fails to take advantage of the prior information in other reconstructed frames. In this work, We propose a learned video coding framework with optical flow based feature prediction module and decomposed context module. Specifically, we employ the previous optical flow to generate a warped frame, and along with other reconstructions, they are used for a more accurate reference forecasting, thereby reducing the bit rate required for motion compression. Moreover, based on the conditional coding framework, our decomposed context module explores conditional context in past decoded frames and further reduces additional spatiotemporal correlations. Experimental results demonstrate that our approach yields better performance than previous learned video compression methods and traditional standard codecs. For example, our neural codec achieves 28.94% coding gain over HEVC in PSNR metric and about 2.00% coding gain over VVC in MS-SSIM metric. Huashan Sun, Qian Huang 0008, Yiming Wang 0008, Ruoyu Hao |
MMAsia | 2 |
| 2023 | Hierarchical Multi-Scale Adaptive Conv-LSTM Network for Human Action Recognition Based on Wearable SensorsabstractRecently, human action recognition has been widely used in the fields of health monitoring, human-robot interaction, medical treatment, and sports. Due to the availability of various wearable devices on the market, we can easily access sensor data for human action recognition. However, it is still a challenge to capture minute action processes as well as extract spatio-temporal motion patterns from serial sensor data. Therefore, we propose a novel hierarchical multi-scale adaptive Conv-LSTM network structure called HMA Conv-LSTM. The finer-grained spatial information in the sensor signals is extracted by hierarchical multi-scale convolution. The multi-channel feature fusion through adaptive channel feature fusion retains important information and improves model efficiency. We capture temporal context information by dynamic channel selection-LSTM based on the attention mechanism. Extensive experiments on the Opportunity and PAMAP2 public datasets show that our proposed model achieves competitive performance compared to several state-of-the-art approaches. Weiliang Xie, Qian Huang 0008, Yanfang Wang 0005, Yanwei Liu 0001 |
MMAsia | 2 |
| 2023 | Dense video captioning based on local attentionabstractAbstract Dense video captioning aims to locate multiple events in an untrimmed video and generate captions for each event. Previous methods experienced difficulties in establishing the multimodal feature relationship between frames and captions, resulting in low accuracy of the generated captions. To address this problem, a novel Dense Video Captioning Model Based on Local Attention (DVCL) is proposed. DVCL employs a 2D temporal differential CNN to extract video features, followed by feature encoding using a deformable transformer that establishes the global feature dependence of the input sequence. Then DIoU and TIoU are incorporated into the event proposal match algorithm and evaluation algorithm during training, to yield more accurate event proposals and hence increase the quality of the captions. Furthermore, an LSTM based on local attention is designed to generate captions, enabling each word in the captions to correspond to the relevant frame. Extensive experimental results demonstrate the effectiveness of DVCL. On the ActivityNet Captions dataset, DVCL performs significantly better than other baselines, with improvements of 5.6%, 8.2%, and 15.8% over the best baseline in BLEU4, METEOR, and CIDEr, respectively. Yong Qian, Yingchi Mao, Olano Teah Bloh, Qian Huang 0008 |
IET Image Process. | 6 |
| 2023 | Video stabilization: A comprehensive survey
Yiming Wang 0008, Qian Huang 0008, Chuanxu Jiang, Jiwen Liu, Mingzhou Shang, Zhuang Miao |
Neurocomputing | 2 |
| 2023 | Spatial and temporal information fusion for human action recognition via Center Boundary Balancing Multimodal Classifier
Xing Li 0005, Qian Huang 0008, Zhijian Wang 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Joint Intent Detection Model for Task-oriented Human-Computer Dialogue System using Asynchronous TrainingabstractHow to accurately understand low-resource languages is the core of the task-oriented human-computer dialogue system. Language understanding consists of two sub-tasks, i.e., intent detection and slot filling. Intent detection still faces challenges due to semantic ambiguity and implicit intentions with users’ input. Moreover, separately modeling intent detection and slot filling significantly decrease the correctness and relevance between questions and answers. To address these issues, we propose a joint intent detection method using asynchronous training strategy. The proposed method firstly encodes local text information extracted by CNN and relationship information among words emphasized by attention structure. Later, a joint intent detection model with asynchronous training strategy is proposed by either fusing hidden states of intent detection and slot filling layers, or adopting the key information to fine-tune the whole network, greatly increasing the relevance of intent detection and slot filling subtasks. The accuracy achieved by the proposed method tested on an open-source airline travel dataset and a self-collected electricity service dataset, i.e., ATIS and ECSF, are 97.49% and 89.68%, respectively, which proves the effectiveness of joint learning and asynchronous training. Yirui Wu, Hao Li 0089, Lilai Zhang, Qian Huang 0008, Shaohua Wan 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Real-Time 3-D Human Action Recognition Based on Hyperpoint SequenceabstractReal-time 3-D human action recognition has broad industrial applications, such as surveillance, human–computer interaction, and healthcare monitoring. By relying on complex spatio-temporal local encoding, most existing point cloud sequence networks capture spatio-temporal local structures to recognize 3-D human actions. To simplify the point cloud sequence modeling task, we propose a lightweight and effective point cloud sequence network referred to as SequentialPointNet for real-time 3-D action recognition. Instead of capturing spatio-temporal local structures, SequentialPointNet encodes the temporal evolution of static appearances to recognize human actions. First, we define a novel type of point data, hyperpoint, to better describe the temporally changing human appearances. A theoretical foundation is provided to clarify the information equivalence property for converting point cloud sequences into hyperpoint sequences. Second, the point cloud sequence modeling task is decomposed into a hyperpoint embedding task and a hyperpoint sequence modeling task. Specifically, for hyperpoint embedding, the static point cloud technology is employed to convert point cloud sequences into hyperpoint sequences, which introduces inherent frame-level parallelism; for hyperpoint sequence modeling, a hyperpoint-mixer module is designed as the basic building block to learning the spatio-temporal features of human actions. Extensive experiments on three widely-used 3-D action recognition datasets demonstrate that the proposed SequentialPointNet achieves a competitive classification performance with up to 10× faster than existing approaches. Xing Li 0005, Qian Huang 0008, Zhijian Wang 0002, Tianjin Yang, Zhenjie Hou, Zhuang Miao |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Hyperpointnet for Point Cloud Sequence-Based 3D Human Action RecognitionabstractPoint cloud sequence-based 3D action recognition achieves impressive performance and efficiency. Conventional approaches for modeling point cloud sequences usually perform cross-frame spatio-temporal local encoding, thus resulting in intensive computation and mutual interference between spatial and temporal information extracting. In this work, to avoid spatio-temporal local encoding, we propose a strong parallelized point cloud sequence network referred to as HyperpointNet for 3D action recognition. HyperpointNet is composed of two serial modules, i.e., a hyperpoint sequence embedding module and a hyperpoint sequence encoding module. In the hyperpoint sequence embedding module, employing static point cloud modeling methods, the point cloud sequence is abstracted into a new point data type named hyper-point sequence. In the hyperpoint sequence encoding module, a temporal PointNet (TPN) layer is designed to model the hyperpoint sequence. Extensive experiments conducted on two public datasets show that HyperpointNet outperforms state-of-the-art approaches. Xing Li 0005, Qian Huang 0008, Tianjin Yang, Qianhan Wu |
ICME | 2 |
| 2022 | Intelligent Video Surveillance Platform Based on FFmpeg and Yolov5abstractWith the development of multimedia, video surveillance systems are becoming more popular. However, the current video surveillance systems have a general function and are unable to provide Intelligent perception. Chuanxu Jiang, Yanfang Wang 0005, Qian Huang 0008, Yiming Wang 0008, Yuhan Dai |
MMAsia | 3 |
| 2022 | Dynamic-scale grid structure with weighted-scoring strategy for fast feature matching
Qian Huang 0008, Xing Li 0005 |
Appl. Intell. | 2 |
| 2022 | VirtualActionNet: A strong two-stream point cloud sequence network for human action recognition
Xing Li 0005, Qian Huang 0008, Zhijian Wang 0002, Tianjin Yang |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | A multi-scale human action recognition method based on Laplacian pyramid depth motion imagesabstractHuman action recognition is an active research area in computer vision. Aiming at the lack of spatial muti-scale information for human action recognition, we present a novel framework to recognize human actions from depth video sequences using multi-scale Laplacian pyramid depth motion images (LP-DMI). Each depth frame is projected onto three orthogonal Cartesian planes. Under three views, we generate depth motion images (DMI) and construct Laplacian pyramids as structured multi-scale feature maps which enhances multi-scale dynamic information of motions and reduces redundant static information in human bodies. We further extract the multi-granularity descriptor called LP-DMI-HOG to provide more discriminative features. Finally, we utilize extreme learning machine (ELM) for action classification. Through extensive experiments on the public MSRAction3D datasets, we prove that our method outperforms state-of-the-art benchmarks. Qian Huang 0008, Xing Li 0005, Qianhan Wu |
MMAsia | 2 |
| 2020 | A review of monocular visual odometry
Chaozheng Zhu, Qian Huang 0008, Baosen Ren |
Vis. Comput. | 3 |
| 2019 | Research on Index Mechanism of HBase Based on Coprocessor for Sensor DataabstractIn order to provide effective management of big data, almost two hundred different NoSQL stores have been developed, among which HBase is one of the best known. When performing data queries, the native HBase supports primary key indexes well. For non-primary key data, only the full table scan can be used, which greatly reduces the multi-condition query speed of HBase. Some additional indexing techniques have been presented to support querying on non-key indexing for HBase. For large-scale sensor data, this research field still requires effective design paradigm, in-depth experimentations and practical implementations. Therefore, we propose and implement a secondary memory index mechanism of HBase based on coprocessor for sensor data. The index mechanism can be automatically updated according to the change of the HBase table. Meanwhile, the index mechanism is persisted and maintained in memory, which can greatly improve the retrieval speed of the index data. By using real sensor datasets with different distribution patterns, experimental results show that the condition retrieval speed of the mechanism is greatly improved compared with the original HBase. In addition, compared with the secondary index mechanism based on Solr and Hibase, the performance of our proposed solution is also improved. Feng Ye 0004, Songjie Zhu, Yuansheng Lou, Yong Chen 0009, Qian Huang 0008 |
COMPSAC (1) | 6 |
| 2018 | The Research of a Lightweight Distributed Crawling SystemabstractNowadays, information on the Internet is growing at an explosive rate. The ability of the stand-alone web crawling system has come to its bottleneck, so more and more companies turn to distributed web crawling techniques. However, existing distributed web crawling systems have some shortcomings. Thread management modules for solving thread synchronization and resource competition are usually designed by using pure multithread asynchronous methods, but the execution of this kind of modules observably reduces the performance. Moreover, the deduplication algorithms lead to low efficiency in dealing with large data sets or the problem of occupying large storage space. To solve the problems mentioned above, this paper proposes a lightweight and practical distributed crawling system, which combines Docker and distributed computing techniques. It can make full use of the computing resources of the cluster and improve the efficiency of the crawling system effectively. Taking the data of Netease news page as an example, the experimental results show that the distributed crawler proposed has higher execution efficiency. Feng Ye 0004, Zongfei Jing, Qian Huang 0008, Yong Chen 0009 |
SERA | 3 |
| 2017 | Light-Field Depth Estimation via Epipolar Plane Image Analysis and Locally Linear EmbeddingabstractIn this paper, we propose a novel method for 4D light-field (LF) depth estimation exploiting the special linear structure of an epipolar plane image (EPI) and locally linear embedding (LLE). Without high computational complexity, depth maps are locally estimated by locating the optimal slope of each line segmentation on the EPIs, which are projected by the corresponding scene points. For each pixel to be processed, we build and then minimize the matching cost that aggregates the intensity pixel value, gradient pixel value, spatial consistency, as well as reliability measure to select the optimal slope from a predefined set of directions. Next, a subangle estimation method is proposed to further refine the obtained optimal slope of each pixel. Furthermore, based on a local reliability measure, all the pixels are classified into reliable and unreliable pixels. For the unreliable pixels, LLE is employed to propagate the missing pixels by the reliable pixels based on the assumption of manifold preserving property maintained by natural images. We demonstrate the effectiveness of our approach on a number of synthetic LF examples and real-world LF data sets, and show that our experimental results can achieve higher performance than the typical and recent state-of-the-art LF stereo matching methods. Yongbing Zhang 0002, Huijin Lv, Yebin Liu, Haoqian Wang, Xingzheng Wang, Qian Huang 0008, Xinguang Xiang, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | A multi-phase sparse probability framework via entropy minimization for single sample face recognitionabstractIn this paper, we propose a robust probability based sparse method to solve single sample face recognition, which harvests the advantages of both local and global representation. Different from previous sparse representation methods that generate sparse coefficients by l1, we produce sparse class probability distribution by proposing a multi-phase sparse probability (MSP) framework. To create class probability distribution, we divide each face image into many local blocks and vote based on the classification results of all blocks. For classifying each block, we propose local similarity assumption that makes many conventional methods feasible to SSPP problem. Moreover, we also propose a heuristic multiphase class selection scheme to solve the entropy minimization problem, which finally provides a higher classification confidence from the global perspective. Experimental results on three popular databases show that our approach not only generalizes well to SSPP problem but also has strong robustness to expression, illumination, occlusion and time variation. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Qian Huang 0008, Feng Xu 0008 |
ICIP | 4 |
| 2010 | A background model based method for transcoding surveillance videos captured by stationary cameraabstractReal-world video surveillance applications require storing videos without neglecting any part of scenarios for weeks or months. To reduce the storage cost, the high bit-rate videos from cameras should be transcoded into a more efficient compressed format with as little quality loss as possible. In this paper, we propose a background model based method to improve the transcoding efficiency for surveillance videos captured by stationary cameras, and objectively measure it. The background model is trained by pre-decoded I frames, and then used to transcode the source stream. Following this method, an H.264/AVC based transcoder employing the background model as long-term reference frame and a difference frame coding based transcoder are implemented and evaluated. Experimental results show that both trancoders save nearly half the used bits while maintaining quality compared with the full-decoding-full-encoding method, and the latter one has slightly better performance. Xianguo Zhang, Luhong Liang, Qian Huang 0008, Tiejun Huang 0001, Wen Gao 0001 |
PCS | 3 |
| 2010 | An efficient coding scheme for surveillance videos captured by stationary camerasabstractIn this paper, a new scheme is presented to improve the coding efficiency of sequences captured by stationary cameras (or namely, static cameras) for video surveillance applications. We introduce two novel kinds of frames (namely background frame and difference frame) for input frames to represent the foreground/background without object detection, tracking or segmentation. The background frame is built using a background modeling procedure and periodically updated while encoding. The difference frame is calculated using the input frame and the background frame. A sequence structure is proposed to generate high quality background frames and efficiently code difference frames without delay, and then surveillance videos can be easily compressed by encoding the background frames and difference frames in a traditional manner. In practice, the H.264/AVC encoder JM 16.0 is employed as a build-in coding module to encode those frames. Experimental results on eight in-door and out-door surveillance videos show that the proposed scheme achieves 0.12 dB~1.53 dB gain in PSNR over the JM 16.0 anchor specially configured for surveillance videos. Xianguo Zhang, Luhong Liang, Qian Huang 0008, Yazhou Liu, Tiejun Huang 0001, Wen Gao 0001 |
VCIP | 3 |
| 2010 | Deinterlacing Using Hierarchical Motion AnalysisabstractA motion-compensated deinterlacing scheme based on hierarchical motion analysis is presented. According to deinterlacing steps, our contribution can be divided into four parts: motion estimation, motion state analysis, motion consistency analysis, and finer-grained interpolation. In motion estimation, we introduce a Gaussian noise model for choosing the best motion vector for each block, and make a tradeoff between utilizing previous de-interlaced frames and avoiding error propagation. A directional interpolation method is also introduced in this part for backward fields. In motion state analysis, we define two motion states for each pixel, thus achieve a compromise between traditional block-based strategies and the extreme pixel-based case. In motion consistency analysis, we propose to measure both the motion vector consistency and the motion state consistency in order to determine whether the previous two parts should be performed again with a different block size. In finer-grained interpolation, we utilize a combination of recursive median filters to generate the final results. Experimental results show that all of the proposed techniques are effective, either objectively or subjectively. As a result, we can achieve much higher image quality, with an average gain of about 1.83 dB in terms of peak signal-to-noise ratio. Moreover, the increased computation complexity is marginal. Qian Huang 0008, Debin Zhao, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | An Efficient Coding Method for Intra Prediction Mode InformationabstractIn H.264/AVC, the bits used for the Intra_4times4 prediction mode information usually occupy a high percentage in intra coding. Towards this issue, we present an efficient coding method for the Intra_4times4 prediction mode information. Firstly, a 3-order Markov random field is introduced to model the correlation among neighboring 4times4 blocks at picture level. Secondly, based on the conditional probabilities learned in this model, we build up a context adaptive coding scheme to code the Intra_4times4 prediction mode information. Although the probabilities and the coding scheme are initialized off-line, they can be revised by automatic adjustments. Thus the proposed algorithm is robust to a variety of video sequences. Experimental results demonstrate that the proposed method can obtain a gain up to 0.3 dB in all I-frames coding without involving any serious computational burden. Kai Zhang 0007, Xiangyang Ji, Qian Huang 0008, Debin Zhao, Wen Gao 0001 |
ISCAS | 3 |
| 2006 | An Adaptive De-Interlacing Algorithm Based on Texture and Motion Vector AnalysisabstractIn this paper, we propose a novel hybrid de-interlacing algorithm, which effectively combines two motion-compensated (MC) de-interlacing techniques: MC median filtering (MCMF) and adaptive recursive (AR) with one spatial approach: line averaging (LA). Despite of its drawbacks, AR is one of the best methods nowadays. MCMF helps reduce flickers and LA is very robust to erroneous motion vectors. The interpolation switches among these methods based on the proposed measurement of texture smoothness and motion vector (MV) reliability. MCMF is adopted when MV is reliable and texture is rich. LA is used when MV is unreliable and texture is smooth. AR is applied to the remaining regions. Experimental results show that the proposed algorithm is superior to the compared algorithms in terms of peak signal-to-noise ratio (PSNR) and that the de-interlaced videos have very high subjective quality Jianguo Du 0001, Songnan Li, Debin Zhao, Qian Huang 0008, Wen Gao 0001 |
ICME | 4 |
| 2006 | An Edge-Based Median Filtering Algorithm with Consideration of Motion Vector Reliability for Adaptive Video DeinterlacingabstractBecause of its ability to preserve signal edges while filtering out impulsive noises, median filtering is widely used in signal processing applications, e.g. deinterlacing. An edge-based median filtering (EMF) algorithm is proposed for adaptive deinterlacing. A criterion of motion vector reliability (MVR) is also introduced for better interpolation. For each motion compensated block, two motion vectors between opposite-parity fields and one between same-parity fields are taken into account. Experiments show that the proposed MVR and EMF are both very efficient. Outputs of the proposed EMF are much more similar to original progressive videos than those of objectively best EMF methods nowadays, without obvious visual distortions. Finally, the proposed EMF and MVR criterion are shown to be suitable for texture-based adaptive deinterlacing Qian Huang 0008, Wen Gao 0001, Debin Zhao, Qingming Huang |
ICME | 1 |