Yongfei Zhang

dblp:16/8496 · DBLP profile ↗
← Back
50ranked-venue papers
19as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 14 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Joint embedding for multi-structural hypergraph based dimensionality reduction in rotor fault diagnosis
Yongfei Zhang, Qibo Liang, Yuqiao Zheng, Rongzhen Zhao, Linfeng Deng, Mingkuan Shi, Kongyuan Wei
Eng. Appl. Artif. Intell.1
2026 Revealing Social Users' Values in Reactions: Modeling User Values in Social Networks
Lianghuan Zhao, Yongfei Zhang, Xiaokang Wang 0001, Linlin Ma
IEEE Trans. Comput. Soc. Syst.4
2025 Multimetric hypergraph embedding for dimensionality reduction in rotor fault diagnosis
Yongfei Zhang, Yuqiao Zheng, Rongzhen Zhao, Linfeng Deng, Mingkuan Shi, Kongyuan Wei
Adv. Eng. Informatics1
2025 Multi-source contrastive cluster center method for cross-domain bearing fault identification
Lizhen Wu, Rongzhen Zhao, Kongyuan Wei, Yuqiao Zheng, Linfeng Deng, Yongfei Zhang, Mingkuan Shi
Eng. Appl. Artif. Intell.7
2025 Relational Clustering-Based Parallel Spaces Construction and Embedding for Dynamic Knowledge Graph
abstract
With the increasing amount of data in various domains, knowledge graphs (KGs) have become powerful tools for representing complex and heterogeneous information in a structured way, and for extracting valuable information from knowledge graphs through embedding techniques to support downstream tasks such as recommendation and Q&A systems. Knowledge graphs consist of triples that are continuously added as knowledge is updated. However, most existing embedding models are designed for static graphs, requiring the entire model to be retrained for each update, which is time-consuming. Existing global dynamic embedding models focus on exploiting the structural and relational information of the whole graph to achieve embedding quality, resulting in reduced dynamic efficiency. To address this problem, we propose a relational clustering-based parallel space model in which knowledge from different domains is embedded in different subspaces, allowing each subspace to focus on the data characteristics of a specific domain, thereby improving the quality of knowledge. Second, the new data only affects some subspaces but not the performance of other spaces, improving the model's adaptability to dynamics. Furthermore, we employ two incremental approaches based on the type of added data to improve the efficiency of dynamic embedding while ensuring that the added data preserves the characteristics of the parallel space. The experimental results show that the dynamic embedding efficiency of our model is improved by an average of 50.3% compared to the SOTA dynamic embedding model for the link prediction task. Particularly on FB15K, our model not only improves the efficiency by 41% but also increases the accuracy by 7.5%, demonstrating the accuracy and efficiency of our model.
Yongfei Zhang
IEEE Trans. Big Data2
2025 CRF Based Pedestrian Trajectory Reconstruction With Spatio-Temporal Feature Embedding and Entropy Constraints
Peiyue Li, Yongfei Zhang
IEEE Trans. Intell. Transp. Syst.2
2024 Heuristic-Driven, Type-Specific Embedding in Parallel Spaces for Enhancing Knowledge Graph Reasoning
abstract
Knowledge Graph Reasoning aims to derive new insights from existing Knowledge Graphs (KGs) and address any missing or incomplete data. Existing models primarily rely on explicit information while neglecting the implicit constraints imposed by entity types on relations types. For example, when the entity type is "person-person," the relation type should be constrained to interpersonal connections like "co-worker." Based on this perspective, we introduce a priori knowledge-based approach for inferring relations types. This approach utilizes the relations type distribution across different entity types in the dataset to guide the inference process. Additionally, recognizing that mapping all different relations types to a single space can decrease inference accuracy due to the diversity of semantics, we propose a parallel spaces KG embedding model that partitions the entire KG into multiple subspaces. Each subspace is dedicated to learning information associated with a specific relation type. Experimental results on three KG reasoning benchmarks demonstrate that our model outperforms other baselines in accuracy. Importantly, our model shows significant advantages when applied to datasets with a substantial number of relations.
Yongfei Zhang, Shan Yang 0003
ICASSP2
2024 SAGS-DynamicBio: Integrating Semantic-Aware and Graph Structure-Aware Embedding for Dynamic Biological Data with Knowledge Graphs
Yongfei Zhang
ECML/PKDD (9)2
2024 A Multi-Attention Feature Distillation Neural Network for Lightweight Single Image Super-Resolution
abstract
In recent years, remarkable performance improvements have been produced by deep convolutional neural networks (CNN) for single image super-resolution (SISR). Nevertheless, a high proportion of CNN-based SISR models are with quite a few network parameters and high computational complexity for deep or wide architectures. How to more fully utilize deep features to make a balance between model complexity and reconstruction performance is one of the main challenges in this field. To address this problem, on the basis of the well-known information multi-distillation model, a multi-attention feature distillation network termed as MAFDN is developed for lightweight and accurate SISR. Specifically, an effective multi-attention feature distillation block (MAFDB) is designed and used as the basic feature extraction unit in MAFDN. With the help of multi-attention layers including pixel attention, spatial attention, and channel attention, MAFDB uses multiple information distillation branches to learn more discriminative and representative features. Furthermore, MAFDB introduces the depthwise over-parameterized convolutional layer (DO-Conv)-based residual block (OPCRB) to enhance its ability without incurring any parameter and computation increase in the inference stage. The results on commonly used datasets demonstrate that our MAFDN outperforms existing representative lightweight SISR models when taking both reconstruction performance and model complexity into consideration. For example, for × 4 SR on Set5, MAFDN (597K/33.79G) obtains 0.21 dB/0.0037 and 0.10 dB/0.0015 PSNR/SSIM gains over the attention-based SR model AFAN (692K/50.90G) and the feature distillation-based SR model DDistill-SR (675K/32.83G), respectively.
Yongfei Zhang, Xinying Lin, Linbo Qing, Xiaohai He, Yi Li 0069, Honggang Chen
Int. J. Intell. Syst.1
2023 PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-Identification
abstract
Although recent studies empirically show that injecting Convolutional Neural Networks (CNNs) into Vision Transformers (ViTs) can improve the performance of person reidentification, the rationale behind it remains elusive. From a frequency perspective, we reveal that ViTs perform worse than CNNs in preserving key high-frequency components (e.g, clothes texture details) since high-frequency components are inevitably diluted by low-frequency ones due to the intrinsic Self-Attention within ViTs. To remedy such inadequacy of the ViT, we propose a Patch-wise High-frequency Augmentation (PHA) method with two core designs. First, to enhance the feature representation ability of high-frequency components, we split patches with high-frequency components by the Discrete Haar Wavelet Transform, then empower the ViT to take the split patches as auxiliary input. Second, to prevent high-frequency components from being diluted by low-frequency ones when taking the entire sequence as input during network optimization, we propose a novel patch-wise contrastive loss. From the view of gradient optimization, it acts as an implicit augmentation to improve the representation ability of key high-frequency components. This benefits the ViT to capture key high-frequency components to extract discriminative person representations. PHA is necessary during training and can be removed during inference, without bringing extra complexity. Extensive experiments on widely-used ReID datasets validate the effectiveness of our method.
Guiwei Zhang, Yongfei Zhang, Bo Li 0006, Shiliang Pu
CVPR2
2023 ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification
abstract
Visible-Infrared person re-identification is challenging due to the large modality gap. To bridge the gap, most studies heavily rely on the correlation of visible-infrared holistic person images, which may perform poorly under severe distribution shifts. In contrast, we find that some cross-modal correlated high-frequency components contain discriminative visual patterns and are less affected by variations such as wavelength, pose, and background clutter than holistic images. Therefore, we are motivated to bridge the modality gap based on such high-frequency components, and propose Prototype-guided High-frequency Patch Enhancement (ProtoHPE) with two core designs. First, to enhance the representation ability of cross-modal correlated high-frequency components, we split patches with such components by Wavelet Transform and exponential moving average Vision Transformer (ViT), then empower ViT to take the split patches as auxiliary input. Second, to obtain semantically compact and discriminative high-frequency representations of the same identity, we propose Multimodal Prototypical Contrast. To be specific, it hierarchically captures comprehensive semantics of different modal instances, facilitating the aggregation of high-frequency representations belonging to the same identity. With it, ViT can capture key high-frequency components during inference without relying on ProtoHPE, thus bringing no extra complexity. Extensive experiments validate the effectiveness of ProtoHPE.
Guiwei Zhang, Yongfei Zhang, Zichang Tan
ACM Multimedia2
2022 CAKE: A Scalable Commonsense-Aware Framework For Multi-View Knowledge Graph Completion
abstract
Knowledge graphs store a large number of factual triples while they are still incomplete, inevitably.The previous knowledge graph completion (KGC) models predict missing links between entities merely relying on fact-view data, ignoring the valuable commonsense knowledge.The previous knowledge graph embedding (KGE) techniques suffer from invalid negative sampling and the uncertainty of fact-view link prediction, limiting KGC's performance.To address the above challenges, we propose a novel and scalable Commonsense-Aware Knowledge Embedding (CAKE) framework to automatically extract commonsense from factual triples with entity concepts.The generated commonsense augments effective selfsupervision to facilitate both high-quality negative sampling (NS) and joint commonsense and fact-view link prediction.Experimental results 1 on the KGC task demonstrate that assembling our framework could enhance the performance of the original KGE models, and the proposed commonsense-aware NS module is superior to other NS techniques.Besides, our proposed framework could be easily adaptive to various KGE models and explain the predicted results.
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Shiliang Pu
ACL (1)3
2022 Perform like an Engine: A Closed-Loop Neural-Symbolic Learning Framework for Knowledge Graph Inference
abstract
Knowledge graph (KG) inference aims to address the natural incompleteness of KGs, including rule learning-based and KG embedding (KGE) models. However, the rule learning-based models suffer from low efficiency and generalization while KGE models lack interpretability. To address these challenges, we propose a novel and effective closed-loop neural-symbolic learning framework EngineKG via incorporating our developed KGE and rule learning modules. KGE module exploits symbolic rules and paths to enhance the semantic association between entities and relations for improving KG embeddings and interpretability. A novel rule pruning mechanism is proposed in the rule learning module by leveraging paths as initial candidate rules and employing KG embeddings together with concepts for extracting more high-quality rules. Experimental results on four real-world datasets show that our model outperforms the relevant baselines on link prediction tasks, demonstrating the superiority of our KG inference model in a neural-symbolic learning fashion. The source code and datasets of this paper are available at https://github.com/ngl567/EngineKG.
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Shiliang Pu
COLING3
2022 Joint semantics and data-driven path representation for knowledge graph reasoning
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Yongpan Sheng, Chuan Shi 0001, Shiliang Pu
Neurocomputing3
2022 Weakly-supervised contrastive learning-based implicit degradation modeling for blind image super-resolution
Yongfei Zhang, Ling Dong, Linbo Qing, Xiaohai He, Honggang Chen
Knowl. Based Syst.1
2021 UnrealPerson: An Adaptive Pipeline Towards Costless Person Re-Identification
abstract
The main difficulty of person re-identification (ReID) lies in collecting annotated data and transferring the model across different domains. This paper presents UnrealPerson, a novel pipeline that makes full use of unreal image data to decrease the costs in both the training and deployment stages. Its fundamental part is a system that can generate synthesized images of high-quality and from controllable distributions. Instance-level annotation goes with the synthesized data and is almost free. We point out some details in image synthesis that largely impact the data quality. With 3,000 IDs and 120,000 instances, our method achieves a 38.5% rank-1 accuracy when being directly transferred to MSMT17. It almost doubles the former record using synthesized data and even surpasses previous direct transfer records using real data. This offers a good basis for unsupervised domain adaption, where our pre-trained model is easily plugged into the state-of-the-art algorithms towards higher accuracy. In addition, the data distribution can be flexibly adjusted to fit some corner ReID scenarios, which widens the application of our pipeline. We publish our data synthesis toolkit and synthesized data in https://github.com/FlyHighest/UnrealPerson.
Lingxi Xie, Longhui Wei, Zijie Zhuang, Yongfei Zhang, Bo Li 0006, Qi Tian 0001
CVPR5
2020 Rule-Guided Compositional Representation Learning on Knowledge Graphs
abstract
Representation learning on a knowledge graph (KG) is to embed entities and relations of a KG into low-dimensional continuous vector spaces. Early KG embedding methods only pay attention to structured information encoded in triples, which would cause limited performance due to the structure sparseness of KGs. Some recent attempts consider paths information to expand the structure of KGs but lack explainability in the process of obtaining the path representations. In this paper, we propose a novel Rule and Path-based Joint Embedding (RPJE) scheme, which takes full advantage of the explainability and accuracy of logic rules, the generalization of KG embedding as well as the supplementary semantic structure of paths. Specifically, logic rules of different lengths (the number of relations in rule body) in the form of Horn clauses are first mined from the KG and elaborately encoded for representation learning. Then, the rules of length 2 are applied to compose paths accurately while the rules of length 1 are explicitly employed to create semantic associations among relations and constrain relation embeddings. Moreover, the confidence level of each rule is also considered in optimization to guarantee the availability of applying the rule to representation learning. Extensive experimental results illustrate that RPJE outperforms other state-of-the-art baselines on KG completion task, which also demonstrate the superiority of utilizing logic rules as well as paths for improving the accuracy and explainability of representation learning.
Guanglin Niu, Yongfei Zhang, Bo Li 0006, Peng Cui 0001, Si Liu 0001, Xiaowei Zhang 0003
AAAI2
2020 Single Camera Training for Person Re-Identification
abstract
Person re-identification (ReID) aims at finding the same person in different cameras. Training such systems usually requires a large amount of cross-camera pedestrians to be annotated from surveillance videos, which is labor-consuming especially when the number of cameras is large. Differently, this paper investigates ReID in an unexplored single-camera-training (SCT) setting, where each person in the training set appears in only one camera. To the best of our knowledge, this setting was never studied before. SCT enjoys the advantage of low-cost data collection and annotation, and thus eases ReID systems to be trained in a brand new environment. However, it raises major challenges due to the lack of cross-camera person occurrences, which conventional approaches heavily rely on to extract discriminative features. The key to dealing with the challenges in the SCT setting lies in designing an effective mechanism to complement cross-camera annotation. We start with a regular deep network for feature extraction, upon which we propose a novel loss function named multi-camera negative loss (MCNL). This is a metric learning loss motivated by probability, suggesting that in a multi-camera system, one image is more likely to be closer to the most similar negative sample in other cameras than to the most similar negative sample in the same camera. In experiments, MCNL significantly boosts ReID accuracy in the SCT setting, which paves the way of fast deployment of ReID systems with good performance on new target scenes.
Lingxi Xie, Longhui Wei, Yongfei Zhang, Bo Li 0006, Qi Tian 0001
AAAI4
2020 Hierarchical Deep Hashing for Fast Large Scale Image Retrieval
abstract
Fast image retrieval is of great importance in many computer vision tasks and especially practical applications. Deep hashing, the state-of-the-art fast image retrieval scheme, introduces deep learning to learn the hash functions and generate binary hash codes, and outperforms the other image retrieval methods in terms of accuracy. However, all the existing deep hashing methods could only generate one level hash codes and require a linear traversal of all the hash codes to figure out the closest one when a new query arrives, which is very time-consuming and even intractable for large scale applications. In this work, we propose a Hierarchical Deep Hashing(HDHash) scheme to speed up the state-of-the-art deep hashing methods. More specifically, hierarchical deep hash codes of multiple levels can be generated and indexed with tree structures rather than linear ones, and pruning irrelevant branches can sharply decrease the retrieval time. To our best knowledge, this is the first work to introduce hierarchical indexed deep hashing for fast large scale image retrieval. Extensive experimental results on three benchmark datasets demonstrate that the proposed HDHash scheme achieves better or comparable accuracy with significantly improved efficiency and reduced memory as compared to state-of-the-art fast image retrieval schemes.
Yongfei Zhang, Xianglong Liu 0001, Shiliang Pu, Changhuai Chen
ICPR1
2020 Background Segmentation for Vehicle Re-identification
Yongfei Zhang
MMM (2)2
2020 Recent Advances on HEVC Inter-Frame Coding: From Optimization to Implementation and Beyond
abstract
High Efficiency Video Coding (HEVC) has doubled the video compression ratio with equivalent subjective quality as compared to its predecessor H.264/AVC. The significant coding efficiency improvement is attributed to many new techniques. Inter-frame coding is one of the most powerful yet complicated techniques therein and has posed high computational burden thus main obstacle in HEVC-based real-time applications. Recently, plenty of research has been done to optimize the inter-frame coding, either to reduce the complexity for real-time applications, or to further enhance the encoding efficiency. In this paper, we provide a comprehensive review of the state-of-the-art techniques for HEVC inter-frame coding from three aspects, namely fast inter coding solutions, implementation on different hardware platforms as well as advanced inter coding techniques. More specifically, different algorithms in each aspect are further subdivided into sub-categories and compared in terms of pros, cons, coding efficiency and coding complexity. To the best of our knowledge, this is the first such comprehensive review of the recent advances of the inter-frame coding for HEVC and hopefully it would help the improvement, implementation and applications of HEVC as well as the ongoing development of the next generation video coding standard.
Yongfei Zhang, Rui Fan 0002, Siwei Ma 0001, Zhibo Chen 0001, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.1
2019 Texture-Classification Accelerated CNN Scheme for Fast Intra CU Partition in HEVC
abstract
High Efficiency Video Coding (HEVC) achieves significant coding performance over H.264. However, the performance gain is achieved at the cost of substantially higher encoding complexity, in which the coding tree unit (CTU) partition is one of the most time-consuming parts due to the rate-distortion optimization-based ergodic search of all possible quad-tree partitions. To address this problem, this paper proposes a texture-classification accelerated convolutional neural network (CNN)-based fast intra CU partition scheme to reduce the encoding complexity for intra-coding in HEVC, by taking into consideration of the heterogeneous texture characteristics into the CNN-based classification. First, a threshold-based texture classification model is developed to identify the heterogeneous and homogeneous CTUs, through jointly consideration of the CU depth, quantization parameter and texture complexity. Second, three different CNN structures are designed and trained to predict the CU partition mode for each CU layer in the heterogeneous CTUs. Finally, extensive experimental results show that the proposed scheme can reduce intra-mode encoding time by 62.13% with negligible BD-rate loss of 2.01%, consistently outperforming two state-of-the-art CNN-based schemes in terms of both coding performance and complexity reduction.
Yongfei Zhang, Gang Wang 0023, Mai Xu, C.-C. Jay Kuo
DCC1
2019 Adaptive intra mode decision for HEVC based on texture characteristics and multiple reference lines
Yongfei Zhang, Miyi Duan
Multim. Tools Appl.2
2019 Motion estimation using maximum sub-image and sub-pixel phase correlation on a DSP platform
Bo Zhai, Yue Wang 0029, Peisong Guo, Yongfei Zhang
Multim. Tools Appl.5
2019 Highly Paralleled Low-Cost Embedded HEVC Video Encoder on TI KeyStone Multicore DSP
abstract
Although HEVC, the emerging video coding standard, has doubled the coding performance of its predecessor H.264/AVC, its significantly increased computational complexity imposes great obstacles for HEVC encoders to be employed in real-time applications with embedded processors, such as digital signal processors (DSPs). In this paper, a TI Keystone multicore TMS320C6678 DSP-based highly paralleled low-cost fast HEVC encoding solution is well designed and implemented. First, the overall structure of HEVC encoder with CTU-level parallelism is re-designed to well support the encoding parallelism, with full consideration of the hardware characteristics. Second, a low-delay and low-memory multicore data transmission mechanism is proposed to reduce the latency of data access between internal L2 memory and external DDR3. Third, the encoding bottlenecks, i.e., the most time-consuming encoding modules, are identified and optimized for acceleration with TI powerful C6000 SIMD instructions. Experimental results show that our proposed HEVC encoder on TI TMS320C6678 DSPs can significantly improve the real-time capacity with tolerable performance loss, 0.93 dB performance loss under on average 465.50 times speedup as compared to CPU-based HM reference software, more specifically, which makes it desirable in power-constrained real-time video applications.
Rui Fan 0002, Yongfei Zhang, Gang Wang 0023, Zhe Li 0015
IEEE Trans. Circuits Syst. Video Technol.3
2018 ReTestDroid: Towards Safer Regression Test Selection for Android Application
abstract
Mobile applications are widely used in our daily life and Android is the most popular open source mobile operating system. Because mobile applications update frequently, it is important developers to perform regression testing to ensure their quality. Modeling the control flow of an android application based on the activity lifecycle model only is imprecise for regression testing. Because many Android applications use asynchronous tasks, fragments, and native code frequently, which must be considered during change impact analysis. Otherwise, regression test selection techniques may miss some failure-revealing test cases, compromising the safety of these techniques. In this work, we propose a novel approach to model asynchronous task invocations, fragment-based activity lifecycle, and native code within the control flow graph of an Android application. Furthermore, we designed a regression test selection tool ReTestDroid based on our graph model. Our experiments on five real-life Android applications showed that our approach could enable much safer regression test selection while significantly saving regression-testing time.
Bo Jiang 0001, Yongfei Zhang, Zhenyu Zhang 0004, Wing Kwong Chan
COMPSAC (1)3
2018 Multi-hypothesis-Based Error Concealment for Whole Frame Loss in HEVC
Yongfei Zhang, Zhe Li 0015
MMM (1)1
2018 A fast and HEVC-compatible perceptual video coding scheme using a transform-domain Multi-Channel JND model
Gang Wang 0023, Yongfei Zhang, Bo Li 0006, Rui Fan 0002, Mingliang Zhou 0001
Multim. Tools Appl.2
2018 Background Modeling and Referencing for Moving Cameras-Captured Surveillance Video Coding in HEVC
abstract
Surveillance video coding is crucial for improving compression efficiency in intelligent video surveillance systems and applications. Plenty of work has been done, which can be roughly divided into two categories: the former mainly focuses on low-complexity background modeling to obtain the clear background, while the latter focuses on an appropriate coding strategy to generate the high-quality background reference picture for effective background prediction. However, almost all existing works focus only on stationary camera scenes, while moving cameras-captured surveillance video coding is left untouched and is still an open problem. In this paper, a background modeling and referencing scheme for moving cameras-captured surveillance video coding in high-efficiency video coding (HEVC) is proposed. First, this paper proposes a low-complexity motion background modeling algorithm for surveillance video coding using the running average based on a global-motion-compensation method. To obtain the global motion vector, we propose a global motion detection method based on character blocks by establishing a low-rank singular value decomposition model for clustering and estimating motion vectors of background character blocks in the cameras movement circumstance. Second, we propose a background referencing coding strategy, in which the motion background coding tree units (MBCTUs) would be selected by anchoring the input video frame on the modeling background frame and coded with the optimized quantization parameter. Then, the reconstructed MBCTU will be used to update the previous coding tree unit in the global compensation location of the background reference picture. Extensive experimental results show that the proposed scheme can achieve significant bit savings of up to 26.6% and, on average, 6.7% with similar subjective quality and negligible encoding complexity, compared to HM12.0. Besides, the proposed scheme consistently outperforms two state-of-the-art surveillance video coding schemes with remarkable bitrate savings.
Gang Wang 0023, Bo Li 0006, Yongfei Zhang, Jinhui Yang
IEEE Trans. Multim.3
2017 Complexity-based intra frame rate control by jointing inter-frame correlation for high efficiency video coding
Mingliang Zhou 0001, Yongfei Zhang, Bo Li 0006, Hai-Miao Hu
J. Vis. Commun. Image Represent.2
2017 Multidirectional parabolic prediction-based interpolation-free sub-pixel motion estimation
Rui Fan 0002, Yongfei Zhang, Bo Li 0006, Gang Wang 0023
Signal Process. Image Commun.2
2017 Motion Classification-Based Fast Motion Estimation for High-Efficiency Video Coding
abstract
High efficiency video coding (HEVC), the latest video coding standard, is becoming popular due to its excellent coding performance. However, the significant gain in performance is achieved at the cost of substantially higher encoding complexity than its precedent H.264/AVC, in which motion estimation (ME) is the most time-consuming module that effectively removes temporal redundancy. Test zone search (TZS) is adopted as the default fast ME method in the reference software of HEVC; however, its computational complexity is still too high for real-time applications. Several fast ME algorithms have been recently proposed to further reduce ME complexity; however, these approaches typically lead to non-negligible performance loss. To address this problem, this paper proposes a motion classification-based fast ME algorithm. By exploring the motion relationship of neighboring blocks and the coding cost characteristic, the prediction unit (PU) is first categorized into one of three classes, namely, motion-smooth PU, motion-medium PU and motion-complex PU. Then different search strategies are carefully designed for PUs of each class according to their respective motion and content characteristics. Furthermore, a fast search priority-based partial internal termination scheme is presented to rapidly skip impossible positions that speeds up cost computation during the ME process. Extensive experimental results demonstrate that the proposed algorithm achieves as much as 12.47% and 20.25% reductions in total encoder complexity when compared with TZS under low delay P and random access configuration, respectively, with negligible rate-distortion degradation; thus, it outperforms state-of-the-art fast ME algorithms in terms of both coding performance and complexity reduction.
Rui Fan 0002, Yongfei Zhang, Bo Li 0006
IEEE Trans. Multim.2
2017 Complexity Correlation-Based CTU-Level Rate Control with Direction Selection for HEVC
abstract
Rate control is a crucial consideration in high-efficiency video coding (HEVC). The estimation of model parameters is very important for coding tree unit (CTU)-level rate control, as it will significantly affect bit allocation and thus coding performance. However, the model parameters in the CTU-level rate control sometimes fails because of inadequate consideration of the correlation between model parameters and complexity characteristic. In this study, we establish a novel complexity correlation-based CTU-level rate control for HEVC. First, we formulate the model parameter estimation scheme as a multivariable estimation problem; second, based on the complexity correlation of the neighbouring CTU, an optimal direction is selected in five directions for reference CTU set selection during model parameter estimation to further improve the prediction accuracy of the complexity of the current CTU. Third, to improve their precision, the relationship between the model parameters and the complexity of the reference CTU set in the optimal direction is established by using least square method (LS), and the model parameters are solved via the estimated complexity of the current CTU. Experimental results show that the proposed algorithm can significantly improve the accuracy of the CTU-level rate control and thus the coding performance; the proposed scheme consistently outperforms HM 16.0 and other state-of-the-art algorithms in a variety of testing configurations. More specifically, up to 8.4% and on average 6.4% BD-Rate reduction is achieved compared to HM 16.0 and up to 4.7% and an average of 3.4% BD-Rate reduction is achieved compared to other algorithms, with only a slight complexity overhead.
Mingliang Zhou 0001, Yongfei Zhang, Bo Li 0006, Xupeng Lin
ACM Trans. Multim. Comput. Commun. Appl.2
2016 Memory-efficient high-speed VLSI implementation of multi-level discrete wavelet transform
Yongfei Zhang, Haiheng Cao, Bo Li 0006
J. Vis. Commun. Image Represent.1
2016 Content-adaptive parameters estimation for multi-dimensional rate control
Mingliang Zhou 0001, Bo Li 0006, Yongfei Zhang
J. Vis. Commun. Image Represent.3
2015 Fast rate distortion optimized quantization for HEVC
abstract
As the new generation video coding standard, the High Efficiency Video Coding (HEVC) achieves significantly better coding efficiency than existing video coding standards, which is however at the cost of a much higher computational complexity. Rate-Distortion Optimized Quantization (RDOQ) is one of the powerful coding tools employed in HEVC to pursue high coding efficiency. Nevertheless, RDOQ requires an exhaustive search over multiple candidates to determine the optimal quantized level through Rate-Distortion Optimization (RDO), which leads to considerable complexity in practice. This paper addresses this issue from a new point of view and presents a fast RDOQ scheme for transformed DCT coefficients quantization in HEVC. More specifically, we first proposed a RDOQ call removal algorithm which can effectively remove the redundant RDOQ calls and thus speed up the RDOQ process. Secondly, an all-zero block-based RDOQ skip algorithm is developed to skip unnecessary RDOQ. Compared to HEVC reference software, the proposed algorithms can save on average 34% of the RDOQ time with only slight performance degradation. Besides, our proposed algorithms can be incorporated with existing fast RDOQ algorithms to further improve the encoding efficiency.
Yongfei Zhang
VCIP1
2014 A high-throughput MQ coder architecture based on dependence extraction method
abstract
MQ is an efficient entropy coder that performs the actual compression in JPEG2000. However, it usually acts as the bottleneck of the hardware architecture due to the feedback loops caused by iterative operations. The current single ejection (SE) architecture achieves higher frequency by adopting more pipeline stages, but the speed is limited by the throughput per cycle. On the other hand, the multiple ejections (ME) usually handles more than one context data (CxD) per cycle, while the frequency is deteriorated due to the longer critical circuit caused by the context dependences. Hence, to enable the MQ arithmetic coder to process more than one sample while running at higher clock frequency, this paper proposes a two-CxDs architecture based on the equality of two adjacent CxDs, in which the two adjacent CxDs with different contexts are processed in a clock. Experiment results illustrate the architecture increases above 30% speed performance.
Haiheng Cao, Yongfei Zhang
ICIP2
2014 An Improved Similarity-Based Fast Coding Unit Depth Decision Algorithm for Inter-frame Coding in HEVC
Rui Fan 0002, Yongfei Zhang, Zhe Li 0015
MMM (1)2
2014 Visual Distortion Sensitivity Modeling for Spatially Adaptive Quantization in Remote Sensing Image Compression
abstract
As remote sensing images are often characterized with strong randomness, weak local correlation, and multiple small targets, the commonly used coarse-granularity subband-level quantization scheme fails to make use of these characteristics; thus, the performance improvements of these methods in literature are often marginal. To address this problem, this letter presents a novel spatially adaptive quantization (SAQ) method for the compression of remote sensing images based on our proposed Visual Distortion Sensitivity (ViDiS) Model. The ViDiS model takes into consideration four ViDiS components, including image luminance, spatial frequency, spatial orientation, and visual masking, to help measure the distortion more consistent to the image quality perceived by human beings. Then, a SAQ scheme is proposed to better exploit the content characteristics of remote sensing images, in which the quantization is conducted on a finer subband block level rather than subband level, with the guidance of the ViDiS model. Experimental results show that the proposed algorithm can preserve better visual quality in low-contrast areas with small targets at a competitive computational cost, which makes it more desirable in compression applications for remote sensing images.
Yongfei Zhang, Haiheng Cao, Bo Li 0006
IEEE Geosci. Remote. Sens. Lett.1
2014 Delay-Bounded Priority-Driven Resource Allocation for Video Transmission Over Multihop Networks
abstract
In this paper we consider the problem of resource allocation for video transmission over mesh networks with delay bound constraints and priority-based packet scheduling. We observe that priority-driven packet scheduling at the intermediate network routers has a direct and significant impact on the queuing behaviors and delay bound violation probabilities of video packets, as well as the overall end-to-end video distortion. Using learning methods, we develop a packet delay bound violation probability model for video transmission over multihop networks with priority-based packet scheduling. With this model, we can successfully predict the probability of packets being dropped due to violation of specified delay bounds. We also observe that the transmission distortion caused by packet drops exhibits a unique exponential behavior with priority-based packet scheduling. With these analysis results, we formulate the resource allocation for multisession video transmission over networks with priority-driven packet scheduling under delay bound constraints as a multiobjective optimization problem. Evolutionary optimization methods based on single- and multiobjective genetic algorithms are proposed to solve the problem and obtain the optimal resource allocation. Extensive experiment results demonstrate the effectiveness of the proposed resource-distortion models and optimization algorithms.
Yongfei Zhang, Shiyin Qin, Bo Li 0006, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.1
2013 Fast Coding Unit Depth Decision Algorithm for Interframe Coding in HEVC
abstract
As the next generation standard of video coding, the High Efficiency Video Coding (HEVC) achieves significantly better coding efficiency than all existing video coding standards. A Coding Unit (CU) quad tree concept is introduced to HEVC to improve the coding efficiency. Each CU node in quad tree will be traversed by depth first search process to find the best Coding Tree Unit (CTU) partition. Although this quad tree search process can obtain the best CTU partition, it is very time consuming, especially in interframe coding. To alleviate the encoder computation load in interframe coding, a fast CU depth decision method is proposed by reducing the depth search range. Based on the depth information correlation between spatio-temporal adjacent CTUs and the current CTU, some depths can be adaptively excluded from the depth search process in advance. Experimental results show that the proposed scheme provides almost 30% encoder time savings on average compared to the default encoding scheme in HM8.0 with only 0.38% bit rate increment in coding performance.
Yongfei Zhang, Zhe Li 0015
DCC1
2013 Region-classification-based rate control for flicker suppression of I-frames in HEVC
abstract
In High Efficiency Video Coding (HEVC), the coding efficiency of I-frames is lower than P-frames and B-frames, which will cause the flicker artifact, especially in low bitrates applications. We propose a region-classification-based rate control for Coding Tree Units (CTUs) in I-frames to improve the reconstructed quality of I-frames to suppress the flicker artifact. The CTUs in I-frame are classified into three regions according to their motion vectors and complexity. When the bit budget of one I-frame is used up, the target bitrates for the remaining CTUs will be adjusted according to the regions they belong to, and the pixel-based unified rate-quantization (URQ) model is then used to calculate the QPs. Experimental results demonstrate that the proposed scheme can efficiently suppress the flicker artifacts and improve both the subjective and objective video quality when compared with the original scheme in HM9.0.
Yongfei Zhang, Hai-Miao Hu, Bo Li 0006
ICIP2
2013 Rate-distortion optimized unequal loss protection for video transmission over packet erasure channels
Yongfei Zhang, Shiyin Qin, Bo Li 0006, Zhihai He
Signal Process. Image Commun.1
2012 Resource-Distortion Modeling for Video Streaming over Mesh Networks with Priority-Based Packet Scheduling
abstract
Video streaming over mesh networks operates under stringent network resource constraints, with a large number of video sessions competing for limited network resources. In this work, we aim to establish so-called resource-distortion models to characterize the inherent relationship between the allocated network resources to each video session and its end-to-end video distortion or presentation quality at the receiver end. We observed that priority-based packet scheduling has significant impact on such resource-distortion relationship. Using ANN-based learning methods, we develop an end-to-end packet delay bound violation (PDBV) probability model for video streaming over multi-hop networks with priority-based packet scheduling. We then derive a quadratic video transmission distortion model to capture the unique behavior of priority-based packet scheduling and its impact on the end-to-end video distortion. Based on these resource-distortion models, we are able to predict the end-to-end distortion for video streaming over multi-hop networks with priority-based packet scheduling. Our extensive experimental results demonstrate that the proposed method is very accurate.
Yongfei Zhang, Shiyin Qin, Zhihai He
ICME1
2012 Gradient-based fast decision for intra prediction in HEVC
abstract
As the next generation standard of video coding, the High Efficiency Video Coding(HEVC) achieves significantly better coding efficiency than all existing video coding standards, which is however at the cost of a much higher computation complexity. To address this issue, this paper presents a gradient-based fast decision algorithm for intra prediction in HEVC. More specifically, the intra prediction in HEVC is divided into two stages: prediction unit(PU) size decision and mode decision. At the PU size decision process, four orientation features are extracted from the coding unit by the intensity gradient filters to decide the texture complexity and texture direction of the coding unit, and then the texture direction is used to exclude impossible prediction modes at the mode decision process. Compared to HEVC reference software, the proposed algorithm saves around 56.7% of the encoding time in intra high efficiency setting and up to 70.86% in intra low complexity setting with slight performance degradation.
Yongfei Zhang, Zhe Li 0015, Bo Li 0006
VCIP1
2011 Transmission Distortion-optimized Unequal Loss Protection for video transmission over packet erasure channels
abstract
In this paper, we study the problem of Transmission Distortion-optimized Unequal Loss Protection (TD-ULP) under rate constraints for non-scalable video transmission over packet erasure channels. Based on a packet-level transmission distortion modeling scheme, we estimate the amount of contribution of each video packet to the reconstructed video quality, which defines the priority level of each packet. Unequal amounts of protections are then allocated to different video packets according to their priority levels as well as the dynamic channel conditions. The optimal ULP resource allocation is formulated as a constrained nonlinear optimization problem. An evolutionary algorithm based on Particle Swarm Optimization (PSO) is developed to obtain the optimal resource allocation. Our extensive experimental results demonstrate the effectiveness of the proposed TD-ULP scheme, which outperforms existing methods by up to 2dB gain in reconstructed video quality. 1*
Yongfei Zhang, Shiyin Qin, Zhihai He
ICME1
2010 Fine-Granularity Transmission Distortion Modeling for Video Packet Scheduling Over Mesh Networks
abstract
Packet scheduling is a critical component in multi-session video streaming over mesh networks. Different video packets have different levels of contribution to the overall video presentation quality at the receiver side. In this work, we develop a fine-granularity transmission distortion model for the encoder to predict the quality degradation of decoded videos caused by lost video packets. Based on this packet-level transmission distortion model, we propose a content-and-deadline-aware scheduling (CDAS) scheme for multi-session video streaming over multi-hop mesh networks, where content priority, queuing delays, and dynamic network transmission conditions are jointly considered for each video packet. Our extensive experimental results demonstrate that the proposed transmission distortion model and the CDAS scheme significantly improve the performance of multi-session video streaming over mesh networks.
Yongfei Zhang, Shiyin Qin, Zhihai He
IEEE Trans. Multim.1
2010 Multihop Packet Delay Bound Violation Modeling for Resource Allocation in Video Streaming Over Mesh Networks
abstract
Resource allocation plays a critical role in multisession video streaming over mesh networks to maximize the overall video presentation quality under transmission delay and network resource constraints. A critical component in efficient resource allocation is to analyze and model the multihop queuing behavior along the transmission path, estimate the packet loss ratio due to delay bound violation, and predict the amount of video quality degradation after multihop video transmission. In this work, we develop a multihop packet delay bound violation model to predict the packet loss probability and end-to-end distortion for video streaming over multihop networks. To this end, we extract salient features to characterize the input source and network conditions of links along the transmission path and construct a learning-based model using artificial neural network (ANN). Based on this model, we then formulate the resource allocation into a nonconvex optimization problem which aims to minimize the overall video distortion while maintaining fairness between sessions. We solve this optimization problem using Lagrangian duality methods. Extensive experimental results demonstrate that, with this widely-used offline-training-online-estimation mechanism, the proposed model is potentially applicable to almost all network conditions and can provide fairly accurate estimation results as compared with other models with a given sample data set. The proposed optimization algorithm achieves more efficient resource allocation than existing schemes.
Yongfei Zhang, Shixin Sun, S. Y. Qin, Zhihai He
IEEE Trans. Multim.2
2009 Packet-level transmission distortion analysis for video streaming over mesh networks
abstract
In video transmission over wired or wireless mesh networks, video packets could be lost due to transmission errors, network congestion or delay bound violation. The compressed video data is highly sensitive to packet loss, and the packet loss will cause decoding failure and more importantly, the error propagation would significantly degrade the reconstructed video quality. In this paper, we extend our previous research and propose a simple yet accurate and robust model to estimate the transmission distortion of each Macro Block (MB) and packet. We then use this transmission distortion estimation results for packet scheduling in video streaming over mesh networks. Extensive simulation results demonstrate the effectiveness and robustness of the proposed model under different experiment settings. More importantly, the estimated transmission distortion, which reveals the unequal importance of each MB and packet, enables important applications in content-aware resource allocation and performance optimization in video communication.
Yongfei Zhang, Shiyin Qin, Zhihai He
PCS1
2007 The object oriented analysis and modeling for obstacle avoidance of a behavior-based robot
abstract
This paper first describes key conceptions about object oriented analysis in software engineering, behavior based robotics and their conceptual similarities. Then, based on these similarities, the paper utilizes object oriented methods of software engineering, such as unified modeling language (UML), to analyze and model the architecture and design behaviors for a behavior-based robot, which is expected to wander with autonomous obstacle avoidance in unknown environment. Object oriented methods permit a translation from conceptual behavior models to computer programming representations, and separate concrete control algorithms from robot modeling. With this approach, the paper also implements a fuzzy algorithm for obstacle avoidance behavior of the constructed behavior models in a physical robot. Finally the paper gives experimental results and points out future directions.
Qian Zhang 0016, Yongfei Zhang, Shi-Yin Qin
SMC2