Fan Yang 0053

dblp:29/3081-53 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-3431-2585ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Three-Stage Progressive Pre-Analysis Framework for VMAF Controllable Image Coding
abstract
To achieve controllable subjective quality in image coding, this paper proposes a Video Multi-method Assessment Fusion (VMAF)-oriented image coding pre-analysis algorithm, enabling the adaptive derivation of quantization parameters corresponding to a specified quality target. First, a$Q-\mathcal{V}$model is constructed to describe the relationship between encoding quantization and VMAF distortion. Then, a Three-stage Progressive Control (TPC) algorithm, shown in Fig. 1(a), is designed to adapt quantization parameters using the discovered$Q-\mathcal{V}$model. The first two stages, based on lightweight feature extraction, iteratively fit the distortion metrics intrinsically calculated by VMAF to predict the VMAF value for a given sample under specified distortion conditions. The final stage fits the$Q-\mathcal{V}$model parameters using multi-point VMAF distortion data and outputs the corresponding quantization step for encoder control. A two-pass refinement algorithm, depicted in Fig. 1(b), further adjusts the quantization parameters based on the first encoding pass, improving quality control accuracy and framework robustness. Experiments on four datasets show that the quality control error remains below 1.293% for various VMAF targets, and the two-pass refinement reduces it further to 0.710%, outperforming existing methods.
Guoqing Xiang, Wenzhao Li, Mingyuan Yang, Fan Yang 0053, Shanghang Zhang, Huizhu Jia
DCC5
2024 VLUReID: Exploiting Vision-Language Knowledge for Unsupervised Person Re-Identification
abstract
The superior performances of pre-trained vision-language models on various downstream tasks demonstrate the effectiveness of integrating cross-modal vision-language knowledge into visual tasks. However, this knowledge is hardly used for visual-based person re-identification (re-ID) because the datasets lack textual descriptions. Existing efforts require manual annotations for training, which can be time-consuming. We propose VLUReID, a framework that improves visual-based person re-ID using vision-language knowledge without requiring manual annotations from datasets. Specifically, the Vision-to-Text Association (VTA) module uses designed textual prompts to prompt the vision-language model in generating pseudo-semantic labels for visual inputs. Subsequently, within the Dual-Branch Asymmetric Training (DBAT) module, we propose an asymmetric training strategy to extract cross-modal knowledge from pseudo-semantic labels and integrate it into the person re-ID model. The experimental results on two widely-used benchmarks for unsupervised video-based person re-ID demonstrate the effectiveness of our framework.
Ray Zhang 0002, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Shanghang Zhang
ICME3
2022 Directed Mix Contrast for Lidar Point Cloud Segmentation
abstract
Comprehensive and real-time scene understanding are crucial for autonomous driving, where LiDAR semantic segmentation plays an indispensable role. However, existing segmentation algorithms are exposed with extremely imbalanced dataset. Besides, due to the sparseness of point cloud, some hard negative samples are intrinsically similar. In this paper, we propose a contrastive learning framework named Directed Mix Contrast (DMC) for LiDAR point cloud segmentation. There are two key contributions in this framework. Firstly, a simple pre-processing step is proposed to maintain the balance between classes. In addition, to deal with the hard negative samples, contrastive learning strategy is introduced to facilitate feature learning. More specifically, we propose a directed mix-up training strategy to take control of the contrasting procedure. To demonstrate the effectiveness of DMC, we conduct experiments on two large-scale LiDAR Segmentation datasets and achieve mIoU of 71.0% on SemanticK-ITTI, which outperforms most existing methods. Furthermore, we extend our method to several backbone networks, and the results show good generalization ability of DMC.
Yuheng Lu, Fangping Chen, Ziwei Zhang 0003, Fan Yang 0053
ICME4
2022 Towards Blind Watermarking: Combining Invertible and Non-invertible Mechanisms
abstract
Blind watermarking provides powerful evidence for copyright protection, image authentication, and tampering identification.However, it remains a challenge to design a watermarking model with high imperceptibility and robustness against strong noise attacks. To resolve this issue, we present a framework Combining the Invertible and Non-invertible (CIN) mechanisms. The CIN is composed of the invertible part to achieve high imperceptibility and the non-invertible part to strengthen the robustness against strong noise attacks. For the invertible part, we develop a diffusion and extraction module (DEM) and a fusion and split module (FSM) to embed and extract watermarks symmetrically in an invertible way. For the non-invertible part, we introduce a non-invertible attention-based module (NIAM) and the noise-specific selection module (NSM) to solve the asymmetric extraction under a strong noise attack. Extensive experiments demonstrate that our framework outperforms the current state-of-the-art methods of imperceptibility and robustness significantly. Our framework can achieve an average of 99.99% accuracy and 67.66 dB PSNR under noise-free conditions, while 96.64% and 39.28 dB combined strong noise attacks. The code will be available in https://github.com/RM1110/CIN.
Rui Ma 0032, Mengxi Guo, Fan Yang 0053, Yuan Li 0014, Huizhu Jia
ACM Multimedia4
2021 Deep Human-Interaction and Association by Graph-Based Learning for Multiple Object Tracking in the Wild
Cong Ma 0006, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Wen Gao 0001
Int. J. Comput. Vis.2
2021 Deep Trajectory Post-Processing and Position Projection for Single & Multiple Camera Multiple Object Tracking
Cong Ma 0006, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Wen Gao 0001
Int. J. Comput. Vis.2
2021 Part-aware Progressive Unsupervised Domain Adaptation for Person Re-Identification
abstract
Unsupervised domain adaptation (UDA) aims to mitigate the domain shift that occurs when transferring knowledge from a labeled source domain to an unlabeled target domain. While it has been studied for application in unsupervised person re-identification (ReID), the relations of feature distribution across the source and target domains remain underexplored, as they either ignore the local relations or omit the in-depth consideration of negative transfer when two domains do not share identical label spaces. In light of the above, this paper presents an innovative part-aware progressive adaptation network (PPAN) that exploits global and local relations for UDA-based ReID across domains. A multi-branch network is developed that explicitly learns discriminative feature representation from both whole-body images and body-part images under the supervision of a labeled source domain. Within each network branch, an independent UDA constraint is designed that aligns the global and local feature distributions from a labeled source domain with those of an unlabeled target domain. In addition, a novel progressive adaptation strategy (PAS) is designed that effectively alleviates the negative influence of outlier source identities. The proposed unsupervised ReID model is evaluated on five widely used datasets (Market-1501, DukeMTMC-reID, CUHK03, VIPeR and PRID), and experimental results demonstrate its superior robustness and effectiveness relative to state-of-the-art approaches.
Fan Yang 0053, Shijian Lu, Huizhu Jia, Don Xie, Zongqiao Yu, Feiyue Huang, Wen Gao 0001
IEEE Trans. Multim.1
2020 BBA-NET: A Bi-Branch Attention Network For Crowd Counting
abstract
In the field of crowd counting, the current mainstream CNNbased regression methods simply extract the density information of pedestrians without finding the position of each person. This makes the output of the network often found to contain incorrect responses, which may erroneously estimate the total number and not conducive to the interpretation of the algorithm. To this end, we propose a Bi-Branch Attention Network (BBA-NET) for crowd counting, which has three innovation points. i) A two-branch architecture is used to estimate the density information and location information separately. ii) Attention mechanism is used to facilitate feature extraction, which can reduce false responses. iii) A new density map generation method combining geometric adaptation and Voronoi split is introduced. Our method can integrate the pedestrian’s head and body information to enhance the feature expression ability of the density map. Extensive experiments performed on two public datasets show that our method achieves a lower crowd counting error compared to other state-of-the-art methods.
Chengyang Li 0001, Fan Yang 0053, Cong Ma 0006, Yuan Li 0014, Huizhu Jia
ICASSP3
2020 Fusion Target Attention Mask Generation Network For Video Segmentation
abstract
Video segmentation aims to segment target objects in a video sequence, which remains a challenge due to the motion and deformation of objects. In this paper, we propose a novel attention-driven hybrid encoder-decoder network that generates object segmentation by fully leveraging spatial and temporal information. Firstly, a multi-branch network is designed to learn feature representation from object appearance, location and motion. Secondly, a target attention module is proposed to further exploit context information from learned representation. In addition, a novel edge loss is designed which constraints the model to generate salient edge features and accurate segmentation. The proposed model has been evaluated over two widely used public benchmarks, and experiments demonstrate its superior robustness and effectiveness as compared with the state of the arts.
Yunyi Li, Fangping Chen, Fan Yang 0053, Yuan Li 0014, Huizhu Jia
ICIP3
2020 Optical Flow-Guided Mask Generation Network for Video Segmentation
abstract
The purpose of video segmentation is to segment foreground objects from a video sequence. In this paper, we propose a CNN based method for the semi-supervised video object segmentation, where a hybrid encoder-decoder network is designed to generate pixel-wise foreground object segmentation in use of both spatial and temporal information. In order to minimize cumulative error of the network as much as possible, we develop a two-stage training scheme: alternate training and back-propagation-through-time training. Then the performances of our method and other state-of-the-art ones are compared on two annotated video segmentation databases. Furthermore, we also run an extensive ablation study to test the effects of different components from our method.
Yunyi Li, Fangping Chen, Fan Yang 0053, Cong Ma 0006, Yuan Li 0014, Huizhu Jia
ISCAS3
2019 Deep Association: End-to-end Graph-Based Learning for Multiple Object Tracking with Conv-Graph Neural Network
abstract
Multiple Object Tracking (MOT) has a wide range of applications in surveillance retrieval and autonomous driving. The majority of existing methods focus on extracting features by deep learning and hand-crafted optimizing bipartite graph or network flow. In this paper, we proposed an efficient end-to-end model, Deep Association Network (DAN), to learn the graph-based training data, which are constructed by spatial-temporal interaction of objects. DAN combines Convolutional Neural Network (CNN), Motion Encoder (ME) and Graph Neural Network (GNN). The CNNs and Motion Encoders extract appearance features from bounding box images and motion features from positions respectively, and then the GNN optimizes graph structure to associate the same object among frames together. In addition, we presented a novel end-to-end training strategy for Deep Association Network. Our experimental results demonstrate the effectiveness of DAN up to the state-of-the-art methods without extra-dataset on MOT16 and DukeMTMCT.
Cong Ma 0006, Yuan Li 0014, Fan Yang 0053, Ziwei Zhang 0003, Yueqing Zhuang, Huizhu Jia
ICMR3
2019 Attention driven person re-identification
Fan Yang 0053, Shijian Lu, Huizhu Jia, Wen Gao 0001
Pattern Recognit.1
2018 Dense Relation Network: Learning Consistent and Context-Aware Representation for Semantic Image Segmentation
abstract
Semantic image segmentation, which aims at assigning pixel-wise category, is one of challenging image understanding problems. Global context plays an important role on local pixel-wise category assignment. To make the best of global context, in this paper, we propose dense relation network (DRN) and context-restricted loss (CRL) to aggregate global and local information. DRN uses Recurrent Neural Network (RNN) with different skip lengths in spatial directions to get context-aware representations while CRL helps aggregate them to learn consistency. Compared with previous methods, our proposed method takes full advantage of hierarchical contextual representations to produce high-quality results. Extensive experiments demonstrate that our method achieves significant state-of-the-art performances on Cityscapes and Pascal Context benchmarks, with mean-IoU of 82.8% and 49.0% respectively.
Yueqing Zhuang, Fan Yang 0053, Cong Ma 0006, Ziwei Zhang 0003, Yuan Li 0014, Huizhu Jia, Wen Gao 0001
ICIP2
2018 Trajectory Factory: Tracklet Cleaving and Re-Connection by Deep Siamese Bi-GRU for Multiple Object Tracking
abstract
Multi-Object Tracking (MOT) is a challenging task in the complex scene such as surveillance and autonomous driving. In this paper, we propose a novel tracklet processing method to cleave and re-connect tracklets on crowd or longterm occlusion by Siamese Bi-Gated Recurrent Unit (GRU). The tracklet generation utilizes object features extracted by CNN and RNN to create the high-confidence tracklet candidates in sparse scenario. Due to mis-tracking in the generation process, the tracklets from different objects are split into several sub-tracklets by a bidirectional GRU. After that, a Siamese GRU based tracklet re-connection method is applied to link the sub-tracklets which belong to the same object to form a whole trajectory. In addition, we extract the track-let images from existing MOT datasets and propose a novel dataset to train our networks. The proposed dataset contains more than 95160 pedestrian images. It has 793 different persons in it. On average, there are 120 images for each person with positions and sizes. Experimental results demonstrate the advantages of our model over the state-of-the-art methods on MOTI6.
Cong Ma 0006, Changshui Yang, Fan Yang 0053, Yueqing Zhuang, Ziwei Zhang 0003, Huizhu Jia
ICME3
2018 RelationNet: Learning Deep-Aligned Representation for Semantic Image Segmentation
abstract
Semantic image segmentation, which assigns labels in pixel level, plays a central role in image understanding. Recent approaches have attempted to harness the capabilities of deep learning. However, one central problem of these methods is that deep convolutional neural network gives little consideration to the correlation among pixels. To handle this issue, in this paper, we propose a novel deep neural network named RelationNet, which utilizes CNN and RNN to aggregate context information. Besides, a spatial correlation loss is applied to train RelationNet to align features of spatial pixels belonging to same category. Importantly, since it is expensive to obtain pixel-wise annotations, we exploit a new training method to combine the coarsely and finely labeled data. Experiments show the detailed improvements of each proposal. Experimental results demonstrate the effectiveness of our proposed method to the problem of semantic image segmentation, which obtains state-of-the-art performance on the Cityscapes benchmark and Pascal Context dataset.
Yueqing Zhuang, Fan Yang 0053, Cong Ma 0006, Ziwei Zhang 0003, Huizhu Jia
ICPR3
2016 Structure preserving single image super-resolution
abstract
In this paper, we present a novel structure preserving method for single image super-resolution to well construct edge structures and small detail structures. In our approach, the sharp edges are recovered via a novel edge preserving interpolation technique based on a well estimated gradient field and the edge preserving method, which incorporate the local and non-local structure information. The gradient of interpolated high-resolution(HR) image is then regarded as an edge preserving constraint to reconstruct the detail structures. Experimental results demonstrate that the new approach can reconstruct faithfully the HR images with sharp edges and texture structures, and annoying artifacts (blurring, jaggies, ringing, etc.) are greatly suppressed. It outperforms the state-of-the-art approaches, based on subjective and objective evaluations.
Fan Yang 0053, Don Xie, Huizhu Jia, Rui Chen 0006, Guoqing Xiang, Wen Gao 0001
ICIP1