Yang Yang 0062

dblp:48/450-62 · DBLP profile ↗
← Back
55ranked-venue papers
15as first author
24since 2021 · last 2026
0000-0003-0559-5464ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 9 first-author · 10 since 2021Security and privacy · 5 · 2 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
abstract
Camera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for transferring rich information from offline models to online models. However, existing KD methods overlook future frames, as they mainly focus on spatial feature distillation under strict frame alignment or on temporal relational distillation, thereby making it challenging for online models to effectively learn future knowledge. To this end, we propose a sparse query-based approach, Future Temporal Knowledge Distillation (FTKD), which effectively transfers future frame knowledge from an offline teacher model to an online student model. Specifically, we present a future-aware feature reconstruction strategy to encourage the student model to capture future features without strict frame alignment. In addition, we further introduce future-guided logit distillation to leverage the teacher's stable foreground and background context. FTKD is applied to two high-performing 3D object detection baselines, achieving up to 1.3 mAP and 1.3 NDS gains on the nuScenes dataset, as well as the most accurate velocity estimation, without increasing inference cost.
Hu Zhu, Weihao Gu, Yang Yang 0062, Yanyan Liang 0001
AAAI5
2026 PlainDrop: Practical Asynchronous Proactive Secret Sharing With Silent Setup
abstract
Dynamic Proactive Secret Sharing (DPSS) is essential for distributed systems, enabling long-term key escrow, BFT protocol reconfiguration, and confidential state machine replication. Yet existing asynchronous schemes, while crucial for realistic settings, suffer from high communication overhead and poor practicality, limiting real-world deployment. We propose PlainDrop, a concise and efficient DPSS protocol designed specifically for asynchronous networks. PlainDrop achieves optimized communication complexity ofO(n2) via commitment–share decoupling combined with homomorphic threshold encryption techniques. PlainDrop also eliminates the need for expensive distributed key generation and complex bivariate polynomial structures by introducing a lightweight silent setup framework and employing direct share processing based on univariate polynomials. We formally prove that PlainDrop provides secrecy, integrity, and termination in asynchronous networks against a mobile adversary corrupting up to one third of the parties. We implement PlainDrop and evaluate it on Amazon EC2 with up to 100 nodes. Our experimental results demonstrate average reductions of 37% and 67% in completion time, and 61% and 89% in communication volume, compared to DyCAPS and LongLive, respectively.
Yang Yang 0062, Bingyu Li 0003, Qin Wang 0008, Qianhong Wu, Willy Susilo
IEEE Internet Things J.1
2026 SA-Person: Text-Based Person Retrieval With Scene-Aware Re-Ranking
Yingjia Xu, Jinlin Wu, Daming Gao, Zhen Chen 0018, Yang Yang 0062, Min Cao 0005, Mang Ye, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.5
2025 RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection
abstract
In radar-camera 3D object detection, the radar point clouds are sparse and noisy, which causes difficulties in fusing camera and radar modalities. To solve this, we introduce a novel query-based detection method named Radar-Camera Transformer (RCTrans). Specifically, we first design a Radar Dense Encoder to enrich the sparse valid radar tokens, and then concatenate them with the image tokens. By doing this, we can fully explore the 3D information of each interest region and reduce the interference of empty tokens during the fusing stage. We then design a Pruning Sequential Decoder to predict 3D boxes based on the obtained tokens and random initialized queries. To alleviate the effect of elevation ambiguity in radar point clouds, we gradually locate the position of the object via a sequential fusion structure. It helps to get more precise and flexible correspondences between tokens and queries. A pruning training strategy is adopted in the decoder, which can save much time during inference and inhibit queries from losing their distinctiveness. Extensive experiments on the large-scale nuScenes dataset prove the superiority of our method, and we also achieve new state-of-the-art radar-camera 3D detection results.
Yang Yang 0062, Zhen Lei 0001
AAAI2
2025 RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images
abstract
In recent years, diffusion models have revolutionized visual generation, outperforming traditional frameworks like Generative Adversarial Networks (GANs). However, generating images of humans with realistic semantic parts, such as hands and faces, remains a significant challenge due to their intricate structural complexity. To address this issue, we propose a novel post-processing solution named RealisHuman. The RealisHuman framework operates in two stages. First, it generates realistic human parts, such as hands or faces, using the original malformed parts as references, ensuring consistent details with the original image. Second, it seamlessly integrates the rectified human parts back into their corresponding positions by repainting the surrounding areas to ensure smooth and realistic blending. The RealisHuman framework significantly enhances the realism of human generation, as demonstrated by notable improvements in both qualitative and quantitative metrics.
Benzhi Wang, Jingkai Zhou, Jingqi Bai, Yang Yang 0062, Fan Wang 0019, Zhen Lei 0001
AAAI4
2025 FlexiADKG: A Flexible Asynchronous Distributed Key Generation Protocol with Constant Round Complexity
Yang Yang 0062, Bingyu Li 0003, Zhenyang Ding, Qianhong Wu, Qin Wang 0008
ACISP (1)1
2025 Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation
abstract
To tackle the threat of fake news, the task of detecting and grounding multi-modal media manipulation (DGM4) has received increasing attention. However, most state-of-the-art methods fail to explore the fine-grained consistency within local content, usually resulting in an inadequate perception of detailed forgery and unreliable results. In this paper, we propose a novel approach named Contextual-Semantic Consistency Learning (CSCL) to enhance the fine-grained perception ability of forgery for DGM4. Two branches for image and text modalities are established, each of which contains two cascaded decoders, i.e., Contextual Consistency Decoder (CCD) and Semantic Consistency Decoder (SCD), to capture within-modality contextual consistency and across-modality semantic consistency, respectively. Both CCD and SCD adhere to the same criteria for capturing fine-grained forgery details. To be specific, each module first constructs consistency features by leveraging additional supervision from the heterogeneous information of each token pair. Then, the forgery-aware reasoning or aggregating is adopted to deeply seek forgery cues based on the consistency features. Extensive experiments on DGM4datasets prove that CSCL achieves new state-of-the-art performance, especially for the results of grounding manipulated content. Codes and weights are avaliable at https://github.com/liyih/CSCL.
Yang Yang 0062, Zichang Tan, Huan Liu 0030, Zhen Lei 0001
CVPR2
2025 Layer-Animate for Transparent Video Generation
abstract
Transparent videos with alpha channels play a crucial role in film production, advertising, and augmented reality fields. However, there is currently no available method for producing transparent videos. Traditional methods are time-consuming and labor-intensive, and employing alternative approaches for this task will result in inaccurate transparent regions, constrained motion, and artifacts. To address these challenges, we propose Layer-Animate, the first method capable of generating transparent videos. Our method comprises two stages: in the first stage, transparent images are generated as the base images to provide content and transparency information for the next stage. In the second stage, Inter-Frame Attention is applied to decouple content from motion, enabling the motion module to focus better on action. Layer-Animate is the first method used to generate transparent videos with accurate transparent regions, sufficient motion, and no artifacts, as demonstrated by notable improvements in qualitative and quantitative metrics.
Jingqi Bai, Jingkai Zhou, Benzhi Wang, Yang Yang 0062, Zhen Lei 0001, Fan Wang 0019
ICASSP5
2025 TrustBlink: A zkSNARK-Powered On-Demand Relay for PoW Cross-Chain Verification With Low Costs
Bohang Wei, Yang Yang 0062, Shihong Xiong, Minghang Li, Qianhong Wu
ICICS (2)2
2025 AccidentBlip: Agent of Accident Warning Based on MA-Former
abstract
In complex transportation systems, accurately sensing the surrounding environment and predicting the risk of potential accidents is crucial. Most existing accident prediction methods are based on temporal neural networks, such as RNN and LSTM. Recent multimodal fusion approaches improve vehicle localization through 3D target detection and assess potential risks by calculating inter-vehicle distances. However, these temporal networks and multimodal fusion methods suffer from limited detection robustness and high economic costs. To address these challenges, we propose AccidentBlip, a vision-only framework that employs our self-designed Motion Accident Transformer (MA-former) to process each frame of video. Unlike conventional self-attention mechanisms, MA-former replaces Q-former's self-attention with temporal attention, allowing the query corresponding to the previous frame to generate the query input for the next frame. Additionally, we introduce a residual module connection between queries of consecutive frames to enhance the model's temporal processing capabilities. For complex V2V and V2X scenarios, AccidentBlip adapts by concatenating queries from multiple cameras, effectively capturing spatial and temporal relationships. In particular, AccidentBlip achieves SOTA performance in both accident detection and prediction tasks on the DeepAccident dataset. It also outperforms current SOTA methods in V2V and V2X scenarios, demonstrating a superior capability to understand complex real-world environments.
Yihua Shao, Yeling Xu, Xinwei Long, Siyu Chen 0021, Ziyang Yan, Haoting Liu, Yan Wang 0068, Hao Tang 0005, Yang Yang 0062
IV9
2025 Privacy Preservation in AI-Driven IoT for Vehicles via Hierarchical Sharding Blockchain
abstract
The AI-driven Internet of Things (AIoT) has been widely applied in the field of Internet of Vehicles (IoV) for vehicular cooperation. Federated learning (FL), due to its ability to protect users’ data privacy, reduce communication overhead, and facilitate real-time decision making, is widely applied in the augmented intelligence of things for vehicles (AIoV). However, integrating FL with AIoV poses challenges, including the absence of fine-grained access control, insufficient safeguards for FL tasks and vehicle identities, inadequate security for data transmission, and shortcomings in protecting data storage. These vulnerabilities may lead to risks such as vehicle tracking, model information theft, and data tampering. To address these challenges, we propose a privacy preservation mechanism for AIoV via cloud–edge–vehicle hierarchical sharding blockchain. First, we propose a hierarchical anonymous authentication scheme for IoV devices with stronger scalability and higher fault tolerance. Vehicles only know the attributes of each other or which shard they belong to. Second, we present a secure FL task assignment scheme for AIoV. Edge nodes utilize attribute-based encryption to deploy fine-grained FL tasks based on vehicle attributes. Only users who meet the attributes can decrypt the content, protecting FL tasks content and participant identities. Third, we present a secure data transmission scheme between AIoV devices to protect the identity and data privacy of both parties, while also achieving noninteractive key agreement. Additionally, we propose a scalable secure data sharing and storage scheme based on hierarchical sharding blockchain, aiming to reduce storage overhead and minimize trust costs.
Mingzhe Zhai, Qianhong Wu, Yizhong Liu, Yang Yang 0062, Muhammad Ghulam, Prayag Tiwari
IEEE Internet Things J.5
2025 RandFlash: Breaking the Quadratic Barrier in Large-Scale Distributed Randomness Beacons
abstract
Random beacons are of paramount importance in distributed systems (e.g., blockchain, electronic voting, governance). The sheer scale of nodes inherent in distributed environments necessitates minimizing communication overhead per node while ensuring protocol availability, particularly under adversarial conditions. Existing solutions have managed to reduce the optimistic overhead to a minimum ofO(n2), wherenrepresents the node count of the system. In this paper, we step further by proposing and implementing RandFlash, a leaderless random beacon protocol that achieves an optimistic communication complexity ofO(nlogn). Evaluation results demonstrate that RandFlash outperforms existing constructions, RandPiper (CCS’21) and OptRand (NDSS’23), in terms of the number of random beacons generated within largescale networks comprising 64 nodes or more (e.g., in sizes of 80 and 128). Furthermore, RandFlash exhibits resilience, capable of withstanding up to one-third of the nodes acting maliciously, all without the need for strongly trusted setups (i.e., embedding a secret trapdoor by trusted third parties). We also provide formal security proofs validating all properties upheld by this lineage.
Yang Yang 0062, Bingyu Li 0003, Qianhong Wu, Qin Wang 0008, Shihong Xiong, Willy Susilo
IEEE Trans. Inf. Forensics Secur.1
2025 Pixel and Feature Transfer Fusion for Unsupervised Cross-Dataset Person Reidentification
abstract
Recently, unsupervised cross-dataset person reidentification (Re-ID) has attracted more and more attention, which aims to transfer knowledge of a labeled source domain to an unlabeled target domain. There are two common frameworks: one is pixel-alignment of transferring low-level knowledge, and the other is feature-alignment of transferring high-level knowledge. In this article, we propose a novel recurrent autoencoder (RAE) framework to unify these two kinds of methods and inherit their merits. Specifically, the proposed RAE includes three modules, i.e., a feature-transfer (FT) module, a pixel-transfer (PT) module, and a fusion module. The FT module utilizes an encoder to map source and target images to a shared feature space. In the space, not only features are identity-discriminative but also the gap between source and target features is reduced. The PT module takes a decoder to reconstruct original images with its features. Here, we hope that the images reconstructed from target features are in the source style. Thus, the low-level knowledge can be propagated to the target domain. After transferring both high-and low-level knowledge with the two proposed modules above, we design another bilinear pooling layer to fuse both kinds of knowledge. Extensive experiments on Market-1501, DukeMTMC-ReID, and MSMT17 datasets show that our method significantly outperforms either pixel-alignment or feature-alignment Re-ID methods and achieves new state-of-the-art results.
Yang Yang 0062, Guan'an Wang, Prayag Tiwari, Hari Mohan Pandey, Zhen Lei 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Euclidean-Distance-Preserved Feature Reduction for efficient person re-identification
Guan'an Wang, Xiaowen Huang 0001, Yang Yang 0062, Prayag Tiwari, Jian Zhang 0018
Neural Networks3
2024 Depth-Aware Unpaired Video Dehazing
abstract
This paper investigates a novel unpaired video dehazing framework, which can be a good candidate in practice by relieving pressure from collecting paired data. In such a paradigm, two key issues including 1) temporal consistency uninvolved in single image dehazing, and 2) better dehazing ability need to be considered for satisfied performance. To handle the mentioned problems, we alternatively resort to introducing depth information to construct additional regularization and supervision. Specifically, we attempt to synthesize realistic motions with depth information to improve the effectiveness and applicability of traditional temporal losses, and thus better regularizing the spatiotemporal consistency. Moreover, the depth information is also considered in terms of adversarial learning. For haze removal, the depth information guides the local discriminator to focus on regions where haze residuals are more likely to exist. The dehazing performance is consequently improved by more pertinent guidance from our depth-aware local discriminator. Extensive experiments are conducted to validate our effectiveness and superiority over other competitors. To the best of our knowledge, this study is the initial foray into the task of unpaired video dehazing. Our code is available at https://github.com/YaN9-Y/DUVD.
Yang Yang 0062, Chunle Guo, Xiaojie Guo 0001
IEEE Trans. Image Process.1
2023 IKE: Threshold Key Escrow Service with Intermediary Encryption
Yang Yang 0062, Bingyu Li 0003, Shihong Xiong, Yan Zhu 0023, Haibin Zheng, Qianhong Wu
ICA3PP (3)1
2023 Visible-Infrared Person Re-Identification via Semantic Alignment and Affinity Inference
abstract
Visible-infrared person re-identification (VI-ReID) focuses on matching the pedestrian images of the same identity captured by different modality cameras. The part-based methods achieve great success by extracting fine-grained features from feature maps. But most existing part-based methods employ horizontal division to obtain part features suffering from misalignment caused by irregular pedestrian movements. Moreover, most current methods use Euclidean or cosine distance of the output features to measure the similarity without considering the pedestrian relationships. Misaligned part features and naive inference methods both limit the performance of existing works. We propose a Semantic Alignment and Affinity Inference framework (SAAI), which aims to align latent semantic part features with the learnable prototypes and improve inference with affinity information. Specifically, we first propose semantic-aligned feature learning that employs the similarity between pixelwise features and learnable prototypes to aggregate the latent semantic part features. Then, we devise an affinity inference module to optimize the inference with pedestrian relationships. Comprehensive experimental results conducted on the SYSU-MM01 and RegDB datasets demonstrate the favorable performance of our SAAI framework. Our code will be released at https://github.com/xiaoye-hhh/SAAI.
Xingye Fang, Yang Yang 0062
ICCV2
2023 Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person Search
abstract
Weakly supervised person search aims to jointly detect and match persons with only bounding box annotations. Existing approaches typically focus on improving the features by exploring the relations of persons. However, scale variation problem is a more severe obstacle and under-studied that a person often owns images with different scales (resolutions). For one thing, small-scale images contain less information of a person, thus affecting the accuracy of the generated pseudo labels. For another, different similarities between cross-scale images of a person increase the difficulty of matching. In this paper, we address it by proposing a novel one-step framework, named Self-similarity driven Scale-invariant Learning (SSL). Scale invariance can be explored based on the self-similarity prior that it shows the same statistical properties of an image at different scales. To this end, we introduce a Multi-scale Exemplar Branch to guide the network in concentrating on the foreground and learning scale-invariant features by hard exemplars mining. To enhance the discriminative power of the learned features, we further introduce a dynamic pseudo label prediction that progressively seeks true labels for training. Experimental results on two standard benchmarks, i.e., PRW and CUHK-SYSU datasets, demonstrate that the proposed method can solve scale variation problem effectively and perform favorably against state-of-the-art methods. Code is available at https://github.com/Wangbenzhi/SSL.git.
Benzhi Wang, Yang Yang 0062, Jinlin Wu, Guo-Jun Qi, Zhen Lei 0001
ICCV2
2023 Reaching consensus for membership dynamic in secret sharing and its application to cross-chain
abstract
The communication efficiency optimization, censorship resilience, and generation of shared randomness are inseparable from the threshold cryptography in the existing Byzantine Fault Tolerant (BFT) consensus. The membership in consensus in a blockchain scenario supports dynamic changes, which effectively prevents the corruption of consensus participants. Especially in cross-chain protocols, the dynamic access to different blockchains will inevitably bring about the demand for member dynamic. Most existing threshold cryptography schemes rely on redefined key shares, leading to a static set of secret sharing participants. In this paper, we propose a general approach to coupling blockchain consensus and dynamic secret sharing. The committee performs consensus confirmation of both dynamic secret sharing and transaction proposals. Our scheme facilitates threshold cryptography membership dynamic, thus underlying support for membership dynamic of threshold cryptography-based BFT consensus schemes. We instantiate a dynamic HotStuff consensus to demonstrate the effectiveness of the scheme. After the correctness and security proof, our scheme achieves the secrecy and integrity of the threshold key shares while ensuring consensus liveness and safety. Experimental results prove that our scheme obtains dynamic membership with negligible overhead.
Yan Zhu 0023, Bingyu Li 0003, Zhenyang Ding, Yang Yang 0062, Qianhong Wu, Haibin Zheng
High Confid. Comput.4
2023 Camera-aware representation learning for person re-identification
Jinlin Wu, Zhen Lei 0001, Yang Yang 0062, Shukai Chen, Stan Z. Li
Neurocomputing4
2022 CAViT: Contextual Alignment Vision Transformer for Video Object Re-identification
Jinlin Wu, Lingxiao He, Wu Liu 0005, Yang Yang 0062, Zhen Lei 0001, Tao Mei 0001, Stan Z. Li
ECCV (14)4
2022 Low-Light Image Enhancement via Feature Restoration
abstract
Besides poor visibility, under-exposed images often suffer from severe noise and color distortion. Most existing Retinex-based methods deal with the noise and color distortion via some careful designs to denoising and/or color correction. In this paper, we propose a simple yet effective network from the perspective of feature map restoration to mitigate such issues without constructing any explicit modules. More concretely, we build an encoder-decoder network to reconstruct images, while a feature restoration subnet is introduced to transform the features of low-light images to those of corresponding clear ones. The enhanced images are consequently acquired through assembling the restored features by the decoder, in which, the noise and possible color distortion can be greatly remedied. Extensive experiments on widely-used datasets are conducted to validate the superiority of our design over other state-of-the-art alternatives both quantitatively and qualitatively. Our code is available at https://github.com/YaN9-Y/FRLIE.
Yang Yang 0062, Xiaojie Guo 0001
ICASSP1
2022 Adversarial Binary Mutual Learning for Semi-Supervised Deep Hashing
abstract
Hashing is a popular search algorithm for its compact binary representation and efficient Hamming distance calculation. Benefited from the advance of deep learning, deep hashing methods have achieved promising performance. However, those methods usually learn with expensive labeled data but fail to utilize unlabeled data. Furthermore, the traditional pairwise loss used by those methods cannot explicitly force similar/dissimilar pairs to small/large distances. Both weaknesses limit existing methods’ performance. To solve the first problem, we propose a novel semi-supervised deep hashing model named adversarial binary mutual learning (ABML). Specifically, our ABML consists of a generative model$G_{H}$and a discriminative model$D_{H}$, where$D_{H}$learns labeled data in a supervised way and$G_{H}$learns unlabeled data by synthesizing real images. We adopt an adversarial learning (AL) strategy to transfer the knowledge of unlabeled data to$D_{H}$by making$G_{H}$and$D_{H}$mutually learn from each other. To solve the second problem, we propose a novel Weibull cross-entropy loss (WCE) by using the Weibull distribution, which can distinguish tiny differences of distances and explicitly force similar/dissimilar distances as small/large as possible. Thus, the learned features are more discriminative. Finally, by incorporating ABML with WCE loss, our model can acquire more semantic and discriminative features. Extensive experiments on four common data sets (CIFAR-10, large database of handwritten digits (MNIST), ImageNet-10, and NUS-WIDE) and a large-scale data set ImageNet demonstrate that our approach successfully overcomes the two difficulties above and significantly outperforms state-of-the-art hashing methods.
Guan'an Wang, Qinghao Hu 0001, Yang Yang 0062, Jian Cheng 0001, Zeng-Guang Hou
IEEE Trans. Neural Networks Learn. Syst.3
2021 Cascaded Split-and-Aggregate Learning with Feature Recombination for Pedestrian Attribute Recognition
Yang Yang 0062, Zichang Tan, Prayag Tiwari, Hari Mohan Pandey, Jun Wan 0001, Zhen Lei 0001, Guodong Guo, Stan Z. Li
Int. J. Comput. Vis.1
2020 Relation-Aware Pedestrian Attribute Recognition with Graph Convolutional Networks
abstract
In this paper, we propose a new end-to-end network, named Joint Learning of Attribute and Contextual relations (JLAC), to solve the task of pedestrian attribute recognition. It includes two novel modules: Attribute Relation Module (ARM) and Contextual Relation Module (CRM). For ARM, we construct an attribute graph with attribute-specific features which are learned by the constrained losses, and further use Graph Convolutional Network (GCN) to explore the correlations among multiple attributes. For CRM, we first propose a graph projection scheme to project the 2-D feature map into a set of nodes from different image regions, and then employ GCN to explore the contextual relations among those regions. Since the relation information in the above two modules is correlated and complementary, we incorporate them into a unified framework to learn both together. Experiments on three benchmarks, including PA-100K, RAP, PETA attribute datasets, demonstrate the effectiveness of the proposed JLAC.
Zichang Tan, Yang Yang 0062, Jun Wan 0001, Guodong Guo, Stan Z. Li
AAAI2
2020 Cross-Modality Paired-Images Generation for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared (IR) person re-identification is very challenging due to the large cross-modality variations between RGB and IR images. The key solution is to learn aligned features to the bridge RGB and IR modalities. However, due to the lack of correspondence labels between every pair of RGB and IR images, most methods try to alleviate the variations with set-level alignment by reducing the distance between the entire RGB and IR sets. However, this set-level alignment may lead to misalignment of some instances, which limits the performance for RGB-IR Re-ID. Different from existing methods, in this paper, we propose to generate cross-modality paired-images and perform both global set-level and fine-grained instance-level alignments. Our proposed method enjoys several merits. First, our method can perform set-level alignment by disentangling modality-specific and modality-invariant features. Compared with conventional methods, ours can explicitly remove the modality-specific features and the modality variation can be better reduced. Second, given cross-modality unpaired-images of a person, our method can generate cross-modality paired images from exchanged images. With them, we can directly perform instance-level alignment by minimizing distances of every pair of images. Extensive experimental results on two standard benchmarks demonstrate that the proposed model favourably against state-of-the-art methods. Especially, on SYSU-MM01 dataset, our model can achieve a gain of 9.2% and 7.7% in terms of Rank-1 and mAP. Code is available at https://github.com/wangguanan/JSIA-ReID.
Guan'an Wang, Tianzhu Zhang 0001, Yang Yang 0062, Jian Cheng 0001, Jianlong Chang, Zeng-Guang Hou
AAAI3
2020 Horizontal Flipping Assisted Disentangled Feature Learning for Semi-supervised Person Re-identification
Gehan Hao, Yang Yang 0062, Guan'an Wang, Zhen Lei 0001
ACCV (3)2
2020 High-Order Information Matters: Learning Relation and Topology for Occluded Person Re-Identification
abstract
Occluded person re-identification (ReID) aims to match occluded person images to holistic ones across dis-joint cameras. In this paper, we propose a novel framework by learning high-order relation and topology information for discriminative features and robust alignment. At first, we use a CNN backbone to learn feature maps and key-points estimation model to extract semantic local features. Even so, occluded images still suffer from occlusion and outliers. Then, we view the extracted local features of an image as nodes of a graph and propose an adaptive direction graph convolutional (ADGC) layer to pass relation information between nodes. The proposed ADGC layer can automatically suppress the message passing of meaningless features by dynamically learning direction and degree of linkage. When aligning two groups of local features, we view it as a graph matching problem and propose a cross-graph embedded-alignment (CGEA) layer to joint learn and embed topology information to local features, and straightly predict similarity score. The proposed CGEA layer can both take full use of alignment learned by graph matching and replace sensitive one-to-one alignment with a robust soft one. Finally, extensive experiments on occluded, partial, and holistic ReID tasks show the effectiveness of our proposed method. Specifically, our framework significantly outperforms state-of-the-art by $6.5\%$ mAP scores on Occluded-Duke dataset.
Guan'an Wang, Shuo Yang 0002, Zhicheng Wang 0001, Yang Yang 0062, Shuliang Wang 0001, Gang Yu 0002, Erjin Zhou, Jian Sun 0001
CVPR5
2020 Towards Fast, Accurate and Stable 3D Dense Face Alignment
Jianzhu Guo, Xiangyu Zhu 0001, Yang Yang 0062, Fan Yang 0062, Zhen Lei 0001, Stan Z. Li
ECCV (19)3
2020 Attentive Hybrid Feature with Two-Step Fusion for Facial Expression Recognition
abstract
Facial expression recognition is inherently a challenging task, especially for the in-the-wild images with various occlusions and large pose variations, which may lead to the loss of some crucial information. To address it, in this paper, we propose an attentive hybrid architecture (AHA) which learns global, local and integrated features based on different face regions. Compared with one type of feature, our extracted features own complementary information and can reduce the loss of crucial information. Specifically, AHA contains three branches, where all sub-networks in those branches employ the attention mechanism to further localize the interested pixels/regions. Moreover, we propose a two-step fusion strategy based on LSTM to deeply explore the hidden correlations among different face regions. Extensive experiments on four popular expression databases (i.e., CK+, FER-2013, SFEW 2.0, RAF-DB) show the effectiveness of the proposed method.
Jun Weng, Yang Yang 0062, Zichang Tan, Zhen Lei 0001
ICPR2
2020 Generative Landmark Guided Face Inpainting
Yang Yang 0062, Xiaojie Guo 0001
PRCV (1)1
2020 Cross-modality paired-images generation and augmentation for RGB-infrared person re-identification
Guan'an Wang, Yang Yang 0062, Tianzhu Zhang 0001, Jian Cheng 0001, Zeng-Guang Hou, Prayag Tiwari, Hari Mohan Pandey
Neural Networks2
2020 An end-to-end exemplar association for unsupervised person Re-identification
Jinlin Wu, Yang Yang 0062, Zhen Lei 0001, Jinqiao Wang, Stan Z. Li, Prayag Tiwari, Hari Mohan Pandey
Neural Networks2
2019 Pedestrian Parsing by Joint Learning from Wholes and Parts
abstract
Parsing pedestrian into different semantic regions makes an important component in pedestrian analysis. In this paper, we propose a novel pedestrian parsing framework that explores the correlated and complementary information by jointly learning global and local features with Fully Connected Networks (FCNs). The proposed network contains two branches: A global-local branch and a main brahch. The global-local branch is learnt under both local and global supervisions, which can discover the discriminative and fine-grained features for assisting the feature learning of the main branch. Comparing with either global or local feature learning, our framework can capture both the contextual and detailed information well. Thus, the learned features are more reliable and discriminative. We also design a multi-scale atrous spatial pyramid pooling (multi-scale ASPP) to reduce the junction effect in local features especially in global-local fusion features. It also helps to capture image context at multiple scales thus parsing pedestrian robustly. We evaluate the proposed method on PPSS and PennFudan datasets and achieve the state-of-the-art performance.
Zhenting Gong, Yang Yang 0062
AVSS2
2019 3DMA: A Multi-modality 3D Mask Face Anti-spoofing Database
abstract
Benefiting from publicly available databases, face anti-spoofing has recently gained extensive attention in the academic community. However, most of the existing databases focus on the 2D object attacks, including photo and video attacks. The only two public 3D mask face anti-spoofing database are very small. In this paper, we release a multi-modality 3D mask face anti-spoofing database named 3DMA, which contains 920 videos of 67 genuine subjects wearing 48 kinds of 3D masks, captured in visual (VIS) and near-infrared (NIR) modalities. To simulate the real world scenarios, two illumination and four capturing distance settings are deployed during the collection process. To the best of our knowledge, the proposed database is currently the most extensive public database for 3D mask face anti-spoofing. Furthermore, we build three protocols for performance evaluation under different illumination conditions and distances. Experimental results with Convolutional Neural Network (CNN) and LBP-based methods reveal that our proposed 3DMA is indeed a challenge for face anti-spoofing. This database is available at http://www.cbsr.ia.ac.cn/english/3DMA.html. We hope our public 3DMA database can help to pave the way for further research on 3D mask face anti-spoofing.
Jinchuan Xiao, Yinhang Tang, Jianzhu Guo, Yang Yang 0062, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li
AVSS4
2019 RGB-Infrared Cross-Modality Person Re-Identification via Joint Pixel and Feature Alignment
abstract
RGB-Infrared (IR) person re-identification is an important and challenging task due to large cross-modality variations between RGB and IR images. Most conventional approaches aim to bridge the cross-modality gap with feature alignment by feature representation learning. Different from existing methods, in this paper, we propose a novel and end-to-end Alignment Generative Adversarial Network (AlignGAN) for the RGB-IR RE-ID task. The proposed model enjoys several merits. First, it can exploit pixel alignment and feature alignment jointly. To the best of our knowledge, this is the first work to model the two alignment strategies jointly for the RGB-IR RE-ID problem. Second, the proposed model consists of a pixel generator, a feature generator and a joint discriminator. By playing a min-max game among the three components, our model is able to not only alleviate the cross-modality and intra-modality variations, but also learn identity-consistent features. Extensive experimental results on two standard benchmarks demonstrate that the proposed model performs favourably against state-of-the-art methods. Especially, on SYSU-MM01 dataset, our model can achieve an absolute gain of 15.4% and 12.9% in terms of Rank-1 and mAP.
Guan'an Wang, Tianzhu Zhang 0001, Jian Cheng 0001, Si Liu 0001, Yang Yang 0062, Zeng-Guang Hou
ICCV5
2019 Unsupervised Graph Association for Person Re-Identification
abstract
In this paper, we propose an unsupervised graph association (UGA) framework to learn the underlying view-invariant representations from the video pedestrian tracklets. The core points of UGA are mining the underlying cross-view associations and reducing the damage of noise associations. To this end, UGA is adopts a two-stage training strategy: (1) intra-camera learning stage and (2) intercamera learning stage. The former learns the intra-camera representation for each camera. While the latter builds a cross-view graph (CVG) to associate different cameras. By doing this, we can learn view-invariant representation for all person. Extensive experiments and ablation studies on seven re-id datasets demonstrate the superiority of the proposed UGA over most state-of-the-art unsupervised and domain adaptation re-id methods.
Jinlin Wu, Yang Yang 0062, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
ICCV3
2019 Clustering and Dynamic Sampling Based Unsupervised Domain Adaptation for Person Re-Identification
abstract
Person Re-Identification (Re-ID) has witnessed great improvements due to the advances of the deep convolutional neural networks (CNN). Despite this, existing methods mainly suffer from the poor generalization ability to unseen scenes because of the different characteristics between different domains. To address this issue, a Clustering and Dynamic Sampling (CDS) method is proposed in this paper, which tries to transfer the useful knowledge of existing labeled source domain to the unlabeled target one. Specifically, to improve the discriminability of CNN model on source domain, we use the commonly shared pedestrian attributes (e.g., gender, hat and clothing color etc.) to enrich the information and resort to the margin-based softmax (e.g., A-Softmax) loss to train the model. For the unlabeled target domain, we iteratively cluster the samples into several centers and dynamically select informative ones from each center to fine-tune the source-domain model. Extensive experiments on DukeMTMC-reID and Market-1501 datasets show that the proposed method greatly improves the state of the arts in unsupervised domain adaptation.
Jinlin Wu, Shengcai Liao, Zhen Lei 0001, Xiaobo Wang 0001, Yang Yang 0062, Stan Z. Li
ICME5
2019 Deeply-learned Hybrid Representations for Facial Age Estimation
abstract
In this paper, we propose a novel unified network named Deep Hybrid-Aligned Architecture for facial age estimation. It contains global, local and global-local branches. They are jointly optimized and thus can capture multiple types of features with complementary information. In each branch, we employ a separate loss for each sub-network to extract the independent features and use a recurrent fusion to explore correlations among those region features. Considering that the pose variations may lead to misalignment in different regions, we design an Aligned Region Pooling operation to generate aligned region features. Moreover, a new large age dataset named Web-FaceAge owning more than 120K samples is collected under diverse scenes and spanning a large age range. Experiments on five age benchmark datasets, including Web-FaceAge, Morph, FG-NET, CACD and Chalearn LAP 2015, show that the proposed method outperforms the state-of-the-art approaches significantly.
Zichang Tan, Yang Yang 0062, Jun Wan 0001, Guodong Guo, Stan Z. Li
IJCAI2
2019 Color-Sensitive Person Re-Identification
abstract
Recent deep Re-ID models mainly focus on learning high-level semantic features, while failing to explicitly explore color information which is one of the most important cues for person Re-ID. In this paper, we propose a novel Color-Sensitive Re-ID to take full advantage of color information. On one hand, we train our model with real and fake images. By using the extra fake images, more color information can be exploited and it can avoid overfitting during training. On the other hand, we also train our model with images of the same person with different colors. By doing so, features can be forced to focus on the color difference in regions. To generate fake images with specified colors, we propose a novel Color Translation GAN (CTGAN) to learn mappings between different clothing colors and preserve identity consistency among the same clothing color. Extensive evaluations on two benchmark datasets show that our approach significantly outperforms state-of-the-art Re-ID models.
Guan'an Wang, Yang Yang 0062, Jian Cheng 0001, Jinqiao Wang, Zeng-Guang Hou
IJCAI2
2019 Attention-Based Pedestrian Attribute Analysis
abstract
Recognizing the pedestrian attributes in surveillance scenes is an inherently challenging task, especially for the pedestrian images with large pose variations, complex backgrounds, and various camera viewing angles. To select important and discriminative regions or pixels against the variations, three attention mechanisms are proposed, including parsing attention, label attention, and spatial attention. Those attentions aim at accessing effective information by considering problems from different perspectives. To be specific, the parsing attention extracts discriminative features by learning not only where to turn attention to but also how to aggregate features from different semantic regions of human bodies, e.g., head and upper body. The label attention aims at targetedly collecting the discriminative features for each attribute. Different from the parsing and label attention mechanisms, the spatial attention considers the problem from a global perspective, aiming at selecting several important and discriminative image regions or pixels for all attributes. Then, we propose a joint learning framework formulated in a multi-task-like way with these three attention mechanisms learned concurrently to extract complementary and correlated features. This joint learning framework is named Joint Learning of Parsing attention, Label attention, and Spatial attention for Pedestrian Attributes Analysis (JLPLS-PAA, for short). Extensive comparative evaluations conducted on multiple large-scale benchmarks, including PA-100K, RAP, PETA, Market-1501, and Duke attribute datasets, further demonstrate the effectiveness of the proposed JLPLS-PAA framework for pedestrian attribute analysis.
Zichang Tan, Yang Yang 0062, Jun Wan 0001, Hanyuan Hang, Guodong Guo, Stan Z. Li
IEEE Trans. Image Process.2
2018 Is Re-ranking Useful for Open-set Person Re-identification?
abstract
Re-ranking algorithms can often boost the performance of close-set person re-identification. However, limited efforts have been devoted to answering whether a similar conclusion could be derived on open-set person re-identification. Considering that open-set scenario is more practical in real applications, in this paper, we try to answer this question and do a benchmark study of re-ranking on open-set person re-identification. Specifically, we evaluate three feature descriptors, namely MB-LBP, LOMO, and IDE, and four distance metrics, namely Euclidean, Cosine, RRDA, and XQDA, with their combinations as baseline algorithms. Then, we evaluate four popular re-ranking algorithms, including k-reciprocal Encoding, ECN-3, ECN-4, and DaF. Through extensive benchmark studies on the OPeRIDv1.0 dataset, the results show that re-ranking algorithms, though useful for closed-set person re-identification, are not generally effective for the open-set person re-identification problem. We argue that this is because re-ranking algorithms change the score distributions per query, and hence disrupt the FAR estimation across all queries. Accordingly, we propose to align the re-ranking scores to the original score via the min-max normalization, which verifies our hypothesis above.
Hongsheng Wang, Shengcai Liao, Zhen Lei 0001, Yang Yang 0062
IEEE BigData4
2018 Dependence-Aware Feature Coding for Person Re-Identification
abstract
In this letter, we focus on how to boost the performance of person re-identification by exploring the discriminative information among person pairs. A novel dependence-aware feature coding framework is proposed for this task. Specifically, we employ the Hilbert–Schmidt independence criterion as the discriminative term, which is to explore the dependence between different kinds of person pairs, i.e., the same person pairs should be dependence maximized, while the different ones should be dependence minimized. Theoretical discussion and analysis on the convexity of the proposed constraint, as well as the convergence of our algorithm, are provided. Experimental results on two benchmark datasets have demonstrated the advantages of our method over the state-of-the-art alternatives.
Xiaobo Wang 0001, Zhen Lei 0001, Shengcai Liao, Xiaojie Guo 0001, Yang Yang 0062, Stan Z. Li
IEEE Signal Process. Lett.5
2017 Unsupervised Learning of Multi-Level Descriptors for Person Re-Identification
abstract
In this paper, we propose a novel coding method named weighted linear coding (WLC) to learn multi-level (e.g., pixel-level, patch-level and image-level) descriptors from raw pixel data in an unsupervised manner. It guarantees the property of saliency with a similarity constraint. The resulting multi-level descriptors have a good balance between the robustness and distinctiveness. Based on WLC, all data from the same region can be jointly encoded. Consequently, when we extract the holistic image features, it is able to preserve the spatial consistency. Furthermore, we apply PCA to these features and compact person representations are then achieved. During the stage of matching persons, we exploit the complementary information resided in multi-level descriptors via a score-level fusion strategy. Experiments on the challenging person re-identification datasets - VIPeR and CUHK 01, demonstrate the effectiveness of our method.
Yang Yang 0062, Longyin Wen, Siwei Lyu, Stan Z. Li
AAAI1
2017 Efficient similarity learning for asymmetric hashing
abstract
Hashing techniques with asymmetric schemes (e.g., only bi-narizing the database points) have recently attracted wide attention in the circle of image retrieval. In comparison with those methods which binarize simultaneously both of the query and database points, they not only enjoy the storage and search efficiencies, but also provide higher accuracy. Gearing to this line, this paper proposes a metric-embedded asymmetric hashing (MEAH) that learns jointly a bilinear similarity measure and binary codes of database points in an unsupervised manner. Technically, the learned similarity measure is able to bridge the gap between the binary codes and the real-valued codes, which are represented possibly with different dimensions. What is more, this measure is capable of preserving the global structure hidden in the database. Extensive experiments on two public image benchmarks demonstrate the superiority of our approach over the several state-of-the-art unsupervised hashing methods.
Cheng Da, Yang Yang 0062, Chunlei Huo, Shiming Xiang, Chunhong Pan
ICIP2
2016 Large Scale Similarity Learning Using Similar Pairs for Person Verification
abstract
In this paper, we propose a novel similarity measure and then introduce an efficient strategy to learn it by using only similar pairs for person verification. Unlike existing metric learning methods, we consider both the difference and commonness of an image pair to increase its discriminativeness. Under a pairconstrained Gaussian assumption, we show how to obtain the Gaussian priors (i.e., corresponding covariance matrices) of dissimilar pairs from those of similar pairs. The application of a log likelihood ratio makes the learning process simple and fast and thus scalable to large datasets. Additionally, our method is able to handle heterogeneous data well. Results on the challenging datasets of face verification (LFW and Pub-Fig) and person re-identification (VIPeR) show that our algorithm outperforms the state-of-the-art methods.
Yang Yang 0062, Shengcai Liao, Zhen Lei 0001, Stan Z. Li
AAAI1
2016 Metric Embedded Discriminative Vocabulary Learning for High-Level Person Representation
abstract
A variety of encoding methods for bag of word (BoW) model have been proposed to encode the local features in image classification. However, most of them are unsupervised and just employ k-means to form the visual vocabulary, thus reducing the discriminative power of the features. In this paper, we propose a metric embedded discriminative vocabulary learning for high-level person representation with application to person re-identification. A new and effective term is introduced which aims at making the same persons closer while different ones farther in the metric space. With the learned vocabulary, we utilize a linear coding method to encode the image-level features (or holistic image features) for extracting high-level person representation. Different from traditional unsupervised approaches, our method can explore the relationship(same or not) among the persons. Since there is an analytic solution to the linear coding, it is easy to obtain the final high-level features. The experimental results on person re-identification demonstrate the effectiveness of our proposed algorithm.
Yang Yang 0062, Zhen Lei 0001, Hailin Shi, Stan Z. Li
AAAI1
2016 Embedding Deep Metric for Person Re-identification: A Study Against Large Variations
Hailin Shi, Yang Yang 0062, Xiangyu Zhu 0001, Shengcai Liao, Zhen Lei 0001, Wei-Shi Zheng 0001, Stan Z. Li
ECCV (1)2
2016 Spatial Constrained Fine-Grained Color Name for Person Re-identification
Yang Yang 0062, Yuhong Yang 0001, Mang Ye, Wenxin Huang, Zheng Wang 0007, Chao Liang 0001, Chunjie Zhang 0001
MMM (1)1
2014 Stacked Deformable Part Model with Shape Regression for Object Part Localization
Zhen Lei 0001, Yang Yang 0062, Stan Z. Li
ECCV (2)3
2014 Salient Color Names for Person Re-identification
Yang Yang 0062, Jimei Yang, Shengcai Liao, Dong Yi, Stan Z. Li
ECCV (1)1
2014 Color Models and Weighted Covariance Estimation for Person Re-identification
abstract
Due to illumination changes, partial occlusions, and object scale differences, person re-identification over disjoint camera views becomes a challenging problem. To address this problem, a variety of image representations have been put forward. In this paper, the illumination invariance and distinctiveness of different color models including the proposed color model are firstly evaluated. Since color distribution is robust to image scales and partial occlusions, color distributions based on different color models are then calculated and fused in the stage of feature extraction. Different color models obtain robustness to different types of illumination and thus fusing them can compensate each other and contribute to better performance. In the stage of feature matching, a weighted KISSME is presented to learn a better distance metric than the original KISSME. Experimental results demonstrate its feasibility and effectiveness. Finally, image pairs are matched based on the learned distance metric. Experiments conducted on two public benchmark datasets (VIPeR and PRID 450S) show that the proposed algorithm outperforms the state-of-the-art methods.
Yang Yang 0062, Shengcai Liao, Zhen Lei 0001, Dong Yi, Stan Z. Li
ICPR1
2014 Low-rank representation based action recognition
abstract
Human action recognition is an important problem in computer vision, which has been applied to many applications. However, how to learn an accurate and discriminative representation of videos based on the features extracted from videos still remains to be a challenging problem. In this paper, we propose a novel method named low-rank representation based action recognition to recognize human actions. Given a dictionary, low-rank representation aims at finding the lowest-rank representation of all data, which can capture the global data structures. According to its characteristics, low-rank representation is robust against noises. Experimental results demonstrate the effectiveness of the proposed approach on several publicly available datasets.
Xiangrong Zhang, Yang Yang 0062, Hanghua Jia, Huiyu Zhou 0001, Licheng Jiao
IJCNN2
2014 Laplacian group sparse modeling of human actions
Xiangrong Zhang, Licheng Jiao, Yang Yang 0062, Feng Dong 0005
Pattern Recognit.4
2013 Manifold-constrained coding and sparse representation for human action recognition
Xiangrong Zhang, Yang Yang 0062, Licheng Jiao, Feng Dong 0005
Pattern Recognit.2