VLDB 2026 Research / reviewers in the wild / expert
Hao Yu 0015
dblp:64/4832-15
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-8298-7181ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CROMBO: Cross-Modality Bootstrapping for Unified Sketch-Photo Representation LearningabstractSketch–photo recognition refers to matching hand-drawn sketches with their corresponding photos, where the performance essentially depends on how well the representations of the two modalities are aligned in the feature spaces. Existing works bluntly force models to reduce the representation discrepancy between the modalities, making the learning less effective. Besides, the current symmetric feature extraction framework prefers the photo modality for richer information while neglecting the sketch modality. Driven by these observations, we argue that, instead of forcefully wiping out the modality discrepancy, we may utilize the discrepancy to enhance model learning. Thus, we propose a Cross-Modality Bootstrapping learning framework (CROMBO) that utilizes the modality discrepancy to bootstrap cross-modality representation learning via a differentiated interaction manner. Specifically, we first present a Sketch Implicit Bootstrapping (SIB) module to magnify the recognizable elements in the photo modality by utilizing the characteristic of sketches having only contours and key details. Second, a Photo-driven Sketch Refinement (PSR) module is developed to guide the sketch representation in the shared feature extraction process by supplementing rich information from the photo modality. Moreover, we design a second-order alignment strategy to dynamically align the latent distribution of two modalities in a Hilbert space. Also, our CROMBO can learn fewer parameters by freezing the weights of shallow layers in the backbone while making no sacrifice in performance. Extensive experiments on six public datasets verify the superior performance of our CROMBO for sketch–photo-based tasks, such as sketch re-identification (Re-ID), sketch–photo face recognition, and sketch-based image retrieval. Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | FreeNet: Liberating Depth-Wise Separable Operations for Building Faster Mobile Vision ArchitecturesabstractIn the pursuit of efficient vision architectures, substantial efforts have been devoted to optimizing operator efficiency. Depth-wise separable operators, such as DWConv, are found cheap in both FLOPs and parameters. As a result, they are increasingly incorporated into efficient backbones, trading for deeper and wider architectures to enhance performance. However, separable operators are not really fast on devices due to the discontinuous memory access requirements. In this paper, we propose FreeNets, a family of simple and efficient backbones that free the separable operation to further accelerate the running speed. We introduce sparse sampling mixers (S2-Mixer) to supersede existing separable token mixers. The S2-Mixer samples multiple segments of partially continuous signals across spatial and channel dimensions for convolutional processing, achieving extremely fast on-device speed. The sparse sampling also enables S2-Mixer to capture long-range pixel relationships from dynamic receptive fields. Furthermore, we introduce a Shift Feed-Forward Network (ShiftFFN) as a faster alternative to existing channel mixers. It utilizes a shift neck architecture that aggregates global information to shift features, enabling faster channel mixing while incorporating global pixel information. Extensive experiments demonstrate that FreeNet offers a superior accuracy-efficiency tradeoff compared to the latest efficient models. On ImageNet-1k, FreeNet-S2 outperforms the StarNet-S4 by 0.4% in top-1 accuracy, while running around 40% faster on desktop GPU and 15% faster on Mobile GPU. Hao Yu 0015, Haoyu Chen 0001, Wei Peng 0009, Xu Cheng 0003, Guoying Zhao 0001 |
AAAI | 1 |
| 2025 | From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-IdentificationabstractAiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities, raising privacy and ownership concerns that make existing centralized training impractical for VI-ReID. To tackle these challenges, we propose L2RW, a benchmark that brings VI-ReID closer to real-world applications. The rationale of L2RW is that integrating decentralized training into VI-ReID can address privacy concerns in scenarios with limited data-sharing regulation. Specifically, we design protocols and corresponding algorithms for different privacy sensitivity levels. In our new benchmark, we ensure the model training is done in the conditions that: 1) data from each camera remains completely isolated, or 2) different data entities (e.g., data controllers of a certain region) can selectively share the data. In this way, we simulate scenarios with strict privacy constraints which is closer to real-world conditions. Intensive experiments with various server-side federated algorithms are conducted, showing the feasibility of decentralized VI-ReID training. Notably, when evaluated in unseen domains (i.e., new data entities), our L2RW, trained with isolated data (privacy-preserved), achieves performance comparable to SOTAs trained with shared data (privacy-unrestricted). We hope this work offers a novel research entry for deploying VI-ReID that fits real-world scenarios and can benefit the community. Hao Yu 0015, Xu Cheng 0003, Haoyu Chen 0001, Zhaodong Sun, Guoying Zhao 0001 |
CVPR | 2 |
| 2025 | Learning Binary-Antithetical Information Bottleneck for Generalizable Face Anti-SpoofingabstractWe investigate generalizable face anti-spoofing (FAS) using information bottleneck theory. As generalizable FAS aims to detect spoofing in unseen scenarios, it has recently gained significant attention. Existing methods often use adversarial strategies or auxiliary modules to learn domain-invariant features by mining data relationships from distinct source domains. However, their learned feature space may still shift for unseen data due to the spurious correlations overfitted from training domains. Our rationale is that the problem of generalized pattern learning in FAS can be framed as a unified binary-antithetical information transition process, grounded in information bottleneck theory. Specifically, we leverage mutual-information optimization to preserve the instance-level spoof-aware information while compressing domain-related information modeled from the antithetical identity distribution. This enables the model to dynamically identify domain-agnostic, minimal sufficient representations that consistently describe the live/spoof distributions while mitigating spurious correlations through cross-identity compression. In light of this, we propose a novel learning framework for FAS, named Binary-Antithetical Information Bottleneck (BIB)-FAS, which is proven to be effectively generalized to unseen scenarios without using auxiliary information (e.g., domain labels) for training. Extensive cross-domain evaluations show that BIB-FAS significantly outperforms state-of-the-art methods. The code is available at: github.com/CV-AC/BIB-FAS. Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
ICASSP | 1 |
| 2025 | DMANet: Dual-modality alignment network for visible-infrared person re-identificationabstractVisible–infrared person re-identification (VI-ReID) is a challenging retrieval task, which aims to match the same pedestrian between visible and infrared modalities. Most existing works achieve performance gains by solving the problem of the inherent cross-modality discrepancies. However, they cannot fully mine the modality information and lead to a poor generalization. In addition, the pedestrian images are unable to align well due to the large inter- and intra- class variations. To tackle the above limitations, we propose a novel dual-modality alignment network (DMANet) for VI-ReID. The core idea of our work is to develop multi-granularity features mutual learning (MGFML) for inadequate perception of modalities information, and to solve modality difference by proposing inter- and intra- modality alignment module (IIMA). Specifically, firstly, an effective multi-granularity features mutual learning module is proposed to mine the multi-granularity features, which combines the domain alignment and self-distillation to relieve modality discrepancy. Further, the maximum mean discrepancy loss and mutual learning loss are presented to enhance the identity-aware ability of the DMANet. Secondly, an effective inter- and intra- modality alignment module is presented to explore the potential alignment relation of inter- and intra- modalities. Finally, joint learning mechanism of multi-granularity features and modality alignment is utilized to improve the VI-ReID accuracy. Extensive experiments on mainstream benchmarks demonstrate that our method is superior to the state-of-the-art methods. Xu Cheng 0003, Shuya Deng, Hao Yu 0015, Guoying Zhao 0001 |
Pattern Recognit. | 3 |
| 2025 | MDANet: Modality-Aware Domain Alignment Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification is a challenging task in video surveillance. Most existing works achieve performance gains by aligning feature distributions or image styles across modalities, whereas the multi-granularity information and domain knowledge are usually neglected. Motivated by these issues, we propose a novel modality-aware domain alignment network (MDANet) for visible-infrared person re-identification (VI-ReID), which utilizes global-local context cues and the generalized domain alignment strategy to solve modal differences and poor generalization. Firstly, modality-aware global-local context attention (MGLCA) is proposed to obtain multi-granularity context features and identity-aware patterns. Secondly, we present a generalized domain alignment learning head (GDALH) to relieve the modality discrepancy and enhance the generalization of MDANet, whose core idea is to enrich feature diversity in the domain alignment procedure. Finally, the entire network model is trained by proposing cross-modality circle, classification, and domain alignment losses in an end-to-end fashion. We conduct comprehensive experiments on two standards and their corrupted VI-ReID datasets to validate the robustness and generalization of our approach. MDANet is obviously superior to the most state-of-the-art methods. Specifically, the proposed method can gain 8.86% and 2.50% in Rank-1 accuracy on SYSU-MM01 (all-search and single-shot mode) and RegDB (infrared to visible mode) datasets, respectively. The source code will be made available soon. Xu Cheng 0003, Hao Yu 0015, Kevin H. M. Cheng, Zitong Yu, Guoying Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | DSAF: Dual Space Alignment Framework for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match visible and infrared pedestrian images across non-overlapped cameras. However, we observe that three crucial challenges remain inadequately addressed by existing methods: (i) limited discriminative capacity for modality-shared representation, (ii) modality misalignment, and (iii) neglect of identity consistency knowledge. To solve the above issues, we propose a novel dual space alignment framework (DSAF) to constrain the modality in two specific spaces. Specifically, for (i), we design a lightweight and plug-and-play modality invariant enhancement (MIE) module to capture fine-grained semantic information and render identity discriminative. This facilitates the establishment of correlations between visible and infrared modalities, enabling the model to learn robust modality-shared features. To tackle (ii), a dual space alignment (DSA) is introduced to conduct the pixel-level alignment in both Euclidean space and Hilbert space. DSA establishes an elastic relationship between these two spaces, remaining invariant knowledge across two spaces. To solve (iii), we propose an adaptive identity-consistent learning (AIL) to discover identity-consistent knowledge between visible and infrared modalities in a dynamic manner. Extensive experiments on mainstream VI-ReID benchmarks show the superiority and flexibility of our proposed method, achieving competitive performance on mainstream datasets. Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Dual-Path Imbalanced Feature Compensation Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) presents significant challenges on account of the substantial cross-modality gap and intra-class variations. Most existing methods primarily concentrate on aligning cross-modality at the feature or image levels and training with an equal number of samples from different modalities. However, in the real world, there exists an issue of modality imbalance between visible and infrared data. Besides, imbalanced samples between train and test impact the robustness and generalization of the VI-ReID. To alleviate this problem, we propose a dual-path imbalanced feature compensation network (DICNet) for VI-ReID, which provides equal opportunities for each modality to learn inconsistent information from different identities of others, enhancing identity discrimination performance and generalization. First, a modality consistency perception (MCP) module is designed to assist the backbone focus on spatial and channel information, extracting diverse and salient features to enhance feature representation. Second, we propose a cross-modality features re-assignment strategy to simulate modality imbalance by grouping and re-organizing the cross-modality features. Third, we perform bidirectional heterogeneous cooperative compensation with cross-modality imbalanced feature interaction modules (CIFIMs), allowing our network to explore the identity-aware patterns from imbalanced features of multiple groups for cross-modality interaction and fusion. Further, we design a feature re-construction difference loss to reduce cross-modality discrepancy and enrich feature diversity within each modality. Extensive experiments on three mainstream datasets show the superiority of the DICNet. Additionally, competitive results in corrupted scenarios verify its generalization and robustness. Xu Cheng 0003, Hao Yu 0015, Jingang Shi, Zitong Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Differentiable Auxiliary Learning for Sketch Re-IdentificationabstractSketch re-identification (Re-ID) seeks to match pedestrians' photos from surveillance videos with corresponding sketches. However, we observe that existing works still have two critical limitations: (i) cross- and intra-modality discrepancies hinder the extraction of modality-shared features, (ii) standard triplet loss fails to constrain latent feature distribution in each modality with inadequate samples. To overcome the above issues, we propose a differentiable auxiliary learning network (DALNet) to explore a robust auxiliary modality for Sketch Re-ID. Specifically, for (i) we construct an auxiliary modality by using a dynamic auxiliary generator (DAG) to bridge the gap between sketch and photo modalities. The auxiliary modality highlights the described person in photos to mitigate background clutter and learns sketch style through style refinement. Moreover, a modality interactive attention module (MIA) is presented to align the features and learn the invariant patterns of two modalities by auxiliary modality. To address (ii), we propose a multi-modality collaborative learning scheme (MMCL) to align the latent distribution of three modalities. An intra-modality circle loss in MMCL brings learned global and modality-shared features of the same identity closer in the case of insufficient samples within each modality. Extensive experiments verify the superior performance of our DALNet over the state-of-the-art methods for Sketch Re-ID, and the generalization in sketch-based image retrieval and sketch-photo face recognition tasks. Xu Cheng 0003, Haoyu Chen 0001, Hao Yu 0015, Guoying Zhao 0001 |
AAAI | 4 |
| 2024 | Domain Shifting: A Generalized Solution for Heterogeneous Cross-Modality Person Re-Identification
Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
ECCV (72) | 3 |
| 2024 | Exploring modality enhancement and compensation spaces for visible-infrared person re-identification
Xu Cheng 0003, Shuya Deng, Hao Yu 0015 |
Image Vis. Comput. | 3 |
| 2024 | Discovering attention-guided cross-modality correlation for visible-infrared person re-identification
Hao Yu 0015, Xu Cheng 0003, Kevin H. M. Cheng, Wei Peng 0009, Zitong Yu, Guoying Zhao 0001 |
Pattern Recognit. | 1 |
| 2024 | ST-Phys: Unsupervised Spatio-Temporal Contrastive Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) is a non-contact method that employs facial videos for measuring physiological parameters. Existing rPPG methods have achieved remarkable performance. However, the success mainly profits from supervised learning over massive labeled data. On the other hand, existing unsupervised rPPG methods fail to fully utilize spatio-temporal features and encounter challenges in low-light or noise environments. To address these problems, we propose an unsupervised contrast learning approach, ST-Phys. We incorporate a low-light enhancement module, a temporal dilated module, and a spatial enhanced module to better deal with long-term dependencies under the random low-light conditions. In addition, we design a circular margin loss, wherein rPPG signals originating from identical videos are attracted, while those from distinct videos are repelled. Our method is assessed on six openly accessible datasets, including RGB and NIR videos. Extensive experiments reveal the superior performance of our proposed ST-Phys over state-of-the-art unsupervised rPPG methods. Moreover, it offers advantages in parameter reduction and noise robustness. Mingyue Cao, Xu Cheng 0003, Hao Yu 0015, Jingang Shi |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | TOPLight: Lightweight Neural Networks with Task-Oriented Pretraining for Visible-Infrared RecognitionabstractVisible-infrared recognition (VI recognition) is a challenging task due to the enormous visual difference across heterogeneous images. Most existing works achieve promising results by transfer learning, such as pretraining on the ImageNet, based on advanced neural architectures like ResNet and ViT. However, such methods ignore the neg-ative influence of the pretrained colour prior knowledge, as well as their heavy computational burden makes them hard to deploy in actual scenarios with limited resources. In this paper, we propose a novel task-oriented pretrained lightweight neural network (TOPLight) for VI recognition. Specifically, the TOPLight method simulates the domain conflict and sample variations with the proposed fake do-main loss in the pretraining stage, which guides the network to learn how to handle those difficulties, such that a more general modality-shared feature representation is learned for the heterogeneous images. Moreover, an effective fine-grained dependency reconstruction module (FDR) is developed to discover substantial pattern dependencies shared in two modalities. Extensive experiments on VI person re-identification and VI face recognition datasets demonstrate the superiority of the proposed TOPLight, which signifi-cantly outperforms the current state of the arts while de-manding fewer computational resources. Hao Yu 0015, Xu Cheng 0003, Wei Peng 0009 |
CVPR | 1 |
| 2023 | Modality Unifying Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging task due to large cross-modality discrepancies and intra-class variations. Existing methods mainly focus on learning modality-shared representations by embedding different modalities into the same feature space. As a result, the learned feature emphasizes the common patterns across modalities while suppressing modality-specific and identity-aware information that is valuable for Re-ID. To address these issues, we propose a novel Modality Unifying Network (MUN) to explore a robust auxiliary modality for VI-ReID. First, the auxiliary modality is generated by combining the proposed cross-modality learner and intra-modality learner, which can dynamically model the modality-specific and modality-shared representations to alleviate both cross-modality and intra-modality variations. Second, by aligning identity centres across the three modalities, an identity alignment loss function is proposed to discover the discriminative feature representations. Third, a modality alignment loss is introduced to consistently reduce the distribution distance of visible and infrared images by modality prototype modeling. Extensive experiments on multiple public datasets demonstrate that the proposed method surpasses the current state-of-the-art methods by a significant margin. Hao Yu 0015, Xu Cheng 0003, Wei Peng 0009, Guoying Zhao 0001 |
ICCV | 1 |