EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyuan Yu
dblp:24/8009
· DBLP profile ↗
21ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Offset-corrected query generation strategies for cross-modality misalignment in 3D object detection: aligning LiDAR and camera
Chak-Fong Cheang, Xiaoyuan Yu, Suigu Tang, Zhaolong Du, Qianxiang Cheng |
Neurocomputing | 3 |
| 2025 | Simplifying complexity: a double-phase detection algorithm for defects of injection molded parts within the limited computer source
Wei Xie 0014, Haorui Wu, Haoming Liang, Langwen Zhang, Xiaoyuan Yu |
Multim. Syst. | 5 |
| 2024 | Rethinking Domain Adaptation and Generalization in the ERA Of ClipabstractIn recent studies on domain adaptation, significant emphasis has been placed on the advancement of learning shared knowledge from a source domain to a target domain. Recently, the large vision-language pre-trained model (i.e., CLIP) has shown strong ability on zero-shot recognition, and parameter efficient tuning can further improve its performance on specific tasks. This work demonstrates that a simple domain prior boosts CLIP’s zero-shot recognition in a specific domain. Besides, CLIP’s adaptation relies less on source domain data due to its diverse pre-training dataset. Furthermore, we create a benchmark for zero-shot adaptation and pseudo-labeling based self-training with CLIP. Last but not least, we propose to improve the task generalization ability of CLIP from multiple unlabeled domains, which is a more practical and unique scenario. We believe our findings motivate a rethinking of domain adaptation benchmarks and the associated role of related algorithms in the era of CLIP. Ruoyu Feng 0001, Tao Yu 0012, Xin Jin 0014, Xiaoyuan Yu, Zhibo Chen 0001 |
ICIP | 4 |
| 2024 | FM-CLIP: Flexible Modal CLIP for Face Anti-SpoofingabstractIn this work, borrowing a solution from the large-scale vision-language models (VLMs) instead of directly removing modality-specific signals from visual features, we propose a novel Flexible Modal CLIP (FM-CLIP) for flexible modal FAS, that can utilize text features to dynamically adjust visual features to be modality independent. In the visual branch, considering the huge visual differences of the same attack in different modalities, which makes it difficult for classifiers to flexibly identify subtle spoofing clues in different test modalities, we propose Cross-Modal Spoofing Enhancer (CMS-Enhancer). It includes a Frequency Extractor (FE) and Cross-Modal Interactor (CMI), aiming to map different modal attacks in a shared frequency space to reduce interference from modality-specific signals and enhance spoofing clues by leveraging cross-modal learning from the shared frequency space. In the text branch, we introduce a Language-Guided Patch Alignment (LGPA) based on prompt learning, which further guides the image encoder to focus on patch-level spoofing representations through dynamic weighting by text features. Thus, our FM-CLIP can flexibly test different modal samples by identifying and enhancing modality-agnostic spoofing cues. Finally, extensive experiments show that FM-CLIP is effective and outperforms state-of-the-art methods on multiple multi-modal datasets. Ajian Liu 0001, Hui Ma 0018, Junze Zheng, Haocheng Yuan, Xiaoyuan Yu, Yanyan Liang 0001, Sergio Escalera, Jun Wan 0001, Zhen Lei 0001 |
ACM Multimedia | 5 |
| 2024 | MBA-Net: multi-branch attention network for occluded person re-identification
Xing Hong, Langwen Zhang, Xiaoyuan Yu, Wei Xie 0014, Yumin Xie |
Multim. Tools Appl. | 3 |
| 2024 | Feature aggregation and modulation network for single image dehazing
Xiaoyuan Yu, Baoquan Ai, Fengguo Li |
Multim. Tools Appl. | 2 |
| 2024 | Local Patch AutoAugment With Multi-Agent CollaborationabstractData augmentation (DA) plays a critical role in improving the generalization of deep learning models. Recent works on automatically searching for DA policies from data have achieved great success. However, existing automated DA methods generally perform the search at the image level, which limits the exploration of diversity in local regions. In this paper, we propose a more fine-grained automated DA approach, dubbed Patch AutoAugment, to divide an image into a grid of patches and search for the joint optimal augmentation policies for the patches. We formulate it as a multi-agent reinforcement learning (MARL) problem, where each agent learns an augmentation policy for each patch based on its content together with the semantics of the whole image. The agents cooperate with each other to achieve the optimal augmentation effect of the entire image by sharing a team reward. We show the effectiveness of our method on multiple benchmark datasets of image classification, fine-grained image recognition and object detection (e.g., CIFAR-10, CIFAR-100, ImageNet, CUB-200-2011, Stanford Cars, FGVC-Aircraft and Pascal VOC 2007). Extensive experiments demonstrate that our method outperforms the state-of-the-art DA methods while requiring fewer computational resources. Shiqi Lin, Tao Yu 0012, Ruoyu Feng 0001, Xin Li 0082, Xiaoyuan Yu, Zhibo Chen 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Multimodal High-order Relation Transformer for Scene Boundary DetectionabstractScene boundary detection breaks down long videos into meaningful story-telling units and plays a crucial role in high-level video understanding. Despite significant advancements in this area, this task remains a challenging problem as it requires a comprehensive understanding of multimodal cues and high-level semantics. To tackle this issue, we propose a multimodal high-order relation transformer, which integrates a high-order encoder and an adaptive decoder in a unified framework. By modeling the mul-timodal cues and exploring similarities between the shots, the encoder is capable of capturing high-order relations between shots and extracting shot features with context semantics. By clustering the shots adaptively, the decoder can discover more universal switch pattern between successive scenes, thus helping scene boundary detection. Extensive experimental results on three standard benchmarks demonstrate that the proposed model performs favorably against state-of-the-art video scene detection methods. Zhangxiang Shi, Tianzhu Zhang 0001, Xiaoyuan Yu |
ICCV | 4 |
| 2023 | Adaptive fractional differential algorithm for image edge enhancement and texture preserve using fuzzy setsabstractAbstract This paper uses a fuzzy set scheme to present an adaptive fractional differential algorithm for image edge enhancement and texture preservation. In the proposed algorithm, an image's membership function and area feature are used to calculate the fuzzy set of images. The function of adaptive fractional differential order (FAFDO) can be constructed by making the linear transformation of the fuzzy set. Then, the fuzzy adaptive fractional differential mask (FAFDM) is obtained by substituting the FAFDO into the fractional differential mask. Finally, the image edge and texture are enhanced and preserved by applying airspace filtering of the FAFDM convolution. The experimental results show that, compared to fractional differential or fuzzy set‐based image enhancement algorithms, the proposed algorithm can adaptively enhance the image edge and preserve the image texture by analysing the fuzziness of the image itself. Wei Xie 0014, Langwen Zhang, Xiaoyuan Yu |
IET Image Process. | 4 |
| 2023 | From low to high: cascade network for restoring low-resolution face image via extracting and transforming edge feature
Xiaoyuan Yu, Wei Xie 0014, Langwen Zhang |
Multim. Tools Appl. | 1 |
| 2023 | A two-stage chaotic encryption algorithm for color face image based on circular diffusion
Jinwei Yu, Xiaoyuan Yu, Langwen Zhang, Wei Xie 0014 |
Multim. Tools Appl. | 2 |
| 2022 | TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionabstractFew-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarity between videos. Recently, it has been observed that directly measuring this similarity is not ideal since different action instances may show distinctive temporal distribution, resulting in severe misalignment issues across query and support videos. In this paper, we arrest this problem from two distinct aspects -- action duration misalignment and action evolution misalignment. We address them sequentially through a Two-stage Action Alignment Network (TA2N). The first stage locates the action by learning a temporal affine transform, which warps each video feature to its action duration while dismissing the action-irrelevant feature (e.g. background). Next, the second stage coordinates query feature to match the spatial-temporal action evolution of support by performing temporally rearrange and spatially offset prediction. Extensive experiments on benchmark datasets show the potential of the proposed method in achieving state-of-the-art performance for few-shot action recognition. Shuyuan Li, Huabin Liu 0001, Rui Qian 0001, Yuxi Li 0009, John See, Mengjuan Fei, Xiaoyuan Yu, Weiyao Lin |
AAAI | 7 |
| 2022 | Frequency Feature Pyramid Network With Global-Local Consistency Loss for Crowd-and-Vehicle Counting in Congested ScenesabstractContext prediction plays a crucial role in implementing autonomous driving applications. As one of important context-prediction tasks, crowd-and-vehicle counting is critical for achieving real-time traffic and crowd analysis, consequently facilitating decision-making processes for autonomous vehicles. However, the completion of crowd-and-vehicle counting also faces challenges, such as large-scale variations, imbalanced data distribution, and insufficient local patterns. To tackle these challenges, we put forth a novel frequency feature pyramid network (FFPNet) in this paper. Our proposed FFPNet extracts the multi-scale information by frequency feature pyramid module, which can tackle the issue of large-scale variations. Meanwhile, the frequency feature pyramid module uses different frequency branches to obtain different scale information. We also adopt the attention mechanism to strength the extraction of different scale information. Moreover, we devise a novel loss function, namely global-local consistency loss, to address the existing problems of imbalanced data distribution and insufficient local patterns. Furthermore, we conduct extensive experiments on six datasets to evaluate our proposed FFPNet. It is worth mentioning that we also construct a novel crowd-and-vehicle dataset (CROVEH), which is the only dataset that contains both crowd-and-vehicle annotations. The experimental results show that FFPNet achieves the best performance on different backbones, e.g., 52.69 mean absolute error (MAE) on P2PNet with FFP module. The codes are available at:https://github.com/MUST-AI-Lab/FFPNet. Xiaoyuan Yu, Yanyan Liang 0001, Xuxin Lin, Jun Wan 0001, Tian Wang 0001, Hongning Dai |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Temporal Alignment via Event Boundary for Few-shot Action Recongnition
Shuyuan Li, Huabin Liu 0001, Mengjuan Fei, Xiaoyuan Yu, Weiyao Lin |
BMVC | 4 |
| 2021 | Uncertainty Guided Collaborative Training for Weakly Supervised Temporal Action DetectionabstractWeakly supervised temporal action detection aims to localize temporal boundaries of actions and identify their categories simultaneously with only video-level category labels during training. Among existing methods, attention based methods have achieved superior performance by separating action and non-action segments. However, without the segment-level ground-truth supervision, the quality of the attention weight hinders the performance of these methods. To alleviate this problem, we propose a novel Uncertainty Guided Collaborative Training (UGCT) strategy, which mainly includes two key designs: (1) The first design is an online pseudo label generation module, in which the RGB and FLOW streams work collaboratively to learn from each other. (2) The second design is an uncertainty aware learning module, which can mitigate the noise in the generated pseudo labels. These two designs work together to promote the model performance effectively and efficiently by imposing pseudo label supervision on attention weight learning. Experimental results on three state-of-the-art attention based methods demonstrate that the proposed training strategy can significantly improve the performance of these methods, e.g., more than 4% for all three methods in terms of mAP@IoU=0.5 on the THUMOS14 dataset. Wenfei Yang, Tianzhu Zhang 0001, Xiaoyuan Yu, Qi Tian 0001, Yongdong Zhang 0001, Feng Wu 0001 |
CVPR | 3 |
| 2021 | Learning Discriminative Features for Adversarial RobustnessabstractDeep Learning models have shown incredible image classification capabilities that extend beyond humans. However, they remain susceptible to image perturbations that a human could not perceive. A slightly modified input, known as an Adversarial Example, will result in drastically different model behavior. The use of Adversarial Machine Learning to generate Adversarial Examples remains a security threat in the field of Deep Learning. Hence, defending against such attacks is a studied field of Deep Learning Security. In this paper, we present the Adversarial Robustness of discriminative loss functions. Such loss functions specialize in either inter-class or intra-class compactness. Therefore, generating an Adversarial Example should be more difficult since the decision barrier between different classes will be more significant. We conducted White-Box and Black-Box attacks on Deep Learning models trained with different discriminative loss functions to test this. Moreover, each discriminative loss function will be optimized with and without Adversarial Robustness in mind. From our experimentation, we found White-Box attacks to be effective against all models, even those trained for Adversarial Robustness, with varying degrees of effectiveness. However, state-of-the-art Deep Learning models, such as Arcface, will show significant Adversarial Robustness against Black-Box attacks while paired with adversarial defense methods. Moreover, by exploring Black-Box attacks, we demonstrate the transferability of Adversarial Examples while using surrogate models optimized with different discriminative loss functions. Ryan Hosler, Tyler Phillips 0001, Xiaoyuan Yu, Agnideven Palanisamy Sundar, Xukai Zou, Feng Li 0001 |
MSN | 3 |
| 2021 | A regional distance regression network for monocular object distance estimation
Lianghui Ding, Yuxi Li 0009, Weiyao Lin, Mingbi Zhao, Xiaoyuan Yu, Yunlong Zhan |
J. Vis. Commun. Image Represent. | 6 |
| 2020 | User-Friendly Design of Cryptographically-Enforced Hierarchical Role-based Access Control ModelsabstractData access control is a critical issue for any organization generating, recording or leveraging sensitive information. The popular Role-based Access Control (RBAC) model is well- suited for large organizations with various groups of personnel, each needing their own set of data access privileges. Unfortunately, the traditional RBAC model does not involve the use of cryptographic keys needed to enforce access control policies and protect data privacy. Cryptography-based Hierarchical Access Control (CHAC) models, on the other hand, have been proposed to facilitate RBAC models and directly enforce data privacy and access controls through the use of key management schemes. Though CHAC models and efficient key management schemes can support large and dynamic organizations, they are difficult to design and maintain without intimate knowledge of symmetric encryption, key management and hierarchical access control models. Therefore, in this paper we propose an efficient algorithm which automatically generates a fine-grained CHAC model based on the input of a highly user-friendly representation of access control policies. The generated CHAC model, the dual-level key management (DLKM) scheme, leverages the collusion-resistant Access Control Polynomial (ACP) and Atallah's Efficient Key Management scheme in order to provide privacy at both the data and user levels. As a result, the proposed model generation algorithm serves to democratize the use of CHAC. We analyze each component of our proposed system and evaluate the resulting performance of the user-friendly CHAC model generation algorithm, as well as the DLKM model itself, along several dimensions. Xiaoyuan Yu, Brandon Haakenson, Tyler Phillips 0001, Xukai Zou |
ICCCN | 1 |
| 2019 | Real-time recovery and recognition of motion blurry QR code image based on fractional order deblurring methodabstractMost image deblurring methods require large algorithm computational cost because multi‐scale blind deconvolution is used for estimating kernel. Furthermore, a moving quick response (QR) code image is regarded as a type of classical blurry image and requires real‐time processing in practical applications. Therefore, this study proposes a new framework of motion blurry QR code image restoration in real‐time based on the fractional‐order deblurring method. The authors perform a trade‐off between algorithm computational cost and quality of the deblurring image. First, a black frame is added around the traditional QR code, which is used for locating QR code and reducing the computational cost. Next, a new image deblurring method is proposed using fractional differential order and is used for improving the quality of the deblurring image. Furthermore, an average grey‐level method is presented to reconstruct the standard QR code images. Comparisons with the existing algorithms demonstrate that the proposed method can achieve favourable deblurring quality and acceptable computational cost. Finally, their framework is validated in a practical platform of an actual conveyor belt system with a low‐cost industrial camera. Experimental results indicate that their framework performs favourably with processing motion blurry QR code images. Xiaoyuan Yu |
IET Image Process. | 1 |
| 2015 | Subcategory-Aware Object DetectionabstractIn this letter, we introduce a subcategory-aware object detection framework to detect generic object classes with high intra-class variance. Motivated by the observation that the object appearance demonstrates some clustering property, we split the training data into subcategories and train a detector for each subcategory. Since the proposed ensemble of detectors relies heavily on subcategory clustering, we propose an effective subcategories generation method that is tuned for the detection task. More specifically, we first initialize subcategories by constrained spectral clustering based on mid-level image features used in object recognition. Then we jointly learn the ensemble detectors and the latent subcategories in an alternative manner. Our performance on the PASCAL VOC 2007 detection challenges and INRIA Person dataset is comparable with state-of-the-art, even with much less computational cost. Xiaoyuan Yu, Jianchao Yang, Zhe Lin 0001, Jiangping Wang, Tianjiang Wang, Thomas S. Huang |
IEEE Signal Process. Lett. | 1 |
| 2015 | Key Point Detection by Max Pooling for TrackingabstractInspired by the recent image feature learning work, we propose a novel key point detection approach for object tracking. Our approach can select mid-level interest key points by max pooling over the local descriptor responses from a set of filters. Linear filters are first learned from targets in first frames. Then max pooling is performed over data driven spatial supporting field to detect discriminant key points, and thus the detected key points bear higher level semantic meanings, which we apply in tracking by structured key point matching. We show that our tracking system is robust to occlusions and cluttered background. Testing on several challenging tracking sequences, we demonstrate that our proposed tracking system can achieve competitive or better performances than the state-of-the-art trackers. Xiaoyuan Yu, Jianchao Yang, Tianjiang Wang, Thomas S. Huang |
IEEE Trans. Cybern. | 1 |