VLDB 2026 Research / reviewers in the wild / expert
Ruiyang Xia
dblp:273/6823
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COMET: An FP32 Matrix Multiplication Accelerator by Extending INT8-Based Arrays
Ruiyang Xia, Yifan Hao 0001, Yongwei Zhao 0001, Zidong Du, Xing Hu 0001, Qi Guo 0001 |
APPT | 1 |
| 2026 | Hardwired-Neuron Language Processing Units as General-Purpose Cognitive SubstratesabstractThe rapid advancement of Large Language Models (LLMs) has established language as a core general-purpose cognitive substrate, driving the demand for specialized Language Processing Units (LPUs) tailored for LLM inference. To overcome the growing energy consumption of LLM inference systems, this paper proposes a Hardwired-Neurons Language Processing Unit (HNLPU), which physically hardwires LLM weight parameters into the computational fabric, achieving several orders of magnitude computational efficiency improvement by extreme specialization. However, a significant challenge still lies in the scale of modern LLMs. A straightforward hardwiring of GPT-OSS-120B would require fabricating photomask sets valued at over 6 billion dollars, rendering this straightforward solution economically impractical. Yang Liu 0466, Yongwei Zhao 0001, Yifan Hao 0001, Zifu Zheng, Weihao Kong, Zhangmai Li, Dongchen Jiang, Ruiyang Xia, Zhihong Ma, Zisheng Liu, Zhaoyong Wan, Yunqi Lu, Hongrui Guo, Zhe Wang 0017, Tianrui Ma, Mo Zou, Rui Zhang 0040, Ling Li 0001, Xing Hu 0001, Zidong Du, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen |
ASPLOS (2) | 9 |
| 2026 | SSD: Making Face Forgery Clues Evident Again With Self-Steganographic DetectionabstractThe rapid development of generative AI techniques enables the synthesis of highly realistic facial images, posing significant challenges for the accurate detection of face forgeries. In contrast to solely elevating detector awareness, proactively reducing the intrinsic difficulty of forgery detection can streamline detector complexity while improving both generalization and robustness. This insight motivates our defense strategy to make face forgery clues more evident. Specifically, a novel proactive approach dubbed Self-Steganographic Detection (SSD) is proposed to imperceptibly embed facial images into themselves as a form of detection evidence. The recovery process is designed to remain robust under normal manipulations while exhibiting deliberate degradation under malicious manipulations, thereby clearly revealing potential forgeries. Unlike embedding bit-level vectors, pixel-level images are informative to ensure the generalization of our approach. Due to the similarity between the protected and embedded images, SSD performs detection without storing any embedded information in advance. To support practical deployment, our approach incorporates a dual detection scheme that aims to identify unprotected images and determine the authenticity of protected images. Extensive experiments using 8 face forgery techniques demonstrate the effectiveness of our approach compared to state-of-the-art methods. Ruiyang Xia, Dawei Zhou 0004, Lin Yuan 0002, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Big Brother Is Watching: Proactive Deepfake Detection via Learnable Hidden FaceabstractAs deepfake technologies continue to evolve, proactive defense techniques have gained increasing attention for their potential to either neutralize deepfake operations or simplify detection through pre-embedded signals. In this paper, inspired by watermark-based forensic methods, we explore a novel detection framework built on the concept of “hiding a learnable face within a face”. Specifically, we use a semi-fragile invertible steganography network to imperceptibly embed a learnable template image within a host face image to be protected. This template serves as an indicator revealing signs of tampering when recovered through the inverse steganography process. Unlike manually designed templates, it is optimized during training to resemble a neutral facial appearance—functioning like a subtle “big brother” hidden within the image. Through a self-blending mechanism and robustness learning strategy with a simulated transmission channel, we develop a robust detector that accurately distinguishes between malicious tampering and benign processing of the steganographic image. Extensive experiments across multiple datasets validate the superiority of the proposed approach over competing passive and proactive detection methods. Shangchao Yang, Ruiyang Xia, Lin Yuan 0002, Xinbo Gao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | A Lightweight Object Counting Network Based on Density Map Knowledge DistillationabstractObject counting aims to count the accurate number of object instances in images, and its operation efficiency is essential. However, most current CNN-based methods rely on complex network architectures, which results in them consuming a significant amount of memory, time, and other resources at runtime. This seriously limits their deployment in practical application scenarios, such as public safety and agriculture planting. Therefore, we propose a lightweight object counting method named EdgeCount to effectively balance inference speed and object counting accuracy. Specifically, we construct a network composed of a student model (EdgeCount) and a teacher model (EdgeCount-T) with the same encoder-decoder structure based on density map knowledge distillation (DMKD), allowing the EdgeCount to learn object density distribution from the EdgeCount-T. After that, we introduce spatial and channel reconstruction convolution (SCConv), composed of a spatial reconstruction unit (SRU) and a channel reconstruction unit (CRU), to decrease spatial and channel redundancy with lower computational costs. Moreover, a low parameter weighted multi-scale feature fusion module (LWMFFM) is designed to further improve the countering ability through segmenting minor structural discrepacies among multi-scale features. Extensive experiments conducted on challenging remote sensing and dense crowd object counting datasets demonstrate the effectiveness and superiority of our method. In particular, under the four NVIDIA Jetson devices, EdgeCount can accurately counter objects with only 0.12M parameters and 19.87M floating-point operations per second (FLOPs) in the size of 128, which achieves the lowest latency and fastest FPS compared with other state-of-the-art object counters. Zhilong Shen, Guoquan Li 0001, Ruiyang Xia, Hongying Meng, Zhengwen Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Advancing Generalized Deepfake Detector with Forgery Perception GuidanceabstractOne of the serious impacts brought by artificial intelligence is the abuse of deepfake techniques. Despite the proliferation of deepfake detection methods aimed at safeguarding the authenticity of media across the Internet, they mainly consider the improvement of detector architecture or the synthesis of forgery samples. The forgery perceptions, including the feature responses and prediction scores for forgery samples, have not been well considered. As a result, the generalization across multiple deepfake techniques always comes with complicated detector structures and expensive training costs. In this paper, we shift the focus to real-time perception analysis in the training process and generalize deepfake detectors through an efficient method dubbed Forgery Perception Guidance (FPG). In particular, after investigating the deficiencies of forgery perceptions, FPG adopts a sample refinement strategy to pertinently train the detector, thereby elevating the generalization efficiently. Moreover, FPG introduces more sample information as explicit optimizations, which makes the detector further adapt the sample diversities. Experiments demonstrate that FPG improves the generality of deepfake detectors with small training costs, minor detector modifications, and the acquirement of real data only. In particular, our approach not only outperforms the state-of-the-art on both the cross-dataset and cross-manipulation evaluation but also surpasses the baseline that needs more than 3× training time. Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Lin Yuan 0002, Shuodi Wang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2024 | MMNet: Multi-Collaboration and Multi-Supervision Network for Sequential Deepfake DetectionabstractAdvanced manipulation techniques have provided criminals with opportunities to make social panic or gain illicit profits through the generation of deceptive media, such as forgery face images. In response, various deepfake detection methods have been proposed to assess image authenticity. Sequential deepfake detection, which is an extension of deepfake detection, aims to identify forged facial regions with the correct sequence for recovery. Nonetheless, due to the different combinations of spatial and sequential manipulations, forgery face images exhibit substantial discrepancies that severely impact detection performance. Additionally, the recovery of forged images requires knowledge of the manipulation model to implement inverse transformations, which is difficult to ascertain as relevant techniques are often concealed by attackers. To address these issues, we propose Multi-Collaboration and Multi-Supervision Network (MMNet) that handles various spatial scales and sequential permutations in forgery face images and achieve recovery without requiring knowledge of the corresponding manipulation method. Furthermore, existing evaluation metrics only consider detection accuracy at a single inferring step, without accounting for the matching degree with ground-truth under continuous multiple steps. To overcome this limitation, we propose a novel evaluation metric called Complete Sequence Matching (CSM), which considers the detection accuracy at multiple inferring steps, reflecting the ability to detect integrally forged sequences. Extensive experiments on several typical datasets demonstrate that MMNet achieves state-of-the-art detection performance and independent recovery performance. Code will be available at https://github.com/xarryon/MMNet. Ruiyang Xia, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Inspector for Face Forgery Detection: Defending Against Adversarial Attacks From Coarse to FineabstractThe emergence of face forgery has raised global concerns on social security, thereby facilitating the research on automatic forgery detection. Although current forgery detectors have demonstrated promising performance in determining authenticity, their susceptibility to adversarial perturbations remains insufficiently addressed. Given the nuanced discrepancies between real and fake instances are essential in forgery detection, previous defensive paradigms based on input processing and adversarial training tend to disrupt these discrepancies. For the detectors, the learning difficulty is thus increased, and the natural accuracy is dramatically decreased. To achieve adversarial defense without changing the instances as well as the detectors, a novel defensive paradigm called Inspector is designed specifically for face forgery detectors. Specifically, Inspector defends against adversarial attacks in a coarse-to-fine manner. In the coarse defense stage, adversarial instances with evident perturbations are directly identified and filtered out. Subsequently, in the fine defense stage, the threats from adversarial instances with imperceptible perturbations are further detected and eliminated. Experimental results across different types of face forgery datasets and detectors demonstrate that our method achieves state-of-the-art performances against various types of adversarial perturbations while better preserving natural accuracy. Code is available on https://github.com/xarryon/Inspector. Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Bi-path Combination YOLO for Real-time Few-shot Object Detection
Ruiyang Xia, Guoquan Li 0001, Zhengwen Huang, Hongying Meng |
Pattern Recognit. Lett. | 1 |
| 2022 | Data augmentation and shadow image classification for shadow detectionabstractAbstract Shadow detection is an important branch of computer vision. Recently, convolutional neural network (CNN)‐based methods for shadow detection have achieved better performance than methods based on manually designed features. However, CNNs are extremely hungry for data and the training of CNN‐based shadow detector requires time‐consuming and expensive pixel‐level annotations. To alleviate this problem in shadow detection, a method of data augmentation based on generative adversarial network (GAN), named ShadowGAN, has been proposed. Given a shadow mask and a shadow‐free image, our ShadowGAN can generate shadow images with labels. To guide the training of ShadowGAN and get more realistic shadow images, loss is further implemented to impose a restriction between real shadow images and generated shadow images. The effectiveness of ShadowGAN is demonstrated by training existing shadow detectors on enlarged dataset. In addition, to better make use of shadow‐free images in shadow detection, shadow image classification task is added for the shadow detectors. Experiments show that this task can guide the feature extraction network to learn more robust shadow features. At last, these two methods are combined and a better performance of shadow detection is achieved. Guoquan Li 0001, Lingyun Wen, Zhengwen Huang, Ruiyang Xia |
IET Image Process. | 4 |
| 2022 | Transformers only look once with nonlinear combination for real-time object detection
Ruiyang Xia, Guoquan Li 0001, Zhengwen Huang, Man Qi |
Neural Comput. Appl. | 1 |
| 2022 | CBASH: Combined Backbone and Advanced Selection Heads With Object Semantic Proposals for Weakly Supervised Object DetectionabstractMost recent object detection methods have achieved growing performance on public datasets. However, enormous efforts are needed for these methods due to the extensive annotations of ground-truth boxes. Weakly Supervised Object Detection (WSOD) methods hence have been proposed to solve this problem as only image-level annotations are required and then output bounding boxes related to the objects. In order to further elevate the weakly supervised detection methods on the extraction of reasonable features, the training of potential positive proposals, and the generation of proposals before training, we propose a new Combined Backbone and Advanced Selection Heads (CBASH) method with the proposals generated from the object semantic information. Specifically, Combined Backbone will make the unobvious object features more noticeable, Advanced Selection Heads promote more potential positive proposals to get training, and the generated object semantic proposals elevate the quality and quantity of positive proposals. The proposed method is evaluated on the challenging PASCAL VOC 2007 and 2012 benchmark datasets. Experimental results show that our proposed method can achieve improved performance on both VOC 2007 and VOC 2012 datasets and outperforms the existing state-of-the-art methods. Ruiyang Xia, Guoquan Li 0001, Zhengwen Huang, Hongying Meng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |