VLDB 2026 Research / reviewers in the wild / expert
Yiwei Ru
dblp:213/8636
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2026
0009-0004-5577-2769ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 10 since 2021Security and privacy · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Symmetric Image-Text Tuning With Entropy-Guided Fusion for Online Continual Learning in Non-Stationary Visual StreamsabstractOnline continual learning studies how models learn from continuous and non-stationary data streams. In this paper, we observe that CLIP models exhibit an asymmetric image-text interaction under online continual learning. Specifically, text features of previously seen classes may introduce unfavorable supervision when paired with visual features of newly observed data, leading to catastrophic forgetting. To alleviate this issue, we propose a simple yet effective symmetric image-text tuning (SIT) strategy that removes such asymmetric text supervision during online learning. We further introduce an entropy-guided fusion (EGF) mechanism that adaptively combines predictions from the pretrained and finetuned branches based on their relative uncertainty. This design allows the model to recover pretrained knowledge when the finetuned branch becomes unreliable, while still preserving plasticity on recently observed classes when confidence is high. In addition, we present MiD-Blurry, an online continual learning benchmark that combines multiple class distribution patterns to better reflect realistic data streams with blurred temporal boundaries. Extensive experiments on standard continual learning benchmarks and the MiD-Blurry setting evaluate inference-at-any-time performance and generalization to future data. The results show that the proposed approach maintains a practical balance between adapting to new data and preserving previously learned information in realistic online learning scenarios. Leyuan Wang, Liuyu Xiang, Yiwei Ru, Yunlong Wang 0003, Zhaofeng He 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Psyche-Wave: Fusing Vector-Quantized Morphology and LLM-Inferred Semantics from Millimeter-Wave SCG for Psychological State DecodingabstractThis paper introduces Psyche-Wave, a novel paradigm for non-contact psychological state assessment, addressing the challenge that existing methods struggle to reconcile signal representation robustness with deep physiological semantic understanding. The proposed framework is built upon high-fidelity Seismocardiogram (SCG) and respiratory signals, captured by a proprietary high-sampling-rate millimeter-wave (mmWave) radar system. Psyche-Wave features a parallel dual-branch architecture for complementary feature extraction. The first, a Data-Driven Morphological Branch, employs Vector Quantization (VQ) to encode the Mel spectrogram of the SCG signal into a codebook-based representation, yielding a noise-resilient morphological embedding. The second, a Knowledge-Driven Semantic Branch, leverages a Large Language Model (LLM) to infer deep contextual relationships from medically significant physiological parameters—including heart rate variability, cardiac time intervals, and cardiopulmonary coupling—outputting a rich semantic embedding. These complementary embeddings are then integrated through a dedicated fusion module and passed to a downstream classifier for precise emotion and personality trait evaluation. Comprehensive evaluations on a newly collected high-fidelity dataset, referred to as mmHeart-Pro, and the public ReMAP dataset demonstrate state-of-the-art performance. This work pioneers a new path that fuses data-driven morphological analysis with knowledge-driven semantic reasoning, significantly advancing the accuracy and interpretability of non-contact psychological sensing. Yiwei Ru, Zhenbo Xu, Yanlin Xu, Huijia Wu, Zhaofeng He 0001, Zhenan Sun |
BIBM | 1 |
| 2025 | Eye Movements as Images: A Multimodal Framework for Eye Movements RepresentationabstractEye movements are increasingly popular for enhancing natural language processing and modeling individual states. Although specialized methods have been developed to represent eye movements for various tasks, effectively modeling the complex dynamics of eye movements and the heterogeneity with stimulus text remains challenging. This paper proposes a text-guided eye movement representation framework that introduces a novel perspective by converting raw eye movement sequences into line graph images and encoding them with a powerful pre-trained vision transformer. To address the disparities between eye movements and text, we guide their temporal alignment using human reading order and combine Canonical Correlation Analysis with Optimal Transport to fuse the two modalities. This approach not only significantly simplifies the design of specialized models but also has the potential to become a universal representation for eye movements. Experimental results on six different domain tasks show that the proposed method achieves state-of-the-art performance. We release the source code at https://github.com/wulalahalala/VLEM. Dongsen Zhang, Peipei Li 0002, Zekun Li 0001, Yiwei Ru, Huijia Wu, Zhaofeng He 0001 |
ICASSP | 4 |
| 2025 | BiommWave: A Non-Visual Approach for Biometric Recognition Using Millimeter-Wave RadarabstractThis paper explores the application of millimeter-wave (mmWave) radar in biometric recognition. As a non-visual human sensing technology, mmWave radar captures reflection properties and micro-movements, providing a complementary modality to visual appearance. We implemented a complete system pipeline to utilize mmWave sensing for individual recognition. Based on physiological mechanisms, we design preprocessing methods to extract intuitive biometric feature maps, concerning body reflections, cardiopulmonary activity, and micro-motion frequencies. To address the data uncertainty, a dynamic pole-based learning strategy is proposed to construct compact and discriminated feature distributions. In real-world evaluations, the system achieves 95.12% accuracy and an Equal Error Rate (EER) of 1.96%. This work leverages the advantages of mmWave radar for flexible, unconstrained, and private biometric systems. From a non-visual sensing perspective, it explores novel modalities as unique biometric cues, demonstrating significant value of research and applications. Mupei Li, Yunlong Wang 0003, Yiwei Ru, Kunbo Zhang, Zhenan Sun |
IJCB | 3 |
| 2025 | LAMAR: LLM-Guided Adaptive Perceptual Modeling for Micro-Action RecognitionabstractWe present LAMAR, a novel framework for Micro-Action Recognition that addresses the challenges of identifying subtle, ephemeral human movements lasting less than one-third of a second. Our approach leverages large language models to estimate semantic complexity of micro-actions and dynamically configure a hierarchical Vision Transformer architecture accordingly. LAMAR introduces: (1) a principled complexity estimation module that quantifies recognition difficulty by analyzing subtlety, noise susceptibility, and intra-class ambiguity; and (2) an adaptive perception pipeline that dynamically adjusts spatiotemporal resolution and attention mechanisms based on estimated complexity. Experiments on the MA-52 benchmark demonstrate that LAMAR outperforms state-of-the-art methods by 7.20% in accuracy while maintaining computational efficiency, establishing a new paradigm for context-aware visual analysis that intelligently allocates resources based on task difficulty. Yiwei Ru, Leyuan Wang, Ma He, Zhaofeng He 0001, Zhenan Sun |
IJCB | 1 |
| 2025 | Beyond Macro-Actions: A Bio-Inspired Framework for Fine-Grained Micro-Action RecognitionabstractHuman Action Recognition (HAR) is pivotal in advancing applications from surveillance to healthcare, but predominantly focuses on easily observable, macro-level actions such as running or jumping. Micro-Action Recognition (MAR), however, delves into the subtle, often involuntary motions like postural shifts, brief gestures, or faint facial twitches, which are critical for revealing underlying emotional states, intentions, or stress levels. MAR presents unique challenges due to the ephemeral nature of micro-actions, their fine-grained inter-class similarities, and significant class imbalance. To overcome these obstacles, our approach draws inspiration from the hierarchical and context-sensitive capabilities of the human visual system. We propose a biologically motivated, multi-pathway framework that cohesively integrates global context analysis, rapid temporal scanning, and meticulous fine-grained scrutiny. This framework combines skeletal dynamics, subtle motion amplitude cues, and RGB-based contextual features to enable a comprehensive and robust recognition of micro-actions, even in unconstrained environments. Our experimental results on the MA-52 dataset demonstrate leading performance, significantly advancing MAR research and broadening the spectrum of applications that require a nuanced understanding of human behavior. Yiwei Ru, Churan Yu, Dongsen Zhang, Mupei Li, Yongji Liu, Zhaofeng He 0001 |
ICME | 1 |
| 2025 | Contextualizing Borderline ECG Analysis via Multi-Modal Feature Extraction and Large Language Model InferenceabstractBorderline electrocardiograms (ECGs) pose a significant diagnostic challenge, as their waveforms often exhibit subtle deviations that overlap with both normal and pathological patterns. Conventional deep learning models, while adept at detecting common arrhythmias, struggle with these ambiguous cases due to sparse annotations and complex signal morphologies. To address this gap, we propose a multi-modal framework that combines structured feature extraction and large language model (LLM) inference for robust ECG classification. First, we employ specialized libraries to derive morphological markers, interval measurements, and heart rate variability (HRV) parameters from raw ECG data. These features, along with demographic metadata, are then seamlessly integrated into prompts for LLM-based few-shot learning. By embedding quantitative signals into textual templates, the model acquires a contextual understanding that transcends static, threshold-based judgments. Our method not only excels at detecting arrhythmias but also demonstrates enhanced performance in classifying borderline ECGs, guided by expert-validated annotations. Experimental results show that the proposed pipeline effectively mitigates class imbalance and improves interpretability, offering a scalable solution that bridges the divide between conventional signal processing and advanced medical-language reasoning. This integration of numeric features with textual prompts paves the way for more accurate, transparent, and adaptable cardiac diagnostics. Yanlin Xu, Yiwei Ru, Dongsen Zhang, Yongji Liu, Zhenan Sun |
ICME | 2 |
| 2025 | CLANet: A Denoising-Driven Framework for Robust mmWave Radar Vital Sign Monitoring
Yiwei Ru, Yongji Liu, Mupei Li, Dongsen Zhang, Zhaofeng He 0001, Zhenan Sun |
PRCV (3) | 1 |
| 2025 | Enhancing Adversarial Transferability With Alignment NetworkabstractDeep neural networks (DNNs) have been confirmed to exhibit vulnerability, as they are susceptible to deception by adversarial examples. Transfer-based attacks perturb a surrogate model and use the transferability of adversarial examples to attack other models. The effectiveness of these attacks relies heavily on the surrogate model, which often focuses on non-critical regions like backgrounds or object edges, leading to poor transferability. The intrinsic properties of the surrogate model fundamentally determine the performance of transfer-based attacks, yet this aspect has rarely been the focus of research. Therefore, we respectively design image masking operations for Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), forcing the model to reallocate attention to the critical regions. The attention of the surrogate model on the masked image and the original image is then aligned by inserting an alignment network inside the model. The modified surrogate model becomes more proficient in capturing the critical regions within the image, thereby generating more powerful adversarial examples. The proposed alignment network can be integrated into existing transfer-based attacks, significantly enhancing their performance. In addition, we also propose a novel feature-level attack based on the aligned attention, demonstrating superior performance compared to existing state-of-the-art feature-level attacks. Qi Li 0005, Yiwei Ru, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | CASIA-PR-V1: A Multi-Ethnic, Multi-Device and Cross-Spectral Dataset and a Multiscale Disentangled Model for Periocular RecognitionabstractPeriocular recognition is regarded as an alternative trait for biometric recognition that can effectively solve the identification problem under large occlusions. However, few datasets are tailored for periocular recognition. For most compromises, iris datasets at near-infrared wavelengths, miss information about the eyebrows or eyelids. In this paper, a challenging dataset for real scenarios named CASIA-PR-V1 with evaluation protocols is released for periocular recognition. It is collected from multiple types of mobile devices with different resolutions or wavelengths. A rich set of attributes, e.g., ethnicities, is tagged to support fine-grained classification tasks. Moreover, we consider a wide range of noisy data in unconstrained environment, especially for glasses. Superior to its counterparts, this periocular dataset is highly valuable for studying cross-device and cross-spectral periocular recognition with occlusions, as well as fine-grained attribute classification. Additionally, a multiscale disentangled model is proposed to extract discriminating representations for periocular recognition with severe occlusions. Extensive experiments are conducted on CASIA-PR-V1, and the results indicate the superiority of our model for unconstraint periocular recognition. Yiwei Ru, Yushan Han, Longteng Kong, Zijian Wang 0009, Yong He 0009, Zhenan Sun |
IEEE Trans. Multim. | 2 |
| 2023 | Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological IntelligenceabstractAffective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP. Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun |
ACM Multimedia | 1 |
| 2021 | Bita-Net: Bi-temporal Attention Network for Facial Video Forgery DetectionabstractDeep forgery detection on video data has attracted remarkable research attention in recent years due to its potential in defending forgery attacks. However, existing methods either only focus on the visual evidence within individual images, or are too sensitive to fluctuations across frames. To address these issues, this paper propose a novel model, named Bita-Net, to detect forgery faces in video data. The network design of Bita-Net is inspired by the mechanism of how human beings detect forgery data, i.e. browsing and scrutinizing, which is reflected by the two-pathway architecture of Bita-Net. Concretely, the browsing pathway scans the entire video at a high frame rate to check the temporal consistency, while the scrutinizing pathway focuses on analyzing key frames of the video at a lower frame rate. Furthermore, an attention branch is introduced to improve the forgery detection ability of the scrutinizing pathway. Extensive experiment results demonstrate the effectiveness and generalization ability of Bita-Net on various popular face forensics detection datasets, including FaceForensics++, CelebDF, DeepfakeTIMIT and UADFV. Yiwei Ru, Yunfan Liu 0001, Jianxin Sun 0003, Qi Li 0005 |
IJCB | 1 |
| 2017 | LivDet iris 2017 - Iris liveness detection competition 2017abstractPresentation attacks such as using a contact lens with a printed pattern or printouts of an iris can be utilized to bypass a biometric security system. The first international iris liveness competition was launched in 2013 in order to assess the performance of presentation attack detection (PAD) algorithms, with a second competition in 2015. This paper presents results of the third competition, LivDet-Iris 2017. Three software-based approaches to Presentation Attack Detection were submitted. Four datasets of live and spoof images were tested with an additional cross-sensor test. New datasets and novel situations of data have resulted in this competition being of a higher difficulty than previous competitions. Anonymous received the best results with a rate of rejected live samples of 3.36% and rate of accepted spoof samples of 14.71%. The results show that even with advances, printed iris attacks as well as patterned contacts lenses are still difficult for software-based systems to detect. Printed iris images were easier to be differentiated from live images in comparison to patterned contact lenses as was also seen in previous competitions. David Yambay, Benedict Becker, Naman Kohli, Daksha Yadav, Adam Czajka, Kevin W. Bowyer, Stephanie Schuckers, Richa Singh 0001, Mayank Vatsa, Afzel Noore, Diego Gragnaniello, Carlo Sansone, Luisa Verdoliva, Lingxiao He, Yiwei Ru, Nianfeng Liu, Zhenan Sun, Tieniu Tan |
IJCB | 15 |