VLDB 2026 Research / reviewers in the wild / expert
Ruijie Quan
dblp:238/0204
· DBLP profile ↗
27ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0003-4077-1398ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 17 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TarPro: Targeted Protection Against Malicious Image Editing
Kaixin Shen, Ruijie Quan, Jiaxu Miao, Jun Xiao 0001 |
AAAI | 2 |
| 2026 | Insert Anything: Image Insertion via In-Context Editing in DiTabstractThis work presents Insert Anything, a unified framework for reference-based image insertion that seamlessly integrates objects from reference images into target scenes under flexible, user-specified control guidance. Instead of training separate models for individual tasks, our approach is trained once on our new AnyInsertion dataset, the first open-source large-scale dataset specifically designed for reference image–based image editing, comprising 136K prompt-image pairs covering diverse tasks such as person, object, and garment insertion--and effortlessly generalizes to a wide range of insertion scenarios. Such a challenging setting requires capturing both identity features and fine-grained details, while allowing versatile local adaptations in style, color, and texture. To this end, we propose to leverage the multimodal attention of the Diffusion Transformer (DiT) to support both mask- and text-guided editing. Furthermore, we introduce an in-context editing mechanism that treats the reference image as contextual information, employing two prompting strategies to harmonize the inserted elements with the target scene while faithfully preserving their distinctive features. Extensive experiments on AnyInsertion, DreamBooth, and VTON-HD benchmarks demonstrate that our method consistently outperforms existing alternatives, underscoring its great potential in real-world applications such as creative content generation, virtual try-on, and scene composition. Wensong Song, Zongxing Yang, Zheqiao Cheng, Ruijie Quan, Yi Yang 0001 |
AAAI | 5 |
| 2026 | Audio-Guided Video Scene Editing
Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 0001, Yi Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Autonomous LLM-Enhanced Adversarial Attack for Text-to-MotionabstractHuman motion generative models have enabled promising applications, but the ability of text-to-motion (T2M) models to produce realistic motions raises security concerns if exploited maliciously. Despite growing interest in T2M, limited research focus on safeguarding these models against adversarial attacks, with existing work on text-to-image models proving insufficient for the unique motion domain. In the paper, we propose ALERT-Motion, an autonomous framework that leverages large language models (LLMs) to generate targeted adversarial attacks against black-box T2M models. Unlike prior methods that modify prompts through predefined rules, ALERT-Motion uses the knowledge of LLMs of human motion to autonomously generate subtle yet powerful adversarial text descriptions. It comprises two key modules: an adaptive dispatching module that constructs an LLM-based agent to iteratively refine and search for adversarial prompts; and a multimodal information contrastive module that extracts semantically relevant motion information to guide the agent's search. Through this LLM-driven approach, ALERT-Motion produces adversarial prompts querying victim models to produce outputs closely matching targeted motions, while avoiding obvious perturbations. Evaluations across popular T2M models demonstrate ALERT-Motion's superiority over previous methods, achieving higher attack success rates with stealthier adversarial prompts. This pioneering work on T2M adversarial attacks highlights the urgency of developing defensive measures as motion generation technology advances, urging further research into safe and responsible deployment. Honglei Miao, Fan Ma, Ruijie Quan, Kun Zhan, Yi Yang 0001 |
AAAI | 3 |
| 2025 | BrainGuard: Privacy-Preserving Multisubject Image Reconstructions from Brain ActivitiesabstractReconstructing perceived images from human brain activity forms a crucial link between human and machine learning through Brain-Computer Interfaces. Early methods primarily focused on training separate models for each individual to account for individual variability in brain activity, overlooking valuable cross-subject commonalities. Recent advancements have explored multisubject methods, but these approaches face significant challenges, particularly in data privacy and effectively managing individual variability. To overcome these challenges, we introduce BrainGuard, a privacy-preserving collaborative training framework designed to enhance image reconstruction from multisubject fMRI data while safeguarding individual privacy. BrainGuard employs a collaborative global-local architecture where personalized models are trained on each subject's data and operate in conjunction with a shared commonality model that captures and leverages cross-subject patterns. This architecture eliminates the need to aggregate fMRI data across subjects, thereby ensuring privacy preservation. To tackle the complexity of fMRI data, BrainGuard integrates a hybrid synchronization strategy, enabling individual models to dynamically incorporate parameters from the global model. By establishing a secure and collaborative training environment, BrainGuard not only protects sensitive brain activity data but also improves the accuracy of image reconstructions. Extensive experiments demonstrate that BrainGuard sets a new benchmark in both high-level and low-level metrics, advancing the state-of-the-art in brain decoding through its innovative design. Zhibo Tian, Ruijie Quan, Fan Ma, Kun Zhan, Yi Yang 0001 |
AAAI | 2 |
| 2025 | ProtChatGPT: Towards Understanding Proteins with Hybrid Representation and Large Language ModelsabstractProtein research is crucial in various scientific disciplines, but understanding their intricate structure-function relationships remains challenging. Recent advancements in Large Language Models (LLMs) have significantly improved the comprehension of task-specific knowledge, suggesting the potential for specialized ChatGPT-like systems in protein research to aid fundamental investigations. In this work, we introduce ProtChatGPT, which aims to learn and understand protein structures using natural language. ProtChatGPT enables users to upload proteins, ask questions, and engage in interactive conversations to produce comprehensive answers. The system comprises multi-level protein encoding, protein-language alignment, and instruction tuning of LLMs. A protein first undergoes multiple protein encoders and PLP-former to produce multi-level hybrid protein embeddings, which are then aligned through a Protein Context Gating (PCG) module with contrastive learning, and projected by an adapter to conform with the LLM. The LLM finally combines user questions with projected protein embeddings to generate informative answers. Experiments show that ProtChatGPT can produce promising responses to proteins and the corresponding user questions. We hope that ProtChatGPT could form the basis for further exploration and application in protein research. Code and our pre-trained model will be publicly available. Chao Wang 0102, Hehe Fan, Ruijie Quan, Lina Yao 0001, Yi Yang 0001 |
SIGIR | 3 |
| 2025 | ExpAvatar: High-Fidelity Avatar Generation of Unseen Expressions with 3D Face PriorsabstractThe reconstruction of dynamic head avatars has gained increasing significance, giving rise to various downstream applications such as visual dubbing and digital human creation. Despite recent advancements, generating novel, unseen expressions for a given identity remains challenging in concurrently achieving (1) accurate expression and consistent appearance and (2) high-quality and realistic faces. This article introduces ExpAvatar, a novel approach crafted to address these challenges. ExpAvatar elaborately leverages the appearance consistency capabilities inherent in 3DMMs-based models along with the robust generalization ability of DDPMs-based models to alleviate appearance drift issues and enhance the generation of unseen expressions. Specifically, ExpAvatar introduces a Face Priors-Conditioned Diffusion (FPDiff) model to inject 3D face priors into generation models through fine-tuning. Furthermore, a Face Priors-Conditioned Catalyst (FPCatalyst) is employed to enhance the inference efficiency and generation quality. Moreover, we propose a unique confidence-based regularizer function to mitigate the effect of imperfect face-tracking estimates, thereby improving the quality of dynamic neural head avatars. Experimental results demonstrate that ExpAvatar surpasses current state-of-the-art solutions in generating unseen expressions, marking an advancement in the realm of dynamic head avatar synthesis. Code: https://github.com/yuangan/ExpAvatar . Yuan Gan, Ruijie Quan, Yawei Luo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Interpretable3D: An Ad-Hoc Interpretable Classifier for 3D Point Cloudsabstract3D decision-critical tasks urgently require research on explanations to ensure system reliability and transparency. Extensive explanatory research has been conducted on 2D images, but there is a lack in the 3D field. Furthermore, the existing explanations for 3D models are post-hoc and can be misleading, as they separate explanations from the original model. To address these issues, we propose an ad-hoc interpretable classifier for 3D point clouds (i.e., Interpretable3D). As an intuitive case-based classifier, Interpretable3D can provide reliable ad-hoc explanations without any embarrassing nuances. It allows users to understand how queries are embedded within past observations in prototype sets. Interpretable3D has two iterative training steps: 1) updating one prototype with the mean of the embeddings within the same sub-class in Prototype Estimation, and 2) penalizing or rewarding the estimated prototypes in Prototype Optimization. The mean of embeddings has a clear statistical meaning, i.e., class sub-centers. Moreover, we update prototypes with their most similar observations in the last few epochs. Finally, Interpretable3D classifies new samples according to prototypes. We evaluate the performance of Interpretable3D on four popular point cloud models: DGCNN, PointNet2, PointMLP, and PointNeXt. Our Interpretable3D demonstrates comparable or superior performance compared to softmax-based black-box models in the tasks of 3D shape classification and part segmentation. Our code is released at: github.com/FengZicai/Interpretable3D. Tuo Feng 0001, Ruijie Quan, Wenguan Wang, Yi Yang 0001 |
AAAI | 2 |
| 2024 | Clustering for Protein Representation LearningabstractProtein representation learning is a challenging task that aims to capture the structure and function of proteins from their amino acid sequences. Previous methods largely ignored the fact that not all amino acids are equally important for protein folding and activity. In this article, we propose a neural clustering framework that can automatically discover the critical components of a protein by considering both its primary and tertiary structure information. Our framework treats a protein as a graph, where each node represents an amino acid and each edge represents a spatial or sequential connection between amino acids. We then apply an iterative clustering strategy to group the nodes into clusters based on their 1D and 3D positions and assign scores to each cluster. We select the highest-scoring clusters and use their medoid nodes for the next iteration of clustering, until we obtain a hierarchical and informative representation of the protein. We evaluate on four protein-related tasks: protein fold classification, enzyme reaction classification, gene ontology term prediction, and enzyme commission number prediction. Experimental results demonstrate that our method achieves state-of-the-art performance. Ruijie Quan, Wenguan Wang, Fan Ma, Hehe Fan, Yi Yang 0001 |
CVPR | 1 |
| 2024 | Psychometry: An Omnifit Model for Image Reconstruction from Human Brain ActivityabstractReconstructing the viewed images from human brain activity bridges human and computer vision through the Brain-Computer Interface. The inherent variability in brain function between individuals leads existing literature to focus on acquiring separate models for each individual using their respective brain signal data, ignoring commonalities between these data. In this article, we devise Psychometry, an omnifit model for reconstructing images from functional Magnetic Resonance Imaging (fMRI) obtained from different subjects. Psychometry incorporates an omni mixture-of-experts (Omni MoE) module where all the experts work together to capture the inter-subject commonalities, while each expert associated with subject-specific parameters copes with the individual differences. Moreover, Psychometry is equipped with a retrieval-enhanced inference strategy, termed Ecphory, which aims to enhance the learned fMRI representation via retrieving from prestored subject-specific memories. These designs collectively render Psychometry omnifit and efficient, enabling it to capture both inter-subject commonality and individual specificity across subjects. As a result, the enhanced fMRI representations serve as conditional signals to guide a generation model to reconstruct high-quality and realistic images, establishing Psychometry as state-of-the-art in terms of both high-level and low-level metrics. Ruijie Quan, Wenguan Wang, Zhibo Tian, Fan Ma, Yi Yang 0001 |
CVPR | 1 |
| 2024 | General and Task-Oriented Video Segmentation
Mu Chen 0001, Liulei Li, Wenguan Wang, Ruijie Quan, Yi Yang 0001 |
ECCV (7) | 4 |
| 2024 | Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
Tuo Feng 0001, Wenguan Wang, Ruijie Quan, Yi Yang 0001 |
ECCV (55) | 3 |
| 2024 | Depth-Aware Blind Image Decomposition for Real-World Adverse Weather Recovery
Chao Wang 0102, Zhedong Zheng, Ruijie Quan, Yi Yang 0001 |
ECCV (82) | 3 |
| 2024 | Neural Interaction Energy for Multi-Agent Trajectory PredictionabstractMaintaining temporal stability is crucial in multi-agent trajectory prediction. Insufficient regularization to uphold this temporal stability often results in fluctuations in kinematic states, leading to inconsistent predictions and the amplification of errors. In this study, we introduce a framework called Multi-Agent Trajectory prediction via neural interaction Energy (MATE). This framework assesses the interactive motion of agents by employing neural interaction energy, which captures the dynamics of interactions and illustrates their influence on the future trajectories of agents. To bolster temporal stability, we introduce two constraints: inter-agent interaction constraint and intra-agent motion constraint. These constraints work together to ensure temporal stability at both the system and agent levels, effectively mitigating prediction fluctuations inherent in multi-agent systems. Comparative evaluations against previous methods on four diverse datasets, including simulated and real-world scenarios, highlight the superior prediction accuracy and generalization capabilities of our model. Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 0001, Yi Yang 0001 |
ACM Multimedia | 2 |
| 2024 | DRIP: Unleashing Diffusion Priors for Joint Foreground and Alpha Prediction in Image MattingabstractRecovering the foreground color and opacity/alpha matte from a single image (i.e., image matting) is a challenging and ill-posed problem where data priors play a critical role in achieving precise results. Traditional methods generally predict the alpha matte and then extract the foreground through post-processing, often failing to produce high-fidelity foreground color. This failure stems from the models' difficulty in learning robust color predictions from limited matting datasets. To address this, we explore the potential of leveraging vision priors embedded in pre-trained latent diffusion models (LDM) for estimating foreground RGBA values in challenging scenarios and rare objects. We introduce Drip, a novel approach for image matting that harnesses the rich prior knowledge of LDM models. Our method incorporates a switcher and a cross-domain attention mechanism to extend the original LDM for joint prediction of the foreground color and opacity. This setup facilitates mutual information exchange and ensures high consistency across both modalities. To mitigate the inherent reconstruction errors of the LDM's VAE decoder, we propose a latent transparency decoder to align the RGBA prediction with the input image, thereby reducing discrepancies. Comprehensive experimental results demonstrate that our approach achieves state-of-the-art performance in foreground and alpha predictions and shows remarkable generalizability across various benchmarks. Zongxin Yang, Ruijie Quan, Yi Yang 0001 |
NeurIPS | 3 |
| 2024 | Semantic Hierarchy-Aware SegmentationabstractHumans are able to recognize structured relations in observation, allowing us to decompose complex scenes into simpler parts and abstract the visual world at multiple levels. However, such hierarchical reasoning ability of human perception remains largely unexplored in current literature of semantic segmentation. Existing works are often aware of flatten labels and distinguish all the semantic categories exclusively for each pixel. In this work, we instead address hierarchical semantic segmentation (HSS), with the aim of providing a structured, pixel-wise description of visual observation in terms of a class hierarchy. We deviseHssn, a general HSS framework that tackles two critical issues in this task:i)how to efficiently adapt existing hierarchy-agnostic segmentation networks to the HSS setting, andii)how to leverage the class hierarchy to regularize HSS network learning. To addressi),Hssndirectly casts HSS as a pixel-wise multi-label classification task, only bringing minimal architecture change to current segmentation models. To solveii),Hssnfirst explores inherent properties of the hierarchy as a training objective, which enforces segmentation predictions to obey the hierarchy structure. Furthermore, with a set of hierarchy-induced margin constraints,Hssnefficiently reshapes the learned pixel embedding space, so as to generate hierarchy-aware pixel representations and facilitate structured segmentation eventually. Building uponHssn, we further exploit the mutual exclusion relation between semantic labels and strengthen the margin based regularization strategy with more meaningful constrains, leading toHssn+, a more effective framework for HSS. We conduct extensive experiments on six semantic segmentation datasets (i.e., Mapillary Vistas 2.0, Cityscapes, LIP, PASCAL-Person-Part, PASCAL-Part-58, and PASCAL-Part-108), with different class hierarchies, network architectures, and backbones, and the results confirm the generalization and superiority of our algorithms. Liulei Li, Wenguan Wang, Tianfei Zhou, Ruijie Quan, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Zero-Shot Video Grounding With Pseudo Query Lookup and VerificationabstractVideo grounding, the process of identifying a specific moment in an untrimmed video based on a natural language query, has become a popular topic in video understanding. However, fully supervised learning approaches for video grounding that require large amounts of annotated data can be expensive and time-consuming. Recently, zero-shot video grounding (ZS-VG) methods that leverage pre-trained object detectors and language models to generate pseudo-supervision for training video grounding models have been developed. However, these approaches have limitations in recognizing diverse categories and capturing specific dynamics and interactions in the video context. To tackle these challenges, we introduce a novel two-stage ZS-VG framework called Lookup-and-Verification (LoVe), which treats the pseudo-query generation procedure as a video-to-concept retrieval problem. Our approach allows for the extraction of diverse concepts from an open-concept pool and employs a verification process to ensure the relevance of the retrieved concepts to the objects or events of interest in the video proposals. Comprehensive experimental results on the Charades-STA, ActivityNet-Captions, and DiDeMo datasets demonstrate the effectiveness of the LoVe framework. Yu Lu 0019, Ruijie Quan, Linchao Zhu, Yi Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Exploiting Unlabeled Videos for Video-Text Retrieval via Pseudo-Supervised LearningabstractLarge-scale pre-trained vision-language models (e.g., CLIP) have shown incredible generalization performance in downstream tasks such as video-text retrieval (VTR). Traditional approaches have leveraged CLIP's robust multi-modal alignment ability for VTR by directly fine-tuning vision and text encoders with clean video-text data. Yet, these techniques rely on carefully annotated video-text pairs, a process that is costly and labor-intensive. In this context, we introduce a new approach, Pseudo-Supervised Selective Contrastive Learning (PS-SCL). PS-SCL minimizes the dependency on manually-labeled text annotations by generating pseudo-supervisions from unlabeled video data for training. We first exploit CLIP's visual recognition capabilities to generate pseudo-texts automatically. These pseudo-texts contain diverse visual concepts from the video and serve as weak textual guidance. Moreover, we introduce Selective Contrastive Learning (SeLeCT), which prioritizes and selects highly correlated video-text pairs from pseudo-supervised video-text pairs. By doing so, SeLeCT enables more effective multi-modal learning under weak pairing supervision. Experimental results demonstrate that our method outperforms CLIP zero-shot performance by a large margin on multiple video-text retrieval benchmarks, e.g., 8.2% R@1 for video-to-text on MSRVTT, 12.2% R@1 for video-to-text on DiDeMo, and 10.9% R@1 for video-to-text on ActivityNet, respectively. Yu Lu 0019, Ruijie Quan, Linchao Zhu, Yi Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | CLIP4STR: A Simple Baseline for Scene Text Recognition With Pre-Trained Vision-Language ModelabstractPre-trained vision-language models (VLMs) are the de-facto foundation models for various downstream tasks. However, scene text recognition methods still prefer backbones pre-trained on a single modality, namely, the visual modality, despite the potential of VLMs to serve as powerful scene text readers. For example, CLIP can robustly identify regular (horizontal) and irregular (rotated, curved, blurred, or occluded) text in images. With such merits, we transform CLIP into a scene text reader and introduce CLIP4STR, a simple yet effective STR method built upon image and text encoders of CLIP. It has two encoder-decoder branches: a visual branch and a cross-modal branch. The visual branch provides an initial prediction based on the visual feature, and the cross-modal branch refines this prediction by addressing the discrepancy between the visual feature and text semantics. To fully leverage the capabilities of both branches, we design a dual predict-and-refine decoding scheme for inference. We scale CLIP4STR in terms of the model size, pre-training data, and training data, achieving state-of-the-art performance on 13 STR benchmarks. Additionally, a comprehensive empirical study is provided to enhance the understanding of the adaptation of CLIP to STR. We believe our method establishes a simple yet strong baseline for future STR research with VLMs. Shuai Zhao 0006, Ruijie Quan, Linchao Zhu, Yi Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Efficient Multimodal Fusion via Interactive PromptingabstractLarge-scale pre-training has brought unimodal fields such as computer vision and natural language processing to a new era. Following this trend, the size of multimodal learning models constantly increases, leading to an urgent need to reduce the massive computational cost of finetuning these models for downstream tasks. In this paper, we propose an efficient and flexible multimodal fusion method, namely PMF, tailored for fusing unimodally pretrained transformers. Specifically, we first present a modular multimodal fusion framework that exhibits high flexibility and facilitates mutual interactions among different modalities. In addition, we disentangle vanilla prompts into three types in order to learn different optimizing objectives for multimodal learning. It is also worth noting that we propose to add prompt vectors only on the deep layers of the unimodal transformers, thus significantly reducing the training memory usage. Experiment results show that our proposed method achieves comparable performance to several other multimodal finetuning methods with less than 3% trainable parameters and up to 66% saving of training memory usage. Ruijie Quan, Linchao Zhu, Yi Yang 0001 |
CVPR | 2 |
| 2023 | Context-Aware Pretraining for Efficient Blind Image DecompositionabstractIn this paper, we study Blind Image Decomposition (BID), which is to uniformly remove multiple types of degradation at once without foreknowing the noise type. There remain two practical challenges: (1) Existing methods typically require massive data supervision, making them infeasible to real-world scenarios. (2) The conventional paradigm usually focuses on mining the abnormal pattern of a superimposed image to separate the noise, which de facto conflicts with the primary image restoration task. Therefore, such a pipeline compromises repairing efficiency and authenticity. In an attempt to solve the two challenges in one go, we propose an efficient and simplified paradigm, called Context-aware Pretraining (CP), with two pretext tasks: mixed image separation and masked image reconstruction. Such a paradigm reduces the annotation demands and explicitly facilitates context-aware feature learning. Assuming the restoration process follows a structure-to-texture manner, we also introduce a Context-aware Pretrained network (CPNet). In particular, CPNet contains two transformer-based parallel encoders, one information fusion module, and one multi-head prediction module. The information fusion module explicitly utilizes the mutual correlation in the spatial-channel dimension, while the multi-head prediction module facilitates texture-guided appearance flow. Moreover, a new sampling loss along with an attribute label constraint is also deployed to make use of the spatial context, leading to high-fidelity image restoration. Extensive experiments on both real and synthetic benchmarks show that our method achieves competitive performance for various BID tasks. Chao Wang 0102, Zhedong Zheng, Ruijie Quan, Yifan Sun 0003, Yi Yang 0001 |
CVPR | 3 |
| 2023 | Action Sensitivity Learning for Temporal Action LocalizationabstractTemporal action localization (TAL), which involves recognizing and locating action instances, is a challenging task in video understanding. Most existing approaches directly predict action classes and regress offsets to boundaries, while overlooking the discrepant importance of each frame. In this paper, we propose an Action Sensitivity Learning framework (ASL) to tackle this task, which aims to assess the value of each frame and then leverage the generated action sensitivity to recalibrate the training procedure. We first introduce a lightweight Action Sensitivity Evaluator to learn the action sensitivity at the class level and instance level, respectively. The outputs of the two branches are combined to reweight the gradient of the two sub-tasks. Moreover, based on the action sensitivity of each frame, we design an Action Sensitive Contrastive Loss to enhance features, where the action-aware frames are sampled as positive pairs to push away the action-irrelevant frames. The extensive studies on various action localization benchmarks (i.e., MultiThumos, Charades, Ego4D-Moment Queries v1.0, Epic-Kitchens 100, Thumos14 and Activi-tyNet1.3) show that ASL surpasses the state-of-the-art in terms of average-mAP under multiple types of scenarios, e.g., single-labeled, densely-labeled and egocentric. Jiayi Shao, Ruijie Quan, Junjun Zheng, Yi Yang 0001 |
ICCV | 3 |
| 2021 | Removing Raindrops and Rain Streaks in One GoabstractExisting rain-removal algorithms often tackle either rain streak removal or raindrop removal, and thus may fail to handle real-world rainy scenes. Besides, the lack of real-world deraining datasets comprising different types of rain and their corresponding rain-free ground-truth also impedes deraining algorithm development. In this paper, we aim to address real-world deraining problems from two aspects. First, we propose a complementary cascaded network architecture, namely CCN, to remove rain streaks and raindrops in a unified framework. Specifically, our CCN removes raindrops and rain streaks in a complementary fashion, i.e., raindrop removal followed by rain streak removal and vice versa, and then fuses the results via an attention based fusion module. Considering significant shape and structure differences between rain streaks and raindrops, it is difficult to manually design a sophisticated network to remove them effectively. Thus, we employ neural architecture search to adaptively find optimal architectures within our specified deraining search space. Second, we present a new real-world rain dataset, namely RainDS, to prosper the development of deraining algorithms in practical scenarios. RainDS consists of rain images in different types and their corresponding rain-free ground-truth, including rain streak only, raindrop only, and both of them. Extensive experimental results on both existing benchmarks and RainDS demonstrate that our method outperforms the state-of-the-art. Ruijie Quan, Xin Yu 0002, Yuanzhi Liang, Yi Yang 0001 |
CVPR | 1 |
| 2021 | Progressive Transfer Learning for Face Anti-SpoofingabstractFace anti-spoofing (FAS) techniques play an important role in defending face recognition systems against spoofing attacks. Existing FAS methods often require a large number of annotated spoofing face data to train effective anti-spoofing models. Considering the attacking nature of spoofing data and its diverse variants, obtaining all the spoofing types in advance is difficult. This would limit the performance of FAS networks in practice. Thus, an online learning FAS method is highly desirable. In this paper, we present a semi-supervised learning based framework to tackle face spoofing attacks with only a few labeled training data (e.g., ∼ 50 face images). Specifically, we progressively adopt the unlabeled data with reliable pseudo labels during training to enrich the variety of training data. We observed that face spoofing data are naturally presented in the format of video streams. Thus, we exploit the temporal consistency to consolidate the reliability of a pseudo label for a selected image. Furthermore, we propose an adaptive transfer mechanism to ameliorate the influence of unseen spoofing data. Benefiting from the progressively-labeling nature of our method, we are able to train our network on not only data of seen spoofing types (i.e., the source domain) but also unlabeled data of unseen attacking types (i.e., the target domain). In this way, our method can reduce the domain gap and is more practical in real-world anti-spoofing scenarios. Extensive experiments in both the intra-database and inter-database scenarios demonstrate that our method is on par with the state-of-the-art methods but employs remarkably less labeled data (less than 0.1% labeled spoofing data in a dataset). Moreover, our method significantly outperforms fully-supervised methods on cross-domain testing scenarios with the help of our progressive learning fashion. Ruijie Quan, Yu Wu 0011, Xin Yu 0002, Yi Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Holistic LSTM for Pedestrian Trajectory PredictionabstractAccurate predictions of future pedestrian trajectory could prevent a considerable number of traffic injuries and improve pedestrian safety. It involves multiple sources of information and real-time interactions, e.g., vehicle speed and ego-motion, pedestrian intention and historical locations. Existing methods directly apply a simple concatenation operation to combine multiple cues while their dynamics over time are less studied. In this paper, we propose a novel Long Short-Term Memory (LSTM), namely, to incorporate multiple sources of information from pedestrians and vehicles adaptively. Different from LSTM, our considers mutual interactions and explores intrinsic relations among multiple cues. First, we introduce extra memory cells to improve the transferability of LSTMs in modeling future variations. These extra memory cells include a speed cell to explicitly model vehicle speed dynamics, an intention cell to dynamically analyze pedestrian crossing intentions and a correlation cell to exploit correlations among temporal frames. These three individual cells uncover the future movement of vehicles, pedestrians and global scenes. Second, we propose a gated shifting operation to learn the movement of pedestrians. The intention of crossing the road or not would significantly affect pedestrian's spatial locations. To this end, global scene dynamics and pedestrian intention information are leveraged to model the spatial shifts. Third, we integrate the speed variations to the output gate and dynamically reweight the output channels via the scaling of vehicle speed. The movement of the vehicle would alter the scale of the predicted pedestrian bounding box: as the vehicle gets closer to the pedestrian, the bounding box is enlarging. Our rescaling process captures the relative movement and updates the size of pedestrian bounding boxes accordingly. Experiments conducted on three pedestrian trajectory forecasting benchmarks show that our achieves state-of-the-art performance. Ruijie Quan, Linchao Zhu, Yu Wu 0011, Yi Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Auto-ReID: Searching for a Part-Aware ConvNet for Person Re-IdentificationabstractPrevailing deep convolutional neural networks (CNNs) for person re-IDentification (reID) are usually built upon ResNet or VGG backbones, which were originally designed for classification. Because reID is different from classification, the architecture should be modified accordingly. We propose to automatically search for a CNN architecture that is specifically suitable for the reID task. There are three aspects to be tackled. First, body structural information plays an important role in reID but it is not encoded in backbones. Second, Neural Architecture Search (NAS) automates the process of architecture design without human effort, but no existing NAS methods incorporate the structure information of input images. Third, reID is essentially a retrieval task but current NAS algorithms are merely designed for classification. To solve these problems, we propose a retrieval-based search algorithm over a specifically designed reID search space, named Auto-ReID. Our Auto-ReID enables the automated approach to find an efficient and effective CNN architecture for reID. Extensive experiments demonstrate that the searched architecture achieves state-of-the-art performance while reducing 50% parameters and 53% FLOPs compared to others. Ruijie Quan, Xuanyi Dong, Yu Wu 0011, Linchao Zhu, Yi Yang 0001 |
ICCV | 1 |
| 2019 | Thinking inside the Box: Differential Fault Localization for SDN Control Plane
Xing Li 0001, Yinbo Yu, Kai Bu, Yan Chen 0004, Ruijie Quan |
IM | 6 |