VLDB 2026 Research / reviewers in the wild / expert
Yanyan Wei
dblp:81/7807
· DBLP profile ↗
27ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0001-8818-6740ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XInsight: Integrative Stage-Consistent Psychological Counseling Support Agents for Digital Well-Being
Fei Wang 0073, Jiangnan Yang, Kun Li 0008, Yanyan Wei, Dan Guo 0001, Meng Wang 0001 |
WWW | 6 |
| 2026 | Modeling Long-Term Emotional Support Through Causal World Modeling With Imitation LearningabstractEmotional support conversation systems have emerged as a promising complement to traditional mental health consultations, offering context-aware dialogue to support seekers’ emotional well-being. Despite their potential, two fundamental challenges remain unresolved: 1) modeling long-term emotional trajectory beyond short-term relief; and 2) adapting support strategies to context in a psychologically coherent manner. To address these challenges, we propose CAIWO, a novel framework that integrates world modeling and causality-enhanced imitation learning to systematically support seekers, alleviate psychological stress, and restore emotional balance. Specifically, CAIWO comprises two core components. The first is an emotional world model, which captures long-term emotional trajectories from historical interactions to inform anticipatory guidance. The second is a causality-enhanced imitation learning module, which infers latent causal dependencies to facilitate coherent strategy transitions and mitigate compounding errors typical of conventional imitation learning. By incorporating the final latent variables into the response decoder, CAIWO dynamically adjusts the strategies and generates emotionally resonant responses. Extensive experiments on the ESConv benchmark demonstrate that CAIWO outperforms state-of-the-art baselines by 8.6%, significantly improving the generation of responses that align with seekers’ emotional development and psychological needs. Mingzheng Li, Fei Wang 0073, Kun Li 0008, Yanyan Wei, Yiqi Nie, Yanbin Hao, Xun Yang 0001, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | IAMAgent: Toward an Interactive and Adaptive Multi-Agent System for Image RestorationabstractExisting image restoration and enhancement (IRE) methods suffer from three fundamental limitations: 1) they present a high technical barrier, requiring expert knowledge and lacking intuitive natural language control; 2) they are inflexible and poorly adaptable, as models are typically designed for single, specific degradations and fail on complex or mixed real-world scenarios; and 3) they lack interactivity and ignore subjectivity, operating as "closed-box" tools that cannot incorporate human feedback or understand nuanced user intentions. To overcome these challenges, we pioneer a novel paradigm: a Multi-Agent System (MAS) for interactive and adaptive image restoration. We design and implement a prototype system, Interactive and Adaptive Multi-Agent System (IAMAgent), which orchestrates a team of specialized agents to collaboratively solve complex IRE tasks. At its core, a Manager Agent, driven by a Large Language Model, interprets user commands, devises strategies, and allocates sub-tasks. It directs a Perception Agent for degradation diagnosis, a suite of specialized Execution Agents that encapsulate various low-level vision models, and a Critique Agent for automated quality assessment. This collaborative framework enables an innovative, language-driven, and human-in-the-loop optimization process. Our work is the first to introduce the MAS paradigm to the IRE domain, transforming it from a collection of static tools into a dynamic, user-centric, and intelligent system. We demonstrate that IAMAgent not only significantly enhances restoration performance and adaptability but also bridges the critical gap between high-level human intention and low-level vision tasks. Yanyan Wei, Yilin Zhang 0012, Jiahuan Ren, Xiaogang Xu 0002, Zenglin Shi, Zhao Zhang 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Exploiting Regional Information Transformer for Single Image DerainingabstractTransformer-based Single Image Deraining (SID) methods have achieved remarkable success, primarily attributed to their robust capability in capturing long-range interactions. However, we've noticed that current methods handle rain-affected and unaffected regions concurrently, overlooking the disparities between these areas, resulting in confusion between rain streaks and background parts, and inabilities to obtain effective interactions, ultimately resulting in suboptimal deraining outcomes. To address the above issue, we introduce the Region Transformer (Regformer), a novel SID method that underlines the importance of independently processing rain-affected and unaffected regions while considering their combined impact for high-quality image reconstruction. The crux of our method is the innovative Region Transformer Block (RTB), which integrates a Region Masked Attention (RMA) mechanism and a Mixed Gate Forward Block (MGFB). Our RTB is used for attention selection of rain-affected and unaffected regions and local modeling of mixed scales. The RMA generates attention maps tailored to these two regions and their interactions, enabling our model to capture comprehensive features essential for rain removal. To better recover high-frequency textures and capture more local details, we develop the MGFB as a compensation module to complete local mixed scale modeling. Extensive experiments demonstrate that our model reaches state-of-the-art performance, significantly improving the image deraining quality. Our code and trained models are publicly available athttps://github.com/ztMotaLee/Regformer. Baiang Li, Zhao Zhang 0001, Xiaogang Xu 0002, Yanyan Wei, Jicong Fan 0001, Meng Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image FusionabstractMultimodal Image Fusion (MMIF) aims to integrate complementary information from different imaging modalities to overcome the limitations of individual sensors. It enhances image quality and facilitates downstream applications such as remote sensing, medical diagnostics, and robotics. Despite significant advancements, current MMIF methods still face challenges such as modality misalignment, high-frequency detail destruction, and task-specific limitations. To address these challenges, we propose AdaSFFuse, a novel framework for task-generalized MMIF through adaptive cross-domain co-fusion learning. AdaSFFuse introduces two key innovations: the Adaptive Approximate Wavelet Transform (AdaWAT) for frequency decoupling, and the Spatial-Frequency Mamba Blocks for efficient multimodal fusion. AdaWAT adaptively separates the high- and low-frequency components of multimodal images from different scenes, enabling fine-grained extraction and alignment of distinct frequency characteristics for each modality. The Spatial-Frequency Mamba Blocks facilitate cross-domain fusion in both spatial and frequency domains, enhancing this process. These blocks dynamically adjust through learnable mappings to ensure robust fusion across diverse modalities. By combining these components, AdaSFFuse improves the alignment and integration of multimodal features, reduces frequency loss, and preserves critical details. Extensive experiments on four MMIF tasks-Infrared-Visible Image Fusion (IVF), Multi-Focus Image Fusion (MFF), Multi-Exposure Image Fusion (MEF), and Medical Image Fusion (MIF)-demonstrate AdaSFFuse's superior fusion performance, ensuring both low computational cost and a compact network, offering a strong balance between performance and efficiency. The code will be publicly available athttps://github.com/Zhen-yu-Liu/AdaSFFuse. Kun Li 0008, Yu Wang 0224, Yuwei Wang 0002, Yanyan Wei, Fei Wang 0073 |
IEEE Trans. Multim. | 6 |
| 2025 | Sign-IDD: Iconicity Disentangled Diffusion for Sign Language ProductionabstractSign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit them, which overlooks the relative positional relationships among joints. To this end, we provide a new perspective, constraining joint associations and gesture details by modeling the limb bones to improve the accuracy and naturalness of the generated poses. In this work, we propose a pioneering iconicity disentangled diffusion framework, termed Sign-IDD, specifically designed for SLP. Sign-IDD incorporates a novel Iconicity Disentanglement (ID) module to bridge the gap between relative positions among joints. The ID module disentangles the conventional 3D joint representation into a 4D bone representation, comprising the 3D spatial direction vector and 1D spatial distance vector between adjacent joints. Additionally, an Attribute Controllable Diffusion (ACD) module is introduced to further constrain joint associations, in which the attribute separation layer aims to separate the bone direction and length attributes, and the attribute control layer is designed to guide the pose generation by leveraging the above attributes. The ACD module utilizes the gloss embeddings as semantic conditions and finally generates sign poses from noise embeddings. Extensive experiments on PHOENIX14T and USTC-CSL datasets validate the effectiveness of our method. Shengeng Tang, Dan Guo 0001, Yanyan Wei, Feng Li 0037, Richang Hong |
AAAI | 4 |
| 2025 | High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity ConsistencyabstractThis paper tackles the challenge of stereoscopic image rain removal by focusing on enhancing texture integrity and disparity consistency. Existing stereoscopic rain removal techniques often fall short due to 1) disruptions in texture coherence caused by complex rain streaks, and 2) inaccuracies in disparity estimation from inadequate feature fusion. To overcome these limitations, we introduce the StereoIRR method, which incorporates: 1) a Long-range and Cross-view Interaction (LCI) framework that preserves texture integrity by mitigating rain’s adverse effects on stereoscopic features, and 2) a Dual-view Mutual Attention mechanism that ensures disparity consistency by generating precise mutual attention maps for cross-view feature fusion. Our approach not only maintains the integrity of stereoscopic textures but also significantly reduces errors in disparity estimation. Extensive experiments demonstrate that StereoIRR consistently outperforms state-of-the-art monocular and stereoscopic methods on multiple benchmark datasets. Yanyan Wei, Zhao Zhang 0001, Zhong-Qiu Zhao, Yang Zhao 0002, Richang Hong, Yi Yang 0001, Meng Wang 0001 |
ICASSP | 1 |
| 2025 | Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable RenderingabstractAccurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has remained a significant challenge. To address these issues, this paper proposes CAPED, an end-to-end robust Category-level Articulated object Pose Estimator integrated differentiable rendering. Given partial point cloud as input, CAPED outputs the per-part 6D pose for articulation. Specifically, with the proposed joint-centric modeling manner, CAPED firstly estimates the pose for the free part. Afterward, we canonicalize the input point cloud to estimate constrained parts’ poses by predicting the joint parameters and states as replacements. For further refinement, we propose a differentiable rendering scheme for pose optimization. Evaluations of the ArtImage and RobotArm datasets demonstrate that CAPED exhibits outstanding effectiveness and generalization in tasks ranging from synthetic data to real-world scenarios. We will publicly release the code. Li Zhang 0104, Yukang Huo, Lin Wu 0001, Yanyan Wei, Harshal Suresh Shende, Liu Liu 0012, Linlin Ou |
ICASSP | 6 |
| 2025 | Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emotion recognition. To overcome this limitation, we propose TF-Mamba, a novel multi-domain framework that captures emotional expressions in both temporal and frequency dimensions. Concretely, we propose a temporal-frequency mamba block to extract temporal- and frequency-aware emotional features, achieving an optimal balance between computational efficiency and model expressiveness. Besides, we design a Complex Metric-Distance Triplet (CMDT) loss to enable the model to capture representative emotional clues for SER. Extensive experiments on the IEMOCAP and MELD datasets show that TF-Mamba surpasses existing methods in terms of model size and latency, providing a more practical solution for future SER applications. Fei Wang 0067, Kun Li 0008, Yanyan Wei, Shengeng Tang, Shu Zhao 0005, Xiao Sun 0003 |
ICASSP | 4 |
| 2025 | Generalizable and Actionable Part Detection and Manipulation with SAM-rectified Segmentation and Iterative Pose RefinementabstractThe ability to perform cross-category object perception and manipulation is highly desirable in building intelligent robots. One promising approach is to define the concept of Generalizable and Actionable Parts (GAParts), such as buttons and handles, on both seen and unseen object categories. However, the accurate cross-category perception of GAParts is still challenging due to the large inter-category object shape variations. To address this issue, we introduce SAMIR, a novel framework using SAM-rectified segmentation and Iterative pose Refinement for GAPart detection and manipulation. Firstly, we introduce a Segment Anything (SAM) segmentation prior to rectify the unconfident, fragmented GAPart instance proposals. Secondly, in addition to the zero-shot generalization of the SAM foundation model, we further finetune it with a lightweight adaptor model on our task dataset. Finally, we propose an iterative pose refinement procedure that improves the accuracy of GAPart pose estimation. Our perception experiments on GAPartNet dataset show that SAMIR consistently outperforms the baseline method on instance segmentation and pose estimation tasks. Our manipulation experiments in Sapien simulator illustrate that SAMIR leads to an improved manipulation success rate. We also deploy our method to a real robot for real-world manipulation. Our code and video are available at sites.google.com/view/samir-gapart. Sucheng Qian, Li Zhang 0104, Yanyan Wei, Liu Liu 0012, Cewu Lu |
IROS | 3 |
| 2025 | Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action RecognitionabstractMicro-Actions (MAs) are an important form of non-verbal communication in social interactions, with potential applications in human emotional analysis. However, existing methods in Micro-Action Recognition often overlook the inherent subtle changes in MAs, which limits the accuracy of distinguishing MAs with subtle changes. To address this issue, we present a novel Motion-guided Modulation Network (MMN) that implicitly captures and modulates subtle motion cues to enhance spatial-temporal representation learning. Specifically, we introduce a Motion-guided Skeletal Modulation module (MSM) to inject motion cues at the skeletal level, acting as a control signal to guide spatial representation modeling. In parallel, we design a Motion-guided Temporal Modulation module (MTM) to incorporate motion information at the frame level, facilitating the modeling of holistic motion patterns in micro-actions. Finally, we propose a motion consistency learning strategy to aggregate the motion cues from multi-scale features for micro-action classification. Experimental results on the Micro-Action 52 and iMiGUE datasets demonstrate that MMN achieves state-of-the-art performance in skeleton-based micro-action recognition, underscoring the importance of explicitly modeling subtle motion cues. The code will be available at https://github.com/momiji-bit/MMN Jihao Gu, Kun Li 0008, Fei Wang 0073, Yanyan Wei, Zhiliang Wu, Hehe Fan, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2025 | From Outline to Detail: An Hierarchical End-to-end Framework for Coherent and Consistent Visual Novel Generation and AssemblyabstractAs a form of multimedia creation, visual novel (VN) conveys engaging narratives through the integrated presentation of text, images, and music, and has shown promise across various application domains. Recent advances in generative AI have fueled interest in automating VN creation using LLMs and other foundation models. However, fully end-to-end VN creation (i.e., from user description to executable VN) remains underexplored and presents several key challenges: 1) the hallucination and limited capacity of LLMs hinder the generation of long and coherent plots; 2) current models lack effective mechanisms for ensuring cross-modal consistency between plot, visual, and audio elements. To address these issues, we propose a hierarchical end-to-end framework for automatic VN generation and assembly, which employs an outline-guided autoregressive generation mechanism that transforms high-level user prompts into coherent plots, while a vision LLM-based self-correction mechanism ensures consistency between multimedia assets and plot content. Additionally, we introduce a script validation mechanism to ensure the executable of the final VN application. Experiments demonstrate that our framework generates high-quality VN applications with coherent storylines and consistent multimedia content. Yilin Zhang 0012, Yanyan Wei, Zhao Zhang 0001, Jicong Fan 0001, Haijun Zhang 0002, Shuicheng Yan |
ACM Multimedia | 2 |
| 2025 | Individual/joint deblurring and low-light image enhancement in one go via unsupervised deblurring paradigm
Suiyi Zhao, Zhao Zhang 0001, Yanyan Wei, Jicong Fan 0001, Yang Zhao 0002, Shuicheng Yan, Meng Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2025 | Leveraging vision-language prompts for real-world image restoration and enhancement
Yanyan Wei, Yilin Zhang 0012, Kun Li 0008, Fei Wang 0067, Shengeng Tang, Zhao Zhang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2025 | Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather RemovalabstractUniversal adverse weather removal (UAWR) seeks to address various weather degradations within a unified framework. Recent methods are inspired by prompt learning using pre-trained vision-language models (e.g., CLIP), leveraging degradation-aware prompts to facilitate weather-free image restoration, yielding significant improvements. In this work, we propose CyclicPrompt, an innovative cyclic prompt approach designed to enhance the effectiveness, adaptability, and generalizability of UAWR. CyclicPrompt comprises two key components: 1) a composite context prompt that integrates weather-related information and context-aware representations into the network to guide restoration. This prompt differs from previous methods by marrying learnable input-conditional vectors with weather-specific knowledge, thereby improving adaptability across various degradations and 2) the erase-and-paste mechanism, after the initial guided restoration, substitutes weather-specific knowledge with constrained restoration priors, inducing high-quality weather-free concepts into the composite prompt to further fine-tune the restoration process. Therefore, we can form a cyclic "Prompt-Restore-Prompt" pipeline that adeptly harnesses weather-specific knowledge, textual contexts, and reliable textures. Extensive experiments on synthetic and real-world datasets validate the superior performance of CyclicPrompt. The code is available at: https://github.com/RongxinL/CyclicPrompt. Rongxin Liao, Feng Li 0037, Yanyan Wei, Zenglin Shi, Le Zhang 0001, Huihui Bai 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Temporal Boundary Awareness Network for Repetitive Action CountingabstractRepetitive Action Counting (RAC) is a critical and challenging task in video analysis, aiming to count the number of repeated actions in videos accurately. Existing methods typically generate a Temporal Self-similarity Matrix (TSM) as an intermediate representation to predict the number of repetitive actions. While this simplifies the process, it often overlooks the variable lengths between action cycles and the phenomenon of motion interruptions. The period inconsistency problem caused by the change in the action period and the motion interruption problem resulting from the motion pause are the two main challenges that affect the accuracy of RAC in complex scenes. To address these challenges, we propose a novel framework. First, we construct a boundary-aware encoder equipped with a temporal pyramid structure to build multi-scale video features, capturing the period information of different lengths of repetitive actions to solve the period inconsistency problem. Next, a cycle and boundary attention module is followed by each layer in the pyramid to enhance these multi-scale features with periodic and event boundary information. Finally, we design a gated density estimator to generate the actionness score for each frame that reflects the probability of the corresponding time point being within the motion cycle. These scores are used to weight features to reduce the impact of noise frames without actions present and solve the motion interruption problem for better density prediction. Extensive experiments conducted on public datasets demonstrate the effectiveness of our method. The source code will be available at https://github.com/zqzhang2023/TBANRAC . Zhenqiang Zhang, Kun Li 0008, Shengeng Tang, Yanyan Wei, Fei Wang 0073, Jinxing Zhou, Dan Guo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Repetitive Action Counting with Feature Interaction Enhancement and Adaptive Gate Fusion
Kun Li 0008, Yanyan Wei, Fei Wang 0073, Jinxing Zhou, Dan Guo 0001 |
MMAsia | 3 |
| 2024 | Progressive Stereo Image Dehazing Network via Cross-View Region InteractionabstractStereo image dehazing aims to restore haze-free images by leveraging the complementary information contained in binocular images. Current methods primarily focus on designing image-level modules and pipelines to utilize complementary information between the left and right-view images. However, these image-level cross-view interactions overlook regional differences in haze concentration and stereo image disparity maps. Consequently, we propose a Progressive Stereo Image Dehazing Network via Cross-view Region Interaction, termed PSIDNet, which fully considers the internal characteristics and external manifestation of haze and disparity, and explicitly addresses the stereo image dehazing task by a regional-aware interactive mechanism. Specifically, we divide hazy images into regions and independently interact with left and right-view information at region levels, meaning weights are not shared across regional patches. This approach allows us to treat different regions with different priorities, i.e., concentrate on regional patches with heavier haze concentration and larger disparities, hence enabling more accurate restoration of hazy images. Furthermore, we introduce an effective cross-view region interactive block that extracts information based on the channel dimension of dual views and later adopts matrix multiplication to generate mutual attention maps based on the fused features. Extensive experiments on synthetic and real-scenario datasets demonstrate the efficacy of our method, compared to other related monocular and stereo image dehazing and restoration methods. Our code will be released publicly at https://github.com/Alvin2112/PSIDNet. Junhu Wang, Yanyan Wei, Zhao Zhang 0001, Jicong Fan 0001, Yang Zhao 0002, Yi Yang 0001, Meng Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Data-Driven single image deraining: A Comprehensive review and new perspectives
Zhao Zhang 0001, Yanyan Wei, Haijun Zhang 0002, Yi Yang 0001, Shuicheng Yan, Meng Wang 0001 |
Pattern Recognit. | 2 |
| 2023 | Low-Latency Dimensional Expansion and Anomaly Detection Empowered Secure IoT NetworkabstractThe Internet of Things (IoT) consists of a myriad of smart devices and offers tremendous innovation opportunities in industry, homes, and businesses to enhance the productivity and the quality of life. However, ecosystem of infrastructures and the services associated with IoT devices have introduced a new set of vulnerabilities and threats, resulting in abnormal values of information collected by sensors, jeopardizing system security. To secure sensor networks, it must be possible to detect such anomalies or sequences of patterns in IoT devices that significantly deviate from normal behavior. To perform this task, this paper proposes a real-time streaming anomaly detection method based on a Bloom filter combined with hashing. This method expands the data dimensions through a hashing algorithm, and then adopts competitive learning (Winner-Take-All) to build a multi-layer Bloom Filter anomaly detection model. The feasibility of the proposed algorithm is verified theoretically using two datasets, KDD (to detect anomalies at the TCP/IP network level) and Credit (to detect anomalies during credit card transactions). The simulation results show that the proposed in this paper can effectively identify anomalies in the simulation data streams, with almost 95% accuracy for both datasets. Wenhao Shao, Yanyan Wei, Praboda Rajapaksha, Dun Li, Zhigang Luo, Noël Crespi |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | COVAD: Content-Oriented Video Anomaly Detection using a Self-Attention based Deep Learning ModelabstractVideo anomaly detection has always been a hot topic and attracting an increasing amount of attention. Much of the existing methods on video anomaly detection depend on processing the entire video rather than considering only the significant context. This paper proposes a novel video anomaly detection method named COVAD, which mainly focuses on the region of interest in the video instead of the entire video. Our proposed COVAD method is based on an auto-encoded convolutional neural network and coordinated attention mechanism, which can effectively capture meaningful objects in the video and dependencies between different objects. Relying on the existing memory-guided video frame prediction network, our algorithm can more effectively predict the future motion and appearance of objects in the video. Our proposed algorithm obtained better experimental results on multiple data sets and outperformed the baseline models considered in our analysis. At the same time we improve a visual test that can provide pixel-level anomaly explanations. Wenhao Shao, Praboda Rajapaksha, Yanyan Wei, Dun Li, Noël Crespi, Zhigang Luo |
Virtual Real. Intell. Hardw. | 3 |
| 2022 | Robust Attention Deraining Network for Synchronous Rain Streaks and Raindrops RemovalabstractSynchronous Rain streaks and Raindrops Removal (SR^3) is a hard and challenging task, since rain streaks and raindrops are two wildly divergent real-world phenomena with different optical properties and mathematical distributions. As such, most existing data-driven deep Singe Image Deraining (SID) methods only focus on one of them. Although there are only a few existing SR^3 methods, they still suffer from blur textures and unknown noise in reality due to weak robustness and generalization ability. In this paper, we will propose a new and universal SID model with novel modules, termed Robust Attention Deraining Network (RadNet), with strong robustness and generalization ability that are reflected in two main aspects. (1) RadNet can restore different rain degenerations, including raindrops, rain streaks, or both; (2) RadNet can adapt to different data strategies, including single-type, superimposed-type, and blended-type. The generalization ability is also demonstrated by the performance of dealing with real rain images. Specifically, we first design a lightweight and robust attention module (RAM) with a universal attention mechanism for coarse rain removal, and then present a new deep refining module (DRM) with multi-scale blocks for precise rain removal. To solve the inconsistent labels of real scenario data, we also introduce a flow & warp module (FWM) into the network, which can greatly improve the performance of real scenario data via optical flow prediction and alignment. The whole process is unified in a network to ensure sufficient robustness and strong generalization ability. We evaluated the performance of our method under a variety of data strategies, and extensive experiments demonstrated that our RadNet could outperform other state-of-the-art SID methods. Yanyan Wei, Zhao Zhang 0001, Mingliang Xu 0001, Richang Hong, Jicong Fan 0001, Shuicheng Yan |
ACM Multimedia | 1 |
| 2022 | SGINet: Toward Sufficient Interaction Between Single Image Deraining and Semantic SegmentationabstractData-driven single image deraining (SID) models have achieved greater progress by simulations, but there is still a large gap between current deraining performance and practical high-level applications, since high-level semantic information is usually neglected in current studies. Although few studies jointly considered high-level tasks (e.g., segmentation) to enable the model to learn more high-level information, there are two obvious shortcomings. First, they require the segmentation labels for training, limiting their operations on other datasets without high-level labels. Second, high- and low-level information are not fully interacted, hence having limited improvement in both deraining and segmentation tasks. In this paper, we propose a Semantic Guided Interactive Network (SGINet), which considers the sufficient interaction between SID and semantic segmentation using a three-stage deraining manner, i.e., coarse deraining, semantic information extraction, and semantics guided deraining. Specifically, a Full Resolution Module (FRM) without down-/up-sampling is proposed to predict the coarse deraining images without context damage. Then, a Segmentation Extracting Module (SEM) is designed to extract accurate semantic information. We also develop a novel contrastive semantic discovery (CSD) loss, which can instruct the process of semantic segmentation without real semantic segmentation labels. Finally, a triple-direction U-net-based Semantic Interaction Module (SIM) takes advantage of the coarse deraining images and semantic information for fully interacting low-level with high-level tasks. Extensive simulations on the newly-constructed complex datasets Cityscapes_syn and Cityscapes_real demonstrated that our model could obtain more promising results. Overall, our SGINet achieved SOTA deraining and segmentation performance in both simulation and real-scenario data, compared with other representative SID methods. Yanyan Wei, Zhao Zhang 0001, Richang Hong, Yi Yang 0001, Meng Wang 0001 |
ACM Multimedia | 1 |
| 2021 | Semi-Deraingan: A New Semi-Supervised Single Image DerainingabstractAlthough supervised single image deraining (SID) have obtained impressive results, they still cannot obtain satisfactory results on real images for the weak generalization of rain removal capacity. In this paper, we mainly discuss the semi-supervised SID and propose a new GAN-based deraining network called Semi-DerainGAN, which can use both synthetic and real data in a uniform network based on two supervised and unsupervised processes. For this task, a semi-supervised rain streak learner termed SSRML sharing the same parameters of both processes is derived, which makes the real images contribute more rain streak information, so that the resulted model has a strong generalization power to the real SID task. We also contribute a new real-world rain image dataset called Real200 to alleviate the difference between both synthetic and real image domains. Extensive results on public datasets show that our model can obtain competitive results, especially on the real rain images. Yanyan Wei, Zhao Zhang 0001, Yang Wang 0023, Haijun Zhang 0002, Ming-Bo Zhao, Mingliang Xu 0001, Meng Wang 0001 |
ICME | 1 |
| 2021 | Low and non-uniform illumination color image enhancement using weighted guided image filteringabstractIn the state of the art, grayscale image enhancement algorithms are typically adopted for enhancement of RGB color images captured with low or non-uniform illumination. As these methods are applied to each RGB channel independently, imbalanced inter-channel enhancements (color distortion) can often be observed in the resulting images. On the other hand, images with non-uniform illumination enhanced by the retinex algorithm are prone to artifacts such as local blurring, halos, and over-enhancement. To address these problems, an improved RGB color image enhancement method is proposed for images captured under non-uniform illumination or in poor visibility, based on weighted guided image filtering (WGIF). Unlike the conventional retinex algorithm and its variants, WGIF uses a surround function instead of a Gaussian filter to estimate the illumination component; it avoids local blurring and halo artifacts due to its anisotropy and adaptive local regularization. To limit color distortion, RGB images are first converted to HSI (hue, saturation, intensity) color space, where only the intensity channel is enhanced, before being converted back to RGB space by a linear color restoration algorithm. Experimental results show that the proposed method is effective for both RGB color and grayscale images captured under low exposure and non-uniform illumination, with better visual quality and objective evaluation scores than from comparator algorithms. It is also efficient due to use of a linear color restoration algorithm. Qi Mu, Yanyan Wei, Zhanli Li |
Comput. Vis. Media | 3 |
| 2021 | DerainCycleGAN: Rain Attentive CycleGAN for Single Image Deraining and RainmakingabstractSingle Image Deraining (SID) is a relatively new and still challenging topic in emerging vision applications, and most of the recently emerged deraining methods use the supervised manner depending on the ground-truth (i.e., using paired data). However, in practice it is rather common to encounter unpaired images in real deraining task. In such cases, how to remove the rain streaks in an unsupervised way will be a challenging task due to lack of constraints between images and hence suffering from low-quality restoration results. In this paper, we therefore explore the unsupervised SID issue using unpaired data, and propose a new unsupervised framework termed DerainCycleGAN for single image rain removal and generation, which can fully utilize the constrained transfer learning ability and circulatory structures of CycleGAN. In addition, we design an unsupervised rain attentive detector (UARD) for enhancing the rain information detection by paying attention to both rainy and rain-free images. Besides, we also contribute a new synthetic way of generating the rain streak information, which is different from the previous ones. Specifically, since the generated rain streaks have diverse shapes and directions, existing derianing methods trained on the generated rainy image by this way can perform much better for processing real rainy images. Extensive experimental results on synthetic and real datasets show that our DerainCycleGAN is superior to current unsupervised and semi-supervised methods, and is also highly competitive to the fully-supervised ones. Yanyan Wei, Zhao Zhang 0001, Yang Wang 0023, Mingliang Xu 0001, Yi Yang 0001, Shuicheng Yan, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | A Coarse-to-Fine Multi-stream Hybrid Deraining Network for Single Image DerainingabstractSingle image deraining task is still a very challenging task due to its ill-posed nature in reality. Recently, researchers have tried to fix this issue by training the CNN-based end-to-end models, but they still cannot extract the negative rain streaks from rainy images precisely, which usually leads to an over de-rained or under de-rained result. To handle this issue, this paper proposes a new coarse-to-fine single image deraining framework termed Multi-stream Hybrid Deraining Network (shortly, MH-DerainNet). To obtain the negative rain streaks during training process more accurately, we present a new module named dual path residual dense block, i.e., Residual path and Dense path. The Residual path is used to reuse com-mon features from the previous layers while the Dense path can explore new features. In addition, to concatenate different scaled features, we also apply the idea of multi-stream with shortcuts between cascaded dual path residual dense block based streams. To obtain more distinct derained images, we combine the SSIM loss and perceptual loss to preserve the per-pixel similarity as well as preserving the global structures so that the deraining result is more accurate. Extensive experi-ments on both synthetic and real rainy images demonstrate that our MH-DerainNet can deliver significant improvements over several recent state-of-the-art methods. Yanyan Wei, Zhao Zhang 0001, Haijun Zhang 0002, Richang Hong, Meng Wang 0001 |
ICDM | 1 |