EDBT 2026 Demo / reviewers in the wild / expert
Zhenyao Wu
dblp:259/5249
· DBLP profile ↗
21ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-5827-7155ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shadow vanishing point detection via combined human/shadow adaptive modulation
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002 |
Signal Process. Image Commun. | 4 |
| 2025 | Acquire and then Adapt: Squeezing out Text-to-Image Model for Image RestorationabstractRecently, pre-trained text-to-image (T2I) models have been extensively adopted for real-world image restoration because of their powerful generative prior. However, controlling these large models for image restoration usually requires a large number of high-quality images and immense computational resources for training, which is costly and not privacy-friendly. In this paper, we find that the well-trained large T2I model (i.e., Flux) is able to produce a variety of high-quality images aligned with real-world distributions, offering an unlimited supply of training samples to mitigate the above issue. Specifically, we proposed a training data construction pipeline for image restoration, namely FluxGen, which includes unconditional image generation, image selection, and degraded image simulation. A novel light-weighted adapter (FluxIR) with squeeze-and-excitation layers is also carefully designed to control the large Diffusion Transformer (DiT)-based T2I model so that reasonable details can be restored. Experiments demonstrate that our proposed method enables the Flux model to adapt effectively to real-world image restoration tasks, achieving superior scores and visual quality on both synthetic and real-world degradation datasets - at only about 8.5% of the training cost compared to current approaches. Junyuan Deng, Xinyi Wu 0002, Yongxing Yang, Congchao Zhu, Song Wang 0002, Zhenyao Wu |
CVPR | 6 |
| 2025 | Coarse-to-fine text injecting for realistic image super-resolution
Chao Bai, Zhenyao Wu, Xinyi Wu 0002, Qi Zou 0001, Song Wang 0002 |
Neurocomputing | 3 |
| 2024 | Estimating intrinsic characteristics of images for shadow removal
Hui Yin 0002, Jin Wan, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002 |
Comput. Graph. | 5 |
| 2024 | CRFormer: A cross-region transformer for shadow removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002 |
Image Vis. Comput. | 3 |
| 2023 | Parametric Surface Constrained Upsampler Network for Point CloudabstractDesigning a point cloud upsampler, which aims to generate a clean and dense point cloud given a sparse point representation, is a fundamental and challenging problem in computer vision. A line of attempts achieves this goal by establishing a point-to-point mapping function via deep neural networks. However, these approaches are prone to produce outlier points due to the lack of explicit surface-level constraints. To solve this problem, we introduce a novel surface regularizer into the upsampler network by forcing the neural network to learn the underlying parametric surface represented by bicubic functions and rotation functions, where the new generated points are then constrained on the underlying surface. These designs are integrated into two different networks for two tasks that take advantages of upsampling layers -- point cloud upsampling and point cloud completion for evaluation. The state-of-the-art experimental results on both tasks demonstrate the effectiveness of the proposed method. The implementation code will be available at https://github.com/corecai163/PSCU. Pingping Cai, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002 |
AAAI | 2 |
| 2023 | Few-Shot 3D Point Cloud Semantic Segmentation via Stratified Class-Specific Attention Based Transformer Networkabstract3D point cloud semantic segmentation aims to group all points into different semantic categories, which benefits important applications such as point cloud scene reconstruction and understanding. Existing supervised point cloud semantic segmentation methods usually require large-scale annotated point clouds for training and cannot handle new categories. While a few-shot learning method was proposed recently to address these two problems, it suffers from high computational complexity caused by graph construction and inability to learn fine-grained relationships among points due to the use of pooling operations. In this paper, we further address these problems by developing a new multi-layer transformer network for few-shot point cloud semantic segmentation. In the proposed network, the query point cloud features are aggregated based on the class-specific support features in different scales. Without using pooling operations, our method makes full use of all pixel-level features from the support samples. By better leveraging the support features for few-shot learning, the proposed method achieves the new state-of-the-art performance, with 15% less inference time, over existing few-shot 3D point cloud segmentation models on the S3DIS dataset and the ScanNet dataset. Our code is available at https://github.com/czzhang179/SCAT. Canyu Zhang 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002 |
AAAI | 2 |
| 2023 | A One-Stage Domain Adaptation Network With Image Alignment for Unsupervised Nighttime Semantic SegmentationabstractIn this paper, we tackle the problem of semantic segmentation for nighttime images that plays an equally important role as that for daytime images in autonomous driving, but is also much more challenging due to very poor illuminations and scarce annotated datasets. It can be treated as an unsupervised domain adaptation (UDA) problem, i.e., applying other labeled dataset taken in the daytime to guide the network training meanwhile reducing the domain shift, so that the trained model can generalize well to the desired domain of nighttime images. However, current general-purpose UDA approaches are insufficient to address the significant appearance difference between the day and night domains. To overcome such a large domain gap, we propose a novel domain adaptation network "DANIA" for nighttime semantic image segmentation by leveraging a labeled daytime dataset (the source domain) and an unlabeled dataset that contains coarsely aligned day-night image pairs (the target daytime and nighttime domains). These three domains are used to perform a multi-target adaptation via adversarial training in the network. Specifically, for the unlabeled day-night image pairs, we use the pixel-level predictions of static object categories on a daytime image as a pseudo supervision to segment its counterpart nighttime image. We also include a step of image alignment to relieve the inaccuracy caused by the misalignment between day-night image pairs by estimating a flow to refine the pseudo supervision produced by daytime images. Finally, a re-weighting strategy is applied to further improve the predictions, especially boosting the prediction accuracy of small objects. The proposed DANIA is a one-stage adaptation framework for nighttime semantic segmentation, which does not train additional day-night image transfer models as a separate pre-processing stage. Extensive experiments on Dark Zurich and Nighttime Driving datasets show that our DANIA achieves state-of-the-art performance for nighttime semantic segmentation. Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Style Mixing and Patchwise Prototypical Matching for One-Shot Unsupervised Domain Adaptive Semantic SegmentationabstractIn this paper, we tackle the problem of one-shot unsupervised domain adaptation (OSUDA) for semantic segmentation where the segmentors only see one unlabeled target image during training. In this case, traditional unsupervised domain adaptation models usually fail since they cannot adapt to the target domain with over-fitting to one (or few) target samples. To address this problem, existing OSUDA methods usually integrate a style-transfer module to perform domain randomization based on the unlabeled target sample, with which multiple domains around the target sample can be explored during training. However, such a style-transfer module relies on an additional set of images as style reference for pre-training and also increases the memory demand for domain adaptation. Here we propose a new OSUDA method that can effectively relieve such computational burden. Specifically, we integrate several style-mixing layers into the segmentor which play the role of style-transfer module to stylize the source images without introducing any learned parameters. Moreover, we propose a patchwise prototypical matching (PPM) method to weighted consider the importance of source pixels during the supervised training to relieve the negative adaptation. Experimental results show that our method achieves new state-of-the-art performance on two commonly used benchmarks for domain adaptive semantic segmentation under the one-shot setting and is more efficient than all comparison approaches. Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002 |
AAAI | 2 |
| 2022 | Style-Guided Shadow Removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002 |
ECCV (19) | 3 |
| 2022 | Is It Necessary to Transfer Temporal Knowledge for Domain Adaptive Video Semantic Segmentation?
Xinyi Wu 0002, Zhenyao Wu, Jin Wan, Lili Ju, Song Wang 0002 |
ECCV (27) | 2 |
| 2022 | SiamDoGe: Domain Generalizable Semantic Segmentation Using Siamese Network
Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Lili Ju, Song Wang 0002 |
ECCV (38) | 1 |
| 2022 | Background-Insensitive Scene Text Recognition with Text Semantic Segmentation
Zhenyao Wu, Xinyi Wu 0002, Greg Wilsbacher, Song Wang 0002 |
ECCV (25) | 2 |
| 2022 | Crossmodal Few-shot 3D Point Cloud Semantic SegmentationabstractRecently, few-shot 3D point cloud semantic segmentation methods have been introduced to mitigate the limitations of existing fully supervised approaches, i.e., heavy dependence on labeled 3D data and poor capacity to generalize to new categories. However, those few-shot learning methods need one or few labeled data as support for testing. In practice, such data labeling usually requires manual annotation of large-scale points in 3D space, which can be very difficult and laborious. To address this problem, in this paper we introduce a novel crossmodal few-shot learning approach for 3D point cloud semantic segmentation. In this approach, the point cloud to be segmented is taken as query while one or few labeled 2D RGB images are taken as support to guide the segmentation of query. This way, we only need to annotate on a few 2D support images for the categories of interest. Specifically, we first convert the 2D support images into 3D point cloud format based on both appearance and the estimated depth information. We then introduce a co-embedding network for extracting the features of support and query, both from 3D point cloud format, to fill their domain gap. Finally, we compute the prototypes of support and employ cosine similarity between the prototypes and the query features for final segmentation. Experimental results on two widely-used benchmarks show that, with one or few labeled 2D images as support, our proposed method achieves competitive results against existing few-shot 3D point cloud semantic segmentation methods. Zhenyao Wu, Xinyi Wu 0002, Canyu Zhang 0002, Song Wang 0002 |
ACM Multimedia | 2 |
| 2021 | Binaural Audio-Visual LocalizationabstractLocalizing sound sources in a visual scene has many important applications and quite a few traditional or learning-based methods have been proposed for this task. Humans have the ability to roughly localize sound sources within or beyond the range of the vision using their binaural system. However most existing methods use monaural audio, instead of binaural audio, as a modality to help the localization. In addition, prior works usually localize sound sources in the form of object-level bounding boxes in images or videos and evaluate the localization accuracy by examining the overlap between the ground-truth and predicted bounding boxes. This is too rough since a real sound source is often only a part of an object. In this paper, we propose a deep learning method for pixel-level sound source localization by leveraging both binaural recordings and the corresponding videos. Specifically, we design a novel Binaural Audio-Visual Network (BAVNet), which concurrently extracts and integrates features from binaural recordings and videos. We also propose a point-annotation strategy to construct pixel-level ground truth for network training and performance evaluation. Experimental results on Fair-Play and YT-Music datasets demonstrate the effectiveness of the proposed method and show that binaural audio can greatly improve the performance of localizing the sound sources, especially when the quality of the visual information is limited. Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002 |
AAAI | 2 |
| 2021 | From Shadow Generation To Shadow RemovalabstractShadow removal is a computer-vision task that aims to restore the image content in shadow regions. While almost all recent shadow-removal methods require shadow-free images for training, in ECCV 2020 Le and Samaras introduces an innovative approach without this requirement by cropping patches with and without shadows from shadow images as training samples. However, it is still laborious and time-consuming to construct a large amount of such unpaired patches. In this paper, we propose a new G2R-ShadowNet which leverages shadow generation for weakly-supervised shadow removal by only using a set of shadow images and their corresponding shadow masks for training. The proposed G2R-ShadowNet consists of three sub-networks for shadow generation, shadow removal and refinement, respectively and they are jointly trained in an end-to-end fashion. In particular, the shadow generation sub-net stylises non-shadow regions to be shadow ones, leading to paired data for training the shadow-removal sub-net. Extensive experiments on the ISTD dataset and the Video Shadow Removal dataset show that the proposed G2R-ShadowNet achieves competitive performances against the current state of the arts and outperforms Le and Samaras’ patch-based shadow-removal method. Hui Yin 0002, Xinyi Wu 0002, Zhenyao Wu, Yang Mi, Song Wang 0002 |
CVPR | 4 |
| 2021 | DANNet: A One-Stage Domain Adaptation Network for Unsupervised Nighttime Semantic SegmentationabstractSemantic segmentation of nighttime images plays an equally important role as that of daytime images in autonomous driving, but the former is much more challenging due to poor illuminations and arduous human annotations. In this paper, we propose a novel domain adaptation network (DANNet) for nighttime semantic segmentation without using labeled nighttime image data. It employs an adversarial training with a labeled daytime dataset and an unlabeled dataset that contains coarsely aligned day-night image pairs. Specifically, for the unlabeled day-night image pairs, we use the pixel-level predictions of static object categories on a daytime image as a pseudo supervision to segment its counterpart nighttime image. We further design a re-weighting strategy to handle the inaccuracy caused by misalignment between day-night image pairs and wrong predictions of daytime images, as well as boost the prediction accuracy of small objects. The proposed DANNet is the first one-stage adaptation framework for nighttime semantic segmentation, which does not train additional day-night image transfer models as a separate pre-processing stage. Extensive experiments on Dark Zurich and Nighttime Driving datasets show that our method achieves state-of-the-art performance for nighttime semantic segmentation. Xinyi Wu 0002, Zhenyao Wu, Hao Guo 0002, Lili Ju, Song Wang 0002 |
CVPR | 2 |
| 2021 | Learning Depth from Single Image Using Depth-Aware Convolution and Stereo KnowledgeabstractEstimating depth from a monocular image has become a very popular task in computer vision for identifying important geometric information of the scene. While its performance has been significantly improved by convolutional neural networks (CNNs) in recent years, depth-estimation accuracy is still unsatisfactory at locations with abrupt depth changes. This is mainly caused by the use of spatially consistent filters in CNNs which directly mix the features of different objects when applied to the pixels near the object borders. Moreover, the performance gap between depth estimation from single image and that from a stereo pair remains quite large due to the ill-posed nature of the former one. In this paper, we propose a new depth-aware convolutional neural network (DACNN) to address these issues. We first design a novel depth-aware convolution operation for DACNN, that can adaptively choose subsets of relevant features for convolutions at each location. Specifically, we compute hierarchical depth features as the guidance, and then estimate the depth map using such depth-aware convolution which can leverage the guidance to adapt the filters. In addition, we also introduce a pre-trained stereo network into DACNN as the teacher to carry out knowledge distillation on the student monocular network with a specially designed loss function. Experimental results on the KITTI online benchmark and Eigen split datasets show that the proposed method achieves the state-of-the-art performance for single-image depth estimation. Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Song Wang 0002, Lili Ju |
ICME | 1 |
| 2020 | SalSAC: A Video Saliency Prediction Model with Shuffled Attentions and Correlation-Based ConvLSTMabstractThe performance of predicting human fixations in videos has been much enhanced with the help of development of the convolutional neural networks (CNN). In this paper, we propose a novel end-to-end neural network “SalSAC” for video saliency prediction, which uses the CNN-LSTM-Attention as the basic architecture and utilizes the information from both static and dynamic aspects. To better represent the static information of each frame, we first extract multi-level features of same size from different layers of the encoder CNN and calculate the corresponding multi-level attentions, then we randomly shuffle these attention maps among levels and multiply them to the extracted multi-level features respectively. Through this way, we leverage the attention consistency across different layers to improve the robustness of the network. On the dynamic aspect, we propose a correlation-based ConvLSTM to appropriately balance the influence of the current and preceding frames to the prediction. Experimental results on the DHF1K, Hollywood2 and UCF-sports datasets show that SalSAC outperforms many existing state-of-the-art methods. Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002 |
AAAI | 2 |
| 2019 | Semantic Stereo Matching With Pyramid Cost VolumesabstractThe accuracy of stereo matching has been greatly improved by using deep learning with convolutional neural networks. To further capture the details of disparity maps, in this paper, we propose a novel semantic stereo network named SSPCV-Net, which includes newly designed pyramid cost volumes for describing semantic and spatial information on multiple levels. The semantic features are inferred by a semantic segmentation subnetwork while the spatial features are derived by hierarchical spatial pooling. In the end, we design a 3D multi-cost aggregation module to integrate the extracted multilevel features and perform regression for accurate disparity maps. We conduct comprehensive experiments and comparisons with some recent stereo matching networks on Scene Flow, KITTI 2015 and 2012, and Cityscapes benchmark datasets, and the results show that the proposed SSPCV-Net significantly promotes the state-of-the-art stereo-matching performance. Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Song Wang 0002, Lili Ju |
ICCV | 1 |
| 2019 | Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular VideosabstractDepth estimation from monocular videos has important applications in many areas such as autonomous driving and robot navigation. It is a very challenging problem without knowing the camera pose since errors in camera-pose estimation can significantly affect the video-based depth estimation accuracy. In this paper, we present a novel SC-GAN network with end-to-end adversarial training for depth estimation from monocular videos without estimating the camera pose and pose change over time. To exploit cross-frame relations, SC-GAN includes a spatial correspondence module which uses Smolyak sparse grids to efficiently match the features across adjacent frames, and an attention mechanism to learn the importance of features in different directions. Furthermore, the generator in SC-GAN learns to estimate depth from the input frames, while the discriminator learns to distinguish between the ground-truth and estimated depth map for the reference frame. Experiments on the KITTI and Cityscapes datasets show that the proposed SC-GAN can achieve much more accurate depth maps than many existing state-of-the-art methods on monocular videos. Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Song Wang 0002, Lili Ju |
ICCV | 1 |