Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ruizheng Wu

dblp:244/2111 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 32% Video understanding and tracking · 26% Image recognition and object detection · 8%
Computer graphics and multimedia
4 papers
Image and video processing · 47% Visual content generation and editing · 36% Computational photography and imaging · 17%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › industrial visual inspection
defect detection
0.812024
An Incremental Unified Framework for Small Defect Inspection · ECCV (31) 2024
Machine learning › Learning paradigms
incremental learning
0.812024
An Incremental Unified Framework for Small Defect Inspection · ECCV (31) 2024
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.812024
GNeRP: Gaussian-guided Neural Reconstruction of Reflective Objects with Noisy Polarization Priors · ICLR 2024
Computer vision › 3D vision
neural radiance field
0.812024
GNeRP: Gaussian-guided Neural Reconstruction of Reflective Objects with Noisy Polarization Priors · ICLR 2024
Computer vision › 3D vision › 3d reconstruction › non-lambertian surface reconstruction
reflective surface reconstruction
0.812024
GNeRP: Gaussian-guided Neural Reconstruction of Reflective Objects with Noisy Polarization Priors · ICLR 2024
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.812024
GNeRP: Gaussian-guided Neural Reconstruction of Reflective Objects with Noisy Polarization Priors · ICLR 2024
Computer vision › Video understanding and tracking
video anomaly detection
0.812024
HAWK: Learning to Understand Open-World Video Anomalies · NeurIPS 2024
Computer vision › Video understanding and tracking
video anomaly understanding
0.812024
HAWK: Learning to Understand Open-World Video Anomalies · NeurIPS 2024
Computer vision › Vision and language
vision-language model
0.812024
HAWK: Learning to Understand Open-World Video Anomalies · NeurIPS 2024
Image and video processing
image restoration
0.812024
Learning to Remove Wrinkled Transparent Film with Polarized Prior · CVPR 2024
Computational photography and imaging
polarization imaging
0.812024
Learning to Remove Wrinkled Transparent Film with Polarized Prior · CVPR 2024
Image and video processing › image restoration › reflection removal
specular highlight removal
0.812024
Learning to Remove Wrinkled Transparent Film with Polarized Prior · CVPR 2024
Machine learning › Generative modeling › video generation
video frame synthesis
0.612022
Video Frame Interpolation with Transformer · CVPR 2022
Image and video processing
video frame interpolation
0.612022
Video Frame Interpolation with Transformer · CVPR 2022
Computer vision › Segmentation and scene understanding
instance segmentation
0.512021
Video Instance Segmentation with a Propose-Reduce Paradigm · ICCV 2021
Computer vision › Video understanding and tracking
video instance segmentation
0.512021
Video Instance Segmentation with a Propose-Reduce Paradigm · ICCV 2021
Computer vision › Video understanding and tracking
video propagation
0.412020
Memory Selection Network for Video Propagation · ECCV (15) 2020
Visual content generation and editing › face editing
identity swapping
0.412020
Particularity Beyond Commonality: Unpaired Identity Transfer with Multiple References · ECCV (4) 2020
Visual content generation and editing
image editing
0.412020
Particularity Beyond Commonality: Unpaired Identity Transfer with Multiple References · ECCV (4) 2020
Visual content generation and editing
image-to-image translation
0.412019
Attribute-Driven Spontaneous Motion in Unpaired Image Translation · ICCV 2019
Visual content generation and editing › image-to-image translation
unpaired image translation
0.412019
Attribute-Driven Spontaneous Motion in Unpaired Image Translation · ICCV 2019

Methods — techniques the papers use, named apart from their topics

transformer · 1.1cross-scale window-based attention · 1.1signed distance function · 0.8reconstruction network · 0.8polarized prior · 0.8polarization priors · 0.8motion modality integration · 0.8large visual language models · 0.8gaussian representation · 0.8consistency loss · 0.8angle estimation network · 0.8sequence propagation head · 0.5propose-reduce paradigm · 0.5generative adversarial network · 0.4refinement module · 0.4deformation learning · 0.4
YearPublicationVenuePosition
2024 Learning to Remove Wrinkled Transparent Film with Polarized Prior
abstract
In this paper, we study a new problem, Film Removal (FR), which attempts to remove the interference of wrinkled transparent films and reconstruct the original information under films for industrial recognition systems. We first physically model the imaging of industrial materials covered by the film. Considering the specular highlight from the film can be effectively recorded by the polarized camera, we build a practical dataset with polarization information containing paired data with and without transparent film. We aim to remove interference from the film (specular highlights and other degradations) with an end-to-end framework. To locate the specular highlight, we use an angle estimation network to optimize the polarization angle with the minimized specular highlight. The image with minimized specular highlight is set as a prior for supporting the reconstruction network. Based on the prior and the polarized images, the reconstruction network can decouple all degradations from the film. Extensive experiments show that our framework achieves SOTA performance in both image reconstruction and industrial downstream tasks. Our code will be released at https://github.com/jqtangust/FilmRemoval.
Jiaqi Tang 0005, Ruizheng Wu, Xiaogang Xu 0002, Sixing Hu, Ying-Cong Chen
CVPR2
2024 An Incremental Unified Framework for Small Defect Inspection
Jiaqi Tang 0005, Hao Lu 0009, Xiaogang Xu 0002, Ruizheng Wu, Sixing Hu, Tong Zhang 0001, Tsz Wa Cheng, Ming Ge, Ying-Cong Chen, Fugee Tsung
ECCV (31)4
2024 GNeRP: Gaussian-guided Neural Reconstruction of Reflective Objects with Noisy Polarization Priors
abstract
Learning surfaces from neural radiance field (NeRF) became a rising topic in Multi-View Stereo (MVS). Recent Signed Distance Function (SDF)-based methods demonstrated their ability to reconstruct exact 3D shapes of Lambertian scenes. However, their results on reflective scenes are unsatisfactory due to the entanglement of specular radiance and complicated geometry. To address the challenges, we propose a Gaussian-based representation of normals in SDF fields. Supervised by polarization priors, this representation guides the learning of geometry behind the specular reflection and capture more details than existing methods. Moreover, we propose a reweighting strategy in optimization process to alleviate the noise issue of polarization priors. To validate the effectiveness of our design, we capture polarimetric information and ground truth meshes in additional reflective scenes with various geometry. We also evaluated our framework on PANDORA dataset. Both qualitative and quantitative comparisons prove our method outperforms existing neural 3D reconstruction methods in reflective scenes by a large margin.
Ruizheng Wu, Ying-Cong Chen
ICLR2
2024 HAWK: Learning to Understand Open-World Video Anomalies
abstract
Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios. In this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk.
Jiaqi Tang 0005, Hao Lu 0009, Ruizheng Wu, Xiaogang Xu 0002, Bin Guo 0001, Jiangbo Lu, Qifeng Chen 0001, Ying-Cong Chen
NeurIPS3
2022 Video Frame Interpolation with Transformer
abstract
Video frame interpolation (VFI), which aims to synthesize intermediate frames of a video, has made remarkable progress with development of deep convolutional networks over past years. Existing methods built upon convolutional networks generally face challenges of handling large motion due to the locality of convolution operations. To overcome this limitation, we introduce a novel framework, which takes advantage of Transformer to model long-range pixel correlation among video frames. Further, our network is equipped with a novel cross-scale window-based attention mechanism, where cross-scale windows interact with each other. This design effectively enlarges the receptive field and aggregates multi-scale information. Extensive quantitative and qualitative experiments demonstrate that our method achieves new state-of-the-art results on various benchmarks.
Liying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu, Jiaya Jia
CVPR2
2021 Video Instance Segmentation with a Propose-Reduce Paradigm
abstract
Video instance segmentation (VIS) aims to segment and associate all instances of predefined classes for each frame in videos. Prior methods usually obtain segmentation for a frame or clip first, and merge the incomplete results by tracking or matching. These methods may cause error accumulation in the merging step. Contrarily, we propose a new paradigm – Propose-Reduce, to generate complete sequences for input videos by a single step. We further build a sequence propagation head on the existing image-level instance segmentation network for long-term propagation. To ensure robustness and high recall of our proposed framework, multiple sequences are proposed where redundant sequences of the same instance are reduced. We achieve state-of-the-art performance on two representative benchmark datasets – we obtain 47.6% in terms of AP on YouTube-VIS validation set and 70.4 % for J&F on DAVIS-UVOS validation set.
Huaijia Lin, Ruizheng Wu, Shu Liu 0005, Jiangbo Lu, Jiaya Jia
ICCV2
2020 Memory Selection Network for Video Propagation
Ruizheng Wu, Huaijia Lin, Xiaojuan Qi 0001, Jiaya Jia
ECCV (15)1
2020 Particularity Beyond Commonality: Unpaired Identity Transfer with Multiple References
Ruizheng Wu, Xin Tao 0001, Ying-Cong Chen, Xiaoyong Shen, Jiaya Jia
ECCV (4)1
2019 Attribute-Driven Spontaneous Motion in Unpaired Image Translation
abstract
Current image translation methods, albeit effective to produce high-quality results in various applications, still do not consider much geometric transform. We in this paper propose the spontaneous motion estimation module, along with a refinement part, to learn attribute-driven deformation between source and target domains. Extensive experiments and visualization demonstrate effectiveness of these modules. We achieve promising results in unpaired-image translation tasks, and enable interesting applications based on spontaneous motion.
Ruizheng Wu, Xin Tao 0001, Xiaoyong Shen, Jiaya Jia
ICCV1