Zhengyan Tong

dblp:281/6780 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
5 papers
Image and video processing · 54% Rendering · 28% Visual content generation and editing · 18%
Artificial intelligence
2 papers
3D vision · 90% Representation and self-supervised learning · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
stroke-based rendering
1.122022
Im2Oil: Stroke-Based Oil Painting Rendering with Linearly Controllable Fineness Via Adaptive Sampling · ACM Multimedia 2022
Sketch Generation with Drawing Process Guided by Vector Flow and Grayscale · AAAI 2021
Image and video processing › super-resolution › image super-resolution
arbitrary-scale super-resolution
0.712023
Deep Arbitrary-Scale Image Super-Resolution via Scale-Equivariance Pursuit · CVPR 2023
Image and video processing › super-resolution › image super-resolution
depth super-resolution
0.712023
Learning Continuous Depth Representation via Geometric Spatial Aggregator · AAAI 2023
Image and video processing
image restoration
0.712023
Learning Continuous Depth Representation via Geometric Spatial Aggregator · AAAI 2023
Image and video processing › super-resolution
image super-resolution
0.712023
Deep Arbitrary-Scale Image Super-Resolution via Scale-Equivariance Pursuit · CVPR 2023
Computer vision › 3D vision
3d human pose estimation
0.612022
OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose Estimation · ACM Multimedia 2022
Computer vision › 3D vision › pose estimation › robust pose estimation
occlusion-aware pose estimation
0.612022
OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose Estimation · ACM Multimedia 2022
Environmental and earth informatics
meteorology
0.612022
RainNet: A Large-Scale Imagery Dataset and Benchmark for Spatial Precipitation Downscaling · NeurIPS 2022
Visual content generation and editing › stylization
image stylization
0.612022
Im2Oil: Stroke-Based Oil Painting Rendering with Linearly Controllable Fineness Via Adaptive Sampling · ACM Multimedia 2022
Rendering
non-photorealistic rendering
0.612022
Im2Oil: Stroke-Based Oil Painting Rendering with Linearly Controllable Fineness Via Adaptive Sampling · ACM Multimedia 2022
Image and video processing
super-resolution
0.612022
RainNet: A Large-Scale Imagery Dataset and Benchmark for Spatial Precipitation Downscaling · NeurIPS 2022
Visual content generation and editing
sketch generation
0.512021
Sketch Generation with Drawing Process Guided by Vector Flow and Grayscale · AAAI 2021
Computer vision › 3D vision
depth estimation
0.212023
Learning Continuous Depth Representation via Geometric Spatial Aggregator · AAAI 2023
Computer vision › 3D vision › depth estimation
depth super-resolution
0.212023
Learning Continuous Depth Representation via Geometric Spatial Aggregator · AAAI 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.212022
OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose Estimation · ACM Multimedia 2022

Methods — techniques the papers use, named apart from their topics

transformer · 2.0implicit neural representation · 1.3geometric spatial aggregator · 1.3distance field · 1.3implicit physical estimation · 1.1deep learning · 1.1pre-training · 0.7neural kriging · 0.7voronoi algorithm · 0.6pose lifting · 0.6pose filling · 0.6iterative training · 0.6contrastive learning · 0.6adaptive sampling · 0.6
YearPublicationVenuePosition
2023 Learning Continuous Depth Representation via Geometric Spatial Aggregator
abstract
Depth map super-resolution (DSR) has been a fundamental task for 3D computer vision. While arbitrary scale DSR is a more realistic setting in this scenario, previous approaches predominantly suffer from the issue of inefficient real-numbered scale upsampling. To explicitly address this issue, we propose a novel continuous depth representation for DSR. The heart of this representation is our proposed Geometric Spatial Aggregator (GSA), which exploits a distance field modulated by arbitrarily upsampled target gridding, through which the geometric information is explicitly introduced into feature aggregation and target generation. Furthermore, bricking with GSA, we present a transformer-style backbone named GeoDSR, which possesses a principled way to construct the functional mapping between local coordinates and the high-resolution output results, empowering our model with the advantage of arbitrary shape transformation ready to help diverse zooming demand. Extensive experimental results on standard depth map benchmarks, e.g., NYU v2, have demonstrated that the proposed framework achieves significant restoration gain in arbitrary scale depth map super-resolution compared with the prior art. Our codes are available at https://github.com/nana01219/GeoDSR.
Xiaohang Wang 0004, Xuanhong Chen, Bingbing Ni, Zhengyan Tong
AAAI4
2023 Deep Arbitrary-Scale Image Super-Resolution via Scale-Equivariance Pursuit
abstract
The ability of scale-equivariance processing blocks plays a central role in arbitrary-scale image super-resolution tasks. Inspired by this crucial observation, this work proposes two novel scale-equivariant modules within a transformer-style framework to enhance arbitrary-scale image super-resolution (ASISR) performance, especially in high upsampling rate image extrapolation. In the feature extraction phase, we design a plug-in module called Adaptive Feature Extractor, which injects explicit scale information in frequency-expanded encoding, thus achieving scale-adaption in representation learning. In the upsampling phase, a learnable Neural Kriging upsampling operator is introduced, which simultaneously encodes both relative distance (i.e., scale-aware) information as well as feature similarity (i.e., with priori learned from training data) in a bilateral manner, providing scale-encoded spatial feature fusion. The above operators are easily plugged into multiple stages of a SR network, and a recent emerging pretraining strategy is also adopted to impulse the model's performance further. Extensive experimental results have demonstrated the outstanding scale-equivariance capability offered by the proposed operators and our learning framework, with much better results than previous SOTA methods at arbitrary scales for SR. Our code is available at https://github.com/neura1chen/EQSR
Xiaohang Wang 0004, Xuanhong Chen, Bingbing Ni, Zhengyan Tong, Yutian Liu 0004
CVPR5
2022 Im2Oil: Stroke-Based Oil Painting Rendering with Linearly Controllable Fineness Via Adaptive Sampling
abstract
This paper proposes a novel stroke-based rendering (SBR) method that translates images into vivid oil paintings. Previous SBR techniques usually formulate the oil painting problem as pixel-wise approximation. Different from this technique route, we treat oil painting creation as an adaptive sampling problem. Firstly, we compute a probability density map based on the texture complexity of the input image. Then we use the Voronoi algorithm to sample a set of pixels as the stroke anchors. Next, we search and generate an individual oil stroke at each anchor. Finally, we place all the strokes on the canvas to obtain the oil painting. By adjusting the hyper-parameter maximum sampling probability, we can control the oil painting fineness in a linear manner. Comparison with existing state-of-the-art oil painting techniques shows that our results have higher fidelity and more realistic textures. A user opinion test demonstrates that people behave more preference toward our oil paintings than the results of other methods. More interesting results and the code are in https://github.com/TZYSJTU/Im2Oil.
Zhengyan Tong, Xiaohang Wang 0004, Shengchao Yuan, Xuanhong Chen, Xiangzhong Fang
ACM Multimedia1
2022 OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose Estimation
abstract
Occlusion is a significant problem in 3D human pose estimation from the 2D counterpart. On one hand, without explicit annotation, the 3D skeleton is hard to be accurately estimated from the occluded 2D pose. On the other hand, one occluded 2D pose might correspond to multiple 3D skeletons with low confidence parts. To address these issues, we decouple the 3D representation feature into view-invariant part termed occlusion-aware feature and view-dependent part termed rotation feature to facilitate subsequent optimization of the former. Then we propose an occlusion-aware contrastive representation based scheme (OCR-Pose) consisting of Topology Invariant Contrastive Learning module (TiCLR) and View Equivariant Contrastive Learning module (VeCLR). Specifically, TiCLR drives invariance to topology transformation, i.e., bridging the gap between an occluded 2D pose and the unoccluded one. While VeCLR encourages equivariance to view transformation, i.e., capturing the geometric similarity of the 3D skeleton in two views. Both modules optimize occlusion-aware constrastive representation with pose filling and lifting networks via an iterative training strategy in an end-to-end manner. OCR-Pose not only achieves superior performance against state-of-the-art unsupervised methods on unoccluded benchmarks, but also obtains significant improvements when occlusion is involved. Our project is available at https://sites.google.com/view/ocr-pose.
Zhenbo Yu, Zhengyan Tong, Jinxian Liu, Wenjun Zhang 0001
ACM Multimedia3
2022 RainNet: A Large-Scale Imagery Dataset and Benchmark for Spatial Precipitation Downscaling
abstract
AI-for-science approaches have been applied to solve scientific problems (e.g., nuclear fusion, ecology, genomics, meteorology) and have achieved highly promising results. Spatial precipitation downscaling is one of the most important meteorological problem and urgently requires the participation of AI. However, the lack of a well-organized and annotated large-scale dataset hinders the training and verification of more effective and advancing deep-learning models for precipitation downscaling. To alleviate these obstacles, we present the first large-scale spatial precipitation downscaling dataset named RainNet, which contains more than 62,400 pairs of high-quality low/high-resolution precipitation maps for over 17 years, ready to help the evolution of deep learning models in precipitation downscaling. Specifically, the precipitation maps carefully collected in RainNet cover various meteorological phenomena (e.g., hurricane, squall), which is of great help to improve the model generalization ability. In addition, the map pairs in RainNet are organized in the form of image sequences (720 maps per month or 1 map/hour), showing complex physical properties, e.g., temporal misalignment, temporal sparse, and fluid properties. Furthermore, two deep-learning-oriented metrics are specifically introduced to evaluate or verify the comprehensive performance of the trained model (e.g., prediction maps reconstruction accuracy). To illustrate the applications of RainNet, 14 state-of-the-art models, including deep models and traditional approaches, are evaluated. To fully explore potential downscaling solutions, we propose an implicit physical estimation benchmark framework to learn the above characteristics. Extensive experiments demonstrate the value of RainNet in training and evaluating downscaling models. Our dataset is available at https://neuralchen.github.io/RainNet/.
Xuanhong Chen, Kairui Feng, Naiyuan Liu, Bingbing Ni, Zhengyan Tong, Ziang Liu 0001
NeurIPS6
2021 Sketch Generation with Drawing Process Guided by Vector Flow and Grayscale
abstract
We propose a novel image-to-pencil translation method that could not only generate high-quality pencil sketches but also offer the drawing process. Existing pencil sketch algorithms are based on texture rendering rather than the direct imitation of strokes, making them unable to show the drawing process but only a final result. To address this challenge, we first establish a pencil stroke imitation mechanism. Next, we develop a framework with three branches to guide stroke drawing: the first branch guides the direction of the strokes, the second branch determines the shade of the strokes, and the third branch enhances the details further. Under this framework's guidance, we can produce a pencil sketch by drawing one stroke every time. Our method is fully interpretable. Comparison with existing pencil drawing algorithms shows that our method is superior to others in terms of texture quality, style, and user evaluation. Our code and supplementary material are now available at: https://github.com/TZYSJTU/Sketch-Generation-withDrawing-Process-Guided-by-Vector-Flow-and-Grayscale
Zhengyan Tong, Xuanhong Chen, Bingbing Ni, Xiaohang Wang 0004
AAAI1