Tianyang Shi

dblp:224/4099 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-4587-7792ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Domain generalization via domain uncertainty shrinkage
Jun-Zheng Chu, Bin Pan, Tianyang Shi, Zhenwei Shi 0001
Pattern Recognit.3
2026 Be Bayesian by attachments to catch more uncertainty
Bin Pan, Tianyang Shi, Tao Li 0022, Zhenwei Shi 0001
Pattern Recognit.3
2024 Joint Variational Inference Network for domain generalization
Jun-Zheng Chu, Bin Pan, Tianyang Shi, Zhenwei Shi 0001, Tao Li 0022
Pattern Recognit.4
2023 ASM: Adaptive Skinning Model for High-Quality 3D Face Modeling
abstract
The research fields of parametric face model and 3D face reconstruction have been extensively studied. However, a critical question remains unanswered: how to tailor the face model for specific reconstruction settings. We argue that reconstruction with multi-view uncalibrated images demands a new model with stronger capacity. Our study shifts attention from data-dependent 3D Morphable Models (3DMM) to an understudied human-designed skinning model. We propose Adaptive Skinning Model (ASM), which redefines the skinning model with more compact and fully tunable parameters. With extensive experiments, we demonstrate that ASM achieves significantly improved capacity than 3DMM, with the additional advantage of model size and easy implementation for new topology. We achieve state-of-the-art performance with ASM for multi-view reconstruction on the Florence MICC Coop benchmark. Our quantitative analysis demonstrates the importance of a high-capacity model for fully exploiting abundant information from multi-view input in reconstruction. Furthermore, our model with physical-semantic parameters can be directly utilized for real-world applications, such as in-game avatar creation. As a result, our work opens up new research direction for parametric face model and facilitates future research on multi-view reconstruction.
Hong Shang, Tianyang Shi, Xinghan Chen, Jingkai Zhou, Zhongqian Sun, Wei Yang 0032
ICCV3
2023 An Imbalanced Discriminant Alignment Approach for Domain Adaptive SAR Ship Detection
abstract
Synthetic aperture radar (SAR) imaging has round-the-clock data acquisition capability regardless of light and climate constraints, so it has been widely used for ship detection. However, SAR images usually suffer lower imaging quality, which may result in indistinct contours and non-negligible noise. Therefore, the manual labeling for SAR images is expensive, leading to a lack of training data in the task of ship detection. In this paper, we propose a route by utilizing domain adaptive methods to transfer information from labeled visible images (source domain) to unlabeled SAR images (target domain) for ship detection. To address the distribution mismatch between domains, we develop a novel imbalanced discriminant alignment (IDA) approach to improve the discriminant ability of the network and prevent negative migration. The core of the IDA approach is applying a new loss function called imbalanced prediction consistency (IPC) loss to describe the domain classifier consistency, and we further provide theoretical analysis for the effectiveness of the IPC loss. IDA ensures consistency at the image level and instance level, and focuses on the consistency of the source domain to enhance the feature extraction capability of the adversarial network. The theoretical discussion has proven that a necessary and sufficient condition for convergence of the IPC loss is that the two discriminant probabilities converge to 0 at the discriminant distance we define. Experimental results have indicated the advantage of IDA when compared with other domain adaptation SAR ship detection methods.
Bin Pan, Zhehao Xu, Tianyang Shi, Tao Li 0022, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Neural Rendering for Game Character Auto-Creation
abstract
Many role-playing games feature character creation systems where players are allowed to edit the facial appearance of their in-game characters. This paper proposes a novel method to automatically create game characters based on a single face photo. We frame this "artistic creation" process under a self-supervised learning paradigm by leveraging the differentiable neural rendering. Considering the rendering process of a typical game engine is not differentiable, an "imitator" network is introduced to imitate the behavior of the engine so that the in-game characters can be smoothly optimized by gradient descent in an end-to-end fashion. Different from previous monocular 3D face reconstruction which focuses on generating 3D mesh-grid and ignores user interaction, our method produces fine-grained facial parameters with a clear physical significance where users can optionally fine-tune their auto-created characters by manually adjusting those parameters. Experiments on multiple large-scale face datasets show that our method can generate highly robust and vivid game characters. Our method has been applied to two games and has now provided over 10 million times of online services.
Tianyang Shi, Zhengxia Zou, Zhenwei Shi 0001, Yi Yuan 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Castle in the Sky: Dynamic Sky Replacement and Harmonization in Videos
abstract
We propose a vision-based framework for dynamic sky replacement and harmonization in videos. Different from previous sky editing methods that either focus on static photos or require real-time pose signal from the camera's inertial measurement units, our method is purely vision-based, without any requirements on the capturing devices, and can be well applied to either online or offline processing scenarios. Our method runs in real-time and is free of manual interactions. We decompose the video sky replacement into several proxy tasks, including motion estimation, sky matting, and image blending. We derive the motion equation of an object at infinity on the image plane under the camera's motion, and propose "flow propagation", a novel method for robust motion estimation. We also propose a coarse-to-fine sky matting network to predict accurate sky matte and design image blending to improve the harmonization. Experiments are conducted on videos diversely captured in the wild and show high fidelity and good generalization capability of our framework in both visual quality and lighting/motion dynamics. We also introduce a new method for content-aware image augmentation and proved that this method is beneficial to visual perception in autonomous driving scenarios. Our code and animated results are available at https://github.com/jiupinjia/SkyAR.
Zhengxia Zou, Rui Zhao 0019, Tianyang Shi, Zhenwei Shi 0001
IEEE Trans. Image Process.3
2021 Stylized Neural Painting
abstract
This paper proposes an image-to-painting translation method that generates vivid and realistic painting artworks with controllable styles. Different from previous image-to-image translation methods that formulate the translation as pixel-wise prediction, we deal with such an artistic creation process in a vectorized environment and produce a sequence of physically meaningful stroke parameters that can be further used for rendering. Since a typical vector render is not differentiable, we design a novel neural renderer which imitates the behavior of the vector renderer and then frame the stroke prediction as a parameter searching process that maximizes the similarity between the input and the rendering output. We explored the zero-gradient problem on parameter searching and propose to solve this problem from an optimal transportation perspective. We also show that previous neural renderers have a parameter coupling problem and we re-design the rendering network with a rasterization network and a shading network that better handles the disentanglement of shape and color. Experiments show that the paintings generated by our method have a high degree of fidelity in both global appearance and local textures. Our method can be also jointly optimized with neural style transfer that further transfers visual style from other images. Our code and animated results are available at https://jiupinjia.github.io/neuralpainter/.
Zhengxia Zou, Tianyang Shi, Yi Yuan 0002, Zhenwei Shi 0001
CVPR2
2021 Multi-view 3D Reconstruction with Transformers
abstract
Deep CNN-based methods have so far achieved the state of the art results in multi-view 3D object reconstruction. Despite the considerable progress, the two core modules of these methods - view feature extraction and multi-view fusion, are usually investigated separately, and the relations among multiple input views are rarely explored. Inspired by the recent great success in Transformer models, we reformulate the multi-view 3D reconstruction as a sequence-to-sequence prediction problem and propose a framework named 3D Volume Transformer. Unlike previous CNN-based methods using a separate design, we unify the feature extraction and view fusion in a single Transformer network. A natural advantage of our design lies in the exploration of view-to-view relationships using self-attention among multiple unordered inputs. On ShapeNet - a large-scale 3D reconstruction benchmark, our method achieves a new state-of-the-art accuracy in multi-view reconstruction with fewer parameters (70% less) than CNN-based methods. Experimental results also suggest the strong scaling capability of our method. Our code will be made publicly available.
Dan Wang 0011, Xinrui Cui, Xun Chen 0001, Zhengxia Zou, Tianyang Shi, Tim Salcudean, Z. Jane Wang 0001, Rabab K. Ward
ICCV5
2021 Automatic Translation of Music-to-Dance for In-Game Characters
abstract
Music-to-dance translation is an emerging and powerful feature in recent role-playing games. Previous works of this topic consider music-to-dance as a supervised motion generation problem based on time-series data. However, these methods require a large amount of training data pairs and may suffer from the degradation of movements. This paper provides a new solution to this task where we re-formulate the translation as a piece-wise dance phrase retrieval problem based on the choreography theory. With such a design, players are allowed to optionally edit the dance movements on top of our generation while other regression-based methods ignore such user interactivity. Considering that the dance motion capture is expensive that requires the assistance of professional dancers, we train our method under a semi-supervised learning fashion with a large unlabeled music dataset (20x than our labeled one) and also introduce self-supervised pre-training to improve the training stability and generalization performance. Experimental results suggest that our method not only generalizes well over various styles of music but also succeeds in choreography for game players. Our project including the large-scale dataset and supplemental materials is available at https://github.com/FuxiCV/music-to-dance.
Yinglin Duan, Tianyang Shi, Zhipeng Hu, Zhengxia Zou, Changjie Fan, Yi Yuan 0002
IJCAI2
2021 Adversarial Training for Solving Inverse Problems in Image Processing
abstract
Inverse problems are a group of important mathematical problems that aim at estimating source data x and operation parameters z from inadequate observations y . In the image processing field, most recent deep learning-based methods simply deal with such problems under a pixel-wise regression framework (from y to x ) while ignoring the physics behind. In this paper, we re-examine these problems under a different viewpoint and propose a novel framework for solving certain types of inverse problems in image processing. Instead of predicting x directly from y , we train a deep neural network to estimate the degradation parameters z under an adversarial training paradigm. We show that if the degradation behind satisfies some certain assumptions, the solution to the problem can be improved by introducing additional adversarial constraints to the parameter space and the training may not even require pair-wise supervision. In our experiment, we apply our method to a variety of real-world problems, including image denoising, image deraining, image shadow removal, non-uniform illumination correction, and underdetermined blind source separation of images or speech signals. The results on multiple tasks demonstrate the effectiveness of our method.
Zhengxia Zou, Tianyang Shi, Zhenwei Shi 0001, Jieping Ye
IEEE Trans. Image Process.2
2020 Fast and Robust Face-to-Parameter Translation for Game Character Auto-Creation
abstract
With the rapid development of Role-Playing Games (RPGs), players are now allowed to edit the facial appearance of their in-game characters with their preferences rather than using default templates. This paper proposes a game character auto-creation framework that generates in-game characters according to a player's input face photo. Different from the previous methods that are designed based on neural style transfer or monocular 3D face reconstruction, we re-formulate the character auto-creation process in a different point of view: by predicting a large set of physically meaningful facial parameters under a self-supervised learning paradigm. Instead of updating facial parameters iteratively at the input end of the renderer as suggested by previous methods, which are time-consuming, we introduce a facial parameter translator so that the creation can be done efficiently through a single forward propagation from the face embeddings to parameters, with a considerable 1000x computational speedup. Despite its high efficiency, the interactivity is preserved in our method where users are allowed to optionally fine-tune the facial parameters on our creation according to their needs. Our approach also shows better robustness than previous methods, especially for those photos with head-pose variance. Comparison results and ablation analysis on seven public face verification datasets suggest the effectiveness of our method.
Tianyang Shi, Zhengxia Zou, Yi Yuan 0002, Changjie Fan
AAAI1
2020 Deep Adversarial Decomposition: A Unified Framework for Separating Superimposed Images
abstract
Separating individual image layers from a single mixed image has long been an important but challenging task. We propose a unified framework named "deep adversarial decomposition" for single superimposed image separation. Our method deals with both linear and non-linear mixtures under an adversarial training paradigm. Considering the layer separating ambiguity that given a single mixed input, there could be an infinite number of possible solutions, we introduce a "Separation-Critic" - a discriminative network which is trained to identify whether the output layers are well-separated and thus further improves the layer separation. We also introduce a "crossroad L1" loss function, which computes the distance between the unordered outputs and their references in a crossover manner so that the training can be well-instructed with pixel-wise supervision. Experimental results suggest that our method significantly outperforms other popular image separation frameworks. Without specific tuning, our method achieves the state of the art results on multiple computer vision tasks, including the image deraining, photo reflection removal, and image shadow removal.
Zhengxia Zou, Sen Lei, Tianyang Shi, Zhenwei Shi 0001, Jieping Ye
CVPR3
2020 Neutral Face Game Character Auto-Creation via PokerFace-GAN
abstract
Game character customization is one of the core features of many recent Role-Playing Games (RPGs), where players can edit the appearance of their in-game characters with their preferences. This paper studies the problem of automatically creating in-game characters with a single photo. In recent literature on this topic, neural networks are introduced to make game engine differentiable and the self-supervised learning is used to predict facial customization parameters. However, in previous methods, the expression parameters and facial identity parameters are highly coupled with each other, making it difficult to model the intrinsic facial features of the character. Besides, the neural network based renderer used in previous methods is also difficult to be extended to multi-view rendering cases. In this paper, considering the above problems, we propose a novel method named "PokerFace-GAN" for neutral face game character auto-creation. We first build a differentiable character renderer which is more flexible than the previous methods in multi-view rendering cases. We then take advantage of the adversarial training to effectively disentangle the expression parameters from the identity parameters and thus generate player-preferred neutral face (expression-less) characters. Since all components of our method are differentiable, our method can be easily trained under a multi-task self-supervised learning paradigm. Experiment results show that our method can generate vivid neutral face game characters that are highly similar to the input photos. The effectiveness of our method is verified by comparison results and ablation studies.
Tianyang Shi, Zhengxia Zou, Xinhui Song, Changjian Gu, Changjie Fan, Yi Yuan 0002
ACM Multimedia1
2020 Unsupervised Learning Facial Parameter Regressor for Action Unit Intensity Estimation via Differentiable Renderer
abstract
Facial action unit (AU) intensity is an index to describe all visually discernible facial movements. Most existing methods learn intensity estimator with limited AU data, while they lack of generalization ability out of the dataset. In this paper, we present a framework to predict the facial parameters (including identity parameters and AU parameters) based on a bone-driven face model (BDFM) under different views. The proposed framework consists of a feature extractor, a generator, and a facial parameter regressor. The regressor can fit the physical meaning parameters of the BDFM from a single face image with the help of the generator, which maps the facial parameters to the game-face images as a differentiable renderer. Besides, identity loss, loopback loss, and adversarial loss can improve the regressive results. Quantitative evaluations are performed on two public databases BP4D and DISFA, which demonstrates that the proposed method can achieve comparable or better performance than the state-of-the-art methods. What's more, the qualitative results also demonstrate the validity of our method in the wild.
Xinhui Song, Tianyang Shi, Zunlei Feng, Mingli Song, Jackie Lin, Chuanjie Lin, Changjie Fan, Yi Yuan 0002
ACM Multimedia2
2019 Face-to-Parameter Translation for Game Character Auto-Creation
abstract
Character customization system is an important component in Role-Playing Games (RPGs), where players are allowed to edit the facial appearance of their in-game characters with their own preferences rather than using default templates. This paper proposes a method for automatically creating in-game characters of players according to an input face photo. We formulate the above "artistic creation" process under a facial similarity measurement and parameter searching paradigm by solving an optimization problem over a large set of physically meaningful facial parameters. To effectively minimize the distance between the created face and the real one, two loss functions, i.e. a "discriminative loss" and a "facial content loss", are specifically designed. As the rendering process of a game engine is not differentiable, a generative network is further introduced as an "imitator" to imitate the physical behavior of the game engine so that the proposed method can be implemented under a neural style transfer framework and the parameters can be optimized by gradient descent. Experimental results demonstrate that our method achieves a high degree of generation similarity between the input face photo and the created in-game character in terms of both global appearance and local details. Our method has been deployed in a new game last year and has now been used by players over 1 million times.
Tianyang Shi, Yi Yuan 0002, Changjie Fan, Zhengxia Zou, Zhenwei Shi 0001, Yong Liu 0007
ICCV1
2019 Generative Adversarial Training for Weakly Supervised Cloud Matting
abstract
The detection and removal of cloud in remote sensing images are essential for earth observation applications. Most previous methods consider cloud detection as a pixel-wise semantic segmentation process (cloud v.s. background), which inevitably leads to a category-ambiguity problem when dealing with semi-transparent clouds. We re-examine the cloud detection under a totally different point of view, i.e. to formulate it as a mixed energy separation process between foreground and background images, which can be equivalently implemented under an image matting paradigm with a clear physical significance. We further propose a generative adversarial framework where the training of our model neither requires any pixel-wise ground truth reference nor any additional user interactions. Our model consists of three networks, a cloud generator G, a cloud discriminator D, and a cloud matting network F, where G and D aim to generate realistic and physically meaningful cloud images by adversarial training, and F learns to predict the cloud reflectance and attenuation. Experimental results on a global set of satellite images demonstrate that our method, without ever using any pixel-wise ground truth during training, achieves comparable and even higher accuracy over other fully supervised methods, including some recent popular cloud detectors and some well-known semantic segmentation frameworks.
Zhengxia Zou, Wenyuan Li 0002, Tianyang Shi, Zhenwei Shi 0001, Jieping Ye
ICCV3
2019 CoinNet: Copy Initialization Network for Multispectral Imagery Semantic Segmentation
abstract
Remote sensing imagery semantic segmentation refers to assigning a label to every pixel. Recently, deep convolutional neural networks (CNNs)-based methods have presented an impressive performance in this task. Due to the lack of sufficient labeled remote sensing images, researchers usually utilized transfer learning (TL) strategies to fine tune networks which were pretrained in huge RGB-scene data sets. Unfortunately, this manner may not work if the target images are multispectral/hyperspectral. The basic assumption of TL is that the low-level features extracted by the former layers are similar in most data sets, hence users only require to train the parameters in the last layers that are specific to different tasks. However, if one should use a pretrained deep model in RGB data for multispectral /hyperspectral imagery semantic segmentation, the structure of the input layer has to be adjusted. In this case, the first convolutional layer has to be trained using the multispectral /hyperspectral data sets which are much smaller. Apparently, the feature representation ability of the first convolutional layer will decrease and it may further harm the following layers. In this letter, we propose a new deep learning model, COpy INitialization Network (CoinNet), for multispectral imagery semantic segmentation. The major advantage of CoinNet is that it can make full use of the initial parameters in the pretrained network's first convolutional layer. Comparison experiments on a challenging multispectral data set have demonstrated the effectiveness of the proposed improvement. The demo and a trained network will be published in our homepage.
Bin Pan, Zhenwei Shi 0001, Tianyang Shi, Xinzhong Zhu
IEEE Geosci. Remote. Sens. Lett.4
2018 Collaborative Sparse Hyperspectral Unmixing Using l0 Norm
abstract
Sparse unmixing has been applied on hyperspectral imagery popularly in recent years. It assumes that every observed signature is a linear combination of just a few spectra (end-members) from a known spectral library. However, solving the sparse unmixing problem directly (using l0norm to control the sparsity of solution at a low level) is NP-hard. Most related works focus on convex relaxation methods, but the sparsity and accuracy of results cannot be well guaranteed. Under these circumstances, this paper proposes a novel algorithm termed collaborative sparse hyperspectral unmixing using l0 norm (CSUnL0), which aims at solving l0problem directly. First, it introduces a row-hard-threshold function. The row-hardthreshold function makes it possible to combine l0 norm, instead of its approximate norms, with alternating direction method of multipliers. Compared with the convex relaxation methods, the l0norm constraint guarantees sparser and more accurate results. Moreover, the antinoise ability of CSUnL0 also gets improved. Second, CSUnL0 uses l2norm of each end-members' abundance across the whole map as a collaborative constraint, which can take advantage of the hyperspectral data's subspace property. The experimental results indicate that l0norm contributes to acquiring a more sparser solution and helps CSUnL0 to enhance calculation accuracy.
Tianyang Shi
IEEE Trans. Geosci. Remote. Sens.2