Jilin Tang

dblp:254/8073 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0001-9478-7489ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 AS-NeRF: Learning Auxiliary Sampling for Generalizable Novel View Synthesis from Sparse Views
abstract
We tackle the problem of novel view synthesis (NVS) which aims to generate realistic images at novel views. Unlike existing works that require either costly per-scene optimization or relatively dense views, we propose AS-NeRF, a neural rendering based approach that can achieve generalizable NVS from only sparse views. Considering the inherent spatial continuity of images, we design a novel sparse-attention based auxiliary sampling module (ASM). Given a 3D point on the ray, the ASM adaptively attends to a sparse set of view-specific 2D auxiliary locations around the 3D point’s original projection pixels, and dynamically computes the attention weights in a cross-attention manner. This enables our model to effectively exploit the local correlation among neighboring pixels, obtaining the enhanced features with more powerful representation. Extensive experiments show that our method outperforms the state-of-the-art on both real and synthetic data.
Jilin Tang, Lincheng Li, Xingqun Qi, Changjie Fan, Xin Yu 0002
ICME1
2024 MarkerNet: A divide-and-conquer solution to motion capture solving from raw markers
abstract
Abstract Marker‐based optical motion capture (MoCap) aims to localize 3D human motions from a sequence of input raw markers. It is widely used to produce physical movements for virtual characters in various games such as the role‐playing game, the fighting game, and the action‐adventure game. However, the conventional MoCap cleaning and solving process is extremely labor‐intensive, time‐consuming, and usually the most costly part of game animation production. Thus, there is a high demand for automated algorithms to replace costly manual operations and achieve accurate MoCap cleaning and solving in the game industry. In this article, we design a divide‐and‐conquer‐based MoCap solving network, dubbed MarkerNet, to estimate human skeleton motions from sequential raw markers effectively. In a nutshell, our key idea is to decompose the task of direct solving of global motion from all markers into first modeling sub‐motions of local parts from the corresponding marker subsets and then aggregating sub‐motions into a global one. In this manner, our model can effectively capture local motion patterns w.r.t. different marker subsets, thus producing more accurate results compared to the existing methods. Extensive experiments on both real and synthetic data verify the effectiveness of the proposed method.
Zhipeng Hu, Jilin Tang, Lincheng Li, Xin Yu 0002, Jiajun Bu
Comput. Animat. Virtual Worlds2
2021 Structure-aware Person Image Generation with Pose Decomposition and Semantic Correlation
abstract
In this paper we tackle the problem of pose guided person image generation, which aims to transfer a person image from the source pose to a novel target pose while maintaining the source appearance. Given the inefficiency of standard CNNs in handling large spatial transformation, we propose a structure-aware flow based method for high-quality person image generation. Specifically, instead of learning the complex overall pose changes of human body, we decompose the human body into different semantic parts (e.g., head, torso, and legs) and apply different networks to predict the flow fields for these parts separately. Moreover, we carefully design the network modules to effectively capture the local and global semantic correlations of features within and among the human parts respectively. Extensive experimental results show that our method can generate high-quality results under large pose discrepancy and outperforms state-of-the-art methods in both qualitative and quantitative comparisons.
Jilin Tang, Yi Yuan 0002, Tianjia Shao, Yong Liu 0007, Mengmeng Wang 0005, Kun Zhou 0001
AAAI1
2021 Vanet: a View Attention Guided Network for 3d Reconstruction from Single and Multi-View Images
abstract
Reconstructing 3D meshes of objects from 2D images is an important but challenging task. Previous 3D reconstruction methods either only focus on generating the mesh from a single image, or multi-view images. Instead of investigating these problems separately, we present a novel view attention guided network called VANet which addresses both single and multi-view 3D reconstruction under a unified frame-work. To explore non-visible parts of an object during the re-construction, a channel-wise view attention mechanism and a dual pathway network architecture are introduced. The proposed network highlights the informative object parts and compensates those non-informative ones with auxiliary views of input.
Yi Yuan 0002, Jilin Tang, Zhengxia Zou
ICME2
2021 Attention guided feature pyramid network for crowd counting
Huanpeng Chu, Jilin Tang, Haoji Hu
J. Vis. Commun. Image Represent.2
2019 Pose Guided Global and Local GAN for Appearance Preserving Human Video Prediction
abstract
We propose a pose-guided approach for appearance preserving video prediction by combining global and local information using Generative Adversarial Networks (GANs). The aim is to predict the subsequent frames based on previous frames of human action videos. Considering that human action videos contain both background scenes which are relatively time-invariant among frames, and human actions which are time-varying components, we use a global GAN to model the time-invariant background and coarse human profiles. Then, a local GAN is utilized to further refine the time-varying human parts. Finally, we use a 3D auto-encoder to fine-tune the frame-by-frame images to obtain the whole predicted video. We evaluate our model on the Penn Action and J-HMDB datasets and demonstrate the superiority of our proposed method over other state-of-the-art methods.
Jilin Tang, Haoji Hu, Hangguan Shan, Chuan Tian, Tony Q. S. Quek
ICIP1