VLDB 2026 Research / reviewers in the wild / expert
Jilin Tang
dblp:254/8073
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0001-9478-7489ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AS-NeRF: Learning Auxiliary Sampling for Generalizable Novel View Synthesis from Sparse ViewsabstractWe tackle the problem of novel view synthesis (NVS) which aims to generate realistic images at novel views. Unlike existing works that require either costly per-scene optimization or relatively dense views, we propose AS-NeRF, a neural rendering based approach that can achieve generalizable NVS from only sparse views. Considering the inherent spatial continuity of images, we design a novel sparse-attention based auxiliary sampling module (ASM). Given a 3D point on the ray, the ASM adaptively attends to a sparse set of view-specific 2D auxiliary locations around the 3D point’s original projection pixels, and dynamically computes the attention weights in a cross-attention manner. This enables our model to effectively exploit the local correlation among neighboring pixels, obtaining the enhanced features with more powerful representation. Extensive experiments show that our method outperforms the state-of-the-art on both real and synthetic data. Jilin Tang, Lincheng Li, Xingqun Qi, Changjie Fan, Xin Yu 0002 |
ICME | 1 |
| 2024 | MarkerNet: A divide-and-conquer solution to motion capture solving from raw markersabstractAbstract Marker‐based optical motion capture (MoCap) aims to localize 3D human motions from a sequence of input raw markers. It is widely used to produce physical movements for virtual characters in various games such as the role‐playing game, the fighting game, and the action‐adventure game. However, the conventional MoCap cleaning and solving process is extremely labor‐intensive, time‐consuming, and usually the most costly part of game animation production. Thus, there is a high demand for automated algorithms to replace costly manual operations and achieve accurate MoCap cleaning and solving in the game industry. In this article, we design a divide‐and‐conquer‐based MoCap solving network, dubbed MarkerNet, to estimate human skeleton motions from sequential raw markers effectively. In a nutshell, our key idea is to decompose the task of direct solving of global motion from all markers into first modeling sub‐motions of local parts from the corresponding marker subsets and then aggregating sub‐motions into a global one. In this manner, our model can effectively capture local motion patterns w.r.t. different marker subsets, thus producing more accurate results compared to the existing methods. Extensive experiments on both real and synthetic data verify the effectiveness of the proposed method. Zhipeng Hu, Jilin Tang, Lincheng Li, Xin Yu 0002, Jiajun Bu |
Comput. Animat. Virtual Worlds | 2 |
| 2021 | Structure-aware Person Image Generation with Pose Decomposition and Semantic CorrelationabstractIn this paper we tackle the problem of pose guided person image generation, which aims to transfer a person image from the source pose to a novel target pose while maintaining the source appearance. Given the inefficiency of standard CNNs in handling large spatial transformation, we propose a structure-aware flow based method for high-quality person image generation. Specifically, instead of learning the complex overall pose changes of human body, we decompose the human body into different semantic parts (e.g., head, torso, and legs) and apply different networks to predict the flow fields for these parts separately. Moreover, we carefully design the network modules to effectively capture the local and global semantic correlations of features within and among the human parts respectively. Extensive experimental results show that our method can generate high-quality results under large pose discrepancy and outperforms state-of-the-art methods in both qualitative and quantitative comparisons. Jilin Tang, Yi Yuan 0002, Tianjia Shao, Yong Liu 0007, Mengmeng Wang 0005, Kun Zhou 0001 |
AAAI | 1 |
| 2021 | Vanet: a View Attention Guided Network for 3d Reconstruction from Single and Multi-View ImagesabstractReconstructing 3D meshes of objects from 2D images is an important but challenging task. Previous 3D reconstruction methods either only focus on generating the mesh from a single image, or multi-view images. Instead of investigating these problems separately, we present a novel view attention guided network called VANet which addresses both single and multi-view 3D reconstruction under a unified frame-work. To explore non-visible parts of an object during the re-construction, a channel-wise view attention mechanism and a dual pathway network architecture are introduced. The proposed network highlights the informative object parts and compensates those non-informative ones with auxiliary views of input. Yi Yuan 0002, Jilin Tang, Zhengxia Zou |
ICME | 2 |
| 2021 | Attention guided feature pyramid network for crowd counting
Huanpeng Chu, Jilin Tang, Haoji Hu |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Pose Guided Global and Local GAN for Appearance Preserving Human Video PredictionabstractWe propose a pose-guided approach for appearance preserving video prediction by combining global and local information using Generative Adversarial Networks (GANs). The aim is to predict the subsequent frames based on previous frames of human action videos. Considering that human action videos contain both background scenes which are relatively time-invariant among frames, and human actions which are time-varying components, we use a global GAN to model the time-invariant background and coarse human profiles. Then, a local GAN is utilized to further refine the time-varying human parts. Finally, we use a 3D auto-encoder to fine-tune the frame-by-frame images to obtain the whole predicted video. We evaluate our model on the Penn Action and J-HMDB datasets and demonstrate the superiority of our proposed method over other state-of-the-art methods. Jilin Tang, Haoji Hu, Hangguan Shan, Chuan Tian, Tony Q. S. Quek |
ICIP | 1 |