VLDB 2026 Research / reviewers in the wild / expert
Chunjin Song
dblp:230/8001
· DBLP profile ↗
6ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0002-7256-3510ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
4 papers |
Geometric modeling and processing · 51% Visual content generation and editing · 33% Image and video processing · 16% | |
| Artificial intelligence
3 papers |
3D vision · 97% Generative modeling · 3% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d human reconstruction
human avatar reconstruction |
1.6 | 2 | 2025 | Locality Sensitive Avatars From Video · ICLR 2025 Pose Modulated Avatars from Video · ICLR 2024 |
Computer vision › 3D vision
neural radiance field |
1.6 | 2 | 2025 | Locality Sensitive Avatars From Video · ICLR 2025 Pose Modulated Avatars from Video · ICLR 2024 |
Geometric modeling and processing
deformation modeling |
1.6 | 2 | 2025 | Locality Sensitive Avatars From Video · ICLR 2025 Pose Modulated Avatars from Video · ICLR 2024 |
Computer vision › 3D vision › 3d shape representation
canonical representation |
0.9 | 1 | 2025 | Locality Sensitive Avatars From Video · ICLR 2025 |
Geometric modeling and processing › shape deformation
non-rigid deformation |
0.9 | 1 | 2025 | Locality Sensitive Avatars From Video · ICLR 2025 |
Visual content generation and editing › style transfer
arbitrary style transfer |
0.8 | 2 | 2020 | EFANet: Exchangeable Feature Alignment Network for Arbitrary Style Transfer · AAAI 2020 ETNet: Error Transition Network for Arbitrary Style Transfer · NeurIPS 2019 |
Visual content generation and editing
style transfer |
0.8 | 2 | 2020 | EFANet: Exchangeable Feature Alignment Network for Arbitrary Style Transfer · AAAI 2020 ETNet: Error Transition Network for Arbitrary Style Transfer · NeurIPS 2019 |
Image and video processing › image matching
feature alignment |
0.4 | 1 | 2020 | EFANet: Exchangeable Feature Alignment Network for Arbitrary Style Transfer · AAAI 2020 |
Image and video processing › image enhancement
image refinement |
0.4 | 1 | 2019 | ETNet: Error Transition Network for Arbitrary Style Transfer · NeurIPS 2019 |
Machine learning › Generative modeling
iterative refinement |
0.1 | 1 | 2019 | ETNet: Error Transition Network for Arbitrary Style Transfer · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
neural radiance field · 3.3graph neural network · 3.3frequency modulation · 1.5progressive refinement · 0.8error transition network · 0.8whitening loss · 0.4feature alignment · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Locality Sensitive Avatars From VideoabstractWe present locality-sensitive avatar, a neural radiance field (NeRF) based network to learn human motions from monocular videos. To this end, we estimate a canonical representation between different frames of a video with a non-linear mapping from observation to canonical space, which we decompose into a skeletal rigid motion and a non-rigid counterpart. Our key contribution is to retain fine-grained details by modeling the non-rigid part with a graph neural network (GNN) that keeps the pose information local to neighboring body parts. Compared to former canonical representation based methods which solely operate on the coordinate space of a whole shape, our locality-sensitive motion modeling can reproduce both realistic shape contours and vivid fine-grained details. We evaluate on ZJU-MoCap, SynWild, ActorsHQ, MVHumanNet and various outdoor videos. The experiments reveal that with the locality sensitive deformation to canonical feature space, we are the first to achieve state-of-the-art results across novel view synthesis, novel pose animation and 3D shape reconstruction simultaneously. Our code is available at https://github.com/ChunjinSong/lsavatar. Chunjin Song, Shih-Yang Su, Bastian Wandt, Leonid Sigal, Helge Rhodin |
ICLR | 1 |
| 2025 | Representing Animatable Avatar via Factorized Neural FieldsabstractAbstract For reconstructing high‐fidelity human 3D models from monocular videos, it is crucial to maintain consistent large‐scale body shapes along with finely matched subtle wrinkles. This paper explores how per‐frame rendering results can be factorized into a pose‐independent component and a corresponding pose‐dependent counterpart to facilitate frame consistency at multiple scales. Pose adaptive texture features are further improved by restricting the frequency bands of these two components. Pose‐independent outputs are expected to be low‐frequency, while high‐frequency information is linked to pose‐dependent factors. We implement this with a dual‐branch network. The first branch takes coordinates in the canonical space as input, while the second one additionally considers features outputted by the first branch and pose information of each frame. A final network integrates the information predicted by both branches and utilizes volume rendering to generate photo‐realistic 3D human images. Through experiments, we demonstrate that our method consistently surpasses all state‐of‐the‐art methods in preserving high‐frequency details and ensuring consistent body contours. Our code is accessible at https://github.com/ChunjinSong/facavatar . Chunjin Song, Bastian Wandt, Leonid Sigal, Helge Rhodin |
Comput. Graph. Forum | 1 |
| 2024 | Pose Modulated Avatars from VideoabstractIt is now possible to reconstruct dynamic human motion and shape from a sparse set of cameras using Neural Radiance Fields (NeRF) driven by an underlying skeleton. However, a challenge remains to model the deformation of cloth and skin in relation to skeleton pose. Unlike existing avatar models that are learned implicitly or rely on a proxy surface, our approach is motivated by the observation that different poses necessitate unique frequency assignments. Neglecting this distinction yields noisy artifacts in smooth areas or blurs fine-grained texture and shape details in sharp regions. We develop a two-branch neural network that is adaptive and explicit in the frequency domain. The first branch is a graph neural network that models correlations among body parts locally, taking skeleton pose as input. The second branch combines these correlation features to a set of global frequencies and then modulates the feature encoding. Our experiments demonstrate that our network outperforms state-of-the-art methods in terms of preserving details and generalization capabilities. Our code is available at https://github.com/ChunjinSong/PM-Avatars. Chunjin Song, Bastian Wandt, Helge Rhodin |
ICLR | 1 |
| 2023 | AudioViewer: Learning to Visualize SoundsabstractA long-standing goal in the field of sensory substitution is enabling sound perception for deaf and hard of hearing (DHH) people by visualizing audio content. Different from existing models that translate to hand sign language, between speech and text, or text and images, we target immediate and low-level audio to video translation that applies to generic environment sounds as well as human speech. Since such a substitution is artificial, with-out labels for supervised learning, our core contribution is to build a mapping from audio to video that learns from unpaired examples via high-level constraints. For speech, we additionally disentangle content from style, such as gender and dialect. Qualitative and quantitative results, including a human study, demonstrate that our unpaired translation approach maintains important audio features in the generated video and that videos of faces and numbers are well suited for visualizing high-dimensional audio features that can be parsed by humans to match and distinguish between sounds and words. Project website: https://chunjinsong.github.io/audioviewer Chunjin Song, Yuchi Zhang, Willis Peng, Parmis Mohaghegh, Bastian Wandt, Helge Rhodin |
WACV | 1 |
| 2020 | EFANet: Exchangeable Feature Alignment Network for Arbitrary Style TransferabstractStyle transfer has been an important topic both in computer vision and graphics. Since the seminal work of Gatys et al. first demonstrates the power of stylization through optimization in the deep feature space, quite a few approaches have achieved real-time arbitrary style transfer with straightforward statistic matching techniques. In this work, our key observation is that only considering features in the input style image for the global deep feature statistic matching or local patch swap may not always ensure a satisfactory style transfer; see e.g., Figure 1. Instead, we propose a novel transfer framework, EFANet, that aims to jointly analyze and better align exchangeable features extracted from the content and style image pair. In this way, the style feature from the style image seeks for the best compatibility with the content information in the content image, leading to more structured stylization results. In addition, a new whitening loss is developed for purifying the computed content features and better fusion with styles in feature space. Qualitative and quantitative experiments demonstrate the advantages of our approach. Chunjin Song, Yang Zhou 0007, Minglun Gong, Hui Huang 0004 |
AAAI | 2 |
| 2019 | ETNet: Error Transition Network for Arbitrary Style TransferabstractNumerous valuable efforts have been devoted to achieving arbitrary style transfer since the seminal work of Gatys et al. However, existing state-of-the-art approaches often generate insufficiently stylized results under challenging cases. We believe a fundamental reason is that these approaches try to generate the stylized result in a single shot and hence fail to fully satisfy the constraints on semantic structures in the content images and style patterns in the style images. Inspired by the works on error-correction, instead, we propose a self-correcting model to predict what is wrong with the current stylization and refine it accordingly in an iterative manner. For each refinement, we transit the error features across both the spatial and scale domain and invert the processed features into a residual image, with a network we call Error Transition Network (ETNet). The proposed model improves over the state-of-the-art methods with better semantic structures and more adaptive style pattern details. Various qualitative and quantitative experiments show that the key concept of both progressive strategy and error-correction leads to better results. Code and models are available at https://github.com/zhijieW94/ETNet. Chunjin Song, Yang Zhou 0007, Minglun Gong, Hui Huang 0004 |
NeurIPS | 1 |