Simian Luo

dblp:317/0715 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 94% Vision and language · 6%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 50% Audio and music processing · 50%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
3d shape generation
0.712023
Learning Versatile 3D Shape Generation with Improved Auto-regressive Models · ICCV 2023
Machine learning › Generative modeling
audio generation
0.712023
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models · NeurIPS 2023
Machine learning › Generative modeling
autoregressive model
0.712023
Learning Versatile 3D Shape Generation with Improved Auto-regressive Models · ICCV 2023
Machine learning › Generative modeling
diffusion model
0.712023
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.712023
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models · NeurIPS 2023
Visual content generation and editing
3d shape generation
0.712023
Learning Versatile 3D Shape Generation with Improved Auto-regressive Models · ICCV 2023
Audio and music processing › sound synthesis
video-to-audio generation
0.712023
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models · NeurIPS 2023
Computer vision › Vision and language › cross-modal alignment › cross-modal feature alignment
audio-visual alignment
0.212023
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

latent vector representation · 1.3latent diffusion · 1.3double guidance · 1.3discrete representation learning · 1.3cross-attention · 1.3contrastive audio-visual pretraining · 1.3
YearPublicationVenuePosition
2023 Learning Versatile 3D Shape Generation with Improved Auto-regressive Models
abstract
Auto-Regressive (AR) models have achieved impressive results in 2D image generation by modeling joint distributions in the grid space. While this approach has been extended to the 3D domain for powerful shape generation, it still has two limitations: expensive computations on volumetric grids and ambiguous auto-regressive order along grid dimensions. To overcome these limitations, we propose the Improved Auto-regressive Model (ImAM) for 3D shape generation, which applies discrete representation learning based on a latent vector instead of volumetric grids. Our approach not only reduces computational costs but also preserves essential geometric details by learning the joint distribution in a more tractable order. Moreover, thanks to the simplicity of our model architecture, we can naturally extend it from unconditional to conditional generation by concatenating various conditioning inputs, such as point clouds, categories, images, and texts. Extensive experiments demonstrate that ImAM can synthesize diverse and faithful shapes of multiple categories, achieving state-of-the-art performance.
Simian Luo, Xuelin Qian, Yanwei Fu 0001, Yinda Zhang 0001, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Xiangyang Xue 0001
ICCV1
2023 Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models
abstract
The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation quality in terms of temporal synchronization and audio-visual relevance. We present Diff-Foley, a synchronized Video-to-Audio synthesis method with a latent diffusion model (LDM) that generates high-quality audio with improved synchronization and audio-visual relevance. We adopt contrastive audio-visual pretraining (CAVP) to learn more temporally and semantically aligned features, then train an LDM with CAVP-aligned visual features on spectrogram latent space. The CAVP-aligned features enable LDM to capture the subtler audio-visual correlation via a cross-attention module. We further significantly improve sample quality with `double guidance'. Diff-Foley achieves state-of-the-art V2A performance on current large scale V2A dataset. Furthermore, we demonstrate Diff-Foley practical applicability and adaptability via customized downstream finetuning. Project Page: https://diff-foley.github.io/
Simian Luo, Chuanhao Yan, Chenxu Hu, Hang Zhao 0021
NeurIPS1
2022 QS-Craft: Learning to Quantize, Scrabble and Craft for Conditional Human Motion Animation
Yuxin Hong, Xuelin Qian, Simian Luo, Guodong Guo, Xiangyang Xue 0001, Yanwei Fu 0001
ACCV (6)3