EDBT 2026 Demo / reviewers in the wild / expert
Can Wang 0007
dblp:71/4716-7
· DBLP profile ↗
15ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-5102-1464ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unifying Multi-Modal Hair Editing via Proxy Feature BlendingabstractHair editing is a long-standing problem in computer vision that demands both fine-grained local control and intuitive user interactions across diverse modalities. Despite the remarkable progress of GANs and diffusion models, existing methods still lack a unified framework that simultaneously supports arbitrary interaction modes (e.g., text, sketch, mask, and reference image) while ensuring precise editing and faithful preservation of irrelevant attributes. In this work, we introduce a novel paradigm that reformulates hair editing as proxy-based hair transfer. Specifically, we leverage the dense and semantically disentangled latent space of StyleGAN for precise manipulation and exploit its feature space for disentangled attribute preservation, thereby decoupling the objectives of editing and preservation. Our framework unifies different modalities by converting editing conditions into distinct transfer proxies, whose features are seamlessly blended to achieve global or local edits. Beyond 2D, we extend our paradigm to 3D-aware settings by incorporating EG3D and PanoHead, where we propose a multi-view boosted hair feature localization strategy together with 3D-tailored proxy generation methods that exploit the inherent properties of 3D-aware generative models. Extensive experiments demonstrate that our method consistently outperforms prior approaches in editing effects, attribute preservation, visual naturalness, and multi-view consistency, while offering unprecedented support for multimodal and mixed-modal interactions. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Can Wang 0007, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Chat2Layout: Interactive 3D Furniture Layout With a Multimodal LLMabstractAutomatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), recent methods address layout generation in a static manner, lacking the feedback-driven refinement essential for interactive user engagement. We introduce Chat2Layout, a novel interactive furniture layout generation system that extends the functionality of MLLMs into the realm of interactive layout design. To achieve this, we establish a unified vision-question paradigm for in-context learning, enabling seamless communication with MLLMs to steer their behavior without altering model weights. Within this framework, we present a novel training-free visual prompting mechanism. This involves a visual-text prompting technique that assist MLLMs in reasoning about plausible layout plans, followed by an Offline-to-Online search (O2O-Search) method, which identifies the minimal set of informative references to provide exemplars for visual-text prompting. By employing an agent system with MLLMs as the core controller, we enable bidirectional interaction. The agent not only comprehends the 3D environment and user requirements through linguistic and visual perception but also plans tasks and reasons about actions to generate and arrange furniture within the virtual space. Furthermore, the agent iteratively updates based on visual feedback from execution results. Experimental results demonstrate that our approach facilitates language-interactive generation and arrangement for diverse and complex 3D furniture. Can Wang 0007, Hongliang Zhong, Menglei Chai, Mingming He, Dongdong Chen 0001, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Animus3D: Text-driven 3D Animation via Motion Score DistillationabstractWe present Animus3D , a text-driven 3D animation framework that generates motion field given a static 3D asset and text prompt. Previous methods mostly leverage the vanilla Score Distillation Sampling (SDS) objective to distill motion from pretrained text-to-video diffusion, leading to animations with minimal movement or noticeable jitter. To address this, our approach introduces a novel SDS alternative, Motion Score Distillation (MSD). Specifically, we introduce a LoRA-enhanced video diffusion model that defines a static source distribution rather than pure noise as in SDS, while another inversion-based noise estimation technique ensures appearance preservation when guiding motion. To further improve motion fidelity, we incorporate explicit temporal and spatial regularization terms that mitigate geometric distortions across time and space. Additionally, we propose a motion refinement module to upscale the temporal resolution and enhance fine-grained details, overcoming the fixed-resolution constraints of the underlying video model. Extensive experiments demonstrate that Animus3D successfully animates static 3D assets from diverse text prompts, generating significantly more substantial and detailed motion than state-of-the-art baselines while maintaining high visual integrity. Code will be released upon acceptance. Qi Sun 0005, Can Wang 0007, Jiaxiang Shang, Wensen Feng, Jing Liao 0001 |
SIGGRAPH Asia | 2 |
| 2025 | StyleRetoucher: Generalized Portrait Image Retouching With GAN PriorsabstractCreating fine-retouched portrait images is tedious and time-consuming even for professional artists. There exist automatic retouching methods, but they either suffer from over-smoothing artifacts or lack generalization ability. To address such issues, we present StyleRetoucher, a novel automatic portrait image retouching framework, leveraging StyleGAN's generation and generalization ability to improve an input portrait image's skin condition while preserving its facial details. Harnessing the priors of pretrained StyleGAN, our method shows superior robustness: a). performing stably with fewer training samples and b). generalizing well on the out-domain data. Moreover, by blending the spatial features of the input image and intermediate features of the StyleGAN layers, our method preserves the input characteristics to the largest extent. We further propose a novel blemish-aware feature selection mechanism to effectively identify and remove the skin blemishes, improving the image skin condition. Qualitative and quantitative evaluations validate the great generalization capability of our method. Further experiments show StyleRetoucher's superior performance to the alternative solutions in the image retouching task. We also conduct a user perceptive study to confirm the superior retouching performance of our method over the existing state-of-the-art alternatives. Wanchao Su, Can Wang 0007, Fangzhou Han, Hongbo Fu 0001, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Generative object insertion in Gaussian splatting with a multi-view diffusion modelabstractGenerating and inserting new objects into 3D content is a compelling approach for achieving versatile scene recreation. Existing methods, which rely on SDS optimization or single-view inpainting, often struggle to produce high-quality results. To address this, we propose a novel method for object insertion in 3D content represented by Gaussian Splatting. Our approach introduces a multi-view diffusion model, dubbed MVInpainter, which is built upon a pre-trained stable video diffusion model to facilitate view-consistent object inpainting. Within MVInpainter, we incorporate a ControlNet-based conditional injection module to enable controlled and more predictable multi-view generation. After generating the multi-view inpainted results, we further propose a mask-aware 3D reconstruction technique to refine Gaussian Splatting reconstruction from these sparse inpainted views. By leveraging these fabricate techniques, our approach yields diverse results, ensures view-consistent and harmonious insertions, and produces better object quality. Extensive experiments demonstrate that our approach outperforms existing methods. Hongliang Zhong, Can Wang 0007, Jingbo Zhang 0002, Jing Liao 0001 |
Vis. Informatics | 2 |
| 2024 | NeRF-Art: Text-Driven Neural Radiance Fields StylizationabstractAs a powerful representation of 3D scenes, the neural radiance field (NeRF) enables high-quality novel view synthesis from multi-view images. Stylizing NeRF, however, remains challenging, especially in simulating a text-guided style with both the appearance and the geometry altered simultaneously. In this paper, we present NeRF-Art, a text-guided NeRF stylization approach that manipulates the style of a pre-trained NeRF model with a simple text prompt. Unlike previous approaches that either lack sufficient geometry deformations and texture details or require meshes to guide the stylization, our method can shift a 3D scene to the target style characterized by desired geometry and appearance variations without any mesh guidance. This is achieved by introducing a novel global-local contrastive learning strategy, combined with the directional constraint to simultaneously control both the trajectory and the strength of the target style. Moreover, we adopt a weight regularization method to effectively suppress cloudy artifacts and geometry noises which arise easily when the density field is transformed during geometry stylization. Through extensive experiments on various styles, we demonstrate that our method is effective and robust regarding both single-view stylization quality and cross-view consistency. Can Wang 0007, Ruixiang Jiang, Menglei Chai, Mingming He, Dongdong Chen 0001, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose ControlabstractNeural implicit fields are powerful for representing 3D scenes and generating high-quality novel views, but it remains challenging to use such implicit representations for creating a 3D human avatar with a specific identity and artistic style that can be easily animated. Our proposed method, AvatarCraft, addresses this challenge by using diffusion models to guide the learning of geometry and texture for a neural avatar based on a single text prompt. We carefully design the optimization framework of neural implicit fields, including a coarse-to-fine multi-bounding box training strategy, shape regularization, and diffusion-based constraints, to produce high-quality geometry and texture. Additionally, we make the human avatar animatable by deforming the neural implicit field with an explicit warping field that maps the target human mesh to a template human mesh, both represented using parametric human models. This simplifies animation and reshaping of the generated avatar by controlling pose and shape parameters. Extensive experiments on various text descriptions show that AvatarCraft is effective and robust in creating human avatars and rendering novel views, poses, and shapes. Our project page is: https://avatar-craft.github.io/. Ruixiang Jiang, Can Wang 0007, Jingbo Zhang 0002, Menglei Chai, Mingming He, Dongdong Chen 0001, Jing Liao 0001 |
ICCV | 2 |
| 2023 | Dual Learning for Joint Facial Landmark Detection and Action Unit RecognitionabstractFacial landmark detection and action unit (AU) recognition are two essential tasks in facial analysis. Previous works rarely consider the relationship between these complementary tasks. In this article, we introduce a novel multi-task dual learning framework to exploit the relationship between facial landmark detection and AU recognition while simultaneously addressing both tasks. When both tasks share middle-level features, common patterns can be exploited and middle- and high-level features can be used to perform facial landmark detection and AU recognition, respectively. In addition, a dual learning mechanism is designed to convert the predicted landmarks and AUs of the label space to the corresponding facial image of the image space, further exploring the strong correlations between the tasks. By jointly training the proposed method at both the feature and label levels, each task improves the other. Experiments on two benchmark databases demonstrate that the proposed method can leverage dependencies to boost the generalization of both tasks. Shangfei Wang, Yanan Chang, Can Wang 0007 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Cross-Domain and Disentangled Face Manipulation With 3D GuidanceabstractFace image manipulation via three-dimensional guidance has been widely applied in various interactive scenarios due to its semantically-meaningful understanding and user-friendly controllability. However, existing 3D-morphable-model-based manipulation methods are not directly applicable to out-of-domain faces, such as non-photorealistic paintings, cartoon portraits, or even animals, mainly due to the formidable difficulties in building the model for each specific face domain. To overcome this challenge, we propose, as far as we know, the first method to manipulate faces in arbitrary domains using human 3DMM. This is achieved through two major steps: 1) disentangled mapping from 3DMM parameters to the latent space embedding of a pre-trained StyleGAN2 [1] that guarantees disentangled and precise controls for each semantic attribute; and 2) cross-domain adaptation that bridges domain discrepancies and makes human 3DMM applicable to out-of-domain faces by enforcing a consistent latent space embedding. Experiments and comparisons demonstrate the superiority of our high-quality semantic manipulation method on a variety of face domains with all major 3D facial attributes controllable - pose, expression, shape, albedo, and illumination. Moreover, we develop an intuitive editing interface to support user-friendly control and instant feedback. Our project page is https://cassiepython.github.io/cddfm3d/index.html. Can Wang 0007, Menglei Chai, Mingming He, Dongdong Chen 0001, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance FieldsabstractWe present CLIP-NeRF, a multi-modal 3D object manipulation method for neural radiance fields (NeRF). By leveraging the joint language-image embedding space of the recent Contrastive Language-Image Pre-Training (CLIP) model, we propose a unified framework that allows manip-ulating NeRF in a user-friendly way, using either a short text prompt or an exemplar image. Specifically, to combine the novel view synthesis capability of NeRF and the controllable manipulation ability of latent representations from generative models, we introduce a disentangled conditional NeRF architecture that allows individual control over both shape and appearance. This is achieved by performing the shape conditioning via applying a learned deformation field to the positional encoding and deferring color conditioning to the volumetric rendering stage. To bridge this disentangled latent representation to the CLIP embedding, we design two code mappers that take a CLIP embedding as input and update the latent codes to reflect the targeted editing. The mappers are trained with a CLIP-based matching loss to ensure the manipulation accuracy. Furthermore, we propose an inverse optimization method that accurately projects an input image to the latent codes for manipulation to enable editing on real images. We evaluate our approach by extensive experiments on a variety of text prompts and exemplar images and also provide an intuitive interface for interactive editing. Can Wang 0007, Menglei Chai, Mingming He, Dongdong Chen 0001, Jing Liao 0001 |
CVPR | 1 |
| 2022 | FDNeRF: Few-shot Dynamic Neural Radiance Fields for Face Reconstruction and Expression EditingabstractWe propose a Few-shot Dynamic Neural Radiance Field (FDNeRF), the first NeRF-based method capable of reconstruction and expression editing of 3D faces based on a small number of dynamic images. Unlike existing dynamic NeRFs that require dense images as input and can only be modeled for a single identity, our method enables face reconstruction across different persons with few-shot inputs. Compared to state-of-the-art few-shot NeRFs designed for modeling static scenes, the proposed FDNeRF accepts view-inconsistent dynamic inputs and supports arbitrary facial expression editing, i.e., producing faces with novel expressions beyond the input ones. To handle the inconsistencies between dynamic inputs, we introduce a well-designed conditional feature warping (CFW) module to perform expression conditioned warping in 2D feature space, which is also identity adaptive and 3D constrained. As a result, features of different expressions are transformed into the target ones. We then construct a radiance field based on these view-consistent features and use volumetric rendering to synthesize novel views of the modeled faces. Extensive experiments with quantitative and qualitative evaluation demonstrate that our method outperforms existing dynamic and few-shot NeRFs on both 3D face reconstruction and expression editing tasks. Code is available at https://fdnerf.github.io . Jingbo Zhang 0002, Xiaoyu Li 0002, Ziyu Wan, Can Wang 0007, Jing Liao 0001 |
SIGGRAPH Asia | 4 |
| 2022 | HITS: Binarizing physiological time series with deep hashing neural network
Zhaoji Fu, Can Wang 0007, Guodong Wei, Shaofu Du, Shenda Hong |
Pattern Recognit. Lett. | 2 |
| 2021 | Deep Portrait Lighting Enhancement with 3D GuidanceabstractAbstract Despite recent breakthroughs in deep learning methods for image lighting enhancement, they are inferior when applied to portraits because 3D facial information is ignored in their models. To address this, we present a novel deep learning framework for portrait lighting enhancement based on 3D facial guidance. Our framework consists of two stages. In the first stage, corrected lighting parameters are predicted by a network from the input bad lighting image, with the assistance of a 3D morphable model and a differentiable renderer. Given the predicted lighting parameter, the differentiable renderer renders a face image with corrected shading and texture, which serves as the 3D guidance for learning image lighting enhancement in the second stage. To better exploit the long‐range correlations between the input and the guidance, in the second stage, we design an image‐to‐image translation network with a novel transformer architecture, which automatically produces a lighting‐enhanced result. Experimental results on the FFHQ dataset and in‐the‐wild images show that the proposed method outperforms state‐of‐the‐art methods in terms of both quantitative metrics and visual quality. Fangzhou Han, Can Wang 0007, Hao Du 0006, Jing Liao 0001 |
Comput. Graph. Forum | 2 |
| 2021 | Video Affective Content Analysis by Exploring Domain KnowledgeabstractFilm grammar is often used to invoke certain emotional experiences from audiences through changing visual, speech, and musical elements of videos. Such film grammar, referred to as domain knowledge, is of great importance for video affective content analysis but has not been thoroughly examined in research. In this paper, we propose an improved method for emotion recognition and regression from videos through exploring domain knowledge. We first investigate the domain knowledge of visual, speech, and musical elements, and infer probabilistic dependencies between elements and emotions from the summarized film grammar. Then, we transfer the summarized dependencies between elements and emotions as constraints, and formulate video affective content analysis, including both emotion recognition and emotion regression from video content, as a constrained optimization problem. Experiments on the LIRIS-ACCEDE database, the FilmStim database, and the DEAP database demonstrate that the proposed video affective content analysis method can successfully leverage well-established film grammar to improve emotion recognition and regression from video content. Shangfei Wang, Can Wang 0007, Tanfang Chen, Yangyang Shu |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | CardioID: learning to identification from electrocardiogram data
Shenda Hong, Can Wang 0007, Zhaoji Fu |
Neurocomputing | 2 |