Jiteng Mu

dblp:249/2990 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 32% 3D vision · 31% Trustworthy machine learning · 11%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 25 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
synthetic-to-real domain adaptation
1.022022
Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic Vehicles · CVPR 2022
Learning From Synthetic Animals · CVPR 2020
Machine learning › Generative modeling
autoregressive model
0.912025
EditAR: Unified Conditional Generation with Autoregressive Models · CVPR 2025
Machine learning › Trustworthy machine learning › interpretability › concept-based explanation
concept attribution
0.912025
IntroStyle: Training-Free Introspective Style Attribution Using Diffusion Features · ICCV 2025
Machine learning › Generative modeling › image generation
conditional image generation
0.912025
EditAR: Unified Conditional Generation with Autoregressive Models · CVPR 2025
Machine learning › Generative modeling
diffusion model
0.912025
IntroStyle: Training-Free Introspective Style Attribution Using Diffusion Features · ICCV 2025
Machine learning › Generative modeling › diffusion model
image editing
0.912025
EditAR: Unified Conditional Generation with Autoregressive Models · CVPR 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
IntroStyle: Training-Free Introspective Style Attribution Using Diffusion Features · ICCV 2025
Machine learning › Generative modeling › image generation
token-based image generation
0.912025
EditAR: Unified Conditional Generation with Autoregressive Models · CVPR 2025
Visual content generation and editing › image generation
controllable image generation
0.812024
Editable Image Elements for Controllable Synthesis · ECCV (2) 2024
Computer vision › 3D vision
human rendering
0.712023
ActorsNeRF: Animatable Few-shot Human Rendering with Generalizable NeRFs · ICCV 2023
Computer vision › 3D vision
neural radiance field
0.712023
ActorsNeRF: Animatable Few-shot Human Rendering with Generalizable NeRFs · ICCV 2023
Computer vision › 3D vision
neural rendering
0.712023
ActorsNeRF: Animatable Few-shot Human Rendering with Generalizable NeRFs · ICCV 2023
Computer vision › 3D vision › correspondence estimation
dense correspondence
0.612022
CoordGAN: Self-Supervised Dense Correspondences Emerge from GANs · CVPR 2022
Machine learning › Generative modeling
generative adversarial network
0.612022
CoordGAN: Self-Supervised Dense Correspondences Emerge from GANs · CVPR 2022
Computer vision › Segmentation and scene understanding
part segmentation
0.612022
Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic Vehicles · CVPR 2022
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised correspondence learning
0.612022
CoordGAN: Self-Supervised Dense Correspondences Emerge from GANs · CVPR 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.612022
Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic Vehicles · CVPR 2022
Computer vision › 3D vision
3d shape representation
0.512021
A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation · ICCV 2021
Computer vision › 3D vision › 3d shape modeling
articulated object modeling
0.512021
A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation · ICCV 2021
Computer vision › 3D vision › 3d shape representation
implicit function
0.512021
A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation · ICCV 2021
Computer vision › 3D vision › 3d shape representation › implicit surface representation
signed distance function
0.512021
A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation · ICCV 2021
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization
0.412020
Learning From Synthetic Animals · CVPR 2020
Computer vision › Vision and language › cross-modal alignment › image-text alignment
text-to-image alignment
0.312025
EditAR: Unified Conditional Generation with Autoregressive Models · CVPR 2025
Visual content generation and editing › style control
artistic style control
0.312025
IntroStyle: Training-Free Introspective Style Attribution Using Diffusion Features · ICCV 2025
Visual content generation and editing › image generation
text-to-image generation
0.312025
IntroStyle: Training-Free Introspective Style Attribution Using Diffusion Features · ICCV 2025

Methods — techniques the papers use, named apart from their topics

synthetic dataset construction · 1.7diffusion features · 1.7knowledge distillation · 0.9autoregressive modeling · 0.9few-shot learning · 0.7canonical space alignment · 0.7warped coordinate frames · 0.6spatial structure guidance · 0.6encoder · 0.6domain adaptation · 0.6
YearPublicationVenuePosition
2025 EditAR: Unified Conditional Generation with Autoregressive Models
abstract
Recent progress in controllable image generation and editing is largely driven by diffusion-based methods. Although diffusion models perform exceptionally well in specific tasks with tailored designs, establishing a unified model is still challenging. In contrast, autoregressive models inherently feature a unified tokenized representation, which simplifies the creation of a single foundational model for various tasks. In this work, we propose EditAR, a single unified autoregressive framework for a variety of conditional image generation tasks, e.g., image editing, depth-to-image, edge-to-image, segmentation-to-image. The model takes both images and instructions as inputs, and predicts the edited images tokens in a vanilla next-token paradigm. To enhance the text-to-image alignment, we further propose to distill the knowledge from foundation models into the autoregressive modeling process. We evaluate its effectiveness across diverse tasks on established benchmarks, showing competitive performance to various state-of-the-art task-specific methods. Project page: https://jitengmu.github.io/EditAR/
Jiteng Mu, Nuno Vasconcelos, Xiaolong Wang 0004
CVPR1
2025 IntroStyle: Training-Free Introspective Style Attribution Using Diffusion Features
abstract
Text-to-image (T2I) models have recently gained widespread adoption. This has spurred concerns about safeguarding intellectual property rights and an increasing demand for mechanisms that prevent the generation of specific artistic styles. Existing methods for style extraction typically necessitate the collection of custom datasets and the training of specialized models. This, however, is resource-intensive, time-consuming, and often impractical for real-time applications. We present a novel, training-free framework to solve the style attribution problem, using the features produced by a diffusion model alone, without any external modules or retraining. This is denoted as Introspective Style attribution (IntroStyle) and is shown to have superior performance to state-of-the-art models for style attribution. We also introduce a synthetic dataset of Artistic Style Split (ArtSplit) to isolate artistic style and evaluate fine-grained style attribution performance. Our experimental results on WikiArt and DomainNet datasets show that \ours is robust to the dynamic nature of artistic styles, outperforming existing methods by a wide margin.
Jiteng Mu, Nuno Vasconcelos
ICCV2
2025 Learning Generalizable Feature Fields for Mobile Manipulation
abstract
An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while understanding fine-grained semantics, whereas the former involves capturing the complexity inherent at an expansive physical scale. In this work, we present GeFF (Generalizable Feature Fields), a scene-level generalizable neural feature field that acts as a unified representation for both navigation and manipulation that performs in real-time. To do so, we treat generative novel view synthesis as a pre-training task, and then align the resulting rich scene priors with natural language via CLIP feature distillation. We demonstrate the effectiveness of this approach by deploying GeFF on a quadrupedal robot equipped with a manipulator. We quantitatively evaluate GeFF’s ability for open-vocabulary object-/part-level manipulation and show that GeFF outperforms point-based baselines in runtime and storage-accuracy trade-offs, with qualitative examples of semantics-aware navigation and articulated object manipulation.
Ri-Zhao Qiu, Yafei Hu, Jianglong Ye, Jiteng Mu, Ruihan Yang, Nikolay Atanasov 0001, Sebastian A. Scherer, Xiaolong Wang 0004
IROS7
2024 Editable Image Elements for Controllable Synthesis
Jiteng Mu, Michaël Gharbi, Richard Zhang 0001, Eli Shechtman, Nuno Vasconcelos, Xiaolong Wang 0004, Taesung Park
ECCV (2)1
2023 ActorsNeRF: Animatable Few-shot Human Rendering with Generalizable NeRFs
abstract
While NeRF-based human representations have shown impressive novel view synthesis results, most methods still rely on a large number of images / views for training. In this work, we propose a novel animatable NeRF called ActorsNeRF. It is first pre-trained on diverse human subjects, and then adapted with few-shot monocular video frames for a new actor with unseen poses. Building on previous generalizable NeRFs with parameter sharing using a ConvNet encoder, ActorsNeRF further adopts two human priors to capture the large human appearance, shape, and pose variations. Specifically, in the encoded feature space, we will first align different human subjects in a category-level canonical space, and then align the same human from different frames in an instance-level canonical space for rendering. We quantitatively and qualitatively demonstrate that ActorsNeRF significantly outperforms the existing state-of-the-art on few-shot generalization to new people and poses on multiple datasets. Project page: https://jitengmu.github.io/ActorsNeRF/.
Jiteng Mu, Shen Sang, Nuno Vasconcelos, Xiaolong Wang 0004
ICCV1
2022 Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic Vehicles
abstract
Part segmentations provide a rich and detailed part-level description of objects. However, their annotation requires an enormous amount of work, which makes it difficult to apply standard deep learning methods. In this paper, we propose the idea of learning part segmentation through unsupervised domain adaptation (UDA) from synthetic data. We first introduce UDA-Part, a comprehensive part segmentation dataset for vehicles that can serve as an adequate benchmark for UDA11https://qliu24.github.io/udapart/. In UDA-Part, we label parts on 3D CAD models which enables us to generate a large set of annotated synthetic images. We also annotate parts on a number of real images to provide a real test set. Secondly, to advance the adaptation of part models trained from the synthetic data to the real images, we introduce a new UDA algorithm that leverages the object's spatial structure to guide the adaptation process. Our experimental results on two real test datasets confirm the superiority of our approach over existing works, and demonstrate the promise of learning part segmentation for general objects from synthetic data. We believe our dataset provides a rich testbed to study UDA for part segmentation and will help to significantly push forward research in this area.
Qing Liu 0017, Adam Kortylewski, Zhishuai Zhang, Zizhang Li, Mengqi Guo, Qihao Liu, Xiaoding Yuan, Jiteng Mu, Weichao Qiu, Alan L. Yuille
CVPR8
2022 CoordGAN: Self-Supervised Dense Correspondences Emerge from GANs
abstract
Recent advances show that Generative Adversarial Networks (GANs) can synthesize images with smooth variations along semantically meaningful latent directions, such as pose, expression, layout, etc. While this indicates that GANs implicitly learn pixel-level correspondences across images, few studies explored how to extract them explicitly. In this work, we introduce Coordinate GAN (CoordGAN), a structure-texture disentangled GAN that learns a dense correspondence map for each generated image. We represent the correspondence maps of different images as warped coordinate frames transformed from a canonical coordinate frame, i.e., the correspondence map, which describes the structure (e.g., the shape of a face), is controlled via a transformation. Hence, finding correspondences boils down to locating the same coordinate in different correspondence maps. In CoordGAN, we sample a transformation to represent the structure of a synthesized instance, while an independent texture branch is responsible for rendering appearance details orthogonal to the structure. Our approach can also extract dense correspondence maps for real images by adding an encoder on top of the generator. We quantitatively demonstrate the quality of the learned dense correspondences through segmentation mask transfer on multiple datasets. We also show that the proposed generator achieves better structure and texture disentanglement compared to existing approaches. Project page: https://jitengmu.github.io/CoordGAN/
Jiteng Mu, Shalini De Mello, Zhiding Yu, Nuno Vasconcelos, Xiaolong Wang 0004, Jan Kautz, Sifei Liu
CVPR1
2021 A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation
abstract
Recent work has made significant progress on using implicit functions, as a continuous representation for 3D rigid object shape reconstruction. However, much less effort has been devoted to modeling general articulated objects. Compared to rigid objects, articulated objects have higher degrees of freedom, which makes it hard to generalize to unseen shapes. To deal with the large shape variance, we introduce Articulated Signed Distance Functions (A-SDF) to represent articulated shapes with a disentangled latent space, where we have separate codes for encoding shape and articulation. With this disentangled continuous representation, we demonstrate that we can control the articulation input and animate unseen instances with unseen joint angles. Furthermore, we propose a Test-Time Adaptation inference algorithm to adjust our model during inference. We demonstrate our model generalize well to out-of-distribution and unseen data, e.g., partial point clouds and real-world depth images. Project page with code: https://jitengmu.github.io/A-SDF/.
Jiteng Mu, Weichao Qiu, Adam Kortylewski, Alan L. Yuille, Nuno Vasconcelos, Xiaolong Wang 0004
ICCV1
2020 Learning From Synthetic Animals
abstract
Despite great success in human parsing, progress for parsing other deformable articulated objects, like animals, is still limited by the lack of labeled data. In this paper, we use synthetic images and ground truth generated from CAD animal models to address this challenge. To bridge the domain gap between real and synthetic images, we propose a novel consistency-constrained semi-supervised learning method (CC-SSL). Our method leverages both spatial and temporal consistencies, to bootstrap weak models trained on synthetic data with unlabeled real images. We demonstrate the effectiveness of our method on highly deformable animals, such as horses and tigers. Without using any real image label, our method allows for accurate keypoint prediction on real images. Moreover, we quantitatively show that models using synthetic data achieve better generalization performance than models trained on real images across different domains in the Visual Domain Adaptation Challenge dataset. Our synthetic dataset contains 10+ animals with diverse poses and rich ground truth, which enables us to use the multi-task learning strategy to further boost models' performance.
Jiteng Mu, Weichao Qiu, Gregory D. Hager, Alan L. Yuille
CVPR1