Muhammad Gohar Javed

dblp:310/5763 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 42% Generative modeling · 28% 3D vision · 15%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling
1.622025
InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling · ICLR 2025
MoMask: Generative Masked Modeling of 3D Human Motions · CVPR 2024
Computer vision › 3D vision › human body modeling › 3d human modeling
3d human motion generation
0.912025
InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling · ICLR 2025
Machine learning › Generative modeling
masked generative modeling
0.912025
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer · ICLR 2025
Machine learning › Deep learning architectures and training › transformer
masked transformer
0.912025
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer · ICLR 2025
Computer animation and physical simulation
motion synthesis
0.912025
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer · ICLR 2025
Machine learning › Generative modeling › diffusion model › human motion generation
text-to-motion generation
0.812024
MoMask: Generative Masked Modeling of 3D Human Motions · CVPR 2024
Machine learning › Representation and self-supervised learning
vector quantization
0.812024
MoMask: Generative Masked Modeling of 3D Human Motions · CVPR 2024
Computer animation and physical simulation
motion editing
0.312025
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer · ICLR 2025

Methods — techniques the papers use, named apart from their topics

vector quantization · 2.5transformer · 1.7sliding window local attention · 1.7masked modeling · 1.7VQ-VAE · 1.7residual transformer · 0.8masked transformer · 0.8
YearPublicationVenuePosition
2025 InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling
abstract
Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often produce results lacking realism and fidelity. In this work, we introduce *InterMask*, a novel framework for generating human interactions using collaborative masked modeling in discrete space. InterMask first employs a VQ-VAE to transform each motion sequence into a 2D discrete motion token map. Unlike traditional 1D VQ token maps, it better preserves fine-grained spatio-temporal details and promotes *spatial awareness* within each token. Building on this representation, InterMask utilizes a generative masked modeling framework to collaboratively model the tokens of two interacting individuals. This is achieved by employing a transformer architecture specifically designed to capture complex spatio-temporal inter-dependencies. During training, it randomly masks the motion tokens of both individuals and learns to predict them. For inference, starting from fully masked sequences, it progressively fills in the tokens for both individuals. With its enhanced motion representation, dedicated architecture, and effective learning strategy, InterMask achieves state-of-the-art results, producing high-fidelity and diverse human interactions. It outperforms previous methods, achieving an FID of $5.154$ (vs $5.535$ of in2IN) on the InterHuman dataset and $0.399$ (vs $5.207$ of InterGen) on the InterX dataset. Additionally, InterMask seamlessly supports reaction generation without the need for model redesign or fine-tuning.
Muhammad Gohar Javed, Chuan Guo 0002, Li Cheng 0001
ICLR1
2025 MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
abstract
Generative masked transformer have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying generative masked modeling to generate diverse instances from a single MoCap reference may lead to overfitting, a challenge that remains unexplored. In this work, we present MotionDreamer, a localized masked modeling paradigm designed to learn motion internal patterns from a given motion with arbitrary topology and duration. By embedding the given motion into quantized tokens with a novel distribution regularization method, MotionDreamer constructs a robust and informative codebook for local motion patterns. Moreover, a sliding window local attention is introduced in our masked transformer, enabling the generation of natural yet diverse animations that closely resemble the reference motion patterns. As demonstrated through comprehensive experiments, MotionDreamer outperforms the state-of-the-art methods that are typically GAN or Diffusion-based in both faithfulness and diversity. Thanks to the consistency and robustness of quantization-based approach, MotionDreamer can also effectively perform downstream tasks such as temporal motion editing, crowd motion synthesis, and beat-aligned dance generation, all using a single reference motion. Our implementation, learned models and results are to be made publicly available upon paper acceptance.
Chuan Guo 0002, Yuxuan Mu, Muhammad Gohar Javed, Xinxin Zuo, Juwei Lu, Hai Jiang 0001, Li Cheng 0001
ICLR4
2024 MoMask: Generative Masked Modeling of 3D Human Motions
abstract
We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In Mo-Mask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity details. Starting at the base layer, with a sequence of motion tokens obtained by vector quan-tization, the residual tokens of increasing orders are de-rived and stored at the subsequent layers of the hierar-chy. This is consequently followed by two distinct bidirectional transformers. For the base-layer motion tokens, a Masked Transformer is designated to predict randomly masked motion tokens conditioned on text input at training stage. During generation (i. e. inference) stage, starting from an empty sequence, our Masked Transformer iteratively fills up the missing tokens; Subsequently, a Residual Transformer learns to progressively predict the next-layer tokens based on the results from current layer. Extensive experiments demonstrate that MoMask outperforms the state-of-art methods on the text-to-motion generation task, with an FID of 0.045 (vs e.g. 0.141 of T2M-GPT) on the HumanML3D dataset, and 0.228 (vs 0.514) on KIT-ML, respectively. MoMask can also be seamlessly applied in related tasks without further model fine-tuning, such as text-guided temporal inpainting.
Chuan Guo 0002, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang 0003, Li Cheng 0001
CVPR3