Weijian Ma

dblp:162/9819 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0009-0004-0650-6448ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Instruct Where the Model Fails: Generative Data Augmentation via Guided Self-contrastive Fine-tuning
abstract
Data augmentation is expected to bring about unseen features of training set, enhancing the model’s ability to generalize in situations where data is limited. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with stereotypes and imperceptible bias when used to augment training data, owing to dataset misalignment and the generator’s ignorance of the downstream model. We improve downstream task awareness in generated images by proposing a task-aware fine-tuning strategy that actively detects failures of downstream task in the target model to fine-tune the generation process between epochs. The dynamic fine-tuning strategy is achieved by (1) inspecting misalignment between generated data and original data via VLM captioners and (2) adjusts both prompts and diffusion model so that the strategy dynamically guides the generator by focusing on the detected bias of VLM. This is done via re-captioning the overfitted data as well as finetuning the diffusion trajectory in a contrastive manner. To co-operate with the VLM captioner, the contrastive fine-tuning process dynamically adjusts different parts of the diffusion trajectory based on detected misalignment, thus shifting the the generated distribution away from making the downstream model overfit. Our experiments on few-shot class incremental learning show that our instruction-guided finetuning strategy consistently assists the downstream model with higher classification accuracy compared to generative data augmentation baselines such as Stable Diffusion and GPT-4o, and state-of-the-art non-generative strategies.
Weijian Ma, Ruoxin Chen, Ke-Yue Zhang, Shuang Wu 0001, Shouhong Ding
AAAI1
2025 CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation
abstract
Recently, Large Language Models (LLMs) have achieved significant success, prompting increased interest in expanding their generative capabilities beyond general text into domain-specific areas. This study investigates the generation of parametric sequences for computer-aided design (CAD) models using LLMs. This endeavor represents an initial step towards creating parametric 3D shapes with LLMs, as CAD model parameters directly correlate with shapes in three-dimensional space. Despite the formidable generative capacities of LLMs, this task remains challenging, as these models neither encounter parametric sequences during their pretraining phase nor possess direct awareness of 3D structures. To address this, we present CAD-Llama, a framework designed to enhance pretrained LLMs for generating parametric 3D CAD models. Specifically, we develop a hierarchical annotation pipeline and a code-like format to translate parametric 3D CAD command sequences into Structured Parametric CAD Code (SPCC), incorporating hierarchical semantic descriptions. Furthermore, we propose an adaptive pretraining approach utilizing SPCC, followed by an instruction tuning process aligned with CAD-specific guidelines. This methodology aims to equip LLMs with the spatial knowledge inherent in parametric sequences. Experimental results demonstrate that our framework significantly outperforms prior autoregressive methods and existing LLM baselines.
Weijian Ma, Yunzhong Lou, Guichun Zhou
CVPR2
2025 MuSeLLM: SDF Generation and Understanding via Multi-Scale Tokenization with Position-Aware Guidance
abstract
The advancement of 3D generation and understanding stands as a critical interface enabling machines to interpret and interact with the physical world. Large language models, as trained on internet-scale text corpus, are well recognized to possess commonsense of real world and perception in real 3D space. However, when it comes to leveraging these abilities into 3D understanding and generation, the difference of data format is a significant gap between the two ends. In this sense, we propose MuSeLLM, a specially designed adapting strategy for finetuning LLMs on a specially selected data format, namely voxelized SDFs with respect to its easiness of tokenization. Based on this, we put forward a multi-scale tokenization strategy, allowing a coarse to fine generation paradigm on different levels of codebook tokens. This serves as a natural fit to prevailing LLM architectures where the generation of tokens at a certain level may refer to the entire shape representations at coarser levels. To avoid overfitting on spatial orders of tokens, we propose a position-aware guidance which perturbs the generation order of tokens in each level during the training stage, serving as a strong data augmentation strategy for adapting LLMs to 3D domains with few 3D data available. Experiment results in text-guided 3D generation and 3D object understanding illustrate that our method achieves superior performance against previous state-of-the-art with the same training data, as the inherent spatial reasoning ability has been triggered by our method design.
Tianwei Ding 0002, Lanshan He, Weijian Ma
ICMR3
2025 CADMorph: Geometry‑Driven Parametric CAD Editing via a Plan-Generate-Verify Loop
abstract
A Computer-Aided Design (CAD) model encodes an object in two coupled forms: a \emph{parametric construction sequence} and its resulting \emph{visible geometric shape}. During iterative design, adjustments to the geometric shape inevitably require synchronized edits to the underlying parametric sequence, called \emph{geometry-driven parametric CAD editing}. The task calls for 1) preserving the original sequence’s structure, 2) ensuring each edit's semantic validity, and 3) maintaining high shape fidelity to the target shape, all under scarce editing data triplets. We present \emph{CADMorph}, an iterative \emph{plan–generate–verify} framework that orchestrates pretrained domain-specific foundation models during inference: a \emph{parameter-to-shape} (P2S) latent diffusion model and a \emph{masked-parameter-prediction} (MPP) model. In the planning stage, cross-attention maps from the P2S model pinpoint the segments that need modification and offer editing masks. The MPP model then infills these masks with semantically valid edits in the generation stage. During verification, the P2S model embeds each candidate sequence in shape-latent space, measures its distance to the target shape, and selects the closest one. The three stages leverage the inherent geometric consciousness and design knowledge in pretrained priors, and thus tackle structure preservation, semantic validity, and shape fidelity respectively. Besides, both P2S and MPP models are trained without triplet data, bypassing the data-scarcity bottleneck. CADMorph surpasses GPT-4o and specialized CAD baselines, and supports downstream applications such as iterative editing and reverse-engineering enhancement.
Weijian Ma, Shizhao Sun, Jiang Bian 0002
NeurIPS1
2024 Draw Step by Step: Reconstructing CAD Construction Sequences from Point Clouds via Multimodal Diffusion
abstract
Reconstructing CAD construction sequences from raw 3D geometry serves as an interface between real-world objects and digital designs. In this paper, we propose CAD-Diffuser, a multimodal diffusion scheme aiming at integrating top-down design paradigm into generative reconstruction. In particular, we unify CAD point clouds and CAD construction sequences at the token level, guiding our proposed multimodal diffusion strategy to understand and link between the geometry and the design intent concentrated in construction sequences. Leveraging the strong decoding abilities of language models, the forward process is modeled as a random walk between the original token and the [MASK] token, while the reverse process naturally fits the masked token modeling scheme. A volume-based noise schedule is designed to encourage outline-first generation, decomposing the top-down design methodology into a machine-understandable procedure. For tokenizing CAD data of multiple modalities, we introduce a tokenizer with a self-supervised face segmentation task to compress local and global geometric information for CAD point clouds, and the CAD construction sequence is transformed into a primitive token string. Experimental results show that our CAD-Diffuser can perceive geometric details and the results are more likely to be reused by human designers.
Weijian Ma, Shuaiqi Chen, Yunzhong Lou
CVPR1
2024 Parametric CAD Primitive Retrieval via Multi-Modal Fusion and Deep Hashing
abstract
In the rapidly evolving field of manufacturing industry, product designers need to be able to quickly and accurately retrieve and reference existing Computer-Aided Design (CAD) primitives to increase efficiency and foster innovation. This paper introduces an innovative deep hashing-based parametric CAD primitive retrieval model DH-CAD, which employs a deep learning framework to capture the profound multi-modal features of parametric CAD primitive (command sequences and point clouds), and utilizes hashing to effectively encode these features into compact binary hash codes, significantly improving retrieval efficiency and accuracy. DH-CAD optimizes the original sequence encodings of Transformer to effectively learn the intricate relationships between CAD tool's command sequences and the corresponding point clouds of CAD primitive. Furthermore, it considers the dependencies between command types and parameter values, combines the output of the command type decoder with the prediction of parameter values to further refine the generated parameters, and employs a binarization network to generate hash codes in an end-to-end manner. Experimental results clearly demonstrate that, compared to traditional feature-matching methods, DH-CAD achieves the state-of-the-art performance on multiple evaluation metrics. Particularly in terms of retrieval speed and accuracy, our proposed model not only retrieves sequences that are most relevant to the query rapidly, but also ensures the high precision of the results, showcasing its efficiency and practicality in supporting CAD design processes.
Minyang Xu, Yunzhong Lou, Weijian Ma
ICMR3
2023 MultiCAD: Contrastive Representation Learning for Multi-modal 3D Computer-Aided Design Models
abstract
CAD models are multimodal data where information and knowledge contained in construction sequences and shapes are complementary to each other and representation learning methods should consider both of them. Such traits have been neglected in previous methods learning unimodal representations. To leverage the information from both modalities, we develop a multimodal contrastive learning strategy where features from different modalities interact via contrastive learning paradigm, driven by a novel multimodal contrastive loss. Two pretext tasks on both geometry and sequence domains are designed along with a two-stage training strategy to make the representation focus on encoding geometric details and decoding representations into construction sequences, thus being more applicable to downstream tasks such as multimodal retrieval and CAD sequence reconstruction. Experimental results show that the performance of our multimodal representation learning scheme has surpassed the baselines and unimodal methods significantly.
Weijian Ma, Minyang Xu
CIKM1
2021 Loading is the Key: A Novel Genetic Quantum Algorithm for SDVRP
abstract
This paper solves Split Demand Vehicle Routing Problem with minimal vehicles and controlled task splits. Our algorithm encodes the mapping between task splits and vehicles into a binary matrix and uses Genetic Quantum Algorithm to control the evolvement process. To convert the binary matrix solution into task loading schemes, we design a novel cost function and successfully convert task assignment problem to Transportation Problem which can be solved by Transportation Simplex Method. Our algorithm uses a simple nearest-neighborhood based heuristic to generate vehicle routes and adopts a local search method tailored for SDVRP to improve solution quality. The experimental results show that our algorithm splits few tasks and can obtain many solutions better than CVRP best-known in TSPLIB 95. Further analysis reveals that savings of SDVRP mostly come from CVRP’s failure to combine tasks geographically close into one route, when the number of vehicles are restricted to minimum.
Weijian Ma, Fei Gao 0001
CEC1