EDBT 2026 Demo / reviewers in the wild / expert
Yunzhong Lou
dblp:358/8138
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0001-2386-0017ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language ModelsabstractWe present ReCAD, a reinforcement learning (RL) framework that bootstraps pretrained large models (PLMs) to generate precise parametric computer-aided design (CAD) models from multimodal inputs by leveraging their inherent generative capabilities. With just access to simple functional interfaces (e.g., point coordinates), our approach enables the emergence of complex CAD operations (e.g., pattern replication and mirror). This stands in contrast to previous methods, which typically rely on knowledge injected through supervised fine-tuning (SFT), offer limited support for editability, and fail to exploit the strong generative priors of PLMs. Specifically, the ReCAD framework begins by fine-tuning vision-language models (VLMs) to equip them with basic CAD model generation capabilities, where we rewrite CAD scripts into parameterized code that is leveraged to generate accurate textual descriptions for supervision. Then, we propose a novel RL strategy that incorporates parameterized code as guidance to enhance the model’s reasoning on challenging questions. Furthermore, we employ a hierarchical primitive learning process to progressively teach structured and compositional skills under a unified reward function that ensures both geometric accuracy and semantic fidelity. ReCAD sets a new state-of-the-art in both text-to-CAD and image-to-CAD tasks, significantly improving geometric accuracy across in-distribution and out-of-distribution settings. In the image-to-CAD task, for instance, it reduces the mean Chamfer Distance from 73.47 to 29.61 (in-distribution) and from 272.06 to 80.23 (out-of-distribution), outperforming existing baselines by a substantial margin. Yusheng Luo, Yunzhong Lou |
AAAI | 3 |
| 2026 | BRep-H: A Data-Centric Framework for Structured Language-Driven BRep Representation and Reconstruction
Yunzhong Lou, Yusheng Luo |
DASFAA (4) | 1 |
| 2025 | Mamba-CAD: State Space Model for 3D Computer-Aided Design Generative ModelingabstractComputer-Aided Design (CAD) generative modeling has a strong and long-term application in the industry. Recently, the parametric CAD sequence as the design logic of an object has been widely mined by sequence models. However, the industrial CAD models, especially in component objects, are fine-grained and complex, requiring a longer parametric CAD sequence to define. To address the problem, we introduce Mamba-CAD, a self-supervised generative modeling for complex CAD models in the industry, which can model on a longer parametric CAD sequence. Specifically, we first design an encoder-decoder framework based on a Mamba architecture and pair it with a CAD reconstruction task for pre-training to model the latent representation of CAD models; and then we utilize the learned representation to guide a generative adversarial network to produce the fake representation of CAD models, which would be finally recovered into parametric CAD sequences via the decoder of Mamba-CAD. To train Mamba-CAD, we further create a new dataset consisting of 77,078 CAD models with longer parametric CAD sequences. Comprehensive experiments are conducted to demonstrate the effectiveness of our model under various evaluation metrics, especially in the generation length of valid parametric CAD sequences. Yunzhong Lou |
AAAI | 2 |
| 2025 | CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model GenerationabstractRecently, Large Language Models (LLMs) have achieved significant success, prompting increased interest in expanding their generative capabilities beyond general text into domain-specific areas. This study investigates the generation of parametric sequences for computer-aided design (CAD) models using LLMs. This endeavor represents an initial step towards creating parametric 3D shapes with LLMs, as CAD model parameters directly correlate with shapes in three-dimensional space. Despite the formidable generative capacities of LLMs, this task remains challenging, as these models neither encounter parametric sequences during their pretraining phase nor possess direct awareness of 3D structures. To address this, we present CAD-Llama, a framework designed to enhance pretrained LLMs for generating parametric 3D CAD models. Specifically, we develop a hierarchical annotation pipeline and a code-like format to translate parametric 3D CAD command sequences into Structured Parametric CAD Code (SPCC), incorporating hierarchical semantic descriptions. Furthermore, we propose an adaptive pretraining approach utilizing SPCC, followed by an instruction tuning process aligned with CAD-specific guidelines. This methodology aims to equip LLMs with the spatial knowledge inherent in parametric sequences. Experimental results demonstrate that our framework significantly outperforms prior autoregressive methods and existing LLM baselines. Weijian Ma, Yunzhong Lou, Guichun Zhou |
CVPR | 4 |
| 2025 | Floorplan-Diffusion: Automatic Floor Plan Generation via Pre-trained Large Latent Diffusion ModelabstractAutomatic floor plan generation is a long-standing goal in the field of engineering and architectural design. It has significant value in practical applications. Although extensive research over several decades has introduced numerous innovative methods, the issue remains a significant challenge. In this study, we introduce a novel approach to floor plan image generation, leveraging the capabilities of the large pre-trained Latent Diffusion Model (LDM). Through improvements in the architecture of the pre-trained Latent Diffusion Model (LDM) and subsequent fine-tuning, we achieve an effective multimodal conditional floor plan image generation with a reduced amount of training data. Specifically, we propose a multi-head self-attention graph convolution-based layout embedding and a novel layout fusion contrast learning module integrating to the pre-trained LDM, which not only enhances the constraint generation under layout instructions, but also preserves the balance and consistency of the floor plan elements, such as different kinds of rooms and walls, as specified in the textual instructions. The experiments were conducted on two commonly used data sets, RPLAN and LIFULL. The experimental results show that our method outperforms the traditional methods and SOTA work and achieves the best generation results. Minyang Xu, Yunzhong Lou, Xiang Gao 0039 |
ICMR | 2 |
| 2024 | Draw Step by Step: Reconstructing CAD Construction Sequences from Point Clouds via Multimodal DiffusionabstractReconstructing CAD construction sequences from raw 3D geometry serves as an interface between real-world objects and digital designs. In this paper, we propose CAD-Diffuser, a multimodal diffusion scheme aiming at integrating top-down design paradigm into generative reconstruction. In particular, we unify CAD point clouds and CAD construction sequences at the token level, guiding our proposed multimodal diffusion strategy to understand and link between the geometry and the design intent concentrated in construction sequences. Leveraging the strong decoding abilities of language models, the forward process is modeled as a random walk between the original token and the [MASK] token, while the reverse process naturally fits the masked token modeling scheme. A volume-based noise schedule is designed to encourage outline-first generation, decomposing the top-down design methodology into a machine-understandable procedure. For tokenizing CAD data of multiple modalities, we introduce a tokenizer with a self-supervised face segmentation task to compress local and global geometric information for CAD point clouds, and the CAD construction sequence is transformed into a primitive token string. Experimental results show that our CAD-Diffuser can perceive geometric details and the results are more likely to be reused by human designers. Weijian Ma, Shuaiqi Chen, Yunzhong Lou |
CVPR | 3 |
| 2024 | CF-CAD: A Contrastive Fusion Network For 3D Computer-Aided Design Generative Modeling
Yunzhong Lou |
DASFAA (3) | 3 |
| 2024 | Parametric CAD Primitive Retrieval via Multi-Modal Fusion and Deep HashingabstractIn the rapidly evolving field of manufacturing industry, product designers need to be able to quickly and accurately retrieve and reference existing Computer-Aided Design (CAD) primitives to increase efficiency and foster innovation. This paper introduces an innovative deep hashing-based parametric CAD primitive retrieval model DH-CAD, which employs a deep learning framework to capture the profound multi-modal features of parametric CAD primitive (command sequences and point clouds), and utilizes hashing to effectively encode these features into compact binary hash codes, significantly improving retrieval efficiency and accuracy. DH-CAD optimizes the original sequence encodings of Transformer to effectively learn the intricate relationships between CAD tool's command sequences and the corresponding point clouds of CAD primitive. Furthermore, it considers the dependencies between command types and parameter values, combines the output of the command type decoder with the prediction of parameter values to further refine the generated parameters, and employs a binarization network to generate hash codes in an end-to-end manner. Experimental results clearly demonstrate that, compared to traditional feature-matching methods, DH-CAD achieves the state-of-the-art performance on multiple evaluation metrics. Particularly in terms of retrieval speed and accuracy, our proposed model not only retrieves sequences that are most relevant to the query rapidly, but also ensures the high precision of the results, showcasing its efficiency and practicality in supporting CAD design processes. Minyang Xu, Yunzhong Lou, Weijian Ma |
ICMR | 2 |
| 2024 | CAD Translator: An Effective Drive for Text to 3D Parametric Computer-Aided Design Generative ModelingabstractComputer-Aided Design (CAD) generative modeling is widely applicable in the fields of industrial engineering. Recently, text-to-3D generation has shown rapid progress in point clouds, mesh, and other non-parametric representations. On the contrary, text to 3D parametric CAD generative modeling is a more appealing task in industry but has not been well explored. The parametric CAD model means the product shape can be defined by using the command sequences of CAD tools. To investigate this, we design an encoder-decoder framework, namely CAD Translator, for incorporating the embedding of parametric CAD sequences into texts appropriately with only one-stage training. We first align texts and parametric CAD sequences via a Cascading Contrastive Strategy in the latent space, and then we propose CT-Mix to conduct the random mask operation on their embeddings separately to further get a fusion embedding via the linear interpolation. This can strengthen the connection between texts and parametric CAD sequences effectively. To train CAD Translator, we build a Text2CAD dataset with the help of Large Multimodal Model (LMM) and conduct thorough experiments to demonstrate the effectiveness of our method. Yunzhong Lou |
ACM Multimedia | 3 |
| 2023 | BRep-BERT: Pre-training Boundary Representation BERT with Sub-graph Node Contrastive LearningabstractObtaining effective entity feature representations is crucial in the field of Boundary Representation (B-Rep), a key parametric representation method in Computer-Aided Design (CAD). However, the lack of labeled large-scale database and the scarcity of task-specific label sets pose significant challenges. To address these problems, we propose an innovative unsupervised neural network approach called BRep-BERT, which extends the concept of BERT to the B-Rep domain. Specifically, we utilize Graph Neural Network (GNN) Tokenizer to generate discrete entity labels with geometric and structural semantic information. We construct new entity representation sequences based on the structural relationships and pre-train the model through the Masked Entity Modeling (MEM) task. To address the attention sparsity issue in large-scale geometric models, we incorporate graph structure information and learnable relative position encoding into the attention module to optimize feature updates. Additionally, we employ geometric sub-graphs and multi-level contrastive learning techniques to enhance the model's ability to learn regional features. Comparisons with previous methods demonstrate that BRep-BERT achieves the state-of-the-art performance on both full-data training and few-shot learning tasks across multiple B-Rep datasets. Particularly, BRep-BERT outperforms previous methods significantly in the few-shot learning scenarios. Comprehensive experiments demonstrate the substantial advantages and potential of BRep-BERT in handling B-Rep data representation. Code will be released at https://github.com/louyz1026/Brep_Bert. Yunzhong Lou |
CIKM | 1 |