Jiantao Lin

dblp:334/6592 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0001-5705-2299ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
abstract
Diffusion models have achieved great success in generating 2D images. However, the quality and generaliz-ability of 3D content generation remain limited. State- of-the-art methods often require large-scale 3D assets for training, which are challenging to collect. In this work, we introduce Kiss3DGen (Keep It Simple and Straightforward in 3D Generation), an efficient framework for generating, editing, and enhancing 3D objects by repurposing a well-trained 2D image diffusion model for 3D generation. Specifically, we fine-tune a diffusion model to generate "3D Bundle Image", a tiled representation composed of multi-view images and their corresponding normal maps. The normal maps are then used to reconstruct a 3D mesh, and the multi-view images provide texture mapping, resulting in a complete 3D model. This simple method effectively transforms the 3D generation problem into a 2D image generation task, maximizing the utilization of knowledge in pretrained diffusion models. Furthermore, we demonstrate that our Kiss3DGen model is compatible with various diffusion model techniques, enabling advanced features such as 3D editing, mesh and texture enhancement, etc. Through extensive experiments, we demonstrate the effectiveness of our approach, showcasing its ability to produce high-quality 3D models efficiently Project page: https://ltt-0.github.io/Kiss3dgen.github.io.
Jiantao Lin, Xin Yang 0020, Meixi Chen, Dongyu Yan, Leyi Wu, Xinli Xu, Lie Xu 0004, Shunsi Zhang, Ying-Cong Chen
CVPR1
2025 PRM: Photometric Stereo Based Large Reconstruction Model
abstract
We propose PRM, a novel photometric stereo based large reconstruction model to reconstruct high-quality meshes with fine-grained local details. Unlike previous large reconstruction models that prepare images under fixed and simple lighting as both input and supervision, PRM renders photometric stereo images by varying materials and lighting for the purposes, which not only improves the precise local details by providing rich photometric cues but also increases the model robustness to variations in the appearance of input images. To offer enhanced flexibility of images rendering, we incorporate a real-time physically-based rendering (PBR) method and mesh rasterization for online images rendering. Moreover, in employing an explicit mesh as our 3D representation, PRM ensures the application of differentiable PBR, which supports the utilization of multiple photometric supervisions and better models the specular color for high-quality geometry optimization. Our PRM leverages photometric stereo images to achieve high-quality reconstructions with fine-grained local details, even amidst sophisticated image appearances. Extensive experiments demonstrate that PRM significantly outperforms other models.
Wenhang Ge, Jiantao Lin, Guibao Shen, Tao Hu 0011, Xinli Xu, Ying-Cong Chen
ICCV2
2025 Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural Rectification
Guibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge, Chaozhe Zhang, Xin Tao 0001, Di Zhang 0026, Pengfei Wan 0001, Guangyong Chen, Yijun Li 0001, Ying-Cong Chen
ICCV3
2025 FlexGen: Flexible Multi-View Generation from Text and Image Inputs
abstract
In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware text annotations. We utilize the strong reasoning capabilities of GPT-4V to generate 3D-aware text annotations. By analyzing four orthogonal views of an object arranged as tiled multi-view images, GPT-4V can produce text annotations that include 3D-aware information with spatial relationship. By integrating the control signal with proposed adaptive dual-control module, our model can generate multi-view images that correspond to the specified text. FlexGen supports multiple controllable capabilities, allowing users to modify text prompts to generate reasonable and corresponding unseen parts. Additionally, users can influence attributes such as appearance and material properties, including metallic and roughness. Extensive experiments demonstrate that our approach offers enhanced multiple controllability, marking a significant advancement over existing multi-view diffusion models. This work has substantial implications for fields requiring rapid and flexible 3D content creation, including game development, animation, and virtual reality. Project page: https://xxu068.github.io/flexgen.github.io/.
Xinli Xu, Wenhang Ge, Jiantao Lin, Lie Xu 0004, HanFeng Zhao, Shunsi Zhang, Ying-Cong Chen
ICCV3
2025 MPPR: Memory-Prior-based Prompt Refinement in Continuous Space for Advanced Text-to-Image Generation
abstract
Refining user-provided natural language prompts allows users to more easily obtain their desired outputs in text-to-image generation. Existing automatic prompt refinement methods predominantly take discrete, human-engineered high-quality prompts as the final optimization target. However, human-engineered prompts are based on human intuition and derived through limited interaction with generative models, which fails to bridge the gap between human preferences and model preferences. Additionally, this discrete optimization target limits the information capacity of the conditional inputs fed to the generative model, leading to suboptimal outcomes. Therefore, we propose an end-to-end prompt optimization method that interacts directly with generative models, eliminating the need for human involvement. The optimization process takes high-quality images as the target and uses the internal states of generative models as optimization signals. This optimizes prompts in a way that aligns more naturally with the model's generation process and producing continuous representations as the final refined prompt. We also introduce a memory module to store common features of high-quality prompts as prior knowledge to guide optimization in continuous space, enabling it to be more efficient. This memory-p rior-based p rompt r efinement in continuous space (MPPR) not only bridges the gap between human preferences and model preferences, but also resolves the issue of insufficient information in the inputs provided to the generative model. Extensive experiments show that our method achieves better performance compared to the state-of-the-art baselines.
Zhibing Zhang, Jiantao Lin, Cangqi Zhou
ACM Multimedia2
2025 ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
abstract
With the rapid advancement of generative models, general-purpose generation has gained increasing attention as a promising approach to unify diverse tasks across modalities within a single system. Despite this progress, existing open-source frameworks often remain fragile and struggle to support complex real-world applications due to the lack of structured workflow planning and execution-level feedback. To address these limitations, we present ComfyMind, a collaborative AI system designed to enable robust and scalable general-purpose generation, built on the ComfyUI platform. ComfyMind introduces two core innovations: Semantic Workflow Interface (SWI) that abstracts low-level node graphs into callable functional modules described in natural language, enabling high-level composition and reducing structural errors; Search Tree Planning mechanism with localized feedback execution, which models generation as a hierarchical decision process and allows adaptive correction at each stage. Together, these components improve the stability and flexibility of complex generative workflows. We evaluate ComfyMind on three public benchmarks: ComfyBench, GenEval, and Reason-Edit, which span generation, editing, and reasoning tasks. Results show that ComfyMind consistently outperforms existing open-source baselines and achieves performance comparable to GPT-Image-1. ComfyMind paves a promising path for the development of open-source general-purpose generative AI systems.
Litao Guo, Xinli Xu, Luozhou Wang, Jiantao Lin, Jinsong Zhou, Bolan Su, Ying-Cong Chen
NeurIPS4
2024 LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching
abstract
The recent advancements in text-to-3D generation mark a significant milestone in generative models, unlocking new possibilities for creating imaginative 3D assets across var-ious real-world scenarios. While recent advancements in text-to-3D generation have shown promise, they often fall short in rendering detailed and high-quality 3D models. This problem is especially prevalent as many methods base themselves on Score Distillation Sampling (SDS). This paper identifies a notable deficiency in SDS, that it brings inconsistent and low-quality updating direction for the 3D model, causing the over-smoothing effect. To address this, we propose a novel approach called Interval Score Matching (ISM). ISM employs deterministic diffusing trajectories and utilizes interval-based score matching to counteract over-smoothing. Furthermore, we incorporate 3D Gaussian Splatting into our text-to-3D generation pipeline. Extensive experiments show that our model largely outperforms the state-of-the-art in quality and training efficiency. Our code is available at: EnVision-Research/LucidDreamer
Yixun Liang, Xin Yang 0020, Jiantao Lin, Xiaogang Xu 0002, Ying-Cong Chen
CVPR3
2024 Graph Representation and Prototype Learning for webly supervised fine-grained image recognition
Jiantao Lin, Tianshui Chen, Ying-Cong Chen, Zhijing Yang, Yuefang Gao
Pattern Recognit. Lett.1
2022 Learning Consistent Global-Local Representation for Cross-Domain Facial Expression Recognition
abstract
Domain shift is one of the knotty problems that seriously restricts the accuracy of cross-domain facial expression recognition. Most existing works mainly focus on learning domain-invariant features by global feature adaption, and little works are conducted using the local features which are more transferable across different domains. In this paper, a consistent global-local feature and semantic learning framework is proposed which can learn domain-invariant global and local feature representation, and generate pseudo labels to facilitate cross-domain facial expression recognition. Specifically, the proposed method first simultaneously learns the domain-invariant global and local features via separately adversarial global and local learning. Once those features are acquired, a global and local semantic consistency is introduced to help generate pseudo labels for unlabeled data of the target dataset. By performing such strategy, more efficiency pseudo labels with high accuracy are produced due to the information diversity in global-local features and do without the image transformation. We conduct extensive experiments and analyses on several public datasets to demonstrate the effectiveness of the proposed model.
Yuefang Gao, Jiantao Lin, Tianshui Chen
ICPR3