EDBT 2026 Demo / reviewers in the wild / expert
Tianrun Chen
dblp:317/5235
· DBLP profile ↗
31ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0003-0177-0157ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Magic3DSketch: Create colorful 3D models from sketch-based 3D modeling guided by text and language-image pre-training
Ying Zang, Yidong Han, Chaotao Ding, Jianqi Zhang, Tianrun Chen |
Neurocomputing | 5 |
| 2026 | SNH-SLAM: Implicit Dense SLAM Based on Scalable Neural-Hash RepresentationabstractWe present SNH-SLAM, a novel expandable dense neural simultaneous localization and mapping (SLAM) method that constructs a neural field in real-time based on run-time observation. To reach this challenging goal without any scene prior, we utilize instant depth supervision to drive the extension of planar convex hulls, where a single hash table maintains multi-level feature units embedded in the planar convex hulls. This design facilitates high-fidelity, hole-free, and low-memory map reconstruction while adding only a tiny time burden to the training process. Our approach performs mapping by minimizing both RGBD-based re-rendering loss and Truncated Signed Distance Field (TSDF) loss. In addition, for camera tracking, our optimization strategy allows SNH-SLAM to converge faster on the pose estimation and maintain robustness. We evaluate our method on common benchmarks and compare it with existing dense neural RGB-D SLAM methods. The evaluation results show the competitiveness of the SNH-SLAM in tracking accuracy, reconstruction quality, memory usage, and frame processing speed. Project page:https://xiaoshumiao123.github.io. Zhenyu Wen, Zhanshuo Dong, Haoran Duan 0001, Tianrun Chen, Zhen Hong |
IEEE Trans. Multim. | 5 |
| 2026 | Let Human Sketches Help: Empowering the Challenging Image Segmentation Task With Freehand SketchesabstractSketches, with their expressive potential, enable humans to convey the essence of an object through a rough contour. This work leverages expressive power for the first time to improve segmentation performance in challenging tasks such as camouflaged object detection (COD). We propose a sketch guided interactive segmentation framework that allows users to intuitively annotate objects with freehand sketches rather than relying on traditional bounding boxes or points commonly used in models such as the SAM. Our method introduces dedicated network architectural enhancements and a novel sketch augmentation strategy to fully exploit sketch input, leading to significant accuracy gains compared with text- or box-based annotations. Furthermore, our model's output can directly train other neural networks, achieving performance comparable to that of pixel-level annotations while reducing the annotation time by up to 120× and thereby lowering the barrier for large-scale dataset creation and model training. To support future research, werelease KOSCamo+, the first freehand sketch dataset for COD, along with code and a labeling tool. These contributions open promising avenues for expanding sketch-based interaction to broader segmentation tasks and exploring multimodal annotation strategies that combine sketches, text, and other lightweight user inputs. Ying Zang, Runlong Cao, Jianqi Zhang, Yidong Han, Ziyue Cao, Didi Zhu, Zejian Li, Lanyun Zhu, Deyi Ji, Tianrun Chen |
IEEE Trans. Multim. | 11 |
| 2026 | From Sketch to Reality: Enabling High-Quality, Cross-Category 3D Model Generation From Free-Hand Sketches With Minimal DataabstractThis paper presents a novel approach for generating high-quality, cross-category 3D models from free-hand sketches with limited training data. We propose the first semi-supervised learning method to our knowledge for sketch-to-3D model conversion. Innovatively, we design a coarse-to-fine pipeline to perform the semi-supervised learning in the coarse stage and train a diffusion-based refiner to get a high-resolution 3D model. We designed a sketch-augmentation method for semi-supervised learning and integrated priors such as CLIP loss, shape prototypes, and adversarial loss to help generate high-quality results even with abstract and imprecise sketches. We also introduce an innovative procedural 3D generation method based on CAD code, which helps pre-train part of the network before fine-tuning with limited real data. Our approach, coupled with a specifically designed curriculum learning, allows us to generate high-quality 3D models across multiple categories with as few as 300 sketch-3D model pairs, marking a significant advancement over previous single-category approaches. In addition, we introduce the KO2D dataset, the largest collection of hand-drawn sketch-3D pairs to support further research in this area. As sketches are a far more intuitive and detailed way for users to express their unique ideas, we believe that this paper can move us closer to democratizing 3D content creation, enabling anyone to transform their ideas into 3D models effortlessly. Ying Zang, Chunan Yu, Jing Li 0145, Shengyuan Zhang, Lanyun Zhu, Chaotao Ding, Renjun Xu, Tianrun Chen |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2026 | DeepSketch2Wear: democratizing 3D garment creation via freehand sketches and text
Jianqi Zhang, Chaotao Ding, Runlong Cao, Lanyun Zhu, Ying Zang, Tianrun Chen |
Vis. Comput. | 8 |
| 2025 | CADCrafter: Generating Computer-Aided Design Models from Unconstrained ImagesabstractCreating CAD digital twins from the physical world is crucial for manufacturing, design, and simulation. However, current methods typically rely on costly 3D scanning with labor-intensive post-processing. To provide a user-friendly design process, we explore the problem of reverse engineering from unconstrained real-world CAD images that can be easily captured by users of all experiences. However, the scarcity of real-world CAD data poses challenges in directly training such models. To tackle these challenges, we propose CADCrafter, an image-to-parametric CAD model generation framework that trains solely on synthetic textureless CAD data while testing on real-world images. To bridge the significant representation disparity between images and parametric CAD models, we introduce a geometry encoder to accurately capture diverse geometric features. Moreover, the texture-invariant properties of the geometric features can also facilitate the generalization to real-world scenarios. Since compiling CAD parameter sequences into explicit CAD models is a non-differentiable process, the network training inherently lacks explicit geometric supervision. To impose geometric validity constraints, we employ direct preference optimization (DPO) to fine-tune our model with the automatic code checker feedback on CAD sequence quality. Furthermore, we collected a real-world dataset, comprised of multi-view images and corresponding CAD command sequence pairs, to evaluate our method. Experimental results demonstrate that our approach can robustly handle real unconstrained CAD images, and even generalize to unseen general objects. Jiacheng Wei, Tianrun Chen, Chi Zhang 0007, Shangzhan Zhang, Bingchen Yang, Chuan-Sheng Foo, Guosheng Lin, Qixing Huang, Fayao Liu |
CVPR | 3 |
| 2025 | POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning SegmentationabstractExisting LVLM-based reasoning segmentation methods often suffer from imprecise segmentation results and hallucinations in their text responses. This paper introduces POPEN, a novel framework designed to address these issues and achieve improved results. POPEN includes a preference-based optimization method to finetune the LVLM, aligning it more closely with human preferences and thereby generating better text responses and segmentation results. Additionally, POPEN introduces a preference-based ensemble method for inference, which integrates multiple outputs from the LVLM using a preference-score-based attention mechanism for refinement. To better adapt to the segmentation task, we incorporate several task-specific designs in our POPEN framework, including a new approach for collecting segmentation preference data with a curriculum learning mechanism, and a novel preference optimization loss to refine the segmentation capability of the LVLM. Experiments demonstrate that our method achieves state-of-the-art performance in reasoning segmentation, exhibiting minimal hallucination in text responses and the highest segmentation accuracy compared to previous advanced methods like LISA and PixelLM. Project page is here. Lanyun Zhu, Tianrun Chen, Qianxiong Xu, Xuanyi Liu, Deyi Ji, De Wen Soh, Jun Liu 0036 |
CVPR | 2 |
| 2025 | Distilling Diffusion Models to Efficient 3D LiDAR Scene CompletionabstractDiffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the practical application of diffusion-based scene completion models since autonomous vehicles require an efficient perception of surrounding environments. This paper proposes a novel distillation method tailored for 3D Li- DAR scene completion models, dubbed ScoreLiDAR, which achieves efficient yet high-quality scene completion. Score- LiDAR enables the distilled model to sample in significantly fewer steps after distillation. To improve completion quality, we also introduce a novel Structural Loss, which encourages the distilled model to capture the geometric structure of the 3D LiDAR scene. The loss contains a scene-wise term constraining the holistic structure and a point-wise term constraining the key landmark points and their relative configuration. Extensive experiments demonstrate that ScoreLiDAR significantly accelerates the completion time from 30.55 to 5.37 seconds per frame (>5x) on SemanticKITTI and achieves superior performance compared to state-of-the-art 3D LiDAR scene completion models. Our model and code are publicly available on https://github.com/happyw1nd/ScoreLiDAR. Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Tianrun Chen, Anyang Wei, Perry Pengyun Gu, Lingyun Sun |
ICCV | 7 |
| 2025 | CPCF: A Cross-Prompt Contrastive Framework for Referring Multimodal Large Language ModelsabstractReferring MLLMs extend conventional multimodal large language models by allowing them to receive referring visual prompts and generate responses tailored to the indicated regions. However, these models often suffer from suboptimal performance due to incorrect responses tailored to misleading areas adjacent to or similar to the target region. This work introduces CPCF, a novel framework to address this issue and achieve superior results. CPCF contrasts outputs generated from the indicated visual prompt with those from contrastive prompts sampled from misleading regions, effectively suppressing the influence of erroneous information outside the target region on response generation. To further enhance the effectiveness and efficiency of our framework, several novel designs are proposed, including a prompt extraction network to automatically identify suitable contrastive prompts, a self-training method that leverages unlabeled data to improve training quality, and a distillation approach to reduce the additional computational overhead associated with contrastive decoding. Incorporating these novel designs, CPCF achieves state-of-the-art performance, as demonstrated by extensive experiments across multiple benchmarks. Project page: https://lanyunzhu.site/CPCF/ Lanyun Zhu, Deyi Ji, Tianrun Chen, De Wen Soh, Jun Liu 0036 |
ICML | 3 |
| 2025 | Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal RetrievalabstractThe success of DeepSeek-R1 demonstrates the immense potential of using reinforcement learning (RL) to enhance LLMs' reasoning capabilities. This paper introduces Retrv-R1, the first R1-style MLLM specifically designed for multimodal universal retrieval, achieving higher performance by employing step-by-step reasoning to produce more accurate retrieval results. We find that directly applying the methods of DeepSeek-R1 to retrieval tasks is not feasible, mainly due to (1) the high computational cost caused by the large token consumption required for multiple candidates with reasoning processes, and (2) the instability and suboptimal results when directly applying RL to train for retrieval tasks. To address these issues, Retrv-R1 introduces an information compression module with a details inspection mechanism, which enhances computational efficiency by reducing the number of tokens while ensuring that critical information for challenging candidates is preserved. Additionally, a new training paradigm is proposed, including an activation stage using a retrieval-tailored synthetic CoT dataset for more effective optimization, followed by RL with a novel curriculum reward to improve both performance and efficiency. Incorporating these novel designs, Retrv-R1 achieves SOTA performance, high efficiency, and strong generalization ability, as demonstrated by extensive experiments across multiple benchmarks and tasks. Lanyun Zhu, Deyi Ji, Tianrun Chen, Shiqi Wang 0001 |
NeurIPS | 3 |
| 2025 | Dyn-E: Local appearance editing of dynamic neural radiance fields
Yinji ShenTu, Shangzhan Zhang, Qing Shuai, Tianrun Chen, Sida Peng, Xiaowei Zhou 0001 |
Comput. Graph. | 5 |
| 2025 | PanopticNeRF-360: Panoramic 3D-to-2D Label Transfer in Urban ScenesabstractTraining perception systems for self-driving cars requires substantial 2D annotations that are labor-intensive to manual label. While existing datasets provide rich annotations on pre-recorded sequences, they fall short in labeling rarely encountered viewpoints, potentially hampering the generalization ability for perception models. In this paper, we present PanopticNeRF-360, a novel approach that combines coarse 3D annotations with noisy 2D semantic cues to generate high-quality panoptic labels and images from any viewpoint. Our key insight lies in exploiting the complementarity of 3D and 2D priors to mutually enhance geometry and semantics. Specifically, we propose to leverage coarse 3D bounding primitives and noisy 2D semantic and instance predictions to guide geometry optimization, by encouraging predicted labels to match panoptic pseudo ground truth. Simultaneously, the improved geometry assists in filtering 3D&2D annotation noise by fusing semantics in 3D space via a learned semantic field. To further enhance appearance, we combine MLP and hash grids to yield hybrid scene features, striking a balance between high-frequency appearance and contiguous semantics. Our experiments demonstrate PanopticNeRF-360's state-of-the-art performance over label transfer methods on the challenging urban scenes of the KITTI-360 dataset. Moreover, PanopticNeRF-360 enables omnidirectional rendering of high-fidelity, multi-view and spatiotemporally consistent appearance, semantic and instance labels. Shangzhan Zhang, Tianrun Chen, Yichong Lu, Xiaowei Zhou 0001, Andreas Geiger 0001, Yiyi Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | LLaFS++: Few-Shot Image Segmentation With Large Language ModelsabstractDespite the rapid advancements in few-shot segmentation (FSS), most of existing methods in this domain are hampered by their reliance on the limited and biased information from only a small number of labeled samples. This limitation inherently restricts their capability to achieve sufficiently high levels of performance. To address this issue, this paper proposes a pioneering framework named LLaFS++, which, for the first time, applies large language models (LLMs) into FSS and achieves notable success. LLaFS++ leverages the extensive prior knowledge embedded by LLMs to guide the segmentation process, effectively compensating for the limited information contained in the few-shot labeled samples and thereby achieving superior results. To enhance the effectiveness of the text-based LLMs in FSS scenarios, we present several innovative and task-specific designs within the LLaFS++ framework. Specifically, we introduce an input instruction that allows the LLM to directly produce segmentation results represented as polygons, and propose a region-attribute corresponding table to simulate the human visual system and provide multi-modal guidance. We also synthesize pseudo samples and use curriculum learning for pretraining to augment data and achieve better optimization, and propose a novel inference method to mitigate potential oversegmentation hallucinations caused by the regional guidance information. Incorporating these designs, LLaFS++ constitutes an effective framework that achieves state-of-the-art results on multiple datasets including PASCAL-$5^{i}$5i, COCO-$20^{i}$20i, and FSS-1000. Our superior performance showcases the remarkable potential of applying LLMs to process few-shot vision tasks. Lanyun Zhu, Tianrun Chen, Deyi Ji, Peng Xu 0023, Jieping Ye, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Replay Master: Automatic Sample Selection and Effective Memory Utilization for Continual Semantic SegmentationabstractContinual Semantic Segmentation (CSS) extends static semantic segmentation by incrementally introducing new classes for training. To alleviate the catastrophic forgetting issue in this task, replay methods can be adopted, constructing a memory buffer that stores a small number of samples from previous classes for future replay. However, existing replay approaches in CSS often lack a thorough exploration of two critical issues: how to find the most suitable memory samples and how to utilize them for replay more effectively. Common strategies either randomly select samples or rely on hand-crafted, single-factor-driven methods that are hard to be optimal, and often employ conventional training techniques for replay that do not account for class imbalance problem resulting from limited memory capacity. In this work, we tackle these challenges by introducing a novel memory sample selection method that leverages a reinforcement learning framework with innovative state representations and a dual-stage action scheme to automatically learn a selection policy. Additionally, we propose an expert mechanism and a dual-phase training method to address the class imbalance issue, thereby enhancing the effectiveness of replay training by making better use of memory samples. Incorporating the proposed automatic sample selection and effective memory utilization methods, we develop a novel and effective replay-based pipeline for CSS. Our extensive experiments on Pascal VOC 2012 and ADE20 K datasets demonstrate the effectiveness of our approach, which achieves state-of-the-art (SOTA) performance and outperforms previous advanced methods significantly. Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, De Wen Soh, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Img2CAD: Conditioned 3-D CAD Model Generation From Single Image With Structured Visual GeometryabstractIn this article, we propose Img2CAD, the first approach to our knowledge that uses 2-D image inputs to generate computer-aided design (CAD) models with editable parameters. Unlike existing artificial intelligence (AI) methods for 3-D model generation using text or image inputs often rely on mesh-based representations, which are incompatible with CAD tools and lack editability and fine control, Img2CAD enables seamless integration between AI-based 3-D reconstruction and CAD software. We have identified an innovative intermediate representation called structured visual geometry, characterized by vectorized wireframes extracted from objects. This representation significantly enhances the performance of generating conditioned CAD models. In addition, we introduce two new datasets to further support research in this area:a big cad model dataset (ABC)-mono, the largest known dataset comprising over 200 000 3-D CAD models with rendered images, andKOCAD, the first dataset featuring real-world captured objects alongside their ground truth CAD models, supporting further research in conditioned CAD model generation. Tianrun Chen, Chunan Yu, Yuanqi Hu, Jing Li 0145, Tao Xu 0048, Runlong Cao, Lanyun Zhu, Ying Zang, Yong Zhang 0030, Zejian Li, Lingyun Sun |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Not Every Patch is Needed: Toward a More Efficient and Effective Backbone for Video-Based Person Re-IdentificationabstractThis paper proposes a new effective and efficient plug-and-play backbone for video-based person re-identification (ReID). Conventional video-based ReID methods typically use CNN or transformer backbones to extract deep features for every position in every sampled video frame. Here, we argue that this exhaustive feature extraction could be unnecessary, since we find that different frames in a ReID video often exhibit small differences and contain many similar regions due to the relatively slight movements of human beings. Inspired by this, a more selective, efficient paradigm is explored in this paper. Specifically, we introduce a patch selection mechanism to reduce computational cost by choosing only the crucial and non-repetitive patches for feature extraction. Additionally, we present a novel network structure that generates and utilizes pseudo frame global context to address the issue of incomplete views resulting from sparse inputs. By incorporating these new designs, our backbone can achieve both high performance and low computational cost. Extensive experiments on multiple datasets show that our approach reduces the computational cost by 74% compared to ViT-B and 28% compared to ResNet50, while the accuracy is on par with ViT-B and outperforms ResNet50 significantly. Lanyun Zhu, Tianrun Chen, Deyi Ji, Jieping Ye, Jun Liu 0036 |
IEEE Trans. Image Process. | 2 |
| 2025 | From Air to Wear: Personalized 3D Digital Fashion With AR/VR Immersive 3D SketchingabstractIn the era of immersive consumer electronics, such as AR/VR headsets and smart devices, people increasingly seek ways to express their identity through virtual fashion. However, existing 3D garment design tools remain inaccessible to everyday users due to steep technical barriers and limited data. In this work, we introduce a 3D sketch-driven 3D garment generation framework that empowers ordinary users - even those without design experience - to create high-quality digital clothing through simple 3D sketches in AR/VR environments. By combining a conditional diffusion model, a sketch encoder trained in a shared latent space, and an adaptive curriculum learning strategy, our system interprets imprecise, free-hand input and produces realistic, personalized garments. To address the scarcity of training data, we also introduce KO3DClothes, a new dataset of paired 3D garments and user-created sketches. Extensive experiments and user studies confirm that our method significantly outperforms existing baselines in both fidelity and usability, demonstrating its promise for democratized fashion design on next-generation consumer platforms. Ying Zang, Yuanqi Hu, Suhui Wang, Yuxia Xu, Chunan Yu, Lanyun Zhu, Deyi Ji, Tianrun Chen |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2024 | Rapid 3D Model Generation with Intuitive 3D InputabstractWith the emergence of AR/VR, 3D models are in tremendous demand. However, conventional 3D modeling with Computer-Aided Design software requires much expertise and is difficult for novice users. We find that AR/VR devices, in addition to serving as effective display mediums, can offer a promising potential as an intuitive 3D model creation tool, especially with the assistance of AI generative models. Here, we propose Deep3DVRSketch, the first 3D model generation network that inputs 3D VR sketches from novice users and generates highly consistent 3D models in multiple categories within seconds, irrespective of the users' drawing abilities. We also contribute KO3D+, the largest 3D sketch-shape dataset. Our method pre-trains a conditional diffusion model on quality 3D data, then fine-tunes an encoder to map 3D sketches onto the generator's manifold using an adaptive curriculum strategy for limited ground truths. In our experiment, our approach achieves state-of-the-art performance in both model quality and fidelity with real-world input from novice users, and users can even draw and obtain very detailed geometric structures. In our user study, users were able to complete the 3D modeling tasks over 10 times faster using our approach compared to conventional CAD software tools. We believe that our Deep3DVRSketch and KO3D+ dataset can offer a promising solution for future 3D modeling in metaverse era. Check the project page at http://research.kokoni3d.com/Deep3DVRSketch. Tianrun Chen, Chaotao Ding, Shangzhan Zhang, Chunan Yu, Ying Zang, Zejian Li, Sida Peng, Lingyun Sun |
CVPR | 1 |
| 2024 | LLaFS: When Large Language Models Meet Few-Shot SegmentationabstractThis paper proposes LLaFS, the first attempt to leverage large language models (LLMs) in few-shot segmentation. In contrast to the conventional few-shot segmentation methods that only rely on the limited and biased information from the annotated support images, LLaFS leverages the vast prior knowledge gained by LLM as an effective supplement and directly uses the LLM to segment images in a few-shot manner. To enable the text-based LLM to handle image-related tasks, we carefully design an input instruction that allows the LLM to produce segmentation results represented as polygons, and propose a region-attribute table to simulate the human visual mechanism and provide multi-modal guidance. We also synthesize pseudo samples and use curriculum learning for pre-training to augment data and achieve better optimization. LLaFS achieves state-of-the-art results on multiple datasets, showing the potential of using LLMs for few-shot computer vision tasks. Lanyun Zhu, Tianrun Chen, Deyi Ji, Jieping Ye, Jun Liu 0036 |
CVPR | 2 |
| 2024 | Addressing Background Context Bias in Few-Shot Segmentation Through Iterative ModulationabstractExisting few-shot segmentation methods usually extract foreground prototypes from support images to guide query image segmentation. However, different background contexts of support and query images can cause their foreground features to be misaligned. This phenomenon, known as background context bias, can hinder the effectiveness of support prototypes in guiding query image segmentation. In this work, we propose a novel framework with an it-erative structure to address this problem. In each iteration of the framework, we first generate a query prediction based on a support foreground feature. Next, we extract background context from the query image to modulate the support foreground feature, thus eliminating the foreground feature misalignment caused by the different backgrounds. After that, we design a confidence-biased attention to eliminate noise and cleanse information. By integrating these components through an iterative structure, we create a novel network that can leverage the synergies between different modules to improve their performance in a mutually reinforcing manner. Through these carefully designed components and structures, our network can effectively elimi-nate background context bias in few-shot segmentation, thus achieving outstanding performance. We conduct extensive experiments on the PASCAL-5iand COCO-20idatasets and achieve state-of-the-art (SOTA) results, which demonstrate the effectiveness of our approach. Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036 |
CVPR | 2 |
| 2024 | Deep3DSketch-im: rapid high-fidelity AI 3D model generation by single freehand sketchesabstractThe rise of artificial intelligence generated content (AIGC) has been remarkable in the language and image fields, but artificial intelligence (AI) generated three-dimensional (3D) models are still under-explored due to their complex nature and lack of training data. The conventional approach of creating 3D content through computer-aided design (CAD) is labor-intensive and requires expertise, making it challenging for novice users. To address this issue, we propose a sketch-based 3D modeling approach, Deep3DSketch-im, which uses a single freehand sketch for modeling. This is a challenging task due to the sparsity and ambiguity. Deep3DSketch-im uses a novel data representation called the signed distance field (SDF) to improve the sketch-to-3D model process by incorporating an implicit continuous field instead of voxel or points, and a specially designed neural network that can capture point and local features. Extensive experiments are conducted to demonstrate the effectiveness of the approach, achieving state-of-the-art (SOTA) performance on both synthetic and real datasets. Additionally, users show more satisfaction with results generated by Deep3DSketch-im, as reported in a user study. We believe that Deep3DSketch-im has the potential to revolutionize the process of 3D modeling by providing an intuitive and easy-to-use solution for novice users. Tianrun Chen, Runlong Cao, Zejian Li, Ying Zang, Lingyun Sun |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2024 | Reality3DSketch: Rapid 3D Modeling of Objects From Single Freehand SketchesabstractThe emerging trend of AR/VR places great demands on 3D content. However, most existing software requires expertise and is difficult for novice users to use. In this paper, we aim to create sketch-based modeling tools for user-friendly 3D modeling. We introduce Reality3DSketch with a novel application of an immersive 3D modeling experience, in which a user can capture the surrounding scene using a monocular RGB camera and can draw a single sketch of an object in the real-time reconstructed 3D scene. A 3D object is generated and placed in the desired location, enabled by our novel neural network with the input of a single sketch. Our neural network can predict the pose of a drawing and can turn a single sketch into a 3D model with view and structural awareness, which addresses the challenge of sparse sketch input and view ambiguity. We conducted extensive experiments synthetic and real-world datasets and achieved state-of-the-art (SOTA) results in both sketch view estimation and 3D modeling performance. According to our user study, our method of performing 3D modeling in a scene is$>$5x faster than conventional methods. Users are also more satisfied with the generated 3D model than the results of existing methods. Tianrun Chen, Chaotao Ding, Lanyun Zhu, Ying Zang, Yiyi Liao, Zejian Li, Lingyun Sun |
IEEE Trans. Multim. | 1 |
| 2023 | Painting 3D Nature in 2D: View Synthesis of Natural Scenes from a Single Semantic MaskabstractWe introduce a novel approach that takes a single semantic mask as input to synthesize multi-view consistent color images of natural scenes, trained with a collection of single images from the Internet. Prior works on 3D-aware image synthesis either require multi-view supervision or learning category-level prior for specific classes of objects, which are inapplicable to natural scenes. Our key idea to solve this challenge is to use a semantic field as the intermediate representation, which is easier to reconstruct from an input semantic mask and then translated to a radiance field with the assistance of off-the-shelf semantic image synthesis models. Experiments show that our method outperforms baseline methods and produces photorealistic and multi-view consistent videos of a variety of natural scenes. The project website is https://zju3dv.github.io/paintingnature/. Shangzhan Zhang, Sida Peng, Tianrun Chen, Linzhan Mou, Haotong Lin, Kaicheng Yu, Yiyi Liao, Xiaowei Zhou 0001 |
CVPR | 3 |
| 2023 | Continual Semantic Segmentation with Automatic Memory Sample SelectionabstractContinual Semantic Segmentation (CSS) extends static semantic segmentation by incrementally introducing new classes for training. To alleviate the catastrophic forgetting issue in CSS, a memory buffer that stores a small number of samples from the previous classes is constructed for replay. However, existing methods select the memory samples either randomly or based on a single-factor-driven handcrafted strategy, which has no guarantee to be optimal. In this work, we propose a novel memory sample selection mechanism that selects informative samples for effective replay in a fully automatic way by considering comprehensive factors including sample diversity and class performance. Our mechanism regards the selection operation as a decision-making process and learns an optimal selection policy that directly maximizes the validation performance on a reward set. To facilitate the selection decision, we design a novel state representation and a dual-stage action space. Our extensive experiments on Pascal-VOC 2012 and ADE 20K datasets demonstrate the effectiveness of our approach with state-of-the-art (SOTA) performance achieved, outperforming the second-place one by 12.54% for the 6-stage setting on Pascal-VOC 2012. Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036 |
CVPR | 2 |
| 2023 | Deep3DSketch: 3D Modeling from Free-Hand Sketches with View- and Structural-Aware Adversarial TrainingabstractThis work aims to investigate the problem of 3D modeling using single free-hand sketches, which is one of the most natural ways we humans express ideas. Although sketch-based 3D modeling can drastically make the 3D modeling process more accessible, the sparsity and ambiguity of sketches bring significant challenges for creating high-fidelity 3D models that reflect the creators’ ideas. In this work, we propose a view-and structural-aware deep learning approach, Deep3DSketch, which tackles the ambiguity and fully uses sparse information of sketches, emphasizing the structural information. Specifically, we introduced random pose sampling on both 3D shapes and 2D silhouettes, and an adversarial training scheme with an effective progressive discriminator to facilitate learning of the shape structures. Extensive experiments demonstrated the effectiveness of our approach, which outperforms existing methods – with state-of-the-art (SOTA) performance on both synthetic and real datasets. Tianrun Chen, Chenglong Fu 0003, Lanyun Zhu, Papa Mao, Ying Zang, Lingyun Sun |
ICASSP | 1 |
| 2023 | Learning Gabor Texture Features for Fine-Grained RecognitionabstractExtracting and using class-discriminative features is critical for fine-grained recognition. Existing works have demonstrated the possibility of applying deep CNNs to exploit features that distinguish similar classes. However, CNNs suffer from problems including frequency bias and loss of detailed local information, which restricts the performance of recognizing fine-grained categories. To address the challenge, we propose a novel texture branch as complimentary to the CNN branch for feature extraction. We innovatively utilize Gabor filters as a powerful extractor to exploit texture features, motivated by the capability of Gabor filters in effectively capturing multi-frequency features and detailed local information. We implement several designs to enhance the effectiveness of Gabor filters, including imposing constraints on parameter values and developing a learning method to determine the optimal parameters. Moreover, we introduce a statistical feature extractor to utilize informative statistical information from the signals captured by Gabor filters, and a gate selection mechanism to enable efficient computation by only considering qualified regions as input for texture extraction. Through the integration of features from the Gabor-filter-based texture branch and CNN-based semantic branch, we achieve comprehensive information extraction. We demonstrate the efficacy of our method on multiple datasets, including CUB-200-2011, NA-bird, Stanford Dogs, and GTOS-mobile. State-of-the-art performance is achieved using our approach. Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036 |
ICCV | 2 |
| 2023 | Deep3DSketch+: Rapid 3D Modeling from Single Free-Hand Sketches
Tianrun Chen, Chenglong Fu 0003, Ying Zang, Lanyun Zhu, Papa Mao, Lingyun Sun |
MMM (2) | 1 |
| 2023 | Novel 3D-Aware Composition Images Synthesis for Object Display with Diffusion ModelabstractDesigning attractive images for object display can be a time-consuming and skill-intensive process. The emergence of advanced algorithms, particularly the Diffusion Model, has made it possible to synthesize attractive images using AI. However, the existing diffusion models are mostly used to generate entire images and lack control over specific objects for object display. Here, to the best of our knowledge, we pioneers to extend the application of the diffusion model to synthesize novel images for specific objects. By encoding the input images of objects into NeRF representation and synthesizing the desired backgrounds using diffusion models with the input of rendered object images and text prompts, our method can generate 3D aware object display images at arbitrary angles and arbitrary backgrounds. We have conducted extensive experiments to demonstrate that our method is capable of generating high-quality and photo-realistic images, which are > 6 times faster than the conventional photomontage approach. Moreover, our generated images have higher compositional scores, image quality scores, and aesthetics scores in our user experiments. By significantly reducing the need for human effort and producing higher quality generated images, our approach opens up exciting possibilities for creating versatile novel images of specific objects. Tianrun Chen, Tao Xu 0048, Yiyu Ye, Papa Mao, Ying Zang, Lingyun Sun |
SMC | 1 |
| 2023 | Deep3DSketch+\+: High-Fidelity 3D Modeling from Single Free-hand SketchesabstractThe rise of AR/VR has led to an increased demand for 3D content. However, the traditional method of creating 3D content using Computer-Aided Design (CAD) is a labor-intensive and skill-demanding process, making it difficult to use for novice users. Sketch-based 3D modeling provides a promising solution by leveraging the intuitive nature of human-computer interaction. However, generating high-quality content that accurately reflects the creator's ideas can be challenging due to the sparsity and ambiguity of sketches. Furthermore, novice users often find it challenging to create accurate drawings from multiple perspectives or follow step-by-step instructions in existing methods. To address this, we introduce a groundbreaking end-to-end approach in our work, enabling 3D modeling from a single free-hand sketch, Deep3DSketch+\+. The issue of sparsity and ambiguity using single sketch is resolved in our approach by leveraging the symmetry prior and structural-aware shape discriminator. We conducted comprehensive experiments on diverse datasets, including both synthetic and real data, to validate the efficacy of our approach and demonstrate its state-of-the-art (SOTA) performance. Users are also more satisfied with results generated by our approach according to our user study. We believe our approach has the potential to revolutionize the process of 3D modeling by offering an intuitive and easy-to-use solution for novice users. Ying Zang, Chaotao Ding, Tianrun Chen, Papa Mao |
SMC | 3 |
| 2023 | Learning-Based Video Compression Framework With Implicit Spatial Transform for Applications in the Internet of ThingsabstractThe rapid development of Big Data and network technology demands more secure and efficient video transmission for surveillance and video analysis applications. Classical video transmission relies on spatial-frequency transformation for compressing with loss but with limited coding efficiencies. The deep learning-based approach exceeds such limitations. In this work, we push the limit further by proposing an implicit spatial transform parameter method, which models the interframe redundancy to efficiently provide information for frame compression. Specifically, our method comprises a transform estimation module, which estimates the conversion from decoded frame to the current frame, and a context generator. The transform compensation and context generator produce a condensed high-dimensional context. Furthermore, we propose a P-frame CoDec for more efficient frame compression by removing the interframe redundancy. The proposed framework is extensible with a flexible context module. We demonstrate experimentally that our method outperforms previous methods by a large margin. Our method brings 34.817% more saved bit rate than H.265/HEVC. We also demonstrate 17.500% more bit rate saving and 0.490 dB gains in peak signal-to-noise ratio (PSNR) compared with the current state-of-the-art learning-based method proposed by Liu et al. (2022). Qinghai Li, Jinxiang Wang 0005, Tianrun Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Panoptic NeRF: 3D-to-2D Label Transfer for Panoptic Urban Scene SegmentationabstractLarge-scale training data with high-quality annotations is critical for training semantic and instance segmentation models. Unfortunately, pixel-wise annotation is labor-intensive and costly, raising the demand for more efficient labeling strategies. In this work, we present a novel 3D-to-2D label transfer method, Panoptic NeRF1, which aims for obtaining per-pixel 2D semantic and instance labels from easy-to-obtain coarse 3D bounding primitives. Our method utilizes NeRF as a differentiable tool to unify coarse 3D annotations and 2D semantic cues transferred from existing datasets. We demonstrate that this combination allows for improved geometry guided by semantic information, enabling rendering of accurate semantic maps across multiple views. Furthermore, this fusion process resolves label ambiguity of the coarse 3D annotations and filters noise in the 2D predictions. By inferring in 3D space and rendering to 2D labels, our 2D semantic and instance labels are multiview consistent by design. Experimental results show that Panoptic NeRF outperforms existing label transfer methods in terms of accuracy and multi-view consistency on challenging urban scenes of the KITTI-360 dataset. Shangzhan Zhang, Tianrun Chen, Yichong Lu, Lanyun Zhu, Xiaowei Zhou 0001, Andreas Geiger 0001, Yiyi Liao |
3DV | 3 |