EDBT 2026 Demo / reviewers in the wild / expert
Ye Chen 0006
dblp:14/3310-6
· DBLP profile ↗
15ranked-venue papers
4as first author
14since 2021 · last 2025
0009-0002-4011-593XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InstantSticker: Realistic Decal Blending via Disentangled Object ReconstructionabstractWe present InstantSticker, a disentangled reconstruction pipeline based on Image-Based Lighting (IBL), which focuses on highly realistic decal blending, simulates stickers attached to the reconstructed surface, and allows for instant editing and real-time rendering. To achieve stereoscopic impression of the decal, we introduce shadow factor into IBL, which can be adaptively optimized during training. This allows the shadow brightness of surfaces to be accurately decomposed rather than baked into the diffuse color, ensuring that the edited texture exhibits authentic shading. To address the issues of warping and blurriness in previous methods, we apply As-Rigid-As-Possible (ARAP) parameterization to pre-unfold a specified area of the mesh and use the local UV mapping combined with a neural texture map to enhance the ability to express high-frequency details in that area. For instant editing, we utilize the Disney BRDF model, explicitly defining material colors with 3-channel diffuse albedo. This enables instant replacement of albedo RGB values during the editing process, avoiding the prolonged optimization required in previous approaches. In our experiment, we introduce the Ratio Variance Warping (RVW) metric to evaluate the local geometric warping of the decal area. Extensive experimental results demonstrate that our method surpasses previous decal blending methods in terms of editing quality, editing speed and rendering speed, achieving the state-of-the-art. Yishun Dou, Ye Chen 0006, Bingbing Ni, Wenjun Zhang 0001 |
AAAI | 6 |
| 2025 | Easy-editable Image Vectorization with Multi-layer Multi-scale Distributed Visual Feature EmbeddingabstractCurrent parameterized image representations embed visual information along the semantic boundaries and struggle to express the internal detailed texture structures of image components, leading to a lack of content consistency after image editing and driving. To address these challenges, this work proposes a novel parameterized representation based on hierarchical image proxy geometry, utilizing multi-layer hierarchically interrelated proxy geometric control points to embed multi-scale long-range structures and fine-grained texture details. The proposed representation enables smoother and more continuous interpolation during image rendering and ensures high-quality consistency within image components during image editing. Additionally, under the layer-wise representation strategy based on semantic-aware image layer decomposition, we enable decoupled image shape/texture editing of the targets of interest within the image. Extensive experimental results on image vectorization and editing tasks demonstrate that our proposed method achieves high rendering accuracy of general images, including natural images, with a significantly higher image parameter compression ratio, facilitating user-friendly editing of image semantic components. Ye Chen 0006, Zhangli Hu, Zhongyin Zhao, Yupeng Zhu, Yuxuan Xiong, Bingbing Ni |
CVPR | 1 |
| 2025 | AMR-Transformer: Enabling Efficient Long-range Interaction for Complex Neural Fluid SimulationabstractAccurately and efficiently simulating complex fluid dynamics is a challenging task that has traditionally relied on computationally intensive methods. Neural network-based approaches, such as convolutional and graph neural networks, have partially alleviated this burden by enabling efficient local feature extraction. However, they struggle to capture long-range dependencies due to limited receptive fields, and Transformer-based models, while providing global context, incur prohibitive computational costs. To tackle these challenges, we propose AMR-Transformer, an efficient and accurate neural CFD-solving pipeline that integrates a novel adaptive mesh refinement scheme with a Navier-Stokes constraint-aware fast pruning module. This design encourages long-range interactions between simulation cells and facilitates the modeling of global fluid wave patterns, such as turbulence and shockwaves. Experiments show that our approach achieves significant gains in efficiency while preserving critical details, making it suitable for high-resolution physical simulations with long-range dependencies. On CFDBench, PDEBench and a new shock wave dataset, our pipeline demonstrates up to an order-of- magnitude improvement in accuracy over baseline models. Additionally, compared to ViT, our approach achieves a reduction in FLOPs of up to 60 times. Zeyi Xu, Jinfan Liu, Kuangxu Chen, Ye Chen 0006, Zhangli Hu, Bingbing Ni |
CVPR | 4 |
| 2025 | SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG GenerationabstractScalable Vector Graphics (SVG) is a code structure used to represent visual information, and with the powerful capabilities of large language models, it holds significant research potential. Current text-to-SVG generation methods lack generalization capabilities and struggle with accurately adhering to input generation instructions. In this paper, we propose a novel approach for generating SVG using large language models, named SVGThinker, which incorporates a reasoning process to align the generation of SVG code with the visualization process, while supporting all SVG primitives. Through sequential rendering of SVG primitives, we first use a multimodal model to annotate the SVG, followed by sequential updates corresponding to the incremental additions of primitives. We then employ a supervised training framework based on Chain-of-Thought reasoning, which enhances the model's robustness and reduces the risk of errors or hallucinations. Through comparisons with state-of-the-art baseline models, our experiments show that our model generates more stable, high-quality, and editable SVG code. In contrast to image-based methods, our approach preserves the structural advantages of SVG and supports precise, hierarchical editing. We believe our work opens new directions for SVG generation, with potential applications in design, content creation, and automated SVG-based graphic generation. Zhongyin Zhao, Ye Chen 0006, Zhujin Liang, Bingbing Ni |
ACM Multimedia | 3 |
| 2025 | AR 2O Painter: An Artistic Oriented Realtime Realistic Oil Painting Agent Powered by Efficient Fluid SimulationabstractWe introduce the AR2 O Painter, an interactive intelligent system designed for real-time, highly realistic oil painting creation. This Agent can faithfully reproduce any portrait image, allowing users to visually enjoy the stroke-by-stroke painting process immersively as the artwork is completed within two minutes. It consists of two modules: the Oil Painting Stroke Sequence Planner, which performs multi-level semantic-based brushstroke sequence decomposition on portrait images, mimicking the logic of artist painting, and the Oil Painting Rendering Engine, which receives the brushstroke sequence, models the pigment via fluid dynamics, simulates its interaction with the canvas and brush, and applies a tailored PBR model with microfacet BRDF, Fresnel effects, and stroke-level geometry, enabling perceptually plausible gloss and fine-grained surface relief. To the best of our knowledge, it is the first real-time intelligent painting system to generate realistic oil paintings with high interactivity and artistic fidelity. The demo video is available at https://youtu.be/aN-W06GmnP8. Jinfan Liu, Zhangli Hu, Ye Chen 0006, Bingbing Ni, Shuicheng Yan |
ACM Multimedia | 4 |
| 2025 | Rig-Reconstruct-Render (R33D): Collaborative Representation for Editable and Skeleton-Drivable 3D Asset Generation
Yuxuan Xiong, Ye Chen 0006, Zhangli Hu, Bingbing Ni |
ACM Multimedia | 2 |
| 2024 | Towards High-fidelity Artistic Image Vectorization via Texture-Encapsulated Shape ParameterizationabstractWe develop a novel vectorized image representation scheme accommodating both shape/geometry and texture in a decoupled way, particularly tailored for reconstruction and editing tasks of artistic/design images such as Emojis and Cliparts. In the heart of this representation is a set of sparsely and unevenly located 2D control points. On one hand, these points constitute a collection of paramet-ric/vectorized geometric primitives (e.g., curves and closed shapes) describing the shape characteristics of the target image. On the other hand, local texture codes, in terms of implicit neural network parameters, are spatially dis-tributed into each control point, yielding local coordinate-to-RGB mappings within the anchored region of each con-trol point. In the meantime, a zero-shot learning algorithm is developed to decompose an arbitrary raster image into the above representation, for the sake of high-fidelity im-age vectorization with convenient editing ability. Extensive experiments on a series of image vectorization and editing tasks well demonstrate the high accuracy offered by our proposed method, with a significantly higher image com-pression ratio over prior art. Ye Chen 0006, Bingbing Ni, Jinfan Liu, Xuanhong Chen |
CVPR | 1 |
| 2024 | Vector Graphics Generation via Mutually Impulsed Dual-Domain DiffusionabstractIntelligent generation of vector graphics has very promising applications in the fields of advertising and logo design, artistic painting, animation production, etc. However, current mainstream vector image generation methods lack the encoding of image appearance information that is associated with the original vector representation and therefore lose valid supervision signal from the strong correlation between the discrete vector parameter (drawing in-struction) sequence and the target shape/structure of the corresponding pixel image. On the one hand, the gener-ation process based on pure vector domain completely ignores the similarity measurement between shape parameter (and their combination) and the paired pixel image appearance pattern; on the other hand, two-stage methods (i.e., generation-and-vectorization) based on pixel diffusion followed by differentiable image-to-vector translation suf-fer from wrong error-correction signal caused by approxi-mate gradients. To address the above issues, we propose a novel generation framework based on dual-domain (vector-pixel) diffusion with cross-modality impulse signals from each other. First, in each diffusion step, the current representation extracted from the other domain is used as a condition variable to constrain the subsequent sampling operation, yielding shape-aware new parameterizations; second, independent supervision signals from both domains avoid the gradient error accumulation problem caused by cross-domain representation conversion. Extensive experimental results on popular benchmarks including font and icon datasets demonstrate the great advantages of our proposed framework in terms of generated shape quality. Zhongyin Zhao, Ye Chen 0006, Zhangli Hu, Xuanhong Chen, Bingbing Ni |
CVPR | 2 |
| 2024 | Towards Artist-Like Painting Agents with Multi-Granularity Semantic Alignment
Zhangli Hu, Ye Chen 0006, Zhongyin Zhao, Jinfan Liu, Bilian Ke, Bingbing Ni |
ACM Multimedia | 2 |
| 2023 | Fast Fluid Simulation via Dynamic Multi-Scale GriddingabstractRecent works on learning-based frameworks for Lagrangian (i.e., particle-based) fluid simulation, though bypassing iterative pressure projection via efficient convolution operators, are still time-consuming due to excessive amount of particles. To address this challenge, we propose a dynamic multi-scale gridding method to reduce the magnitude of elements that have to be processed, by observing repeated particle motion patterns within certain consistent regions. Specifically, we hierarchically generate multi-scale micelles in Euclidean space by grouping particles that share similar motion patterns/characteristics based on super-light motion and scale estimation modules. With little internal motion variation, each micelle is modeled as a single rigid body with convolution only applied to a single representative particle. In addition, a distance-based interpolation is conducted to propagate relative motion message among micelles. With our efficient design, the network produces high visual fidelity fluid simulations with the inference time to be only 4.24 ms/frame (with 6K fluid particles), hence enables real-time human-computer interaction and animation. Experimental results on multiple datasets show that our work achieves great simulation acceleration with negligible prediction error increase. Jinxian Liu, Ye Chen 0006, Bingbing Ni, Zhenbo Yu |
AAAI | 2 |
| 2023 | Editable Image Geometric Abstraction via Neural Primitive AssemblyabstractThis work explores a novel image geometric abstraction paradigm based on assembly out of a pool of pre-defined simple parametric primitives (i.e., triangle, rectangle, circle and semicircle), facilitating controllable shape editing in images. While cast as a mixed combinatorial and continuous optimization problem, the above task is approximately reformulated within a token translation neural framework that simultaneously outputs primitive assignments and corresponding transformation and color parameters in an image-to-set manner, thus bypassing complex/non-differentiable graph-matching iterations. To relax the searching space and address the vanishing gradient issue, a novel Neural Soft Assignment scheme that well explores the quasi-equivalence between the assignment in Bipartite b-Matching and opacity-aware weighted multiple rasterization combination is introduced, drastically reducing the optimization complexity. Without ground-truth image abstraction labeling (i.e., vectorized representation), the whole pipeline is end-to-end trainable in a self-supervised manner, based on the linkage of differentiable rasterization techniques. Extensive experiments on several datasets well demonstrate that our framework is able to predict highly compelling vectorized geometric abstraction results with a combination of ONLY four simple primitives, also with VERY straightforward shape editing capability by simple replacement of primitive type, compared to previous image abstraction and image vectorization methods. Ye Chen 0006, Bingbing Ni, Xuanhong Chen, Zhangli Hu |
ICCV | 1 |
| 2023 | Learning by Restoring Broken 3D GeometryabstractThe key point for an experienced craftsman to repair broken objects effectively is that he must know about them deeply. Similarly, we believe that a model can capture rich geometry information from a shape/scene and generate discriminative representations if it is able to find distorted parts of shapes/scenes and restore them. Inspired by this observation, we propose a novel self-supervised 3D learning paradigm named learning by restoring broken shapes/scenes (collectively called 3D geometry). We first develop a destroy-method cluster, from which we sample methods to break some local parts of an object. Then the destroyed object and the normal object are both sent into a point cloud network to get representations, which are employed to segment points that belong to distorted parts and further reconstruct/restore them to normal. To perform better in these two associated pretext tasks, the model is constrained to capture useful object features, such as rich geometric and contextual information. The object representations learned by this self-supervised paradigm transfer well to different datasets and perform well on downstream classification, segmentation and detection tasks. Experimental results on shape datasets and scene datasets demonstrate that our method achieves state-of-the-art performance among unsupervised methods. We also show experimentally that pre-training with our framework significantly boosts the performance of supervised models. Jinxian Liu, Bingbing Ni, Ye Chen 0006, Zhenbo Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Joint Global and Dynamic Pseudo Labeling for Semi-Supervised Point Cloud Sequence SegmentationabstractSupervised learning is a mainstay for large discriminative models in 3D computer vision, while large amounts of human-annotated data are the key to achieve state-of-the-art performance. This limitation is particularly notable for large-scale point cloud sequence segmentation tasks, because point-level annotations are very time-consuming and especially expensive. To overcome this challenge, we develop a novel semi-supervised framework for point cloud sequences segmentation. Specifically, we develop two kinds of pseudo labeling methods with extracting global semantic information from labeled frames and dynamic information from each sequence respectively. Then the two kinds of generated labels are combined as more robust pseudo labels (GD-Pseudo labels) for unlabeled frames. We finally apply an efficient iterative learning scheme to train a model with a small quantity of human-annotated data and large-scale pseudo-labeled data. Equipped with our framework, the model achieves significant performance improvement (+12—25 mIoU) on SemanticKITTI and Synthia when compared with frameworks that do not utilize large amounts of unlabeled data. Moreover, our method achieves comparable performance with only 20% annotated frames on SemanticKITTI to state-of-the-art models trained with 100% human-annotated frames. Jinxian Liu, Ye Chen 0006, Bingbing Ni, Zhenbo Yu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Shape Self-Correction for Unsupervised Point Cloud UnderstandingabstractWe develop a novel self-supervised learning method named Shape Self-Correction for point cloud analysis. Our method is motivated by the principle that a good shape representation should be able to find distorted parts of a shape and correct them. To learn strong shape representations in an unsupervised manner, we first design a shape-disorganizing module to destroy certain local shape parts of an object. Then the destroyed shape and the normal shape are sent into a point cloud network to get representations, which are employed to segment points that belong to distorted parts and further reconstruct them to restore the shape to normal. To perform better in these two associated pretext tasks, the network is constrained to capture useful shape features from the object, which indicates that the point cloud network encodes rich geometric and contextual information. The learned feature extractor transfers well to downstream classification and segmentation tasks. Experimental results on ModelNet, ScanNet and ShapeNetPart demonstrate that our method achieves state-of-the-art performance among unsupervised methods. Our framework can be applied to a wide range of deep learning networks for point cloud analysis and we show experimentally that pre-training with our framework significantly boosts the performance of supervised models. Ye Chen 0006, Jinxian Liu, Bingbing Ni, Jiancheng Yang, Teng Li 0001, Qi Tian 0001 |
ICCV | 1 |
| 2020 | Self-Prediction for Joint Instance and Semantic Segmentation of Point Clouds
Jinxian Liu, Minghui Yu, Bingbing Ni, Ye Chen 0006 |
ECCV (22) | 4 |