EDBT 2026 Demo / reviewers in the wild / expert
Zhanpeng Huang
dblp:135/5933
· DBLP profile ↗
22ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-1888-7490ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Computer networks · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Head Convolution Module With Dynamic Feature Fusion for Medical Image SegmentationabstractABSTRACT Transformers have achieved remarkable success in medical image segmentation due to their large receptive fields and global context extraction capabilities. However, convolutional neural networks (CNNs) face limitations with fixed‐size kernels, hindering their ability to capture multi‐scale features. To address these challenges, we propose two novel modules: the multi‐head channel mixed convolution (MHCMC) module, which enhances feature extraction by expanding the receptive field and utilizing channel attention, and the dynamic feature aggregation (DFA) module, which adaptively prioritizes crucial spatial features based on global information. Additionally, we introduce the evolutionary hybrid network (EHN) to simulate the transition from local to global dependency capture. The MHCMC and DFA modules are first fused to construct the multi‐head DFA (MHDFA) module. Subsequently, the MHDFA and EHN modules are sequentially stacked to form the encoder, which we denote as the multi‐head dynamic aggregation transformer (MDAT). By further integrating MDAT into the U‐Net architecture, we obtain the proposed MDAT‐Net. Experimental results show that MDAT‐Net outperforms other state‐of‐the‐art models in three medical image segmentation tasks: the liver tumor segmentation benchmark (LiTs2017), CVC LinicDB, and the automated cardiac diagnosis challenge (ACDC). Jiangwei Qin, Zhaohan Cai, Zhanpeng Huang |
IET Image Process. | 4 |
| 2026 | Incomplete multi-view clustering based on kernel graph fusion tensor under self-expression framework
Muke Chen, Yongming Cai, Zhanpeng Huang |
Neurocomputing | 3 |
| 2025 | Shining Yourself: High-Fidelity Ornaments Virtual Try-on with Diffusion ModelabstractWhile virtual try-on for clothes and shoes with diffusion models has gained attraction, virtual try-on for ornaments, such as bracelets, rings, earrings, and necklaces, remains largely unexplored. Due to the intricate tiny patterns and repeated geometric sub-structures in most ornaments, it is much more difficult to guarantee identity and appearance consistency under large pose and scale variances between ornaments and models. This paper proposes the task of virtual try-on for ornaments and presents a method to improve the geometric and appearance preservation of ornament virtual try-ons. Specifically, we estimate an accurate wearing mask to improve the alignments between ornaments and models in an iterative scheme alongside the denoising process. To preserve structure details, we further regularize attention layers to map the reference ornament mask to the wearing mask in an implicit way. Experimental results demonstrate that our method successfully wears ornaments from reference images onto target models, handling substantial differences in scale and pose while preserving identity and achieving realistic visual effects. Yingmao Miao, Zhanpeng Huang, Zibin Wang, Chenhao Lin, Chao Shen 0001 |
CVPR | 2 |
| 2025 | From One to More: Contextual Part Latents for 3D GenerationabstractTo generate 3D objects, early research focused on multi-view-driven approaches relying solely on 2D renderings. Recently, the 3D native latent diffusion paradigm has demonstrated superior performance in 3D generation, because it fully leverages the geometric information provided in ground truth 3D data. Despite its fast development, 3D diffusion still faces three challenges. First, the majority of these methods represent a 3D object by one single latent, regardless of its complexity. This may lead to detail loss when generating 3D objects with multiple complicated parts. Second, most 3D assets are designed parts by parts, yet the current holistic latent representation overlooks the independence of these parts and their interrelationships, limiting the model's generative ability. Third, current methods rely on global conditions (e.g., text, image, point cloud) to control the generation process, lacking detailed controllability. Therefore, motivated by how 3D designers create a 3D object, we present a new part-based 3D generation framework, CoPart, which represents a 3D object with multiple contextual part latents and simultaneously generates coherent 3D parts. This part-based framework has several advantages, including: i) reduces the encoding burden of intricate objects by decomposing them into simpler parts, ii) facilitates part learning and part relationship modeling, and iii) naturally supports part-level control. Furthermore, to ensure the coherence of part latents and to harness the powerful priors from foundation models, we propose a novel mutual guidance strategy to fine-tune pre-trained diffusion models for joint part latent denoising. Benefiting from the part-based representation, we demonstrate that CoPart can support various applications including part-editing, articulated object generation, and mini-scene generation. Moreover, we collect a new large-scale 3D part dataset named Partverse from Objaverse through automatic mesh segmentation and subsequent human post-annotations. By training on the proposed dataset, CoPart achieves promising part-based 3D generation with high controllability. Project page: https://hkdsc.github.io/project/copart. Shaocong Dong, Lihe Ding, Yaokun Li, Jaehyeok Kim, Chenjian Gao, Zhanpeng Huang, Zibin Wang, Tianfan Xue |
ICCV | 10 |
| 2025 | From Gallery to Wrist: Realistic 3D Bracelet Insertion in VideosabstractInserting 3D objects into videos is a longstanding challenge in computer graphics with applications in augmented reality, virtual try-on, and video composition. Achieving both temporal consistency, or realistic lighting remains difficult, particularly in dynamic scenarios with complex object motion, perspective changes, and varying illumination. While 2D diffusion models have shown promise for producing photorealistic edits, they often struggle with maintaining temporal coherence across frames. Conversely, traditional 3D rendering methods excel in spatial and temporal consistency but fall short in achieving photorealistic lighting. In this work, we propose a hybrid object insertion pipeline that combines the strengths of both paradigms. Specifically, we focus on inserting bracelets into dynamic wrist scenes, leveraging the high temporal consistency of 3D Gaussian Splatting (3DGS) for initial rendering and refining the results using a 2D diffusion-based enhancement model to ensure realistic lighting interactions. Our method introduces a shading-driven pipeline that separates intrinsic object properties (albedo, shading, reflectance) and refines both shading and sRGB images for photorealism. To maintain temporal coherence, we optimize the 3DGS model with multi-frame weighted adjustments. This is the first approach to synergize 3D rendering and 2D diffusion for video object insertion, offering a robust solution for realistic and consistent video editing. Project Page: https://cjeen.github.io/BraceletPaper/ Chenjian Gao, Lihe Ding, Zhanpeng Huang, Zibin Wang, Tianfan Xue |
ICCV | 4 |
| 2025 | Multi-view subspace clustering via double-constrained matrix factorization and graph filter
Zhaohan Cai, Muke Chen, Zhanpeng Huang |
Appl. Intell. | 6 |
| 2025 | High-order tensor based multi-view clustering via enhanced adaptive graph propagation
Muke Chen, Yongming Cai, Zhanpeng Huang |
Pattern Anal. Appl. | 3 |
| 2024 | Text-to-3D Generation with Bidirectional Diffusion Using Both 2D and 3D PriorsabstractMost 3D generation research focuses on up-projecting 2D foundation models into the 3D space, either by minimizing 2D Score Distillation Sampling (SDS) loss or fine-tuning on multi-view datasets. Without explicit 3D priors, these methods often lead to geometric anomalies and multi-view inconsistency. Recently, researchers have attempted to improve the genuineness of 3D objects by directly training on 3D datasets, albeit at the cost of low-quality texture generation due to the limited texture diversity in 3D datasets. To harness the advantages of both approaches, we propose Bidirectional Diffusion (BiDiff), a unified framework that incorporates both a 3D and a 2D diffusion process, to preserve both 3D fidelity and 2D texture richness, respectively. Moreover, as a simple combination may yield inconsistent generation results, we further bridge them with novel bidirectional guidance. In addition, our method can be used as an initialization of optimization-based models to further improve the quality of 3D models and the efficiency of optimization, reducing the process from 3.4 hours to 20 minutes. Experimental results have shown that our model achieves high-quality, diverse, and scalable 3D generation. Project website https://bidiff.github.io/. Lihe Ding, Shaocong Dong, Zhanpeng Huang, Zibin Wang, Kaixiong Gong, Dan Xu 0002, Tianfan Xue |
CVPR | 3 |
| 2024 | Interactive3D: Create What You Want by Interactive 3D Generationabstract3D object generation has undergone significant advancements, yielding high-quality results. However, fall short of achieving precise user control, often yielding results that do not align with user expectations, thus limiting their applicability. User-envisioning 3D object generation faces significant challenges in realizing its concepts using current generative models due to limited interaction capabilities. Existing methods mainly offer two approaches: (i) interpreting textual instructions with constrained controllability, or (ii) reconstructing 3D objects from 2D images. Both of them limit customization to the confines of the 2D reference and potentially introduce undesirable artifacts during the 3D lifting process, restricting the scope for direct and versatile 3D modifications. In this work, we introduce Interactive3D, an innovative framework for interactive 3D generation that grants users precise control over the generative process through extensive 3D interaction capabilities. Interactive3D is constructed in two cascading stages, utilizing distinct 3D representations. The first stage employs Gaussian Splatting for direct user interaction, allowing modifications and guidance of the generative direction at any intermediate step through (i) Adding and Removing components, (ii) Deformable and Rigid Dragging, (iii) Geometric Transformations, and (iv) Semantic Editing. Subsequently, the Gaussian splats are transformed into InstantNGP. We introduce a novel (v) Interactive Hash Refinement module to further add details and extract the geometry in the second stage. Our experiments demonstrate that proposed Interactive3D markedly improves the controllability and quality of 3D generation. Our project webpage is available at https://interactive-3d.github.io/. Shaocong Dong, Lihe Ding, Zhanpeng Huang, Zibin Wang, Tianfan Xue, Dan Xu 0002 |
CVPR | 3 |
| 2022 | A Multiview Clustering Method With Low-Rank and Sparsity Constraints for Cancer SubtypingabstractMultiomics data clustering is one of the major challenges in the field of precision medicine. Integration of multiomics data for cancer subtyping can improve the understanding on cancer and reveal systems-level insights. How to integrate multiomics data for accurate cancer subtyping is an interesting and challenging research problem. To capture the global and the local structure of omics data, a novel framework for integrating multiomics data is proposed for cancer subtyping. Multiview clustering with low-rank and sparsity constraints (MVCLRS) can measure the local similarities of samples in each omics data and obtain global consensus structures by integrating the multiomics data. The main insight provided by MVCLRS is that low-rank sparse subspace clustering for the construction of an affinity matrix can best capture the local similarities in omics data. Extensive testing is conducted on 10 real world cancer datasets with multiomics from The Cancer Genome Atlas. Compared with 10 state-of-the-art multiomics clustering algorithms, the MVCLRS performs better in the 10 cancer datasets by providing its clustering results with at least one enriched clinical label in nine of ten cancer subtypes, the most of any method. Zhanpeng Huang, Jiekang Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Automatically Generate Rigged Character from Single ImageabstractAnimation plays an important role in virtual reality and augmented reality applications. However, it requires great efforts for non-professional users to create animation assets. In this paper, we propose a systematic pipeline to generate ready-to-used characters from images for real-time animation without user intervention. Rather than per-pixel mapping or synthesis in image space using optical flow or generative models, we employ an approximate geometric embodiment to undertake 3D animation without large distortion. The geometry structure is generated from a type-agnostic character. A skeleton adaption is then adopted to guarantee semantic motion transfer to the geometry proxy. The generated character is compatible with standard 3D graphics engines and ready to use for real-time applications. Experiments show that our method works on various images (e.g. sketches, cartoons, and photos) of most object categories (e.g. human, animals, and non-creatures). We develop an AR demo to show its potential usage for fast prototyping. Zhanpeng Huang, Rui Han 0005, Jianwen Huang, Zipeng Qin, Zibin Wang |
MMAsia | 1 |
| 2019 | Dandelion: A Unified Code Offloading System for Wearable ComputingabstractExecution speed seriously bothers application developers and users for wearable devices such as Google Glass. Intensive applications like 3D games suffer from significant delays when CPU is busy. Energy is another concern when the devices are in low battery level, but users need them for urgency use. To ease such pains, one approach is to expand the computational power by cloud offloading. This paradigm works well when the available Internet access has enough bandwidth. Another way is to leverage nearby devices for computation-offloading, which is known as device-to-device (D2D) offloading. In this paper, we present Dandelion, a unified code offloading system for wearable computing. Such applications can leverage both the nearby devices and cloud for performance acceleration and energy efficiency. Dandelion is a novel generic code offloading system for wearable computing with a reference implementation on Google Glass. Dandelion includes a programmer-friendly framework based on Java annotation, a lightweight offloading service, and a runtime task scheduler to make offloading decisions. We design some wearable applications and several parallel execution benchmark methods for Dandelion performance evaluation. Extensive experiments on a testbed of Google Glass and Android phones demonstrate that Dandelion generally achieves over 5X execution speedup for local execution and can quickly recover from errors caused by network disruption. Morteza Golkarifard, Zhanpeng Huang, Ali Movaghar-Rahimabadi, Pan Hui 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2017 | Hyperion: A Wearable Augmented Reality System for Text Extraction and Manipulation in the AirabstractWe develop Hyperion a Wearable Augmented Reality (WAR) system based on Google Glass to access text information in the ambient environment. Hyperion is able to retrieve text content from users' current view and deliver the content to them in different ways according to their context. We design four work modalities for different situations that mobile users encounter in their daily activities. In addition, user interaction interfaces are provided to adapt to different application scenarios. Although Google Glass may be constrained by its poor computational capabilities and its limited battery capacity, we utilize code-level offloading to companion mobile devices to improve the runtime performance and the sustainability of WAR applications. System experiments show that Hyperion improves users ability to be aware of text information around them. Our prototype indicates promising potential of converging WAR technology and wearable devices such as Google Glass to improve people's daily activities. Dimitris Chatzopoulos, Carlos Bermejo 0001, Zhanpeng Huang, Arailym Butabayeva, Morteza Golkarifard, Pan Hui 0001 |
MMSys | 3 |
| 2017 | Ubii: Physical World Interaction Through Augmented RealityabstractWe describe a new set of interaction techniques that allow users to interact with physical objects through augmented reality (AR). Previously, to operate a smart device, physical touch is generally needed and a graphical interface is normally involved. These become limitations and prevent the user from operating a device out of reach or operating multiple devices at once. Ubii (Ubiquitous interface and interaction) is an integrated interface system that connects a network of smart devices together, and allows users to interact with the physical objects using hand gestures. The user wears a smart glass which displays the user interface in an augmented reality view. Hand gestures are captured by the smart glass, and upon recognizing the right gesture input, Ubii will communicate with the connected smart devices to complete the designated operations. Ubii supports common inter-device operations such as file transfer, printing, projecting, as well as device pairing. To improve the overall performance of the system, we implement computation offloading to perform the image processing computation. Our user test shows that Ubii is easy to use and more intuitive than traditional user interfaces. Ubii shortens the operation time on various tasks involving operating physical devices. The novel interaction paradigm attains a seamless interaction between the physical and digital worlds. Sikun Lin, Hao Fei Cheng, Weikai Li 0001, Zhanpeng Huang, Pan Hui 0001, Christoph Peylo |
IEEE Trans. Mob. Comput. | 4 |
| 2016 | Vortex particle smoke simulation with an octree data structureabstractAbstract We propose an octree‐based presentation of vortex particles to simulate smoke and gaseous phenomena in a physical way. Vortex particle method prevails over grid‐based method in terms of less numerical dissipation and more detail features, but it suffers from heavy computational overhead due to per‐particle Biot–Savart integration over the entire simulation space. To alleviate this problem, we employ an octree background grid to separate the vortex particles into individual groups. Particles in groups are aggregated as a single super vortex particle to reduce computational cost. The proposed method produces comparable visual result as previous methods with much less computational overhead. Copyright © 2014 John Wiley & Sons, Ltd. Zhanpeng Huang, Guanghong Gong |
Comput. Animat. Virtual Worlds | 1 |
| 2015 | Ubii: Towards Seamless Interaction between Digital and Physical WorldsabstractWe present Ubii (Ubiquitous interface and interaction), an interface system that aims to expand people's perception and interaction from the digital space to the physical world. The centralized user interface is broken into pieces woven in the domain environment. Augmented user interface is paired to the physical objects, where physical and digital presentations are displayed in the same context. The augmented interface and physical affordance respond as one control to provide seamless interaction. By connecting digital interface with physical objects, the system presents a nearby embodiment to afford users sense of awareness to interact with domain objects. Integrated on wearable devices as Google Glass, a less intrusive and more convenient interaction is afforded. Our research illustrates the great potential of direct mapping of interaction between digital interfaces and physical affordance by converging wearable devices and augmented reality (AR) technology. Zhanpeng Huang, Weikai Li 0001, Pan Hui 0001 |
ACM Multimedia | 1 |
| 2015 | Offloading Guidelines for Augmented Reality Applications on Wearable DevicesabstractAs Augmented Reality (AR) gets popular on wearable devices such as Google Glass, various AR applications have been developed by leveraging synergetic benefits beyond the single technologies. However, the poor computational capability and limited power capacity of current wearable devices degrade runtime performance and sustainability. Computational offloading strategy has been proposed to outsource computation to remote cloud for improving performance. Nevertheless, comparing with mobile devices, the wearable devices have their specific limitations, which induce additional problems and require new thoughts of computational offloading. In this paper, we propose several guidelines of computational offloading for AR applications on wearable devices based on our practical experiences of designing and developing AR applications on Google Glass. The guidelines have been adopted and proved by our application prototypes. Zhanpeng Huang, Pan Hui 0001 |
ACM Multimedia | 3 |
| 2015 | Reducing numerical dissipation in smoke simulation
Zhanpeng Huang, Ladislav Kavan, Weikai Li 0001, Pan Hui 0001, Guanghong Gong |
Graph. Model. | 1 |
| 2015 | A local adaptive Catmull-Rom to reduce numerical dissipation of semi-Lagrangian advectionabstractAbstract We propose an adaptive Catmull–Rom interpolation to improve accuracy of semi‐Lagrangian advection for smoke simulation. Original Catmull–Rom improves numerical accuracy but overshoot violates global stability. Monotonic Catmull–Rom is unconditionally stable, whereas it sweeps out detail features due to overly suppression operations. Our method modifies original Catmull–Rom to obtain second‐order accuracy and unconditional stability. It flattens locations where interpolations might break through global bounds but maintains local overshoots to conserve diversity of fluid flow. The scheme is easy to collaborate with existing fluid simulators to improve small features. Copyright © 2013 John Wiley & Sons, Ltd. Zhanpeng Huang, Guanghong Gong |
Comput. Animat. Virtual Worlds | 1 |
| 2015 | Physically-based smoke simulation for computer graphics: a survey
Zhanpeng Huang, Guanghong Gong |
Multim. Tools Appl. | 1 |
| 2014 | Physically-based modeling, simulation and rendering of fire for computer animation
Zhanpeng Huang, Guanghong Gong |
Multim. Tools Appl. | 1 |
| 2009 | A Multi-view Nonlinear Active Shape Model Based on 3D Transformation Shape SearchabstractActive shape model (ASM) is an efficient method for locating face feature points. When the face poses vary largely, it is difficult for the two-dimension (2D) transformation shape search method in previous works to cope with this kind of nonlinear shape variations. We propose a novel shape search method of the three-dimension (3D) transformation; and the variation of the face pose can simulated by 3D transformation shape search correctly. The method includes the following steps: firstly, constructing an average face 3D model; secondly, by the average face 3D model, making 2D initial shape get the third dimension coordinate in the face images of ASM training set; thirdly, tracking the object shape by 3D transforming and projecting to 2D view-plane based on the 3D initial shape. In the third step, 3D transformation shape search is implemented by the two-step transformation. The 10 pose parameters of 3D transformation are calculated after the two step transformations are iterated into the pre-defining precision. The test data shows the 3D transformation shape search method has better performance than the current standard 2D transformation while the face poses vary largely. Faling Yi, Zhanpeng Huang |
IAS | 3 |