Jianchi Sun

dblp:272/2403 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HumanPro: Single-view 3D Clothed Human Reconstruction with Progressive Normal Guidance
abstract
Reconstructing fine-grained geometry of clothed human from single-view image is a challenging task, particularly in accurately recovering complex shapes and generating clothes details. To address these limitations, we propose a novel approach named HumanPro, which estimates high-quality human normals via a generative model, and progressively deforms a parametric body into the final clothed human mesh guided by normals. First, we propose a geometry-aware latent diffusion model with a normal enhancer to estimate high-quality human normals from four views. Then, we propose a progressive mesh optimization consisting of shape-aware deformation alignment and global-to-patch detail refinement for human mesh reconstruction. The shape-aware deformation alignment applies image morphing to learn the shape-level gap of normals, addressing large-scale deformation of complex clothes. It can recover the overall silhouette of a clothed human, and serves as an initialization for the global-to-patch detail refinement. Our detail refinement combines global and patch-wise optimization strategies to iteratively produce the clothed human mesh by minimizing the pixel-level difference of normals. This way effectively recovers fine-grained details while avoiding local minima. Extensive experiments demonstrate that HumanPro can deal with various challenging scenarios and outperforms state-of-the-art methods.
Jianchi Sun, Fei Luo 0004, Wenzhuo Fan, Yu Jiang 0007, Chunxia Xiao
AAAI1
2026 HumanFlow: Controllable Human Image Generation via Flow Matching
abstract
We present HumanFlow, a unified flow-matching-based framework that enables high-fidelity and controllable full-body human image generation under diverse human-centric control conditions. Despite recent progress, controllable human image generation poses a fundamental challenge in balancing high visual fidelity with strict adherence to human-centric control conditions. HumanFlow formulates human image generation as a conditional flow-matching process with deterministic generation dynamics. To incorporate such human-centric control conditions into the pretrained model, we introduce a unified control framework with Control Encoder and Token-ControlNet. A Control Encoder maps diverse conditions into a unified latent representation that is spatially aligned with the image latent space. Token-ControlNet is a lightweight control network architecturally aligned with the FLUX double-stream design. To address accurate structural control over human bodies, we further propose the Human Topology Consistency Loss (HTCL). HTCL regularizes conditional flow matching by constraining generated human configurations to a union of statistically grounded topology manifolds defined by normalized bone ratios and joint angles. To support large-scale training and systematic evaluation, we construct MiCoGen, a multi-condition human image dataset comprising over one million full-body human images with aligned text descriptions and rich human-centric control conditions. Extensive quantitative and qualitative evaluations on the MiCoGen dataset show that HumanFlow consistently achieves improved structural consistency than the existing diffusion-based and flow-matching-based methods, while maintaining high visual fidelity.
Wenzhuo Fan, Hongsheng Zheng, Jianchi Sun, Chunxia Xiao
ACM Trans. Graph.3
2026 GUIDE: Dual-Gated semantic conditioning with correctable priors for single-image human material estimation
Yu Jiang 0007, Jianchi Sun, Xiangqian Shen, Chunxia Xiao
Vis. Comput.2
2026 DuoLit: dual-level priors and gated lighting injection for human material estimation
Yu Jiang 0007, Jianchi Sun, Xiangqian Shen, Chunxia Xiao
Vis. Comput.2
2026 HAFMat: Hybrid priors guided adaptive fusion for single-image human material estimation
Yu Jiang 0007, Jiahao Xia 0001, Jiongming Qin, Jianchi Sun, Chunxia Xiao
Vis. Comput.4
2025 I 2HDiffuser: Image Illumination Harmonization Meets the Diffusion Model
abstract
Recently, since diffusion models show great potential in image generation, many pretrained diffusion models based image composition methods have been proposed for image illumination harmonization. However, they mainly face two key challenges: 1) the effective preservation of foreground appearance (i.e., content structure and texture details, etc); 2) Reasonable generation of the foreground casting shadow. To this end, we propose a novel Image Illumination Harmonization Diffusion model called I 2 HDiffuser to achieve image illumination harmonization with high-fidelity foreground appearance and reasonable cast shadows. I 2 HDiffuser mainly consists of frequency domain feature enhancement branch (FDFEB) and illumination-shadow consistency generation branch (ISCGB). Specifically, FDFEB first introduces the Wavelet Transform Module (WTM) for decomposing composite image features into low-frequency (i.e., illumination features, etc) and high-frequency (i.e., texture and content structure features, etc) components using the Haar wavelet transform. Then the Multi-Condition Guidance Mechanism (M-CGM) is proposed to interact these components as prior conditions, which are further injected into the ISCGB with a noise-to-denoise process for guiding high-fidelity content and background illumination-aware foreground regeneration. Meanwhile, a shadow mask step-wise iterative optimization strategy is introduced to the ISCGB to explicitly provide a reasonable shadow generation space for foreground objects. Extensive experiments on public image harmonization datasets DESOBAv2 and iHarmony4 and real illumination harmonization dataset IH-SG show that the I 2HDiffuser achieves the superiority.
Zhongyun Bao, Gang Fu 0003, Jianchi Sun, Chunxia Xiao
ACM Multimedia3
2025 EE-Head: emotion estimation for precise facial expression in NeRF head avatars
Enxu Zhao, Jianchi Sun, Fei Luo 0004, Chunxia Xiao
Vis. Comput.2
2024 Evaluation of the gross motor abilities of autistic children with a computerised evaluation method
abstract
To effectively evaluate the gross motor ability of autistic children, we proposed a method of computerised evaluation of gross motor skills (CEGM). The CEGM integrates Dynamic Time Warping (DTW) method and OpenPose technology to automatically detect key joints and return a score. Ten items were selected for evaluation based on the gross motor subtest of the Psychoeducational Profile – Third Edition (PEP-3) scale, including upper limb movement, lower limb movement, and body coordination performance. 30 autistic participants (males: 23, female: 7) with an average age of 5.00 years were recruited in this study. Then we compared the results of evaluation using CEGM and the original PEP-3 gross motor subtest in autistic children. The results showed that in the evaluations using CEGM and PEP-3, Cronbach’s α coefficients and Spearman-rank correlation coefficients were all greater than 0.80, intraclass correlation coefficient (ICC) were all greater than 0.90, indicating good agreement in evaluating the gross motor ability of autistic children. Moreover, compared to the PEP-3, the evaluation using CEGM provided precise quantitative indicators (trajectory, velocity, and angle of joint). Therefore, our findings demonstrate that CEGM can be used in the initial evaluation of the gross motor ability of autistic children.
Xiaodi Liu, Jingying Chen 0001, Guangshuai Wang, Kun Zhang 0031, Jianchi Sun, Pianpian Ma, Rujing Zhang
Behav. Inf. Technol.5
2024 Efficient Data Extraction Circuit for Posit Number System: LDD-Based Posit Decoder
abstract
Since being proposed in 2017, the posit number system has attracted much attention due to its advantages over the IEEE 754 standard floating-point format for better dynamic range and higher accuracy, which are crucial to many applications such as neural networks. Those advantages are yielded from a varying-length segment, regime bits, which lead to the size variations for all rest components except the sign bit. Consequently, it requires an extra decoding process to extract the numerical value of a posit number. The state-of-the-art posit decoder is designed based on a leading one/zero detector. However, we find that this conventional method holds implicit redundancy when dealing with binary numbers. In this paper, we design a novel hardware architecture, i.e., the leading difference detector, to optimize the circuit operation by eliminating the redundancy. The experimental results show that the proposed architecture can decrease the delay and power consumption by over 41% compared to the conventional designs for 8-bit, 16-bit, 32-bit, and 64-bit posit decoders.
Jianchi Sun, Yingjie Lao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 SeIF: Semantic-Constrained Deep Implicit Function for Single-Image 3D Head Reconstruction
abstract
Various applications require realistic, artifact-free, and animatable 3D avatars. However, traditional 3D morphable models (3DMMs) produce animatable 3D heads but fail to capture accurate geometries and details, while existing deep implicit functions have been shown to achieve realistic reconstructions but suffer from artifacts and struggle to yield 3D heads that are easy to animate. To reconstruct high-fidelity, artifact-less, and animatable 3D heads from single-view images, we leverage semantics to bridge the best properties of 3DMMs and deep implicit functions and propose SeIF—a semantic-constrained deep implicit function. First, SeIF derives fine-grained semantics from a standard 3DMM (e.g., FLAME) and samples a semantic code for each query point in the query space to provide a soft constraint to the deep implicit function. The reconstruction results show that this semantic constraint does not weaken the powerful representation ability of the deep implicit function while significantly suppressing artifacts. Second, SeIF predicts a more accurate semantic code for each query point and utilizes the semantic codes to uniformize the structure of reconstructed 3D head meshes with the standard 3DMM. Since our reconstructed 3D head meshes have the same structure as the 3DMM, 3DMM-based animation approaches can be easily transferred to animate our reconstructed 3D heads. As a result, SeIF can reconstruct high-fidelity, artifact-less, and animatable 3D heads from single-view images of individuals with diverse ages, genders, races, and facial expressions. Quantitative and qualitative experimental results on seven datasets show that SeIF outperforms existing state-of-the-art methods by a large margin. The code and models of SeIF are available for research purposes athttps://github.com/starVisionTeam/SeIF.
Leyuan Liu 0001, Jianchi Sun, Changxin Gao, Jingying Chen 0001
IEEE Trans. Multim.3
2023 Single-image clothed 3D human reconstruction guided by a well-aligned parametric body model
Leyuan Liu 0001, Yunqi Gao, Jianchi Sun, Jingying Chen 0001
Multim. Syst.3
2022 Hardware Acceleration for Postdecision State Reinforcement Learning in IoT Systems
abstract
Reinforcement learning (RL) is increasingly being used to optimize resource-constrained wireless Internet of Things (IoT) devices. However, existing RL algorithms that are lightweight enough to be implemented on these devices, such as$Q$-learning, converge too slowly to effectively adapt to the experienced information source and channel dynamics, while deep RL algorithms are too complex to be implemented on these devices. By integrating basic models of the IoT system into the learning process, the so-called postdecision state (PDS)-based RL can achieve faster convergence speeds than these alternative approaches at lower complexity than deep RL; however, its complexity may still hinder the real-time and energy-efficient operations on IoT devices. In this article, we develop efficient hardware accelerators for PDS-based RL. We first develop an arithmetic hardware acceleration architecture and then propose a stochastic computing (SC)-based reconfigurable hardware architecture. By using simple bitwise computations enabled by SC, we eliminate costly multiplications involved in PDS learning, which simultaneously reduces the hardware area and power consumption. We show that the computational efficiency can be further improved by using extremely short stochastic representations without sacrificing learning performance. We demonstrate our proposed approach on a simulated wireless IoT sensor that must transmit delay-sensitive data over a fading channel while minimizing its energy consumption. Our experimental results show that our arithmetic accelerator is$5.3\times $faster than$Q$-learning and$2.6\times $faster than a baseline hardware architecture, while the proposed SC-based architecture further reduces the critical path of the arithmetic accelerator by 87.9%.
Jianchi Sun, Nikhilesh Sharma, Jacob Chakareski, Nicholas Mastronarde, Yingjie Lao
IEEE Internet Things J.1
2021 HEI-Human: A Hybrid Explicit and Implicit Method for Single-View 3D Clothed Human Reconstruction
Leyuan Liu 0001, Jianchi Sun, Yunqi Gao, Jingying Chen 0001
PRCV (2)2