VLDB 2026 Research / reviewers in the wild / expert
Yuhan Li 0003
dblp:116/8661-3
· DBLP profile ↗
12ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0001-7740-8202ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Gaze Synthesizer via 3D-Eye Controlled Diffusion and Cross-Domain Feature AlignmentabstractAppearance-based supervised methods with full-face image input have made tremendous advances in recent gaze estimation tasks. However, intensive human annotation requirement inhibits current methods from achieving industrial level accuracy and robustness. Although methods based on generative AI can synthesize eye images to expand self-annotated eye data, these methods usually have limited model capacity and cannot effectively inject gaze information, resulting in poor quality, monotonous texture, and inaccurate gaze direction of the generated eye images. To alleviate the above challenge, we propose a novel gaze data synthesizer framework, in which a 3D-eye model that can flexibly manipulate the gaze direction is used to finely control eye image synthesis based on a stable diffusion large generative model, so that high-quality eye images with arbitrary gaze angle can be synthesized. At the same time, when training the gaze feature extractor, we propose a cross-domain feature alignment module to minimize the feature distribution discrepancy between real samples and synthetic ones, to pursue domain-invariant (shape, texture, etc.) gaze representation. Both qualitative and quantitative experimental results demonstrate that the proposed scheme generates high-quality gaze images and also achieves superior gaze estimation performances over state-of-the-art. Muchun Chen, Yinxin Lin, Yuhan Li 0003, Yugang Chen, Xuanhong Chen, Bilian Ke, Bingbing Ni |
IEEE Trans. Image Process. | 3 |
| 2025 | SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic PriorsabstractDespite significant advances in accurately estimating geometry in contemporary single-image 3D human reconstruction, creating a high-quality, efficient, and animatable 3D avatar remains an open challenge. Two key obstacles persist: incomplete observation and inconsistent 3D priors. To address these challenges, we propose SinGS, aiming to achieve high-quality and efficient animatable 3D avatar reconstruction. At the heart of SinGS are two key components: Kinematic Human Diffusion and Geometry-Preserving 3D Gaussain Splatting. The former is a foundational human model that samples within pose space to generate a highly 3D-consistent and high-quality sequence of human images, inferring unseen viewpoints and providing kinematic priors. The latter is a system that reconstructs a compact, high-quality 3D avatar even under imperfect priors, achieved through a novel semantic Laplacian regularization and a geometry-preserving density control strategy that enable precise and compact assembly of 3D primitives. Extensive experiments demonstrate that SinGS enables lifelike, animatable human reconstructions, maintaining both high quality and inference efficiency (up to 70FPS). Xuanhong Chen, Shunran Jia, Hualiang Wei, Kairui Feng, Yuhan Li 0003, Ang He, Bingbing Ni, Wenjun Zhang 0001 |
CVPR | 8 |
| 2025 | RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation
Yuhan Li 0003, Xianfeng Tan, Wenxiang Shang, Yubo Wu, Xuanhong Chen, Hangcheng Zhu, Bingbing Ni |
ICCV | 1 |
| 2025 | BrokerAS: Towards Fault-tolerant Atomic Cross-chain Swaps
Gaowei Shi, Xiulong Liu 0001, Yuhan Li 0003, Hao Xu 0025, Keqiu Li |
INFOCOM | 4 |
| 2025 | ShoeFit: A New Dataset and Dual-image-stream DiT Framework for Virtual Footwear Try-OnabstractVirtual footwear try-on (VFTON), a critical yet underexplored area in virtual try-on (VTON), aims to synthesize faithful try-on results given diverse footwear and model images while maintaining 3D consistency and texture authenticity.
Unlike conventional garment-focused VTON methods, VFTON presents unique challenges due to (1) Data Scarcity, which arises from the difficulty of perfectly matching product shoes with models wearing the identical ones, (2) Viewpoint Misalignment, where the target foot pose and source shoe views are always misaligned, leading to incomplete texture information and detail distortion, and (3) Background-induced Color Distortion, where complex material of footwear interacts with environmental lighting, causing unintended color contamination.
To address these challenges, we introduce MVShoes, a multi-view shoe try-on dataset consisting of 7305 well-annotated image triplets, covering diverse footwear categories and challenging try-on scenarios. Furthermore, we propose a dual-stream DiT architecture, ShoeFit, designed to mitigate viewpoint misalignment through Multi-View Conditioning with 3D Rotary Position Embedding, and alleviate background-induced distortion using the LayeredRefAttention which leverages background features to modulate footwear latents. The proposed framework effectively decouples shoe appearance from environmental interferences while preserving high-quality texture detail through decoupled denoising and conditioning branches.
Extensive quantitative and qualitative experiments demonstrate that our method substantially improves rendering fidelity and robustness under challenging real-world product shoes, establishing a new benchmark in high-fidelity footwear try-on synthesis. The dataset and benchmark will be publicly available upon acceptance of the paper. Yuhan Li 0003, Zhiyu Jin, Yifan Tong, Wenxiang Shang, Benlei Cui, Xuanhong Chen, Ran Lin, Bingbing Ni |
NeurIPS | 1 |
| 2024 | FocalDreamer: Text-Driven 3D Editing via Focal-Fusion AssemblyabstractWhile text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape with editable parts according to text prompts for fine-grained editing within desired regions. Specifically, equipped with geometry union and dual-path rendering, FocalDreamer assembles independent 3D parts into a complete object, tailored for convenient instance reuse and part-wise control. We propose geometric focal loss and style consistency regularization, which encourage focal fusion and congruent overall appearance. Furthermore, FocalDreamer generates high-fidelity geometry and PBR textures which are compatible with widely-used graphics engines. Extensive experiments have highlighted the superior editing capabilities of FocalDreamer in both quantitative and qualitative evaluations. Yuhan Li 0003, Yishun Dou, Xuanhong Chen, Peng Zhou 0010, Bingbing Ni |
AAAI | 1 |
| 2024 | Differentiable Micro-Mesh ConstructionabstractMicro-mesh (μ-mesh.) is a new graphics primitive for compact representation of extreme geometry, consisting of a low-polygon base mesh enriched by per micro-vertex displacement. A new generation of GPUs supports this structure with hardware evolution on μ-mesh ray tracing, achieving real-time rendering in pixel level geometric details. In this article, we present a differentiable framework to convert standard meshes into this efficient format, offering a holistic scheme in contrast to the previous stage-based methods. In our construction context, a μ-mesh is defined where each base triangle is a parametric primitive, which is then reparameterized with Laplacian operators for efficient geometry optimization. Our framework offers numerous advantages for high-quality μ-mesh production: (i) end-to-end geometry optimization and displacement baking; (ii) enabling the differentiation of renderings with respect to μ-mesh for faithful reprojectability; (iii) high scalability for integrating useful features for μ-mesh production and rendering, such as minimizing shell volume, maintaining the isotropy of the base mesh, and visual-guided adaptive level of detail. Extensive experiments on μ-mesh construction for a large set of high-resolution meshes demonstrate the superior quality achieved by the proposed scheme. Yishun Dou, Qiaoqiao Jin, Yuhan Li 0003, Bingbing Ni |
CVPR | 5 |
| 2024 | Asynchronous Complete Secret Sharing with Linear Communication CostabstractAsynchronous Complete Secret Sharing (ACSS) in Byzantine fault-tolerant systems has become one of the essential building blocks in multiple threshold cryptosystems. However, current ACSS schemes scale poorly due to high communication costs, which are quadratic in the number of participants n. In this paper, we propose a new scheme ALCES to reduce such communication costs from O(n2) to O(cn) with a negligible probability of failure ${e^{ - \frac{c}{{18}}}}$, while guaranteeing completeness and agreement properties. The key point of ALCES is to sample c parties to construct a committee, which then verifies and distributes the encrypted shares to other parties. Additionally, we introduce a new mechanism, referred to as secret labels in ALCES, by encoding the information of labels in polynomial coefficients. This mechanism allows an arbitrary string to act as the label, binding it to a specific secret while efficiently ensuring security and privacy with minimal communication cost. Experimental results show that our technique reduces the overall communication cost in a single sharing process by 66% and 83% for very large quantities, such as 4096 and 8192 parties, respectively, when compared with prior work. Yuhan Li 0003, Xiulong Liu 0001, Gaowei Shi, Hao Xu 0025, Keqiu Li |
HPCC | 1 |
| 2024 | AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any ScenarioabstractWhile image-based virtual try-on has made significant strides, emerging approaches still fall short of delivering high-fidelity and robust fitting images across various scenarios, as their models suffer from issues of ill-fitted garment styles and quality degrading during the training process, not to mention the lack of support for various combinations of attire. Therefore, we first propose a lightweight, scalable, operator known as Hydra Block for attire combinations. This is achieved through a parallel attention mechanism that facilitates the feature injection of multiple garments from conditionally encoded branches into the main network. Secondly, to significantly enhance the model's robustness and expressiveness in real-world scenarios, we evolve its potential across diverse settings by synthesizing the residuals of multiple models, as well as implementing a mask region boost strategy to overcome the instability caused by information leakage in existing models.
Equipped with the above design, AnyFit surpasses all baselines on high-resolution benchmarks and real-world data by a large gap, excelling in producing well-fitting garments replete with photorealistic and rich details. Furthermore, AnyFit’s impressive performance on high-fidelity virtual try-ons in any scenario from any image, paves a new path for future research within the fashion community. Yuhan Li 0003, Wenxiang Shang, Ran Lin, Xuanhong Chen, Bingbing Ni |
NeurIPS | 1 |
| 2024 | GFBE: A Generalized and Fine-Grained Blockchain Evaluation FrameworkabstractMulti-dimensional performance evaluation is crucial for blockchain systems as it enables appropriate blockchain choosing for a given scenario and helps to pinpoint the bottleneck module of a blockchain system to optimize its performance. However, the existing evaluation frameworks for blockchain suffer from low system generality, inefficient workload execution, and incomprehensible evaluation metrics. In order to overcome their limitations, we design and implement the Generalized and Fine-grained Blockchain Evaluation (GFBE) framework. Specifically, we abstract 3 types of Universal Evaluation Interface (UEI) via the dynamic proxying approach to enable generalized evaluation of heterogeneous blockchain systems. Through the design of Lua-based workloads plugin with high flexibility and reusability, GFBE improves the efficiency of workload execution. To achieve comprehensive measurement, we define 15 key performance metrics across hierarchical layers of blockchain architecture. We also implement and deploy GFBE on 16 machines each with 8 CPUs and 16GB RAM, and evaluate three open-source blockchain systems namely Ethereum, ChainMaker, and Haihe smart chain. The experimental results demonstrate that GFBE efficiently and accurately measure 15 key performance metrics such as Contract Execution Efficiency at the contract layer, Consensus Agreement Time Ratio at the consensus layer, and State Query Time at the data layer. Compared with state-of-the-art frameworks such as BLOCKBENCH, Log-based, and Caliper, GFBE distinguishes itself as the only framework that encompasses the appealing features of universal interface, reusable workload, and all-layer metrics. Xiulong Liu 0001, Yuhan Li 0003, Chenyu Zhang 0008, Gaowei Shi, Keqiu Li |
IEEE Trans. Computers | 3 |
| 2023 | Generalized Deep 3D Shape Prior via Part-Discretized Diffusion ProcessabstractWe develop a generalized 3D shape generation prior model, tailored for multiple 3D tasks including unconditional shape generation, point cloud completion, and cross-modality shape generation, etc. On one hand, to precisely capture local fine detailed shape information, a vector quantized variational autoencoder (VQ-VAE) is utilized to index local geometry from a compactly learned code-book based on a broad set of task training data. On the other hand, a discrete diffusion generator is introduced to model the inherent structural dependencies among different tokens. In the meantime, a multi-frequency fusion module (MFM) is developed to suppress high-frequency shape feature fluctuations, guided by multi-frequency contextual information. The above designs jointly equip our proposed 3D shape prior model with high-fidelity, diverse features as well as the capability of cross-modality alignment, and extensive experiments have demonstrated superior performances on various 3D shape generation tasks. Yuhan Li 0003, Yishun Dou, Xuanhong Chen, Bingbing Ni, Yilin Sun, Yutian Liu 0004, Fuzhen Wang |
CVPR | 1 |
| 2019 | PalmGAN for Cross-Domain Palmprint RecognitionabstractNowadays, many efficient palmprint recognition algorithms have emerged. However, previous algorithms can only be used in a single domain. Furthermore, they also require a large amount of labeled data, which is difficult and costly to obtain. In order to solve these problems, we proposed PalmGAN for cross-domain palmprint recognition. Firstly, the labeled fake images were generated to reduce domain gaps, whose styles are similar to the target domain, and at the same time, the identity information remains unchanged. Based on these fake images, supervised Deep Hash Network (DHN) can be trained and directly used for unsupervised identification in the target domain. Moreover, we established semi-uncontrolled and uncontrolled databases, which were collected in uncontrolled environments. Experiments on several popular databases and self-built databases obtained satisfactory performances. PalmGAN can effectively achieve up to 5.08% improvement for cross-domain recognition, and Equal Error Rate (EER) can decrease to 0% for cross-domain recognition between Blue and Green databases. Huikai Shao, Dexing Zhong, Yuhan Li 0003 |
ICME | 3 |