Hui Li 0035

dblp:66/3387-35 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-9198-3951ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 MCMC with adaptive principal-component transformation: rotation-invariant universal samplers for bayesian structural system identification
Xianghao Meng, James L. Beck, Kui Jiang, Hui Li 0035
Adv. Eng. Informatics5
2026 Deep Graph Learning for Spatial-Temporal Group-Correlated In-Service Performance Prediction of Transportation Infrastructure System
Yang Xu 0053, Chenglong Lin, Yuequan Bao, Hui Li 0035
Adv. Eng. Informatics4
2026 Bridging geometry-coherent text-to-3D generation with multiview diffusion priors and Gaussian Splatting
Wenliang Qian, Wangmeng Zuo, Hui Li 0035
Neural Networks4
2026 Denoised Semantic Features for Local Consistent No-Reference Image Quality Assessment
abstract
Multi-dataset no-reference image quality assessment (NR-IQA) aims to deliver consistent image quality evaluation across a variety of contexts, empowering platform developers to optimize image processing pipelines while maintaining acceptable visual quality. Human vision, when observing images, tends to prioritize local semantics, for example, a blurry sky is perceived differently than a blurry face. This insight forms the basis of many multi-dataset NR-IQA models, which commonly rely on pretrained deep networks to extract semantic information that is crucial for assessing perceptual quality. Vision Transformer-based pre-trained models often exhibit persistent noise artifacts, as demonstrated by previous studies such as Denoising Vision Transformers; many existing IQA approaches fail to appropriately address these local semantic artifacts, leading to inconsistent local IQA score maps, even when overall performance appears satisfactory. To tackle this, we introduce DINO-IQA, a novel dual-branch network architecture designed for NR-IQA to multi-dataset. The first branch focuses on extracting local distortion features, effectively capturing image degradation, while the second branch utilizes denoised DINOv2 from ViT decomposition to extract refined semantic features, free from local artifacts. By enabling visual interaction between distortion and semantic features, our method generates locally consistent quality maps that align more closely with human perception. This approach achieves remarkable accuracy and sets a new benchmark for state-of-the-art multi-dataset NR-IQA performance. Our findings underscore the critical need to address semantic noise in pre-trained networks for enhancing NR-IQA, demonstrating that our dual-branch framework offers a robust solution to this previously underexplored challenge.
Hui Li 0035, Chaofeng Chen, Xiaopeng Fan 0001, Wangmeng Zuo, Weisi Lin
IEEE Trans. Multim.1
2025 DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors
abstract
Dynamic 3D interaction has been attracting a lot of attention recently. However, creating such 4D content remains challenging. One solution is to animate 3D scenes with physics-based simulation, which requires manually assigning precise physical properties to the object or the simulated results would become unnatural. Another solution is to learn the deformation of 3D objects with the distillation of video generative models, which, however, tends to produce 3D videos with small and discontinuous motions due to the inappropriate extraction and application of physics priors. In this work, to combine the strengths and complementing shortcomings of the above two solutions, we propose to learn the physical properties of a material field with video diffusion priors, and then utilize a physics-based Material-Point-Method (MPM) simulator to generate 4D content with realistic motions. In particular, we propose motion distillation sampling to emphasize video motion information during distillation. In addition, to facilitate the optimization, we further propose a KAN-based material field with frame boosting. Experimental results demonstrate that our method enjoys more realistic motions than state-of-the-arts do.
Haoze Zhang, Yihan Zeng, Zhilu Zhang 0001, Hui Li 0035, Wangmeng Zuo, Rynson W. H. Lau
AAAI5
2025 FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
abstract
Interactive image editing allows users to modify images through visual interaction operations such as drawing, clicking, and dragging. Existing methods construct such supervision signals from videos, as they capture how objects change with various physical interactions. However, these models are usually built upon text-to-image diffusion models, so necessitate (i) massive training samples and (ii) an additional reference encoder to learn real-world dynamics and visual consistency. In this paper, we reformulate this task as an image-to-video generation problem, so that inherit powerful video diffusion priors to reduce training costs and ensure temporal consistency. Specifically, we introduce FramePainter as an efficient instantiation of this formulation. Initialized with Stable Video Diffusion, it only uses a lightweight sparse control encoder to inject editing signals. Considering the limitations of temporal attention in handling large motion between two frames, we further propose matching attention to enlarge the receptive field while encouraging dense correspondence between edited and source image tokens. We highlight the effectiveness and efficiency of FramePainter across various of editing signals: it domainantly outperforms previous state-of-the-art methods with far less training data, achieving highly seamless and coherent editing of images, \eg, automatically adjust the reflection of the cup. Moreover, FramePainter also exhibits exceptional generalization in scenarios not present in real-world videos, \eg, transform the clownfish into shark-like shape. Our code will be available at https://github.com/YBYBZhang/FramePainter.
Yabo Zhang, Xinpeng Zhou, Yihan Zeng, Hang Xu 0004, Hui Li 0035, Wangmeng Zuo
ICCV5
2025 Lie group convolution neural networks with scale-rotation equivariance
Weidong Qiao, Yang Xu 0053, Hui Li 0035
Neural Networks3
2025 Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
abstract
Real-world applications like video gaming and virtual reality often demand the ability to model 3D scenes that users can explore along custom camera trajectories. While significant progress has been made in generating 3D objects from text or images, creating long-range, 3D-consistent, explorable 3D scenes remains a complex and challenging problem. In this work, we present Voyager , a novel video diffusion framework that generates world-consistent 3D point-cloud sequences from a single image with user-defined camera path. Unlike existing approaches, Voyager achieves end-to-end scene generation and reconstruction with inherent consistency across frames, eliminating the need for 3D reconstruction pipelines (e.g., structure-from-motion or multi-view stereo). Our method integrates three key components: 1) World-Consistent Video Diffusion : A unified architecture that jointly generates aligned RGB and depth video sequences, conditioned on existing world observation to ensure global coherence 2) Long-Range World Exploration : An efficient world cache with point culling and an auto-regressive inference with smooth video sampling for iterative scene extension with context-aware consistency, and 3) Scalable Data Engine : A video reconstruction pipeline that automates camera pose estimation and metric depth prediction for arbitrary videos, enabling large-scale, diverse training data curation without manual 3D annotations. Collectively, these designs result in a clear improvement over existing methods in visual quality and geometric accuracy, with versatile applications. Code for this paper are at https://github.com/Tencent-Hunyuan/HunyuanWorld-Voyager.
Wangguandong Zheng, Tengfei Wang 0002, Yuhao Liu 0001, Zhenwei Wang 0003, Junta Wu, Jie Jiang 0008, Hui Li 0035, Rynson W. H. Lau, Wangmeng Zuo, Chunchao Guo
ACM Trans. Graph.8
2024 Few-shot learning for structural health diagnosis of civil infrastructure
abstract
The successful development of deep learning and computer vision techniques has recently revolutionized structural health diagnosis (SHD) during life-cycle construction, inspection, maintenance, and disaster prevention for civil infrastructure. Multi-source big data of structural health monitoring (SHM) and inspection are projected into a high-level feature space, and data-driven models are established to map the relationships between inputs and outputs and dig out embedded structural behaviors and implicit physical mechanisms. However, the model performance highly relies on the extensive amount and diversity, intra-class completeness, and inter-class balance of training data, and the generalization ability on scarce data with specific features and particular patterns is challenging under real-world scenarios. To address the above challenges, few-shot learning (FSL) has emerged as a cutting-edge machine learning technique that designs training strategies using only a small amount of annotated data under a limited supervision regime to enhance effectiveness and generalization ability. This article systematically summarizes recent advances in FSL algorithms and the corresponding applications in SHD for civil infrastructure. A unified mathematical framework of FSL is formulated, and an FSL taxonomy is summarized according to intrinsic learning mechanisms and implementation principles, including metric learning-based, optimization-based, transfer learning-based, and generative model-based methods. Various applications of SHD for civil infrastructure under real-world scenarios are reviewed, including remote sensing monitoring, structural damage recognition, post-disaster safety evaluation, and construction risk assessment. Error analyses of approximation, generalization, and optimization errors corresponding to the four paradigms mentioned above are formulated for FSL-based data-driven modeling, which exactly acts as the logical connection with how new thoughts of FSL-based SHD should be designed. Finally, potential prospects of FSL-based SHD are outlined to design weakly supervised FSL considering low-quality data from SHM systems , develop multi-modal FSL for multi-source SHM data , establish cross-task FSL with high commonality and generality for various SHD tasks, and construct lightweight FSL framework with model compression for real-world applications of SHD considering the super-large stream of monitoring data.
Yang Xu 0053, Yunlei Fan, Yuequan Bao, Hui Li 0035
Adv. Eng. Informatics4
2024 Multi-dataset learning with channel modulation loss for blind image quality assessment
Hui Li 0035, Zhaoyi Yan, Xiaopeng Fan 0001
Multim. Tools Appl.1
2024 Continual Learning of No-Reference Image Quality Assessment With Channel Modulation Kernel
abstract
No-Reference Image Quality Assessment (NR-IQA), a subset of IQA techniques, is critical in scenarios where reference images are unavailable. With advancements in camera technology and computer vision, IQA datasets have evolved significantly in distortion types, image contents, and domains. This highlights the need for a broad study of NR-IQA continual learning, optimizing on a sequence of tasks, in both in-domain and domain-transfer settings. In this paper, we introduce the Channel Modulation Kernel (CMKernel) as a solution to enhance NR-IQA continual learning from two perspectives. Firstly, CMKernel encodes channel attention information for both in-domain and domain-transfer scenarios. By imposing constraints on CMKernels of successive models, the channel attention distillation loss effectively mitigates the divergence between old and new models. Secondly, in the context of the domain-transfer setting, a significant challenge lies in training a robust and transferable base model from the general domain for subsequent continual learning across specific domains. To tackle this, we introduce CMKernel-based multi-dataset learning to acquire a generative model. By dynamically weighting convolutional channels, the base model learns more equally from mixed datasets, enhancing its performance for subsequent incremental tasks. Comprehensive experiments validate the superiority of CMKernel in both in-domain and domain-transfer continual learning settings, showcasing its efficacy in addressing the evolving challenges of NR-IQA in diverse image contexts.
Hui Li 0035, Chaofeng Chen, Xiaopeng Fan 0001, Wangmeng Zuo, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1