Hongwei Hu

dblp:85/7623 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Text-Driven Relation Manipulation of Diffusion Imagery
abstract
Text-guided image manipulation has recently attracted significant attention. Prevailing algorithms predominantly focus on modifying the appearances of existing instances, such as texture and attribute editing, while they often fail to address the interactions between different instances or achieve fundamental structural changes, such as multi-object editing. This paper introduces a novel text-guided manipulation task named "relation manipulation", aimed at fundamentally altering the structure of images. This task is capable of modifying the quantity of instances and, more importantly, enhancing the understanding and editing of interactions among diverse instances. Our approach comprises two main components: relation customization and multi-region guided diffusion. Relation customization fine-tunes specific relationships using a compact dataset of exemplary relations, facilitating nuanced understanding and implementation of instance interactions. Multi-region guided diffusion employs gradient optimization to update the generation process across multiple regions, integrating a fine-grained attention control strategy to minimize regional interference and conflict. Additionally, the demonstrated applications of our method in multi-region inversion underline its potential in practical scenarios, such as relation manipulation of real images and consecutive image manipulation. Compatible with different variants of Stable Diffusion models, our approach seamlessly integrates into the Stable Diffusion WebUI, enabling high-quality image generation and exceptional control over extensive manipulation. This makes it a robust tool for both academic research and creative industries. Code is available at https://github.com/liyiming09/RMD.
Peng Zhou 0010, Hongwei Hu, Xiaokang Qin, Jun Sun 0005, Yi Xu 0001
IEEE Trans. Image Process.3
2026 vSoC: Efficient and Debug-Friendly Virtual System-on-Chip for Mobile Emulation
abstract
Emerging heavy-load mobile apps like UHD video and AR/VR access diverse high-throughput hardware devices, e.g., video codecs and cameras. However, today's mobile emulators exhibit poor performance when emulating these devices. We pinpoint the major reason to be the discrepancy between the guest's (system-on-chip) and host's (PC or cloud server) memory architectures for these devices, which makes the shared virtual memory (SVM) architecture of mobile emulators highly inefficient. To address this, we introduce vSoC, the first virtual mobile SoC featuring aunifiedSVM framework that enables efficient and secure data sharing among virtual devices, as well as anintelligentprefetch engine that effectively eliminates the vast majority of coherence maintenance overhead. While vSoC addresses runtime performance issues, app developers face a severe debugging challenge due to the inability of traditional tools to capture complete system states. Therefore, we devise an adaptive VM (virtual machine) snapshot-based approach that dynamically selects the optimal resource loading strategy to make vSoC debug-friendly. Compared to state-of-the-art emulators, vSoC brings 1.8–9.0$\times$frame rates, 35%-62% lower motion-tophoton latency, and 6.5–14.4$\times$bug reproduction rates for heavyload apps. It is applicable to a variety of scenarios like end-user high-performance emulation and cloud/web-based rendering.
Jiaxing Qiu, Zhenhua Li 0001, Feng Qian 0001, Yunhao Liu 0001, Hongwei Hu
IEEE Trans. Mob. Comput.7
2025 Linear Attention Modeling for Learned Image Compression
abstract
Recent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low complexity design. In this paper, we propose LALIC, a linear attention modeling for learned image compression. Specially, we propose to use Bi-RWKV blocks, by utilizing the Spatial Mix and Channel Mix modules to achieve more compact feature extraction, and apply the Conv based Omni-Shift module to adapt to two-dimensional latent representation. Furthermore, we propose a RWKV-based Spatial-Channel ConTeXt model (RWKV-SCCTX), that leverages the Bi-RWKV to modeling the correlation between neighboring features effectively. To our knowledge, our work is the first work to utilize efficient Bi-RWKV models with linear attention for learned image compression. Experimental results demonstrate that our method achieves competitive RD performances by outperforming VTM-9.1 by -15.26%, -15.41%, -17.63% in BD-rate on Kodak, CLIC and Tecnick datasets. The code is available at https://github.com/sjtu-medialab/RwkvCompress.
Donghui Feng 0003, Zhengxue Cheng, Shen Wang 0013, Ronghua Wu, Hongwei Hu, Guo Lu, Li Song 0001
CVPR5
2025 Position-LoRA: Enhanced Relation Customization through Structural Prior in Initial Latent Noise
abstract
Recent advancements in concept customization via diffusion models have significantly enhanced controllability and quality. However, precise relation customization, which controls the position of interactions among multiple instances, remains challenging due to unpredictable initial latent noise. Existing methods primarily rely on conditional prompts and attention control, overlooking the structured potential of initial noise. This paper introduces Position-LoRA, a novel framework leveraging structural prior in initial noise to improve relation customization and layout control. Position-LoRA employs a differential fine-tuning scheme and a latent noise encoder. The guided fine-tuning enhances generation tendencies from structured initial noise, embedding explicit relationship-specific spatial information. The latent noise encoder dynamically manipulates latent noises, enabling precise spatial control and flexibility in relational image generation. Furthermore, a fine-grained guidance and control strategy is employed during generation to enhance the image-text alignment and layout alignment. Experiments demonstrate that Position-LoRA improves stability, controllability, and fidelity in relational image generation with layout control, surpassing existing concept customization and layout-to-image methods in qualitative and quantitative evaluations. Code is available at https://github.com/liyiming09/Position-LoRA.
Peng Zhou 0010, Xiaokang Qin, Hongwei Hu, Jun Sun 0005, Yi Xu 0001
ACM Multimedia4
2025 Cross-Modal Collaborative Recovery: Breaking Multimodal Unlearnable Examples Protection
abstract
With the rapid development of multimodal machine learning, data security has become a critical bottleneck constraining multimodal artificial intelligence advancement. Multimodal unlearnable examples, as an emerging data protection paradigm, prevent unauthorized model training by disrupting cross-modal semantic consistency through adversarial perturbations and trigger identifiers. However, existing multimodal protection mechanisms suffer from significant robustness deficiencies and security vulnerabilities. This paper conducts a systematic security evaluation of multimodal unlearnable examples and proposes a Cross-Modal Collaborative Recovery (CMCR) attack framework. CMCR integrates Joint Conditional Diffusion Purification for image denoising, Adaptive Text Purification for trigger removal, and Visual Cross-modal BART for cross-modal reconstruction. Experiments demonstrate that CMCR successfully restores image-text retrieval performance to levels approaching clean data across different model architectures, achieving an average recovery rate of 89.0%. Our results reveal that current multimodal unlearnable examples have security vulnerabilities and insufficient robustness against coordinated cross-modal attacks.
Ruijia Li, Zijiao Zhang, Ximin Huang, Hongwei Hu, Xining Gao
TrustCom4
2025 Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video Conference
abstract
The demand for immersive and interactive communication has driven advancements in 3D video conferencing, yet achieving high-fidelity 3D talking face representation at low bitrates remains a challenge. Traditional 2D video compression techniques fail to preserve fine-grained geometric and appearance details, while implicit neural rendering methods like NeRF suffer from prohibitive computational costs. To address these challenges, we propose a lightweight, high-fidelity, low-bitrate 3D talking face compression framework that integrates FLAME-based parametric modeling with 3DGS neural rendering. Our approach transmits only essential facial metadata in real time, enabling efficient reconstruction with a Gaussian-based head model. Additionally, we introduce a compact representation and compression scheme, including Gaussian attribute compression and MLP optimization, to enhance transmission efficiency. Experimental results demonstrate that our method achieves superior rate-distortion performance, delivering high-quality facial rendering at extremely low bitrates, making it well-suited for real-time 3D video conferencing applications.
Jianglong Li, Bingcong Lu, Zhengxue Cheng, Hongwei Hu, Ronghua Wu, Li Song 0001
VCIP5
2023 Novel motor fault detection scheme based on one-class tensor hyperdisk
Yuting Zeng, Haidong Shao, Hongwei Hu, Xiaoqiang Xu
Knowl. Based Syst.4
2023 EEG pattern identification for motor imagery based on 1DCNN-GRU
Hongwei Hu, Guangxu Li, Zixi Chang
Multim. Tools Appl.3
2020 Regularized matrix completion with partial side information
Kefu Yi, Hongwei Hu, Yang Yu 0002, Wei Hao 0002
Neurocomputing2
2019 Robust Object Tracking Using Manifold Regularized Convolutional Neural Networks
abstract
In visual tracking, usually only a small number of samples are labeled, and most existing deep learning based trackers ignore abundant unlabeled samples that could provide additional information for deep trackers to boost their tracking performance. An intuitive way to explain unlabeled data is to incorporate manifold regularization into the common classification loss functions, but the high computational cost may prohibit those deep trackers from practical applications. To overcome this issue, we propose a two-stage approach to a deep tracker that takes into account both labeled and unlabeled samples. The annotation of unlabeled samples is propagated from its labeled neighbors first by exploring the manifold space that these samples are assumed to lie in. Then, we refine it by training a deep convolutional neural network using both labeled and unlabeled data in a supervised manner. Online visual tracking is further carried out under the framework of particle filters with the presented manifold regularized deep model being updated every few frames. Experimental results on different tracking datasets demonstrate that our tracker outperforms most existing tracking approaches. The source code and results are available at: https://github.com/shenjianbing/MRCNNTracking.
Hongwei Hu, Bo Ma 0001, Jianbing Shen, Hanqiu Sun, Ling Shao 0001, Fatih Porikli
IEEE Trans. Multim.1
2018 Manifold Regularized Correlation Object Tracking
abstract
In this paper, we propose a manifold regularized correlation tracking method with augmented samples. To make better use of the unlabeled data and the manifold structure of the sample space, a manifold regularization-based correlation filter is introduced, which aims to assign similar labels to neighbor samples. Meanwhile, the regression model is learned by exploiting the block-circulant structure of matrices resulting from the augmented translated samples over multiple base samples cropped from both target and nontarget regions. Thus, the final classifier in our method is trained with positive, negative, and unlabeled base samples, which is a semisupervised learning framework. A block optimization strategy is further introduced to learn a manifold regularization-based correlation filter for efficient online tracking. Experiments on two public tracking data sets demonstrate the superior performance of our tracker compared with the state-of-the-art tracking approaches.
Hongwei Hu, Bo Ma 0001, Jianbing Shen, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Robust Object Tracking by Nonlinear Learning
abstract
We propose a method that obtains a discriminative visual dictionary and a nonlinear classifier for visual tracking tasks in a sparse coding manner based on the globally linear approximation for a nonlinear learning theory. Traditional discriminative tracking methods based on sparse representation learn a dictionary in an unsupervised way and then train a classifier, which may not generate both descriptive and discriminative models for targets by treating dictionary learning and classifier learning separately. In contrast, the proposed tracking approach can construct a dictionary that fully reflects the intrinsic manifold structure of visual data and introduces more discriminative ability in a unified learning framework. Finally, an iterative optimization approach, which computes the optimal dictionary, the associated sparse coding, and a classifier, is introduced. Experiments on two benchmarks show that our tracker achieves a better performance compared with some popular tracking algorithms.
Bo Ma 0001, Hongwei Hu, Jianbing Shen, Ling Shao 0001, Fatih Porikli
IEEE Trans. Neural Networks Learn. Syst.2
2016 Randomized Canonical Correlation Discriminant Analysis for Face Recognition
abstract
As an important technique in multivariate statistical analysis, Canonical Correlation Analysis (CCA) has been widely used in face recognition. But existing CCA based face recognition methods need two kinds of expression for the same face sample, and usually suffers high computational complexity in dealing with large samples. In this paper, we present a supervised method called Randomized Canonical Correlation Discriminant Analysis (RCCDA) based on Randomized non-linear Canonical Correlation Analysis (RCCA) to make up for the shortage of CCA based face recognition methods. We first obtain basis vectors approximately with random features instead of the calculation of kernel matrix to improve the efficiency of computation, then we use these basis vectors to compute random optimal discriminant features which can reduce the dimension of face features while preserving as much discriminatory information as possible. The result of experiments on Extended Yale B, AR, ORL and FERET face databases demonstrates that the performance of our method compares favorably with some state-of-the-art algorithms.
Bo Ma 0001, Hongwei Hu, Meili Wei
ECAI3
2016 Generalized Pooling for Robust Object Tracking
abstract
Feature pooling in a majority of sparse coding-based tracking algorithms computes final feature vectors only by low-order statistics or extreme responses of sparse codes. The high-order statistics and the correlations between responses to different dictionary items are neglected. We present a more generalized feature pooling method for visual tracking by utilizing the probabilistic function to model the statistical distribution of sparse codes. Since immediate matching between two distributions usually requires high computational costs, we introduce the Fisher vector to derive a more compact and discriminative representation for sparse codes of the visual target. We encode target patches by local coordinate coding, utilize Gaussian mixture model to compute Fisher vectors, and finally train semi-supervised linear kernel classifiers for visual tracking. In order to handle the drifting problem during the tracking process, these classifiers are updated online with current tracking results. The experimental results on two challenging tracking benchmarks demonstrate that the proposed approach achieves a better performance than the state-of-the-art tracking algorithms.
Bo Ma 0001, Hongwei Hu, Jianbing Shen, Yangbiao Liu, Ling Shao 0001
IEEE Trans. Image Process.2
2015 Linearization to Nonlinear Learning for Visual Tracking
abstract
Due to unavoidable appearance variations caused by occlusion, deformation, and other factors, classifiers for visual tracking are nonlinear as a necessity. Building on the theory of globally linear approximations to nonlinear functions, we introduce an elegant method that jointly learns a nonlinear classifier and a visual dictionary for tracking objects in a semi-supervised sparse coding fashion. This establishes an obvious distinction from conventional sparse coding based discriminative tracking algorithms that usually maintain two-stage learning strategies, i.e., learning a dictionary in an unsupervised way then followed by training a classifier. However, the treating dictionary learning and classifier training as separate stages may not produce both descriptive and discriminative models for objects. By contrast, our method is capable of constructing a dictionary that not only fully reflects the intrinsic manifold structure of the data, but also possesses discriminative power. This paper presents an optimization method to obtain such an optimal dictionary, associated sparse coding, and a classifier in an iterative process. Our experiments on a benchmark show our tracker attains outstanding performance compared with the state-of-the-art algorithms.
Bo Ma 0001, Hongwei Hu, Jianbing Shen, Fatih Porikli
ICCV2
2015 Transfer Metric Learning for Kinship Verification with Locality-Constrained Sparse Features
Bo Ma 0001, Lianghua Huang, Hongwei Hu
ICONIP (1)4
2015 Multi-task l0 gradient minimization for visual tracking
Hongwei Hu, Bo Ma 0001, Yunde Jia
Neurocomputing1
2015 Visual Tracking Using Strong Classifier and Structural Local Sparse Descriptors
abstract
Sparse coding methods have achieved great success in visual tracking, and we present a strong classifier and structural local sparse descriptors for robust visual tracking. Since the summary features considering the sparse codes are sensitive to occlusion and other interfering factors, we extract local sparse descriptors from a fraction of all patches by performing a pooling operation. The collection of local sparse descriptors is combined into a boosting-based strong classifier for robust visual tracking using a discriminative appearance model. Furthermore, a structural reconstruction error based weight computation method is proposed to adjust the classification score of each candidate for more precise tracking results. To handle appearance changes during tracking, we present an occlusion-aware template update scheme. Comprehensive experimental comparisons with the state-of-the-art algorithms demonstrated the better performance of the proposed method.
Bo Ma 0001, Jianbing Shen, Yangbiao Liu, Hongwei Hu, Ling Shao 0001, Xuelong Li 0001
IEEE Trans. Multim.4
2014 Boosting-Based Visual Tracking Using Structural Local Sparse Descriptors
Yangbiao Liu, Bo Ma 0001, Hongwei Hu, Yin Han
ACCV (5)3
2014 Nonlinear learning using LCC for online visual tracking
abstract
In this paper, we propose to address online visual tracking on the basis of Local Coordinate Coding (LCC), which integrates the advantages of the discriminative method and the generative method. In the discriminative module, a nonlinear function is trained using the local coordinate codes of image patches to identify the foreground patches from background. In the generative module, we introduce a similarity function that takes the spatial structures of local patches in the target into account between the candidate and holistic templates by reconstruction error. To deal with appearance change during tracking, an online update method is introduced. The proposed tracking method is evaluated on different challenging video sequences with center location error, and experimental results demonstrate the good performance of our method.
Hongwei Hu, Bo Ma 0001, Junbiao Pang
ICME1
2013 Covariance based local salient descriptors for visual tracking
abstract
When visual tracking is performed by human, we typically pay attention to some salient regions or points of the target instead of the whole target. Inspired by this visual saliency property of human visual system, the paper proposes a novel salient regions extraction method to model target appearance. In order to capture the salient and spatial information within this model, the method extracts a set of local salient descriptors based on covariance features from the target. Afterwards, an optimization problem is constructed with respect to the features of these salient regions, and the optimal target state is obtained by solving this problem using a gradient descent algorithm. Experiments on several challenging video sequences demonstrate the good performance of the proposed method compared with four state-of-art tracking methods.
Hongwei Hu, Bo Ma 0001, Qiaofeng Ma
ICME1
2013 Robust Visual Tracking Using Local Sparse Covariance Descriptor and Matching Pursuit
Bo Ma 0001, Hongwei Hu, Jianglong Chen
ICONIP (3)2
2013 PCA-Based Appearance Template Learning for Contour Tracking
Bo Ma 0001, Hongwei Hu, Yin Han
ICONIP (3)2
2010 Improved modelling of speech dynamics using non-linear formant trajectories for HMM-based speech synthesis
Hongwei Hu, Martin J. Russell
INTERSPEECH1
2008 Speech recognition using non-linear trajectories in a formant-based articulatory layer of a multiple-level segmental HMM
Hongwei Hu, Martin J. Russell
INTERSPEECH1