Haoji Hu

dblp:65/11145 · also Haoji Roland Hu, Roland Hu · DBLP profile ↗
← Back
85ranked-venue papers
10as first author
45since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 2 first-author · 24 since 2021Artificial intelligence and machine learning · 40 · 3 first-author · 29 since 2021Databases, data management, data science and information retrieval · 8 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Q Cache: Visual Attention Is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
abstract
Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual tokens engenders a substantial computational load and key-value (KV) cache footprint bottleneck. Existing approaches focus on token-wise optimization, leveraging diverse intricate token pruning techniques to eliminate non-crucial visual tokens. Nevertheless, these methods often unavoidably undermine the integrity of the KV cache, resulting in failures in long-text generation tasks. To this end, we conduct an in-depth investigation towards the attention mechanism of the model from a new perspective, and discern that attention within more than half of all decode layers are semantic similar. Upon this finding, we contend that the attention in certain layers can be streamlined by inheriting the attention from their preceding layers. Consequently, we propose Lazy Attention, an efficient attention mechanism that enables cross-layer sharing of similar attention patterns. It ingeniously reduces layer-wise redundant computation in attention. In Lazy Attention, we develop a novel layer-shared cache, Q Cache, tailored for MLLMs, which facilitates the reuse of queries across adjacent layers. In particular, Q Cache is lightweight and fully compatible with existing inference frameworks, including Flash Attention and KV cache. Additionally, our method is highly flexible as it is orthogonal to existing token-wise techniques and can be deployed independently or combined with token pruning approaches. Empirical evaluations on multiple benchmarks demonstrate that our method can reduce KV cache usage by over 35% and achieve 1.5x throughput improvement, while sacrificing only approximately 1% of performance on various MLLMs. Compared with SOTA token-wise methods, our technique achieves superior accuracy preservation.
Jiedong Zhuang, Haoji Hu
AAAI7
2026 Enhanced 3D tumor synthesis and segmentation framework using multiscale diffusion and hybrid SAM-Swin models across diverse anatomical datasets
Jincao Yao, Mudassar Ali 0002, Wenjie Zheng 0003, Jiaqi Hu 0005, Jing Wang 0234, Xingze Zou, Haoji Hu, Weizeng Zheng, Neng Jin, Dong Xu 0006
Neurocomputing7
2026 SyncQ-DiT: Synchronizing activation dynamics for quantizing diffusion transformers
Wenjie Zheng 0003, Haoji Hu, Xingze Zou, Jing Wang 0234, Lianrui Mu
Neurocomputing2
2026 PAGS-relight: Position-aware Gaussian Splatting scene relighting with multi-view diffusion models
Jiangnan Ye 0002, Jiedong Zhuang, Jianhong Bai, Lianrui Mu, Wuhao Tan, Syed Abdul Rahman Abu-Bakar, Haoji Hu
Pattern Recognit.8
2025 ST3: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming
abstract
Multimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual tokens incurs a significant computational cost. Existing analysis of the MLLM attention mechanisms remains shallow, leading to coarse-grain token pruning strategies that fail to effectively balance speed and accuracy. In this paper, we conduct a comprehensive investigation of MLLM attention mechanisms with LLaVA. We find that numerous visual tokens and partial attention computations are redundant during the decoding process. Based on this insight, we propose Spatial-Temporal Visual Token Trimming (ST3), a framework designed to accelerate MLLM inference without retraining. ST3 consists of two primary components: 1) Progressive Visual Token Pruning (PVTP), which eliminates inattentive visual tokens across layers, and 2) Visual Token Annealing (VTA), which dynamically reduces the number of visual tokens in each layer as the generated tokens grow. Together, these techniques deliver around 2x faster inference with only about 30% KV cache memory compared to the original LLaVA, while maintaining consistent performance across various datasets. Crucially, ST3 can be seamlessly integrated into existing pre-trained MLLMs, providing a plug-and-play solution for efficient inference.
Jiedong Zhuang, Haoji Hu
AAAI7
2025 UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation
abstract
Photorealistic 3D vehicle models with high controllability are essential for autonomous driving simulation and data augmentation. While handcrafted CAD models provide flexible controllability, free CAD libraries often lack the high-quality materials necessary for photorealistic rendering. Conversely, reconstructed 3D models offer high-fidelity rendering but lack controllability. In this work, we introduce UrbanCAD, a framework that generates highly controllable and photorealistic 3D vehicle digital twins from a single urban image, leveraging a large collection of free 3D CAD models and handcrafted materials. To achieve this, we propose a novel pipeline that follows a retrieval-optimization manner, adapting to observational data while preserving fine-grained expert-designed priors for both geometry and material. This enables vehicles’ realistic 360° rendering, background insertion, material transfer, relighting, and component manipulation. Furthermore, given multi-view background perspective and fisheye images, we approximate environment lighting using fisheye images and reconstruct the background with 3DGS, enabling the photorealistic insertion of optimized CAD models into rendered novel view backgrounds. Experimental results demonstrate that UrbanCAD outperforms baselines in terms of photorealism. Additionally, we show that various perception models maintain their accuracy when evaluated on UrbanCAD with in-distribution configurations but degrade when applied to realistic out-of-distribution data generated by our method. This suggests that UrbanCAD is a significant advancement in creating photorealistic, safety-critical driving scenarios for downstream applications.
Yichong Lu, Yichi Cai, Shangzhan Zhang, Haoji Hu, Andreas Geiger 0001, Yiyi Liao
CVPR5
2025 Recammaster: Camera-Controlled Generative Rendering From a Single Video
abstract
Camera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the extra constraints of maintaining multiple-frame appearance and dynamic synchronization. To address this, we present ReCamMaster, a camera-controlled generative video re-rendering framework that reproduces the dynamic scene of an input video at novel camera trajectories. The core innovation lies in harnessing the generative capabilities of pre-trained text-to-video models through a simple yet powerful video conditioning mechanism--its capability is often overlooked in current research. To overcome the scarcity of qualified training data, we construct a comprehensive multi-camera synchronized video dataset using Unreal Engine 5, which is carefully curated to follow real-world filming characteristics, covering diverse scenes and camera movements. It helps the model generalize to in-the-wild videos. Lastly, we further improve the robustness to diverse inputs through a meticulously designed training strategy. Extensive experiments show that our method substantially outperforms existing state-of-the-art approaches. Our method also finds promising applications in video stabilization, super-resolution, and outpainting. Our code and dataset are publicly available at: https://github.com/KwaiVGI/ReCamMaster.
Jianhong Bai, Menghan Xia, Xintao Wang 0002, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan 0001, Di Zhang 0026
ICCV8
2025 SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
abstract
Recent advancements in video diffusion models demonstrate remarkable capabilities in simulating real-world dynamics and 3D consistency. This progress motivates us to explore the potential of these models to maintain dynamic consistency across diverse viewpoints, a feature highly sought after in applications like virtual filming. Unlike existing methods focused on multi-view generation of single objects for 4D reconstruction, our interest lies in generating open-world videos from arbitrary viewpoints, incorporating six degrees of freedom (6 DoF) camera poses. To achieve this, we propose a plug-and-play module that enhances a pre-trained text-to-video model for multi-camera video generation, ensuring consistent content across different viewpoints. Specifically, we introduce a multi-view synchronization module designed to maintain appearance and geometry consistency across these viewpoints. Given the scarcity of high-quality training data, we also propose a progressive training scheme that leverages multi-camera images and monocular videos as a supplement to Unreal Engine-rendered multi-camera videos. This comprehensive approach significantly benefits our model. Experimental results demonstrate the superiority of our proposed method over existing competitors and several baselines. Furthermore, our method enables intriguing extensions, such as re-rendering a video from multiple novel viewpoints. Project webpage: https://jianhongbai.github.io/SynCamMaster/
Jianhong Bai, Menghan Xia, Xintao Wang 0002, Ziyang Yuan, Zuozhu Liu, Haoji Hu, Pengfei Wan 0001, Di Zhang 0026
ICLR6
2025 RAG with Visual Alert: Boosting Multimodal Language Models for Enhanced Visual Question Answering
Hongze Ou, Lianrui Mu, Haoji Hu
KSEM (5)4
2025 UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
abstract
Recent advances in text-guided video editing have showcased promising results in appearance editing (e.g., stylization). However, video motion editing in the temporal dimension (e.g., from eating to waving), which distinguishes video editing from image editing, is underexplored. In this work, we present UniEdit, a tuning-free framework that supports both video motion and appearance editing by harnessing the power of a pre-trained text-to-video generator within an inversion-then-generation framework. To realize motion editing while preserving source video content, based on the insights that temporal and spatial self-attention layers encode inter-frame and intra-frame dependency, we introduce auxiliary motion-reference and reconstruction branches to produce text-guided motion and source features respectively. The obtained features are then injected into the main editing path via temporal and spatial self-attention layers. We also validate the effectiveness and flexibility of UniEdit by deploying it on three T2V generative models with different architectures. Experiments demonstrate that UniEdit covers video motion editing and various appearance editing scenarios, and surpasses the state-of-the-art methods. Our code is publicly available.
Jianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo, Haoji Hu, Zuozhu Liu, Jiang Bian 0002
ACM Multimedia5
2025 Orientation Matters: Making 3D Generative Models Orientation-Aligned
abstract
Humans intuitively perceive object shape and orientation from a single image, guided by strong priors about canonical poses. However, existing 3D generative models often produce misaligned results due to inconsistent training data, limiting their usability in downstream tasks. To address this gap, we introduce the task of orientation-aligned 3D object generation: producing 3D objects from single images with consistent orientations across categories. To facilitate this, we construct Objaverse-OA, a dataset of 14,832 orientation-aligned 3D models spanning 1,008 categories. Leveraging Objaverse-OA, we fine-tune two representative 3D generative models based on multi-view diffusion and 3D variational autoencoder frameworks to produce aligned objects that generalize well to unseen objects across various categories. Experimental results demonstrate the superiority of our method over post-hoc alignment approaches. Furthermore, we showcase downstream applications enabled by our aligned object generation, including zero-shot object orientation estimation via analysis-by-synthesis and efficient arrow-based object rotation manipulation.
Yichong Lu, Yuzhuo Tian, Zijin Jiang, Hao Ouyang, Haoji Hu, Yujun Shen, Yiyi Liao
NeurIPS7
2025 LTB-Solver: Long-tailed Bias Solver for image synthesis of diffusion models
abstract
Though diffusion models have shown the merits of generating high-quality visual data while preserving better diversity in recent studies, they do not generalize well on long-tailed datasets due to the minority classes lacking of diversity and semantic information . To overcome the aforementioned challenges, we first take a closer look at the collapse of tail category patterns under long-tail distributed data and propose an alternative but easy-to-use and effective solution, a L ong- T ailed B ias Solver in diffusion model image synthesis ( LTB-Solver ), which thereby enhances the overall diversity and quality of synthetic samples building upon the properties of the long-tailed distribution training data. Especially, we extract rich generative distribution knowledge of ‘head’ categories within proxy model and transfer the head-tail consistency distance to ‘tail’ categories, enabling the target diffusion model to learn diverse generation preserving inter-sample variation during the diffusion training process. Moreover, we incorporate the minority guidance loss function that better aligns training objectives with sampling behaviors and adjust the loss values for different classes by multiplying them with different weights. Extensive experiments are conducted on various datasets and several state-of-the-art diffusion model frameworks to verify the effectiveness of the proposed method. The results show that our method significantly improves the performance of diffusion models on long-tailed datasets by a large margin.
Siming Fu, Xiaoxuan He, Haoji Hu
Neurocomputing3
2025 PHiD: Preserving human identity in pose-guided character animation
Wenjie Zheng 0003, Xingze Zou, Lianrui Mu, Jing Wang 0234, Jiaqi Hu 0005, Jiangnan Ye 0002, Jiedong Zhuang, Mudassar Ali 0002, Olumayowa Idowu, Haoji Hu
Neurocomputing10
2025 SemiGMMPoint: Semi-supervised point cloud segmentation based on Gaussian mixture models
Xianwei Zhuang, Hualiang Wang, Xiaoxuan He, Siming Fu, Haoji Hu
Pattern Recognit.5
2025 Segmentation of MRI tumors and pelvic anatomy via cGAN-synthesized data and attention-enhanced U-Net
Mudassar Ali 0002, Haoji Hu, Maryam Mansoor, Weizeng Zheng, Neng Jin
Pattern Recognit. Lett.2
2025 L2H-NeRF: low- to high-frequency-guided NeRF for 3D reconstruction with a few input scenes
Taoqi Bao, Jiangnan Ye 0002, Zhankong Bao, Chee Siang Leow, Haoji Hu, Issei Fujishiro
Vis. Comput.5
2024 Robustness-Guided Image Synthesis for Data-Free Quantization
abstract
Quantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternative to real training data. Existing methods use classification loss to ensure the reliability of the synthesized images. Unfortunately, even if these images are well-classified by the pre-trained model, they still suffer from low semantics and homogenization issues. Intuitively, these low-semantic images are sensitive to perturbations, and the pre-trained model tends to have inconsistent output when the generator synthesizes an image with low semantics. To this end, we propose Robustness-Guided Image Synthesis (RIS), a simple but effective method to enrich the semantics of synthetic images and improve image diversity, further boosting the performance of data-free compression tasks. Concretely, we first introduce perturbations on input and model weight, then define the inconsistency metrics at feature and prediction levels before and after perturbations. On the basis of inconsistency on two levels, we design a robustness optimization objective to eliminate low-semantic images. Moreover, we also make our approach diversity-aware by forcing the generator to synthesize images with small correlations. With RIS, we achieve state-of-the-art performance for various settings on data-free quantization and can be extended to other data-free compression tasks.
Jianhong Bai, Huanpeng Chu, Hualiang Wang, Zuozhu Liu, Ruizhe Chen, Xiaoxuan He, Lianrui Mu, Chengfei Cai, Haoji Hu
AAAI10
2024 Unified Medical Image Pre-training in Language-Guided Common Semantic Space
Xiaoxuan He, Yifan Yang 0004, Xinyang Jiang, Xufang Luo, Haoji Hu, Siyun Zhao, Dongsheng Li 0002, Yuqing Yang 0001, Lili Qiu
ECCV (81)5
2024 FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance
Jiedong Zhuang, Jiaqi Hu 0005, Lianrui Mu, Jiangnan Ye 0002, Haoji Hu
ECCV (10)7
2024 Perceptual Image Compression with Text-Guided Multi-level Fusion
Jiaqi Hu 0005, Jiedong Zhuang, Lu Yu 0003, Haoji Hu
PRCV (5)6
2024 Mitigating Hallucination in Visual-Language Models via Re-balancing Contrastive Decoding
Jiayuan Yu, Lianrui Mu, Jiedong Zhuang, Jiaqi Hu 0005, Jiangnan Ye 0002, Haoji Hu
PRCV (5)10
2024 Trans-DONeRF for Transparent Object Rendering with Mixed Depth Prior
Jiangnan Ye 0002, Taoqi Bao, Lianrui Mu, Jiedong Zhuang, Haoji Hu
PRCV (6)8
2024 Zero-Shot Referring Image Segmentation with Hierarchical Prompts and Frequency Domain Fusion
Jiedong Zhuang, Jiaqi Hu 0005, Haoji Hu
PRICAI (4)4
2024 Data-Free Quantization of Vision Transformers Through Perturbation-Aware Image Synthesis
Lianrui Mu, Jiedong Zhuang, Jiangnan Ye 0002, Haoji Hu
PRICAI (3)6
2024 Application of Tswin-F network based on multi-scale feature fusion in tomato leaf lesion recognition
Yuanbo Ye, Houkui Zhou, Haoji Hu, Guangqun Zhang, Junguo Hu, Tao He 0011
Pattern Recognit.4
2023 Learning Dynamic Graphs from All Contextual Information for Accurate Point-of-Interest Visit Forecasting
abstract
Forecasting the number of visits to Points-of-Interest (POI) in an urban area is critical for planning and decision making in various application domains, from urban planning and transportation management to public health and social studies. Although this forecasting problem can be formulated as a multivariate time-series forecasting task, current approaches cannot fully exploit the ever-changing multi-context correlations among POIs. Therefore, we propose Busyness Graph Neural Network (BysGNN), a temporal graph neural network designed to learn and uncover the underlying multi-context correlations between POIs for accurate visit forecasting. Unlike other approaches where only time-series data is used to learn a dynamic graph, BysGNN utilizes all contextual information and time-series data to learn an accurate dynamic graph representation. By incorporating all contextual, temporal, and spatial signals, we observe a significant improvement in our forecasting accuracy over state-of-the-art forecasting models in our experiments with real-world datasets across the United States.
Arash Hajisafi, Haowen Lin, Sina Shaham, Haoji Hu, Maria Despoina Siampou, Yao-Yi Chiang, Cyrus Shahabi
SIGSPATIAL/GIS4
2023 Video Surveillance on Mobile Edge Networks: Exploiting Multi-Exit Network
abstract
Video surveillance systems are playing increasingly important roles in our everyday lives. To get meaningful surveillance information in a timely and accurate manner, it is vital to optimally allocate computation and communication resources for image classification tasks. In this paper, taking face recognition as an example, we propose a novel end-to-edge collaborative computing system based on a multi-exit network to dynamically allocate computation at the front end (the camera sensor) and back end (the mobile edge computing server). With the ∊-greedy algorithm for reinforcement learning, the decision module decides whether to obtain recognition results from earlier exits at the front end or transmit the feature maps to the back end to obtain more accurate results. The module balances recognition accuracy and time overhead under different channel conditions. Experimental results show that the proposed system can significantly save inference time and maintain competitive accuracy in various communication channel conditions.
Yuchen Cao 0005, Siming Fu, Xiaoxuan He, Haoji Hu, Hangguan Shan, Lu Yu 0003
ICC4
2023 On the Effectiveness of Out-of-Distribution Data in Self-Supervised Long-Tail Learning
Jianhong Bai, Zuozhu Liu, Hualiang Wang, Jin Hao, Yang Feng 0011, Huanpeng Chu, Haoji Hu
ICLR7
2023 Towards Distribution-Agnostic Generalized Category Discovery
abstract
Data imbalance and open-ended distribution are two intrinsic characteristics of the real visual world. Though encouraging progress has been made in tackling each challenge separately, few works dedicated to combining them towards real-world scenarios. While several previous works have focused on classifying close-set samples and detecting open-set samples during testing, it's still essential to be able to classify unknown subjects as human beings. In this paper, we formally define a more realistic task as distribution-agnostic generalized category discovery (DA-GCD): generating fine-grained predictions for both close- and open-set classes in a long-tailed open-world setting. To tackle the challenging problem, we propose a Self-**Ba**lanced **Co**-Advice co**n**trastive framework (BaCon), which consists of a contrastive-learning branch and a pseudo-labeling branch, working collaboratively to provide interactive supervision to resolve the DA-GCD task. In particular, the contrastive-learning branch provides reliable distribution estimation to regularize the predictions of the pseudo-labeling branch, which in turn guides contrastive learning through self-balanced knowledge transfer and a proposed novel contrastive loss. We compare BaCon with state-of-the-art methods from two closely related fields: imbalanced semi-supervised learning and generalized category discovery. The effectiveness of BaCon is demonstrated with superior performance over all baselines and comprehensive analysis across various datasets. Our code is publicly available.
Jianhong Bai, Zuozhu Liu, Hualiang Wang, Ruizhe Chen, Lianrui Mu, Xiaomeng Li 0001, Joey Tianyi Zhou, Yang Feng 0011, Jian Wu 0001, Haoji Hu
NeurIPS10
2023 Repdistiller: Knowledge Distillation Scaled by Re-parameterization for Crowd Counting
Tian Ni, Yuchen Cao 0005, Haoji Hu
PRCV (10)4
2023 Dynamic connection pruning for densely connected convolutional neural networks
Hangxiang Fang, Ling Zhang 0011, Howard H. Yang, Dongxiao Yang, Zheyang Li, Haoji Hu
Appl. Intell.9
2023 Dual cross knowledge distillation for image super-resolution
Hangxiang Fang, Yongwen Long, Yangtao Ou, Yuanjia Huang, Haoji Hu
J. Vis. Commun. Image Represent.6
2023 Class semantic enhancement network for semantic segmentation
Siming Fu, Hualiang Wang, Haoji Hu, Xiaoxuan He, Yongwen Long, Jianhong Bai, Yangtao Ou, Yuanjia Huang, Mengqiu Zhou
J. Vis. Commun. Image Represent.3
2023 AuxBranch: Binarization residual-aware network design via auxiliary branch search
Siming Fu, Huanpeng Chu, Lu Yu 0003, Zheyang Li, Wenming Tan, Haoji Hu
Pattern Recognit.7
2023 Hierarchical Self-Supervised Learning for 3D Tooth Segmentation in Intra-Oral Mesh Scans
abstract
Accurately delineating individual teeth and the gingiva in the three-dimension (3D) intraoral scanned (IOS) mesh data plays a pivotal role in many digital dental applications, e.g., orthodontics. Recent research shows that deep learning based methods can achieve promising results for 3D tooth segmentation, however, most of them rely on high-quality labeled dataset which is usually of small scales as annotating IOS meshes requires intensive human efforts. In this paper, we propose a novel self-supervised learning framework, named STSNet, to boost the performance of 3D tooth segmentation leveraging on large-scale unlabeled IOS data. The framework follows two-stage training, i.e., pre-training and fine-tuning. In pre-training, three hierarchical-level, i.e., point-level, region-level, cross-level, contrastive losses are proposed for unsupervised representation learning on a set of predefined matched points from different augmented views. The pretrained segmentation backbone is further fine-tuned in a supervised manner with a small number of labeled IOS meshes. With the same amount of annotated samples, our method can achieve an mIoU of 89.88%, significantly outperforming the supervised counterparts. The performance gain becomes more remarkable when only a small amount of labeled samples are available. Furthermore, STSNet can achieve better performance with only 40% of the annotated samples as compared to the fully supervised baselines. To the best of our knowledge, we present the first attempt of unsupervised pre-training for 3D tooth segmentation, demonstrating its strong potential in reducing human efforts for annotation and verification.
Zuozhu Liu, Xiaoxuan He, Hualiang Wang, Huimin Xiong, Yan Zhang 0004, Gaoang Wang, Jin Hao, Yang Feng 0011, Fudong Zhu, Haoji Hu
IEEE Trans. Medical Imaging10
2023 Deep Residual Weight-Sharing Attention Network With Low-Rank Attention for Visual Question Answering
abstract
The attention-based networks have become prevailing recently in visual question answering (VQA) due to their high performances. However, the extensive memory consumption of attention-based models poses excessive-high demand for the implementation equipment, raising concerns about their future application scenarios. Therefore, designing an efficient and lightweight VQA model is central to expanding possible application areas. Our work presents a novel lightweight attention-based VQA model, namely residual weight-sharing attention network (RWSAN), consisting of residual weight-sharing attention (RWSA) layers cascaded in depth. Each RWSA layer models the textual representation with self residual weight-sharing attention (SRWSA) and captures question features and question-image interactions with self-guided residual weight-sharing attention (SGRWSA). Inside each RWSA layer, the proposed low-rank attention units perform residual learning with learned connection patterns and shared parameters, and every stacked RWSA layer also uses the same parameters. Extensive ablation experiments with quantitative and qualitative analysis are conducted to illustrate the effectiveness and generality of RWSA. Experiments on VQA-v2, GQA, and CLEVR datasets show that the RWSAN achieves competitive performance with much fewer parameters over the state-of-the-art methods.
Bosheng Qin, Haoji Hu, Yueting Zhuang
IEEE Trans. Multim.2
2022 Renovate Yourself: Calibrating Feature Representation of Misclassified Pixels for Semantic Segmentation
abstract
Existing image semantic segmentation methods favor learning consistent representations by extracting long-range contextual features with the attention, multi-scale, or graph aggregation strategies. These methods usually treat the misclassified and correctly classified pixels equally, hence misleading the optimization process and causing inconsistent intra-class pixel feature representations in the embedding space during learning. In this paper, we propose the auxiliary representation calibration head (RCH), which consists of the image decoupling, prototype clustering, error calibration modules and a metric loss function, to calibrate these error-prone feature representations for better intra-class consistency and segmentation performance. RCH could be incorporated into the hidden layers, trained together with the segmentation networks, and decoupled in the inference stage without additional parameters. Experimental results show that our method could significantly boost the performance of current segmentation methods on multiple datasets (e.g., we outperform the original HRNet and OCRNet by 1.1% and 0.9% mIoU on the Cityscapes test set). Codes are available at https://github.com/VipaiLab/RCH.
Hualiang Wang, Huanpeng Chu, Siming Fu, Zuozhu Liu, Haoji Hu
AAAI5
2022 Meta-prototype Decoupled Training for Long-Tailed Learning
Siming Fu, Huanpeng Chu, Xiaoxuan He, Hualiang Wang, Haoji Hu
ACCV (6)6
2022 Clustering Human Mobility with Multiple Spaces
abstract
Human mobility clustering is an important problem for understanding human mobility behaviors (e.g., work and school commutes). Existing methods typically contain two steps: choosing/learning a mobility representation and applying a clustering algorithm to the representation. However, these methods rely on strict visiting orders in trajectories and cannot take advantage of multiple types of mobility representations. This paper proposes a novel mobility clustering method for mobility behavior detection. First, the proposed method contains a permutation-equivalent operation to handle sub-trajectories that might have different visiting orders but similar impacts on mobility behaviors. Second, the proposed method utilizes a variational autoencoder architecture to simultaneously perform clustering in both latent and original spaces. Also, in order to handle the bias of a single latent space, our clustering assignment prediction considers multiple learned latent spaces at different epochs. This way, the proposed method produces accurate results and can provide reliability estimates of each trajectory’s cluster assignment. The experiment shows that the proposed method outperformed state-of-the-art methods in mobility behavior detection from trajectories with better accuracy and more interpretability.
Haoji Hu, Haowen Lin, Yao-Yi Chiang
IEEE Big Data1
2022 Towards Calibrated Hyper-Sphere Representation via Distribution Overlap Coefficient for Long-Tailed Learning
Hualiang Wang, Siming Fu, Xiaoxuan He, Hangxiang Fang, Zuozhu Liu, Haoji Hu
ECCV (24)6
2022 Meta-BNS FOR Adversarial Data-Free Quantization
abstract
Data-free quantization has recently been a promising method to perform quantization without access to the original data. However, the drawback of such approaches is the homogenization of synthetic data due to low efficiency for diverse data generation and the performance collapse of the generator. To alleviate the above issue, we propose a novel Meta-BNS for adversarial data-free quantization scheme which consists of Meta-BNS module and adversarial exploration module. Meta-BNS module automatically learns an enhancement coefficient matrix function for BN loss module to provide a suitable constrain on the generator. Adversarial exploration module leverages minimax game between the generator and quantized model via input gradient to encourage the generator to learn high-dimensional and complex real data distribution. The experimental results show that our method achieves state-of-the-art performance for various settings on data-free quantization.
Siming Fu, Hualiang Wang, Yuchen Cao 0005, Haoji Hu, Wenming Tan, Tingqun Ye
ICIP4
2022 Residual Swin Transformer Unet with Consistency Regularization for Automatic Breast Ultrasound Tumor Segmentation
abstract
Automatic Breast Ultrasound (ABUS) image segmentation is of great significance for breast cancer diagnosis and treatment. However, similar to most medical datasets, ABUS image datasets are often small-scale and seriously imbalanced, which makes ABUS image segmentation become a challenge. To solve this problem, we propose the Residual Swin Transformer Unet with Consistency Regularization (RSTUnet-CR) which can make full use of non-lesion and unlabeled images for high-precision tumor segmentation on ABUS images. We design a consistency-regularization decoder to reconstruct the input image, which can learn well from non-lesion and unlabeled data. The reconstruction task makes the model more suitable for the imbalanced medical image datasets. In addition, observing that the ABUS images have global semantic correlation, we establish long-distance dependence of images by the residual Swin Transformer block to improve segmentation performance. We evaluate our method on the ABUS dataset collected from 256 subjects and demonstrate the superiority of the proposed method over other state-of-the-art methods in this imbalanced dataset.
Xianwei Zhuang, Xiner Zhu, Haoji Hu, Jincao Yao, Na Feng, Dong Xu 0006
ICIP3
2022 Dynamic Feature Pyramid Networks for Detection
abstract
Feature Pyramid Network (FPN) has been a generic feature extractor in computer vision tasks, which utilizes multi-level features to generate discriminative pyramidal representations. However, the way simply using Sum or Concatenate operation on features to integrate multi-scale information is not sufficient to obtain discriminative semantic representations. In this paper, we propose a dynamic feature pyramid network (DyFPN) to merge multi-scale information in both features and weights. DyFPN uses both high-level context features and low-level spatial structural features to obtain dynamic convolution kernel that contains multi-scale information. In this manner, each resolution in the pyramid performs unique and adaptive convolution directly, meanwhile strengthening the information flow. Specially, DyFPN can be regarded as a complementary enhancement to existing feature pyramid networks. We analyze the effective receptive field and attention map of DyFPN. It proves that our method contains more local information and global information compared with merging multi-scale information only on feature level. Benefit from multi-ways of integrating multi-scale information, our method outperforms other existing feature pyramid methods on COCO detection tasks by a large margin.
Kai Zhang 0055, Zheyang Li, Haoji Hu, Bin Li 0025, Wenming Tan, Haixian Lu, Jun Xiao 0001, Ye Ren, Shiliang Pu
ICME3
2021 Binarizing Super-Resolution Networks by Pixel-Correlation Knowledge Distillation
abstract
Convolutional neural networks (CNNs) have been widely used in single image super-resolution (SR) and obtained remarkable performance. However, most CNN-based SR models require heavy computation, which limits their real-world applications. In this paper, we address the computation problem of SR by network binarization, which converts the full-precision network into the binary network, thus intensively reducing computation. We propose the pixel-correlation distillation for SR network binarization, which distills the knowledge of pixel relationship from the original full-precision network to the binary network. In addition, we further reduce the quantization errors of the binary network by introducing trainable scaling factors to replace the fixed scaling factors in most existing binarization methods. We carry out extensive experiments on SRResNet [1] and VDSR [2], which are two commonly used SR networks. It is shown that the proposed method generates more visually pleasing SR images, and consistently outperforms other state-of-the-art methods in PSNR and SSIM.
Qiu Huang, Haoji Hu, Yongdong Zhu, Zhifeng Zhao
ICIP3
2021 Attention guided feature pyramid network for crowd counting
Huanpeng Chu, Jilin Tang, Haoji Hu
J. Vis. Commun. Image Represent.3
2020 Appearance and Motion Enhancement for Video-Based Person Re-Identification
abstract
In this paper, we propose an Appearance and Motion Enhancement Model (AMEM) for video-based person re-identification to enrich the two kinds of information contained in the backbone network in a more interpretable way. Concretely, human attribute recognition under the supervision of pseudo labels is exploited in an Appearance Enhancement Module (AEM) to help enrich the appearance and semantic information. A Motion Enhancement Module (MEM) is designed to capture the identity-discriminative walking patterns through predicting future frames. Despite a complex model with several auxiliary modules during training, only the backbone model plus two small branches are kept for similarity evaluation which constitute a simple but effective final model. Extensive experiments conducted on three popular video-based person ReID benchmarks demonstrate the effectiveness of our proposed model and the state-of-the-art performance compared with existing methods.
Shuzhao Li, Haoji Hu
AAAI3
2020 Collaborative Distillation for Ultra-Resolution Universal Style Transfer
abstract
Universal style transfer methods typically leverage rich representations from deep Convolutional Neural Network (CNN) models (e.g., VGG-19) pre-trained on large collections of images. Despite the effectiveness, its application is heavily constrained by the large model size to handle ultra-resolution images given limited memory. In this work, we present a new knowledge distillation method (named Collaborative Distillation) for encoder-decoder based neural style transfer to reduce the convolutional filters. The main idea is underpinned by a finding that the encoder-decoder pairs construct an exclusive collaborative relationship, which is regarded as a new kind of knowledge for style transfer models. Moreover, to overcome the feature size mismatch when applying collaborative distillation, a linear embedding loss is introduced to drive the student network to learn a linear embedding of the teacher’s features. Extensive experiments show the effectiveness of our method when applied to different universal style transfer approaches (WCT and AdaIN), even if the model size is reduced by 15.5 times. Especially, on WCT with the compressed models, we achieve ultra-resolution (over 40 megapixels) universal style transfer on a 12GB GPU for the first time. Further experiments on optimization-based stylization scheme show the generality of our algorithm on different stylization paradigms. Our code and trained models are available at https://github.com/mingsun-tse/collaborative-distillation.
Huan Wang 0014, Yijun Li 0001, Yuehai Wang, Haoji Hu, Ming-Hsuan Yang 0001
CVPR4
2020 Triplet Distillation For Deep Face Recognition
abstract
Convolutional neural networks (CNNs) have achieved great successes in face recognition, which unfortunately comes at the cost of massive computation and storage consumption. Many compact face recognition networks are thus proposed to resolve this problem, and triplet loss is effective to further improve the performance of these compact models. However, it normally employs a fixed margin to all the samples, which neglects the informative similarity structures between different identities. In this paper, we borrow the idea of knowledge distillation and define the informative similarity as the transferred knowledge. Then, we propose an enhanced version of triplet loss, named triplet distillation, which exploits the capability of a teacher model to transfer the similarity information to a student model by adaptively varying the margin between positive and negative pairs. Experiments on the LFW, AgeDB and CPLFW datasets show the merits of our method compared to the original triplet loss.
Yushu Feng, Huan Wang 0014, Haoji Hu, Lu Yu 0003, Wei Wang 0118, Shiyan Wang
ICIP3
2020 Affine Deformation Model Based Intra Block Copy for Intra Frame Coding
abstract
In intra frame coding, many methods are proposed to reduce spatial redundancy, including angular intra prediction and intra block copy. However, these two methods cannot deal with complex structure redundancy, such as scaling and rotation relationship where spatial structures are the same but sizes and orientations are different, leading to limited coding efficiency. In this paper, we first mathematically prove that 4-parameter affine deformation model can describe this complex relationship. Then we propose an affine deformation model based intra block copy (ADMIBC) scheme. In the scheme, we focus on several problems, including a fast displacement vector validity judgement algorithm to reduce encoding time, a padding method for unavailable reference pixels during sub-pixel interpolation to improve prediction accuracy and candidate list building methods of control point displacement vector predictor to reduce coding bits. Experimental results show that ADMIBC achieves up to -2.16% BD-rate reduction and -0.99% on average in all intra configuration, compared to the latest video coding standard Versatile Video Coding (VVC).
Daowen Li, Kaitian Qiu, Yaqing Pan, Yingming Li, Haoji Hu, Lu Yu 0003
ISCAS6
2020 Modeling Personalized Item Frequency Information for Next-basket Recommendation
abstract
Next-basket recommendation (NBR) is prevalent in e-commerce and retail industry. In this scenario, a user purchases a set of items (a basket) at a time. NBR performs sequential modeling and recommendation based on a sequence of baskets. NBR is in general more complex than the widely studied sequential (session-based) recommendation which recommends the next item based on a sequence of items. Recurrent neural network (RNN) has proved to be very effective for sequential modeling, and thus been adapted for NBR. However, we argue that existing RNNs cannot directly capture item frequency information in the recommendation scenario.
Haoji Hu, Xiangnan He 0001, Jinyang Gao, Zhi-Li Zhang
SIGIR1
2020 Compressing Facial Makeup Transfer Networks by Collaborative Distillation and Kernel Decomposition
abstract
Although the facial makeup transfer network has achieved high-quality performance in generating perceptually pleasing makeup images, its capability is still restricted by the massive computation and storage of the network architecture. We address this issue by compressing facial makeup transfer networks with collaborative distillation and kernel decomposition. The main idea of collaborative distillation is underpinned by a finding that the encoder-decoder pairs construct an exclusive collaborative relationship, which is regarded as a new kind of knowledge for low-level vision tasks. For kernel decomposition, we apply the depth-wise separation of convolutional kernels to build a light-weighted Convolutional Neural Network (CNN) from the original network. Extensive experiments show the effectiveness of the compression method when applied to the state-of-the-art facial makeup transfer network - BeautyGAN [1].
Bianjiang Yang, Zi Hui, Haoji Hu, Lu Yu 0003
VCIP3
2020 Video Surveillance on Mobile Edge Networks - A Reinforcement-Learning-Based Approach
abstract
Video surveillance systems or Internet of Multimedia Things are playing a more and more important role in our daily life. To obtain useful surveillance information timely and accurately, not only image recognition algorithms but also computing and communication resources can be bottlenecks of the whole system. In this article, taking face recognition application as an example, we study how to build video surveillance systems by utilizing mobile edge computing (MEC), one of the 5G's key technologies. Specifically, to achieve high recognition accuracy and low recognition time, we design image recognition algorithms for both the camera sensor and MEC server, and utilize the action-value methods to train actions of the system by jointly optimizing offloading decision and image compression parameters. The experimental results show the advantages of the proposed system for enabling communication environment-adaptive, efficient, and intelligent video surveillance.
Haoji Hu, Hangguan Shan, Chuankun Wang, Tengxu Sun, Xiaojian Zhen, Kunpeng Yang, Lu Yu 0003, Zhaoyang Zhang 0001, Tony Q. S. Quek
IEEE Internet Things J.1
2020 Attributes-aided part detection and refinement for person re-identification
Shuzhao Li, Haoji Hu
Pattern Recognit.3
2019 Face Alignment by Discriminative Feature Learning
abstract
In this paper, we study the effect of discriminative feature learning for face alignment. We claim that features of the same facial landmarks on different images should share similarities at the feature-level. Thus, we propose the Discriminative Feature Learning method (DFL) for face alignment based on the Fully Convolutional Network (FCN). First, the face image is aligned frontal to eliminate the effect of pose and scale. Second, landmark-specific features are extracted from feature maps of the FCN. A distance constraint on landmark features is added to learn discriminative landmark features. Our experiment results show that DFL can effectively improve the performance of face alignment.
Haoji Hu
ICIP3
2019 Pose Guided Global and Local GAN for Appearance Preserving Human Video Prediction
abstract
We propose a pose-guided approach for appearance preserving video prediction by combining global and local information using Generative Adversarial Networks (GANs). The aim is to predict the subsequent frames based on previous frames of human action videos. Considering that human action videos contain both background scenes which are relatively time-invariant among frames, and human actions which are time-varying components, we use a global GAN to model the time-invariant background and coarse human profiles. Then, a local GAN is utilized to further refine the time-varying human parts. Finally, we use a 3D auto-encoder to fine-tune the frame-by-frame images to obtain the whole predicted video. We evaluate our model on the Penn Action and J-HMDB datasets and demonstrate the superiority of our proposed method over other state-of-the-art methods.
Jilin Tang, Haoji Hu, Hangguan Shan, Chuan Tian, Tony Q. S. Quek
ICIP2
2019 Hepatic Lesion Segmentation by Combining Plain and Contrast-Enhanced CT Images with Modality Weighted U-Net
abstract
We propose the Modality Weighted U-Net (MW-UNet) to combine plain Computed Tomography (CT) and Contrast-Enhanced Computed Tomography (CECT) images for hepatic lesion segmentation. Observing that CT and CECT images provide complimentary but different amount of information for the segmentation task, we propose to fuse their features at specific layers of the U-Net by the weighted sum rule. The weight parameters are updated through backpropagation during training. Compared with most combination methods which concatenate feature maps at the last or intermediate layers, the proposed method obtains feature level fusion with very simple combination rules. Thus, great amount of parameters and computation can be saved. We evaluate our model on the MCGHD database and demonstrate the superiority of the proposed method over other state-of-the-arts both in accuracy and computation.
Yichao Wu, Haoji Hu, Guanghua Rong, Yongwu Li, Shiyan Wang
ICIP3
2019 Three-Dimensional Convolutional Neural Network Pruning with Regularization-Based Method
abstract
Despite enjoying extensive applications in video analysis, three-dimensional convolutional neural networks (3D CNNs) are restricted by their massive computation and storage consumption. To solve this problem, we propose a three-dimensional regularization-based neural network pruning method to assign different regularization parameters to different weight groups based on their importance to the network. Further we analyze the redundancy and computation cost for each layer to determine the different pruning ratios. Experiments show that pruning based on our method can lead to 2× theoretical speedup with only 0.41% accuracy loss for 3D-ResNet18 and 3.28% accuracy loss for C3D. The proposed method performs favorably against other popular methods for model compression and acceleration.
Huan Wang 0014, Lu Yu 0003, Haoji Hu, Hangguan Shan, Tony Q. S. Quek
ICIP5
2019 Structured Pruning for Efficient ConvNets via Incremental Regularization
abstract
Parameter pruning is a promising approach for CNN compression and acceleration by eliminating redundant model parameters with tolerable performance degrade. Despite its effectiveness, existing regularization-based parameter pruning methods usually drive weights towards zero with large and constant regularization factors, which neglects the fragility of the expressiveness of CNNs, and thus calls for a more gentle regularization scheme so that the networks can adapt during pruning. To achieve this, we propose a new and novel regularization-based pruning method, named IncReg, to incrementally assign different regularization factors to different weights based on their relative importance. Empirical analysis on CIFAR-10 dataset verifies the merits of IncReg. Further extensive experiments with popular CNNs on CIFAR-10 and ImageNet datasets show that IncReg achieves comparable to even better results compared with state-of-the-arts. Our source codes and trained models are available here: https://github.com/mingsun-tse/caffe_increg.
Huan Wang 0014, Qiming Zhang 0001, Yuehai Wang, Lu Yu 0003, Haoji Hu
IJCNN5
2019 Sets2Sets: Learning from Sequential Sets with Neural Networks
abstract
Given past sequential sets of elements, predicting the subsequent sets of elements is an important problem in different domains. With the past orders of customers given, predicting the items that are likely to be bought in their following orders can provide information about the future purchase intentions. With the past clinical records of patients at each visit to the hospitals given, predicting the future clinical records in the subsequent visits can provide information about the future disease progression. These useful information can help to make better decisions in different domains. However, existing methods have not studied this problem well. In this paper, we formulate this problem as a sequential sets to sequential sets learning problem. We propose an end-to-end learning approach based on an encoder-decoder framework to solve the problem. In the encoder, our approach maps the set of elements at each past time step into a vector. In the decoder, our method decodes the set of elements at each subsequent time step from the vectors with a set-based attention mechanism. The repeated elements pattern is also considered in our method to further improve the performance. In addition, our objective function addresses the imbalance and correlation existing among the predicted elements. The experimental results on three real-world data sets showthat our method outperforms the best performance of the compared methods with respect to recall and person-wise hit ratio by 2.7-20.6% and 2.1-26.3%, respectively. Our analysis also shows that our decoder has good generalization to output sequential sets that are even longer than the output of training instances.
Haoji Hu, Xiangnan He 0001
KDD1
2019 Analyzing multiple types of behaviors from traffic videos via nonparametric topic model
Houkui Zhou, Haoji Hu, Guangqun Zhang, Junguo Hu, Tao He 0011
J. Vis. Commun. Image Represent.3
2018 Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
Huan Wang 0014, Qiming Zhang 0001, Yuehai Wang, Haoji Hu
BMVC4
2018 Face Alignment by Combining Residual Features in Cascaded Hourglass Network
abstract
Fully Convolutional Networks (FCN) are popular in face alignment thanks to its capacity to retain accurate spatial information. In this work we study the effect of kernel functions of FCN for face alignment. We claim that neither the cross entropy nor the pixel-wise L2 losses can reflect the alignment error accurately if we generate the ground truth probability matrix with kernel functions. Based on this analysis, firstly, we develop a Cascaded Hourglass Network (CHN) as our baseline, and then regress the residual face shape via features obtained from the middle layer of the network, which are called Residual Features (RF). The proposed RF-CHN method obtains Normalized Mean Error (NME) of 6.84, which gives an error reduction of 0.21 compared to the current state-of-the-art on the challenging 300-W database.
Haoji Hu
ICIP3
2018 Find Who to Look at: Turning From Action to Saliency
abstract
The past decade has witnessed the use of highlevel features in saliency prediction for both videos and images. Unfortunately, the existing saliency prediction methods only handle high-level static features, such as face. In fact, high-level dynamic features (also called actions), such as speaking or head turning, are also extremely attractive to visual attention in videos. Thus, in this paper, we propose a data-driven method for learning to predict the saliency of multiple-face videos, by leveraging both static and dynamic features at high-level. Specifically, we introduce an eye-tracking database, collecting the fixations of 39 subjects viewing 65 multiple-face videos. Through analysis on our database, we find a set of high-level features that cause a face to receive extensive visual attention. These high-level features include the static features of face size, center-bias and head pose, as well as the dynamic features of speaking and head turning. Then, we present the techniques for extracting these high-level features. Afterwards, a novel model, namely multiple hidden Markov model (M-HMM), is developed in our method to enable the transition of saliency among faces. In our MHMM, the saliency transition takes into account both the state of saliency at previous frames and the observed high-level features at the current frame. The experimental results show that the proposed method is superior to other state-of-the-art methods in predicting visual attention on multiple-face videos. Finally, we shed light on a promising implementation of our saliency prediction method in locating the region-of-interest (ROI), for video conference compression with high efficiency video coding (HEVC).
Mai Xu, Yufan Liu 0001, Haoji Hu, Feng He 0007
IEEE Trans. Image Process.3
2017 Topic evolution based on the probabilistic topic model: a review
Houkui Zhou, Haoji Hu
Frontiers Comput. Sci.3
2017 Topic discovery and evolution in scientific literature based on content and citations
abstract
Researchers across the globe have been increasingly interested in the manner in which important research topics evolve over time within the corpus of scientific literature. In a dataset of scientific articles, each document can be considered to comprise both the words of the document itself and its citations of other documents. In this paper, we propose a citation- content-latent Dirichlet allocation (LDA) topic discovery method that accounts for both document citation relations and the con-tent of the document itself via a probabilistic generative model. The citation-content-LDA topic model exploits a two-level topic model that includes the citation information for ‘father’ topics and text information for sub-topics. The model parameters are estimated by a collapsed Gibbs sampling algorithm. We also propose a topic evolution algorithm that runs in two steps: topic segmentation and topic dependency relation calculation. We have tested the proposed citation-content-LDA model and topic evolution algorithm on two online datasets, IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI) and IEEE Computer Society (CS), to demonstrate that our algorithm effectively discovers important topics and reflects the topic evolution of important research themes. According to our evaluation metrics, citation-content-LDA outperforms both content-LDA and citation-LDA.
Houkui Zhou, Haoji Hu
Frontiers Inf. Technol. Electron. Eng.3
2017 A survey on trends of cross-media topic evolution map
Houkui Zhou, Haoji Hu, Junguo Hu
Knowl. Based Syst.3
2017 A new sparse representation-based object segmentation framework
Jincao Yao, Haoji Hu
Vis. Comput.3
2016 Implicit kernel presentation aware object segmentation framework
abstract
Given a set of training shapes and an input image with a shape similar to some of the elements in the training set, this paper introduces a new implicit kernel sparse model with a twofold goal. First, to obtain an implicit kernel sparse neighbor based combination that best represents the object. Second, to accurately segment the object taking into accounts both the high-level implicit kernel presentation and the low-level image information. A new energy function that combines the variational image segmentation with the implicit kernel presentation is introduced to accomplish both goals simultaneously. The experimental results on the public datasets show the superior capabilities of the proposed model.
Jincao Yao, Haoji Hu
ICASSP3
2015 Subjective rate-distortion optimization in HEVC with perceptual model of multiple faces
abstract
This paper proposes a novel perceptual video coding approach with a perceptual model of multiple faces, to improve the coding efficiency of HEVC in video conferencing scenarios. For the perceptual model, a latest active appearance model (AAM) is used to detect multiple faces in a video frame. Then, the perceptual model of multiple faces can be established on the basis of the detected multiple faces. With the established perceptual model of multiple faces, all faces in a video frame can be taken into account for subjective rate-distortion optimization, which is based on the state-of-the-art r-λ rate control scheme of HEVC. As such, the perceptual video coding can be achieved for HEVC of video conferencing scenarios. Finally, the experimental results validate the effectiveness of the proposed perceptual video coding approach, in terms of subjective quality.
Yufan Liu 0001, Haoji Hu, Mai Xu
VCIP2
2015 Unsupervised regions based segmentation using object discovery
Bai Yang, Haoji Hu
J. Vis. Commun. Image Represent.3
2015 Robust sparse kernel density estimation by inducing randomness
Fei Chen 0012, Jincao Yao, Haoji Hu
Pattern Anal. Appl.4
2015 GFilter: A General Gram Filter for String Similarity Search
abstract
Numerous applications such as data integration, protein detection, and article copy detection share a similar core problem: given a string as the query, how to efficiently find all the similar answers from a large scale string collection. Many existing methods adopt a prefix-filter-based framework to solve this problem, and a number of recent works aim to use advanced filters to improve the overall search performance. In this paper, we propose a gram-based framework to achieve near maximum filter performance. The main idea is to judiciously choose the high-quality grams as the prefix of query according to their estimated ability to filter candidates. As this selection process is proved to be NP-hard problem, we give a cost model to measure the filter ability of grams and develop efficient heuristic algorithms to find high-quality grams. Extensive experiments on real datasets demonstrate the superiority of the proposed framework in comparison with the state-of-art approaches.
Haoji Hu, Kai Zheng 0001, Xiaoling Wang 0004, Aoying Zhou
IEEE Trans. Knowl. Data Eng.1
2014 Probabilistic hypergraph based hash codes for social image search
abstract
With the rapid development of the Internet, recent years have seen the explosive growth of social media. This brings great challenges in performing efficient and accurate image retrieval on a large scale. Recent work shows that using hashing methods to embed high-dimensional image features and tag information into Hamming space provides a powerful way to index large collections of social images. By learning hash codes through a spectral graph partitioning algorithm, spectral hashing (SH) has shown promising performance among various hashing approaches. However, it is incomplete to model the relations among images only by pairwise simple graphs which ignore the relationship in a higher order. In this paper, we utilize a probabilistic hypergraph model to learn hash codes for social image retrieval. A probabilistic hypergraph model offers a higher order representation among social images by connecting more than two images in one hyperedge. Unlike a normal hypergraph model, a probabilistic hypergraph model considers not only the grouping information, but also the similarities between vertices in hyperedges. Experiments on Flickr image datasets verify the performance of our proposed approach.
Haoji Hu
J. Zhejiang Univ. Sci. C3
2014 A unified framework for semi-supervised PU learning
Haoji Hu, Chaofeng Sha, Xiaoling Wang 0004, Aoying Zhou
World Wide Web1
2013 Community-based user recommendation in uni-directional social networks
abstract
Advances in Web 2.0 technology has led to the rising popularity of many social network services. For example, there are over 500 million active users in Twitter. Given the huge number of users, user recommendation has gained importance where the goal is to find a set of users whom a target user is likely to follow. Content-based approaches that rely on tweet content for user recommendation have low precision as tweet contents are typically short and noisy, while collaborative filtering approaches that utilize follower-followee relationships lead to higher precision but data sparsity remains a challenge. In this work, we propose a community-based approach to user recommendation in Twitter-style social networks. Forming communities enables us to reduce data sparsity as the focus is on discover the latent characteristics of communities instead of individuals. We employ an LDA-based method on the follower-followee relationships to discover communities before applying the state-of-the-art matrix factorization method on each of the communities. This approach proves effective in improving the conversion rate (by as much as 20%) as demonstrated by the results of extensive experiments on two real world data sets Twitter and Weibo. In addition, the community-based approach is scalable as the individual community can be analyzed separately.
Mong-Li Lee, Wynne Hsu, Wei Chen 0025, Haoji Hu
CIKM5
2013 Deep Learning Shape Priors for Object Segmentation
abstract
In this paper we introduce a new shape-driven approach for object segmentation. Given a training set of shapes, we first use deep Boltzmann machine to learn the hierarchical architecture of shape priors. This learned hierarchical architecture is then used to model shape variations of global and local structures in an energetic form. Finally, it is applied to data-driven variational methods to perform object extraction of corrupted data based on shape probabilistic representation. Experiments demonstrate that our model can be applied to dataset of arbitrary prior shapes, and can cope with image noise and clutter, as well as partial occlusions.
Fei Chen 0012, Haoji Hu, Xunxun Zeng
CVPR3
2013 Shape Sparse Representation for Joint Object Classification and Segmentation
abstract
In this paper, a novel variational model based on prior shapes for simultaneous object classification and segmentation is proposed. Given a set of training shapes of multiple object classes, a sparse linear combination of training shapes in a low-dimensional representation is used to regularize the target shape in variational image segmentation. By minimizing the proposed variational functional, the model is able to automatically select the reference shapes that best represent the object by sparse recovery and accurately segment the image, taking into account both the image information and the shape priors. For some applications under an appropriate size of training set, the proposed model allows artificial enlargement of the training set by including a certain number of transformed shapes for transformation invariance, and then the model remains jointly convex and can handle the case of overlapping or multiple objects presented in an image within a small range. Numerical experiments show promising results and the potential of the method for object classification and segmentation.
Fei Chen 0012, Haoji Hu
IEEE Trans. Image Process.3
2012 Estimate Unlabeled-Data-Distribution for Semi-supervised PU Learning
Haoji Hu, Chaofeng Sha, Xiaoling Wang 0004, Aoying Zhou
APWeb1
2012 Pooling Search: Serum Samples Test Simulated Video Fingerprint Search
abstract
Inspired by the serum pooling strategy in medical area, this paper presents a new approach for video fingerprint search. The proposed method has adopted the serum pooling strategy to reduce the unnecessary matching calculation during the search process. Two observations about random vectors are given, which enable us to obtain a general similarity measure and accelerate the search speed without losing search precisions. Simulations on public database indicate the reduction of unnecessary matching and significant improvements in search speed.
Jincao Yao, Haoji Hu
ICME3
2012 Reduced set density estimator for object segmentation based on shape probabilistic representation
Fei Chen 0012, Haoji Hu, Shiyan Wang
J. Vis. Commun. Image Represent.2
2010 Applying Spread Transform Dither Modulation for 3D-mesh watermarking by using perceptual models
abstract
A 3D-mesh watermarking method is proposed in this paper. The watermarking occurs in the spherical coordinates system, where only the vertex norm of each point is modified. The embedding method is based on Spread Transform Dither Modulation (STDM). A perceptual model based on the curvature and the roughness of the 3D mesh is defined and is used to modulate the STDM method. Our results show that the proposed method has high tolerance to noise and is robust against mesh smoothing.
Rony Darazi, Haoji Hu, Benoît Macq
ICASSP2
2010 Simultaneous variational image segmentation and object recognition via shape sparse representation
abstract
In this paper, we propose a novel model for simultaneous image segmentation and object recognition. Our model is different from previous prior-based level set variatioinal image segmentation in two aspects. The first is the use of the shape sparse representation, which is able to integrate shape priors by linear combination into variational image segmentation. The second is that segmentation and recognition procedures are carried out automatically. The sparsest solution will determine the identity of the target. In addition, our model can handle more general shape priors. Numerical experiments show promising results on synthetic and real images.
Fei Chen 0012, Haoji Hu
ICIP3
2010 Incorporating Watson's perceptual model into patchwork watermarking for digital images
abstract
This paper presents a modified patchwork watermarking scheme for digital images by incorporating the Watson's perceptual model into the watermarking process. Watermarking occurs in the DCT domain. Watson's perceptual model provides a measure of distortions that each DCT coefficient can resist based on the human visual system (HVS). Perceptual information is incorporated into the watermarking scheme by minimizing the Watson's distance between the host image and watermarked image using quadratic programming. Compared with patchwork schemes which do not consider perceptual models, experiments indicate that the proposed method has increased robustness and fidelity of the watermarking system.
Haoji Hu, Fei Chen 0012
ICIP1
2009 Constrained optimisation of 3D polygonal mesh watermarking by quadratic programming
abstract
In this paper, we propose a blind and robust watermarking method for 3D polygonal meshes by minimising the mean square error between the original mesh and the watermarked mesh under several constraints. We have formulated the problem of assigning distortions to points in a 3D mesh to a quadratic programming problem, so it can be solved reliably and efficiently. Comparing with similar approaches in [1], experiments indicate the advantages of our method in resisting Gaussian noise.
Haoji Hu, Patrice Rondao-Alface, Benoît Macq
ICASSP1
2008 A 'No Panacea Theorem' for classifier combination
Haoji Hu, Robert I. Damper
Pattern Recognit.1