VLDB 2026 Research / reviewers in the wild / expert
Wei Huang 0013
dblp:81/6685-13
· DBLP profile ↗
77ranked-venue papers
27as first author
45since 2021 · last 2026
0000-0002-0541-8612ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 15 first-author · 18 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 9 first-author · 6 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prior Refinement Is Better: Diffusion-Driven Graph Harmonization for Federated Graph LearningabstractFederated Graph Learning (FGL) has emerged as a compelling paradigm for collaboratively training a global model while preserving the privacy of multi-source graphs. Nonetheless, FGL faces a critical challenge of data heterogeneity, where semantic and structural discrepancies across clients significantly degrade its performance. Although existing methods attempt to calibrate client-specific graph distributions during federated training, they inevitably fall short in aligning the optimization behaviors across clients due to dynamic parameter updates, thereby inducing a bottleneck in generalization improvement. To tackle this challenge, we propose a solution from a new perspective of prior refinement, which seeks to proactively harmonize client graph distributions before the federated training. In particular, we propose a Federated Graph Harmonization (FedGH) framework that exploits the generative strengths of graph diffusion models to perform prior refinement of local graphs. In a nutshell, FedGH designs a conditional diffusion mechanism on each client that synthesizes pseudo-graphs encapsulating both feature and structural priors, thereby facilitating explicit correction of inter-client distributional bias. On the server side, we employ the graph contrastive learning between various client-specific pseudo-graphs to incorporate the global information, subsequently guiding local data reconstruction. Importantly, model-agnostic FedGH can be seamlessly deployed as a plug-and-play module to be easily integrated with existing FGL architectures. Extensive experiments demonstrate that FedGH consistently outperforms state-of-the-art FGL baselines. Shuman Zhuang, Zhihao Wu 0003, Wei Huang 0013, Luojun Lin, Jiali Yin, Lele Fu, Hongning Dai |
AAAI | 3 |
| 2026 | HFTS: Time-Span-Aware Historical-Future Modeling for Temporal Knowledge Graph Completion
Wei Huang 0013, Tianyong Hao, Fu Lee Wang, Jing He 0004, Hai Liu 0006 |
DASFAA (5) | 1 |
| 2026 | Discovering new intents via spatio-temporal pseudo-label denoising
Yuming Shang, Wei Huang 0013, Sanchuan Guo, Jinhu Chen, Xi Zhang 0008, Philip S. Yu |
Inf. Process. Manag. | 3 |
| 2026 | LoTTA: Low-rank test-time adaptation for unsupervised tabular anomaly detection
Yuming Shang, Wenzhi Peng, Wei Huang 0013, Ninglun Gu, Kailai Zhang |
Inf. Process. Manag. | 3 |
| 2026 | A Tri-Factor Adaptive Federated Learning Framework for Parkinson's Disease Diagnosis via Multi-Source Facial Expression AnalysisabstractEarly diagnosis of Parkinson's disease (PD) is crucial for timely treatment and disease management. Recent studies link PD to impaired facial muscle control, manifesting as "masked face" symptoms, offering a novel diagnostic approach through facial expression analysis. However, data privacy concerns and legal restrictions have resulted in significant "data silos", hindering data sharing and limiting the accuracy and generalizability of existing diagnostic models due to small, localized datasets. To address these challenges, we propose an innovative Tri-Factor Adaptive Federated Learning (TriAFL) framework, designed to collaboratively analyze facial expression data across multiple medical institutions while ensuring robust data privacy protection. TriAFL introduces a comprehensive evaluation mechanism that assesses client contributions across three dimensions: gradient, data, and learning efficiency, effectively addressing Non-IID issues arising from data size variations and heterogeneity. To validate the real-world applicability of our method, we collaborate with a hospital to build the largest known facial expression dataset of PD patients. Furthermore, we explore the integration of local data augmentation strategy to further enhance diagnostic accuracy. Comprehensive experimental results demonstrate TriAFL's superior performance over conventional FL methods in classification task, as well as confirms TriAFL's efficacy in PD diagnosis, delivering a rapid, non-invasive screening tool while driving advancements in AI-powered healthcare. Houwei Xu, Yintao Zhou, Shengbo Chen, Binghui Wang, Wei Huang 0013 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Breaking Data Silos in Parkinson's Disease Diagnosis: An Adaptive Federated Learning Approach for Privacy-Preserving Facial Expression AnalysisabstractThe early diagnosis of Parkinson’s disease (PD) is crucial for potential patients to receive timely treatment and prevent disease progression. Recent studies have shown that PD is closely linked to impairments in facial muscle control, resulting in characteristic “masked face” symptoms. This discovery offers a novel perspective for PD diagnosis by leveraging facial expression recognition and analysis techniques to capture and quantify these features, thereby distinguishing between PD patients and non-PD individuals based on their facial expressions. However, concerns about data privacy and legal restrictions have led to significant “data silos”, posing challenges to data sharing and limiting the accuracy and generalization of existing diagnostic models due to small, localized datasets. To address this issue, we propose an innovative adaptive federated learning approach that aims to jointly analyze facial expression data from multiple medical institutions while preserving data privacy. Our proposed approach comprehensively evaluates each client's contributions in terms of gradient, data, and learning efficiency, overcoming the non-IID issues caused by varying data sizes or heterogeneity across clients. To demonstrate the real-world impact of our approach, we collected a new facial expression dataset of PD patients in collaboration with a hospital. Extensive experiments validate the effectiveness of our proposed method for PD diagnosis and facial expression recognition, offering a promising avenue for rapid, non-invasive initial screening and advancing healthcare intelligence. Houwei Xu, Yintao Zhou, Wei Huang 0013, Binghui Wang |
AAAI | 5 |
| 2025 | Dynamic Localisation of Spatial-Temporal Graph Neural NetworkabstractSpatial-temporal data, fundamental to many intelligent applications, reveals dependencies indicating causal links between present measurements at specific locations and historical data at the same or other locations. Within this context, adaptive spatial-temporal graph neural networks (ASTGNNs) have emerged as valuable tools for modelling these dependencies, especially through a data-driven approach rather than pre-defined spatial graphs. While this approach offers higher accuracy, it presents increased computational demands. Addressing this challenge, this paper delves into the concept of localisation within ASTGNNs, introducing an innovative perspective that spatial dependencies should be dynamically evolving over time. We introduce DynAGS, a localised ASTGNN framework aimed at maximising efficiency and accuracy in distributed deployment. This framework integrates dynamic localisation, time-evolving spatial graphs, and personalised localisation, all orchestrated around the Dynamic Graph Generator, a light-weighted central module leveraging cross attention. The central module can integrate historical information in a node-independent manner to enhance the feature representation of nodes at the current moment. This improved feature representation is then used to generate a dynamic sparse graph without the need for costly data exchanges, and it supports personalised localisation. Performance assessments across two core ASTGNN architectures and nine real-world datasets from various applications reveal that DynAGS outshines current benchmarks, underscoring that the dynamic modelling of spatial dependencies can drastically improve model expressibility, flexibility, and system efficiency, especially in distributed settings. © 2025 Owner/Author. Wenying Duan, Shujun Guo, Zimu Zhou, Wei Huang 0013, Hong Rao, Xiaoxi He |
KDD (1) | 4 |
| 2025 | Audio-visual correspondences based joint learning for instrumental playing source separation
Peng Zhang 0005, Siliang Wang, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001 |
Neurocomputing | 4 |
| 2025 | Flexible Temperature Parallel Distillation for Dense Object Detection: Make Response-Based Knowledge Distillation Great AgainabstractFeature-based approaches have been the focal point of previous research on knowledge distillation (KD) for dense object detection. These methods employ feature imitation and result in competitive performance. Despite being able to achieve comparable performance in image recognition, response-based KD methods can not reach the same level in dense object detection. Inspired by improving distillation performance from two key aspects: where to distill and how to distill, in this paper, a parallel distillation (PD) is introduced to fully utilize the sophisticated detection head and transfer all the output responses from the teacher to the student efficiently. In particular, the proposed PD takes an important consideration of the specific location of distillation, which is crucial for effective knowledge transfer. Regarding the discrepancies in output responses between the localization branch and the classification branch, we propose a novel Dynamic Localization Temperature (DLT) module to enhance the precision of distilling localization information. As for the classification branch, a Classification Temperature-Free (CTF) module is also designed to increase the robustness of distillation in heterogeneous networks. By incorporating the DLT and CTF into the PD framework to avoid setting temperature values manually, the Flexible Temperature Parallel Distillation (FTPD) is proposed to achieve a state-of-the-art (SOTA) performance, which can also be further combined with mainstream feature-based methods for better results. In terms of accuracy and robustness with extensive experiments, the proposed FTPD outperforms other KD methods in the task of dense object detection. Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Toward Unifying Saliency Transformer for Video Saliency Prediction and DetectionabstractVideo saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceive dynamic scenes. While many approaches have crafted task-specific training paradigms for either video saliency prediction or video salient object detection tasks, few attention has been devoted to devising a generalized saliency modeling framework that seamlessly bridges both these distinct tasks. In this study, we introduce the Unified Saliency Transformer (UniST) framework, which comprehensively utilizes the essential attributes of video saliency prediction and video salient object detection. In addition to extracting representations of frame sequences, a saliency-aware transformer is designed to learn the spatio-temporal representations at progressively increased resolutions, while incorporating effective cross-scale saliency information to produce a robust representation. Furthermore, task-specific decoders are proposed to perform the final prediction for each task. To the best of our knowledge, this is the first work to explore the design of a unified framework for both saliency modeling tasks. Convincible experiments demonstrate that the proposed UniST achieves superior performance across eight challenging benchmarks for two tasks, outperforming other state-of-the-art methods in most metrics. The project page ishttps://junwenxiong.github.io/UniST. Junwen Xiong, Chuanyue Li, Peng Zhang 0005, Yue Huo, Wei Huang 0013, Yufei Zha |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | M2FE-YOLO: Multibranch and Multilevel Feature Enhancement Network for Remote Sensing Object DetectionabstractThe detection of remote sensing (RS) objects plays a crucial role in various earth observation tasks. Current RS object detection methods tend to face great challenges due to large-scale variations of object sizes and limited representation of semantic features, which affect the final bounding box regression results in practice. To address the challenges, we propose a novel Multi-branch and Multi-level Feature Enhancement framework (M2FE-YOLO) to refine feature representation learning for improving the performance of object detection. The proposed M2FE-YOLO primarily comprises three components, i.e., Multi-Branch Feature-awareness based CSP (MBFA-CSP) module, Multi-Level Feature Fusion (MLFF) module, and Shape-IoU as a loss function. MBFA-CSP builds on a dual-scale feature-aware mechanism and a serial convolution pathway to dynamically adjust receptive fields, capturing critical contextual patterns in RS images. MLFF resolves the persistent semantic discrepancy between low-level texture features and high-level abstract representations, enabling precise localization of objects in cluttered RS landscapes. Shape-IoU instead of C-IoU incorporates a geometric compatibility factor (i.e., aspect ratio consistency), and is more crucial for bounding box regression of elongated or irregular RS objects. Compared with existing state-of-the-art methods, extensive experiments on RSOD, NWPU VHR-10, and DOTA datasets quantitatively and qualitatively demonstrate the superiority of the proposed M2FE-YOLO method, achieving up to 92.1%, 93.3%, and 72.2% mAP, respectively. Meanwhile, M2FE-YOLO-OBB achieves an excellent detection result of 73.9% mAP on DOTA dataset for oriented object detection task. Qinggang Wu, Xiaotian You, Wei Huang 0013, Le Sun 0002, Yang Xu 0006, Xinnian Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | In Vitro Diagnosis of Parkinson's Disease Based on Facial Expression and Behavioral Gait DataabstractParkinson's disease (PD) is characterized by incurable, rapid progression, and severe disability, severely impacting the lives of patients and their families. With an aging population, the need for early detection of PD is increasing. In vitro diagnosis has attracted attention because of its non-invasiveness and low cost, but there are some problems with the existing methods: 1) facial expression diagnosis has little training data; 2) gait diagnosis requires specialized equipment and acquisition environment, which is poorly generalizable; 3) a single modality is easy to miss the diagnosis; and 4) multimodal diagnostic methods are not universally applicable. To address the above issues, we propose a novel multimodal in vitro diagnostic method for PD based on facial expression and behavioral gait. The method uses a lightweight deep learning model for feature extraction and feature fusion to improve diagnostic accuracy and ease of use. Meanwhile, we have established the largest multimodal PD data set in collaboration with hospitals and conducted a large number of experiments to verify the effectiveness of the method. Yinxuan Xu, Yintao Zhou, Wei Huang 0013 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | DynImpt: A Dynamic Data Selection Method for Improving Model Training EfficiencyabstractSelecting key data subsets for model training is an effective way to improve training efficiency. Existing methods generally utilize a well-trained model to evaluate samples and select crucial subsets, ignoring the fact that the sample importance changes dynamically during model training, resulting in the selected subset only being critical in a specific training epoch rather than a changing training phase. To address this issue, we attempt to evaluate the significant changes in sample importance during dynamic training and propose a novel data selection method to improve model training efficiency. Specifically, the temporal changes in sample importance are considered from three perspectives: (i) loss, the difference between the predicted labels and the true labels of samples in the current training epoch; (ii) instability, the dispersion of sample importance in the recent training phase; and (iii) inconsistency, the comparison of the changing trend in the importance of an individual sample relative to the average importance of all samples in the recent training phase. Extensive experiments demonstrate that dynamic data selection can reduce computational costs and improve model training efficiency. Additionally, we find that the difficulty level of the training task influences the data selection strategy. Wei Huang 0013, Shangmin Guo, Yuming Shang, Xiangling Fu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Heterogeneous Prototype Learning From Contaminated Faces Across Domains via Disentangling Latent FactorsabstractThis article studies an emerging practical problem called heterogeneous prototype learning (HPL). Unlike the conventional heterogeneous face synthesis (HFS) problem that focuses on precisely translating a face image from a source domain to another target one without removing facial variations, HPL aims at learning the variation-free prototype of an image in the target domain while preserving the identity characteristics. HPL is a compounded problem involving two cross-coupled subproblems, that is, domain transfer and prototype learning (PL), thus making most of the existing HFS methods that simply transfer the domain style of images unsuitable for HPL. To tackle HPL, we advocate disentangling the prototype and domain factors in their respective latent feature spaces and then replacing the source domain with the target one for generating a new heterogeneous prototype. In doing so, the two subproblems in HPL can be solved jointly in a unified manner. Based on this, we propose a disentangled HPL framework, dubbed DisHPL, which is composed of one encoder-decoder generator and two discriminators. The generator and discriminators play adversarial games such that the generator embeds contaminated images into a prototype feature space only capturing identity information and a domain-specific feature space, while generating realistic-looking heterogeneous prototypes. Experiments on various heterogeneous datasets with diverse variations validate the superiority of DisHPL. Binghui Wang, Mang Ye, Yiu-Ming Cheung, Yintao Zhou, Wei Huang 0013, Bihan Wen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | DiffSal: Joint Audio and Video Learning for Diffusion Saliency PredictionabstractAudio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising diffusion models have shown more promising in unifying task frameworks owing to their inherent ability of generalization. Following this motivation, a novel Diffusion architecture for generalized audio-visual Saliency prediction (DiffSal) is proposed in this work, which formulates the prediction problem as a conditional generative task of the saliency map by utilizing input audio and video as the conditions. Based on the spatiotemporal audio-visual features, an extra network Saliency-UNet is designed to perform multimodal attention modulation for progressive refinement of the ground-truth saliency map from the noisy map. Extensive experiments demonstrate that the proposed DiffSal can achieve excellent performance across six challenging audio-visual benchmarks, with an average relative improvement of 6.3% over the previous state-of-the-art results by six metrics. The project url is htt ps: //junwenxiong. github.io/DiffSal. Junwen Xiong, Peng Zhang 0005, Tao You, Chuanyue Li, Wei Huang 0013, Yufei Zha |
CVPR | 5 |
| 2024 | Early Diagnosing Parkinson's Disease Via a Deep Learning Model Based on Augmented Facial Expression DataabstractIt is crucial to promptly diagnose potential Parkinson's disease (PD) patients in order to facilitate early treatment and prevent disease progression. In recent years, there has been growing interest in using facial expressions for in-vitro PD diagnosis due to the distinct "masked face" characteristics of PD patients and the cost-effectiveness of this approach. However, current facial expression-based PD diagnosis methods are hindered by limited training data on PD patients' facial expressions and weak prediction models. To address these issues, we propose a new PD diagnosis method that utilizes facial expression data augmentation and deep neural network prediction. Our approach involves two stages: 1) generating virtual facial expression images depicting six basic emotions (anger, disgust, fear, happiness, sadness, and surprise) through multi-domain adversarial learning to expand the original training data; 2) training a deep neural network prediction model using a combination of the augmented training data from PD patients and facial expression images of normal individuals from public datasets. Qualitative and quantitative experiments confirm the efficacy of our multi-domain adversarial learning-based facial expression synthesis and demonstrate the promising performance of our proposed approach for PD diagnosis. Yintao Zhou, Wei Huang 0013, Binghui Wang |
ICASSP | 3 |
| 2024 | Reconstructing Prototype From Contaminated Face With Variations Across Heterogeneous DomainsabstractThis paper focuses on a new heterogeneous prototype learning (HPL) problem, which aims at reconstructing the variation-free and identity-preserved prototype in the target domain from a contaminated input image in the source domain. Most existing heterogeneous face synthesis (HFS) methods are unsuitable for HPL, as these methods focus on performing accurate image-to-image translation with facial details unaltered, but cannot effectively remove the input facial variations. In this paper, we propose an identity-aware cycle-consistent network, dubbed IAC2N, for image-to-prototype transformation across domains. To address HPL, IAC2N designs three effective losses, i.e., prototype adversarial loss, label information guided identity loss, and prototype learning cycle loss, in its objective. The first loss is used for transferring the domain style as well as removing the universal facial variations. The latter two losses are used for maintaining the identity consistency during HPL from an explicit and an implicit perspectives, respectively. Furthermore, IAC2N is a joint learning framework that is able to learn the identity feature for the contaminated image via its encoder-decoder structural generator in order to perform heterogeneous face recognition (HFR). Extensive experiments on various heterogeneous face datasets demonstrate the effectiveness of IAC2N in both tasks of HPL and HFR. Binghui Wang, Nanrun Zhou, Yintao Zhou, Wei Huang 0013 |
ICME | 5 |
| 2024 | Enhancing Multi-view Graph Neural Network with Cross-view Confluent Message PassingabstractWith the growing diversity of data sources, multi-view learning methods have attracted considerable attention. Among these, by modeling the multi-view data as multi-view graphs, multi-view Graph Neural Networks (GNNs) have shown encouraging performance on various multi-view learning tasks. The message passing is the critical mechanism empowering GNNs with superior capacity to process complex graph data. However, most multi-view GNNs are designed on the well-established overall framework, overlooking the intrinsic challenges of the message passing on multi-view scenarios. To clarify this, we first revisit the message passing mechanism from a Laplacian smoothing perspective, revealing the key to designing a multi-view message passing. Following the analysis, in this paper, we propose an enhanced GNN framework termed Confluent Graph Neural Networks (CGNN), with Cross-view Confulent Message Pssing (CCMP) tailored for multi-view learning. Inspired by the optimization of an improved multi-view Laplacian smoothing problem, CCMP contains three sub-modules that enable the interaction between graph structures and consistent representations, which makes it aware of consistency and complementarity information across views. Extensive experiments on four types of data including multi-modality data demonstrate that our proposed model exhibits superior effectiveness and robustness. The code is available at https://github.com/shumanzhuang/CGNN. Shuman Zhuang, Sujia Huang, Wei Huang 0013, Zhihao Wu 0003, Ximeng Liu |
ACM Multimedia | 3 |
| 2024 | A fragmentation-aware redundancy elimination scheme for inline backup systems
Wenxuan Zhu, Dan Feng 0001, Wei Huang 0013, Nan Jiang 0013, Meng Chen 0024, Renxin Xia |
Future Gener. Comput. Syst. | 4 |
| 2024 | How does Layer Normalization improve Batch Normalization in self-supervised sound source localization?
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2024 | Closed-loop unified knowledge distillation for dense object detection
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2024 | Applying Delta Compression to Packed Datasets for Efficient Data ReductionabstractBackup systems often adopt deduplication techniques for data reduction. Real-world backup products often group files into larger units (called packed files) before deduplicating them. The grouping entails inserting metadata immediately before the contents of each file in the packed file. Some metadata change with every backup, producing substantial similar (non-duplicate) chunks. Delta compression can remove redundancy among those similar chunks but cannot be applied to HDD-based backup storage because I/Os required for fetching base chunks result in severe throughput loss. For packed datasets, some duplicate chunks, called persistent fragmented chunks (PFCs), are rewritten every backup. We observe that corresponding chunk pairs surrounding identical PFCs are non-identical due to different metadata but similar to each other. In this article, we propose PFC-delta to perform high-performance delta compression for the aforementioned similar chunks on top of deduplication. PFC-delta identifies and prefetches potential base chunks stored along with PFCs by piggybacking on the routine I/Os during deduplication, thus avoiding extra I/Os. We also propose a hash-less delta encoding approach to reduce extra computational overheads. Evaluation results with four real-world datasets show that PFC-delta improves both compression ratio and restore performance, while increasing the backup throughput on all but one datasets. Hong Jiang 0001, Wei Huang 0013, Meng Chen 0024, Yongxuan Zhang |
IEEE Trans. Computers | 4 |
| 2024 | Auto Diagnosis of Parkinson's Disease Via a Deep Learning Model Based on Mixed Emotional Facial ExpressionsabstractParkinson's disease (PD) is a common degenerative disease of the nervous system in the elderly. The early diagnosis of PD is very important for potential patients to receive prompt treatment and avoid the aggravation of the disease. Recent studies have found that PD patients always suffer from emotional expression disorder, thus forming the characteristics of "masked faces". Based on this, we thus propose an auto PD diagnosis method based on mixed emotional facial expressions in the paper. Specifically, the proposed method is cast into four steps: Firstly, we synthesize virtual face images containing six basic expressions (i.e., anger, disgust, fear, happiness, sadness, and surprise) via generative adversarial learning, in order to approximate the premorbid expressions of PD patients; Secondly, we design an effective screening scheme to assess the quality of the above synthesized facial expression images and then shortlist the high-quality ones; Thirdly, we train a deep feature extractor accompanied with a facial expression classifier based on the mixture of the original facial expression images of the PD patients, the high-quality synthesized facial expression images of PD patients, and the normal facial expression images from other public face datasets; Finally, with the well-trained deep feature extractor, we thus adopt it to extract the latent expression features for six facial expression images of a potential PD patient to conduct PD/non-PD prediction. To show real-world impacts, we also collected a new facial expression dataset of PD patients in collaboration with a hospital. Extensive experiments are conducted to validate the effectiveness of the proposed method for PD diagnosis and facial expression recognition. Wei Huang 0013, Renjie Wan, Peng Zhang 0005, Yufei Zha |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveabstractIncorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalities but ignoring the negative effects due to the temporal inconsistency of audio-visual intrinsics. Inspired by the biological inconsistency-correction within multi-sensory information, in this study, a consistency-aware audio-visual saliency prediction network (CASP-Net) is proposed, which takes a comprehensive consideration of the audio-visual semantic interaction and consistent perception. In addition a two-stream encoder for elegant association between video frames and corresponding sound source, a novel consistency-aware predictive coding is also designed to improve the consistency within audio and visual representations iteratively. To further aggregate the multi-scale audio-visual information, a saliency decoder is introduced for the final saliency map generation. Substantial experiments demonstrate that the proposed CASP-Net outperforms the other state-of-the-art methods on six challenging audio-visual eye-tracking datasets. For a demo of our system please see our project webpage. Junwen Xiong, Ganglai Wang, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Guangtao Zhai |
CVPR | 4 |
| 2023 | Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source LocalizationabstractSelf-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and sound sources in visual scenarios. Unfortunately, the insufficient attention to the heterogeneity influence in the different modality features still limits this scheme to be further improved, which also becomes the motivation of our work. In this study, an Induction Network is proposed to bridge the modality gap more effectively. By decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the audio modality to be aligned with the visual modality consistently. In addition to a visual weighted contrastive loss, an adaptive threshold selection strategy is introduced to enhance the robustness of the Induction Network. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have demonstrated a superior performance compared to other state-of-the-art works in different challenging scenarios. The code is available at https://github.com/Tahy1/AVIN. Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | LoopDelta: Embedding Locality-aware Opportunistic Delta Compression in Inline Deduplication for Highly Efficient Data Reduction
Hong Jiang 0001, Dan Feng 0001, Nan Jiang 0013, Taorong Qiu, Wei Huang 0013 |
USENIX ATC | 6 |
| 2023 | Efficient thermal infrared tracking with cross-modal compress distillation
Hangfei Li, Yufei Zha, Huanyu Li 0003, Peng Zhang 0005, Wei Huang 0013 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Object detection based on cortex hierarchical activation in border sensitive mechanism and classification-GIou joint representation
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2023 | Conditional invertible image re-scaling
Yufei Zha, Peng Zhang 0005, Wei Huang 0013 |
Pattern Recognit. | 4 |
| 2023 | Facial Expression Guided Diagnosis of Parkinson's Disease via High-Quality Data AugmentationabstractParkinson's disease (PD) is a neurodegenerative disease which is prevalent among the elder population and severely affects the life quality of patients and their families. Therefore, it is important to conduct an early diagnosis for potential patients with PD, so as to promote prompt treatment and avoid the aggravation of the disease. Recently, the in-vitro PD diagnosis based on facial expressions has received increasing attention because of its distinguishability (i.e., PD patients always possess the characteristics of “masked face”) and affordability. However, the performance of the existing facial expression-based PD diagnosis approaches is limited by: 1) the small-scale training data on PD patients' facial expressions, and 2) the weak prediction model. To address these two problems, we propose a new facial expression guided PD diagnosis method based on high-quality training data augmentation and deep neural network prediction. Specifically, the proposed method consists of three stages: Firstly, we synthesize virtual facial expression images with 6 basic emotions (i.e., anger, disgust, fear, happiness, sadness, and surprise) based on multi-domain adversarial learning to approximate the premorbid expressions of PD patients. Secondly, we introduce three facial image quality assessment (FIQA) criteria to measure the quality of these synthesized facial expression images and design a fusion screening strategy that shortlists the high-quality ones to augment the training data. Finally, we train a deep neural network prediction model based on the original and synthesized high-quality facial expression images for PD diagnosis. To show real-world impacts and evaluate the proposed method under different facial expressions, we also create a (currently largest) multiple facial expressions-based PD face dataset in collaboration with a hospital. Extensive experiments are performed to demonstrate the effectiveness of the multi-domain adversarial learning-based facial expression synthesis and the fusion screening strategy, particularly the superior performance of the proposed method for PD diagnosis. Wei Huang 0013, Yintao Zhou, Yiu-Ming Cheung, Peng Zhang 0005, Yufei Zha |
IEEE Trans. Multim. | 1 |
| 2023 | Look&listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech EnhancementabstractActive speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been widely used in correspondence to each single task. This may lead to the representation learned by the model being task-specific, and inevitably result in the lack of generalization ability of the feature based on multi-modal modeling. More recent studies have shown that establishing cross-modal relationship between auditory and visual stream is a promising solution for the challenge of audio-visual multi-task learning. Therefore, as a motivation to bridge the multi-modal associations in audio-visual tasks, a unified framework is proposed to achieve target speaker detection and speech enhancement with joint learning of audio-visual modeling in this study. With the assistance of audio-visual channels of videos in challenging real-world scenarios, the proposed method is able to exploit inherent correlations in both audio and visual signals, which is used to further anticipate and model the temporal audio-visual relationships across spatial-temporal space via a cross-modal conformer. In addition, a plug-and-play multi-modal layer normalization is introduced to alleviate the distribution misalignment of multi-modal features. Based on cross-modal circulant fusion, the proposed model is capable to learned all audio-visual representations in a holistic process. Substantial experiments demonstrate that the correlations between different modalities and the associations among diverse tasks can be learned by the optimized model more effectively. In comparison to other state-of-the-art works, the proposed work shows a superior performance for active speaker detection and audio-visual speech enhancement on three benchmark datasets, also with a favorable generalization in diverse challenges. Junwen Xiong, Peng Zhang 0005, Lei Xie 0001, Wei Huang 0013, Yufei Zha |
IEEE Trans. Multim. | 5 |
| 2022 | Cross-domain Prototype Learning from Contaminated Faces via Disentangling Latent FactorsabstractThis paper focuses on an emerging challenging problem called heterogeneous prototype learning (HPL) across face domains-It aims to learn the variation-free target domain prototype for a contaminated input image from the source domain and meanwhile preserve the personal identity. HPL involves two coupled subproblems, i.e., domain transfer and prototype learning. To address the two subproblems in a unified manner, we advocate disentangling the prototype and domain factors in their respected latent feature spaces, and replace the latent source domain features with the target domain ones to generate the heterogeneous prototype. To this end, we propose a disentangled heterogeneous prototype learning framework, dubbed DisHPL, which consists of one encoder-decoder generator and two discriminators. The generator and discriminators play adversarial games such that the generator learns to embed the contaminated image into a prototype feature space only capturing identity information and a domain-specific feature space, as well as generating a realistic-looking heterogeneous prototype. The two discriminators aim to predict personal identities and distinguish between real prototypes versus fake generated prototypes in the source/target domain. Experiments on various heterogeneous face datasets validate the effectiveness of DisHPL. Binghui Wang, Shengbo Chen, Yiu-Ming Cheung, Wei Huang 0013 |
CIKM | 6 |
| 2022 | Multi-scale feature learning and temporal probing strategy for one-stage temporal action localizationabstractThe aim of temporal action localization (TAL) is to determine the start and end frames of an action in a video. In recent years, TAL has attracted considerable attention because of its increasing applications in video understanding and retrieval. However, precisely estimating the duration of an action in the temporal dimension is still a challenging problem. In this paper, we propose an effective one-stage TAL method based on a self-defined motion data structure, called a dense joint motion matrix (DJMM), and a novel temporal detection strategy. Our method provides three main contributions. First, compared with mainstream motion images, DJMMs can preserve more pre-processed motion features and provides more precise detail representations. Furthermore, DJMMs perfectly solve the temporal information loss problem caused by motion trajectory overlaps within a certain time period. Second, a spatial pyramid pooling (SPP) layer, which is widely used in the object detection and tracking fields, is innovatively incorporated into the proposed method for multi-scale feature learning. Moreover, the SPP layer enables the backbone convolutional neural network (CNN) to receive DJMMs of any size in the temporal dimension. Third, a large-scale-first temporal detection strategy inspired by a well-developed Chinese text segmentation algorithm is proposed to address long-duration videos. Our method is evaluated on two benchmark data sets and one self-collected data set: Florence-3D, UTKinect-Action3D and HanYue-3D. The experimental results show that our method achieves competitive action recognition accuracy and high TAL precision, and its time efficiency and few-shot learning capabilities enable it to be utilized for real-time surveillance. Leiyue Yao, Wei Huang 0013, Nan Jiang 0013, Bingbing Zhou |
Int. J. Intell. Syst. | 3 |
| 2022 | One-shot Video Graph Generation for Explainable Action Reasoning
Tao Zhuo, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001, Mohan Kankanhalli |
Neurocomputing | 4 |
| 2022 | A novel locally-constrained GAN-based ensemble to synthesize arterial spin labeling images
Wei Huang 0013, Mingyuan Luo, Jing Li 0027, Peng Zhang 0005, Yufei Zha |
Inf. Sci. | 1 |
| 2022 | MAFI: GNN-Based Multiple Aggregators and Feature Interactions Network for Fraud Detection Over Heterogeneous GraphabstractRecently, Graph Neural Networks (GNNs) have been widely used for fraud detection. GNNs first generate node embedding by aggregating neighboring information under different relations, and then use the final node embedding to detect the node’s suspiciousness. However, traditional GNNs employing only a single type of aggregator fail to capture neighbor information from multiple perspectives and treating different relations equally inevitably weakens the semantic information of heterogeneous graphs. Meanwhile, expressive ability of GNNs is limited by using conventional concatenating or averaging operations to update the center node. Also, camouflaged entities could damage GNN-based models. To handle these problems, a novel heterogeneous GNN model calledMultiple Aggregators and Feature Interactions Network(MAFI) is proposed in this paper to conduct fraud detection tasks. Concretely, multiple types of aggregators are applied on different relations to aggregate neighbor information and aggregator-level attention is utilized to learn the importance of different aggregators. Also, relation-level attention is leveraged to learn the importance of each relation. Besides, conventional update operations are replaced with vector-wise implicit and explicit feature interactions. Moreover, a trainable neighbor sampler is employed to filter camouflaged fraudsters. Comprehensive experiments on two real-world fraud datasets indicate that the proposed MAFI outperforms existing GNN-based fraud detectors. Nan Jiang 0013, Fuxian Duan, Honglong Chen, Wei Huang 0013, Ximeng Liu |
IEEE Trans. Big Data | 4 |
| 2022 | Multi-Structure KELM With Attention Fusion Strategy for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification refers to accurately corresponding each pixel in an HSI to a land-cover label. Recently, the successful application of multiscale and multifeature methods has greatly improved the performance of HSI classification due to their enhanced utilization of the available spectral–spatial information. However, as the number of scales and the number of features increases, it becomes more difficult to achieve an optimal degree of fusion for multiple classifiers [e.g., kernel extreme learning machine (KELM)]. On the other hand, a limited sample size of the HSI may cause overfitting problems, which seriously affects the classification accuracy. Therefore, in this article, a novel multi-structure KELM with attention fusion strategy (MSAF-KELM) is proposed to achieve accurate fusion of multiple classifiers for effective HSI classification with ultrasmall sample rates. First, a multi-structure network is built, which combines multiple scales and multiple features to extract abundant spectral–spatial information. Second, a fast and efficient KELM is employed to enable rapid classification. Finally, a weighted self-attention fusion strategy (WSAFS) is introduced, which combines the output weights of each KELM subbranch and the self-attention mechanism to achieve an efficient fusion result on multi-structure networks. We conducted experiments on four types of HSI datasets with different evaluation methods and compared them with several classical and state-of-the-art methods, which demonstrate the excellent performance of our method on ultrasmall sample rates. The code is available athttps://github.com/Fang666666/MSAF-KELMfor reproducibility. Le Sun 0002, Yu Fang 0012, Yuwen Chen 0001, Wei Huang 0013, Zebin Wu 0001, Byeungwoo Jeon |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Identity-Aware Facial Expression Recognition Via Deep Metric Learning Based on Synthesized ImagesabstractPerson-dependent facial expression recognition has received considerable research attention in recent years. Unfortunately, different identities can adversely influence recognition accuracy, and the recognition task becomes challenging. Other adverse factors, including limited training data and improper measures of facial expressions, can further contribute to the above dilemma. To solve these problems, a novel identity-aware method is proposed in this study. Furthermore, this study also represents the first attempt to fulfill the challenging person-dependent facial expression recognition task based on deep metric learning and facial image synthesis techniques. Technically, a StarGAN is incorporated to synthesize facial images depicting different but complete basic emotions for each identity to augment the training data. Then, a deep-convolutional-neural-network-based network is employed to automatically extract latent features from both real facial images and all synthesized facial images. Next, a Mahalanobis metric network trained based on extracted latent features outputs a learned metric that measures facial expression differences between images, and the recognition task can thus be realized. Extensive experiments based on several well-known publicly available datasets are carried out in this study for performance evaluations. Person-dependent datasets, including CK+, Oulu (all 6 subdatasets), MMI, ISAFE, ISED, etc., are all incorporated. After comparing the new method with several popular or state-of-the-art facial expression recognition methods, its superiority in person-dependent facial expression recognition can be proposed from a statistical point of view. Wei Huang 0013, Peng Zhang 0005, Yufei Zha, Yuming Fang 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Unsupervised Cross-Modal Distillation for Thermal Infrared TrackingabstractThe target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model trained on the RGB data. However, this representation does not take into account the information in the TIR modality itself, limiting the performance of TIR tracking. Jingxian Sun 0003, Lichao Zhang 0001, Yufei Zha, Abel Gonzalez-Garcia, Peng Zhang 0005, Wei Huang 0013, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2021 | Multiple object tracking based on multi-task learning with strip attentionabstractAbstract Multiple object tracking (MOT) framework based on bifurcate strategy was usually challenged by data association of different model path, which work for object localisation and appearance embedding independently. By incorporating the re‐identification (re‐ID) as appearance embedding model, more recent studies on task combination of a single network have made a great progress in tracking performance. Unfortunately, the contributive improvement from re‐ID model is hard to balance the accuracy and efficiency for the whole framework. For more effective enhancement of the overall tracking performance, a real‐time detection needs to be taken into consideration with other auxiliary means for MOT modelling. Therefore, in this study, a one‐shot multiple object tracking is proposed based on multi‐task learning to obtain satisfactory performance in both speed and robustness. With updated re‐training strategy for the backbone model of detection, a D2LA network is proposed to achieve more characteristic fine‐grained feature extraction in branching task of pedestrian recognition. Additionally, a strip attention module is also introduced to further strengthen the feature discriminative capability of the tracking framework in occlusion. Experiments on the 2DMOT15, MOT16, MOT17, and MOT20 benchmark data sets have shown a superior performance in comparison to other state‐of‐the‐art tracking approaches. Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
IET Image Process. | 3 |
| 2021 | Learning spatial-channel regularization jointly with correlation filter for visual tracking
Yufei Zha, Zhuling Qiu, Jingxian Sun 0003, Peng Zhang 0005, Wei Huang 0013 |
Neurocomputing | 5 |
| 2021 | Full-scaled deep metric learning for pedestrian re-identification
Wei Huang 0013, Mingyuan Luo, Peng Zhang 0005, Yufei Zha |
Multim. Tools Appl. | 1 |
| 2021 | A novel multi-loss-based deep adversarial network for handling challenging cases in semi-supervised image semantic segmentation
Wei Huang 0013, Zhanfei Shao, Mingyuan Luo, Peng Zhang 0005, Yufei Zha |
Pattern Recognit. Lett. | 1 |
| 2021 | TSLRLN: Tensor subspace low-rank learning with non-local prior for hyperspectral image mixed denoising
Chengxun He, Le Sun 0002, Wei Huang 0013, Jianwei Zhang 0005, Yuhui Zheng, Byeungwoo Jeon |
Signal Process. | 3 |
| 2021 | Multiple Instance Models Regression for Robust Visual TrackingabstractIn comparison to single-model based trackers, the model-ensembled tracking strategy has shown a substantial adaptivity in handling various tracking challenges. As the performance of the tracker has been improved by combining different model outputs linearly, the insufficient consideration of each ensemble member’s contribution still limits the tracking performance to be further enhanced. As the performance of the tracker has been improved by combining different model outputs linearly, the insufficient consideration of each ensemble member’s contribution still limits the tracking performance to be further enhanced. In this paper, a tracking strategy based on multiple instance models regression (MIMRT) is proposed with a unified ensembling scheme. By formulating the tracking initialization with an instance model, the encoding process for an object’s specific detail is performed corresponding to the samples in each frame. The advantage of this operation is to guarantee the model frame-wise discrimination of short-term training, as well as to evaluate the reliability of each instance model by utilizing the long-lifetime samples obtained throughout the whole tracking procedure. To finalize the proposed tracking, all the independent instance models attached to the learned regression coefficients are ensembled with respect to the long-lifetime samples. This also effectively bridges the instance model as a latent variable to investigate a semantic association between the tracking model and the overall samples. A comprehensive experiment has shown that the proposed tracker is able to achieve superior performances compared to the state-of-art tracking approaches on both short-term datasets (e.g., OTB2013, OTB100, VOT2016, UAV123) and long-term dataset (UAV20L). Yufei Zha, Yuanqiang Zhang, Tao Ku, Hanqiao Huang, Wei Huang 0013, Peng Zhang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Compressive Hyperspectral Image Reconstruction Based on Spatial-Spectral Residual Dense NetworkabstractA spatial–spectral residual dense network-based compressive hyperspectral image (HSI) reconstruction method is proposed in this letter. The proposed method contains two networks: residual dense network for hyperspectral image reconstruction (RDNHIR) and spectral difference reconstruction network (SDRN). The RDNHIR network can extract the local features and global hierarchical features by cascading features of all residual dense blocks (RDBs). Then, SDRN takes full advantage of the strong correlation between spectral adjacent bands to better preserve the spectral feature of HSI. Finally, the adjacent spectral difference regularization is introduced into the loss function to further improve the performance. The experimental results show that the proposed method has better reconstruction quality than other state-of-the-art reconstruction methods, especially in the spectral domain. Wei Huang 0013, Yang Xu 0006, Zhihui Wei |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Robust Visual Tracking based on Adversarial Unlabeled Instance Generation with Label Smoothing Loss Regularization
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Garth Douglas Cooper, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2020 | Ensemble Tracking Based on Diverse Collaborative Framework With Multi-Cue Dynamic FusionabstractTracking with deep neural networks has been verified to arrive at a new level accuracy in many challenging scenarios, but the tracking robustness has been still challenged by model singularity and self-learning loop mechanism. As a promising solution for the limitations, to ensemble diverse tracking strategies into a highly-interactive framework has shown a potential effectiveness in recent studies. In this work, a collaborative tracking framework is proposed by exploiting both discriminative correlation filters and deep classifiers into an ensembling framework. With a multi-cue dynamic fusion scheme performed on all the ensembled members’ outputs, a robust long-term tracking can be achieved by calculating the optimal robustness scores based on a dynamic weighted sum of multi-cue metrics. Meanwhile, the obtained reliable and diverse training samples are also utilized to adaptively update the tracker in each branch with heuristic frequency, which is able to alleviate the training samples’ contamination and model corruption. Experiments on the OTB-2015, Temple color 128, UAV123, VOT2016, and VOT2018 benchmark datasets have shown superior performance in comparison to other state-of-the-art tracking approaches. Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Arterial Spin Labeling Images Synthesis via Locally-Constrained WGAN-GP Ensemble
Wei Huang 0013, Mingyuan Luo, Xi Liu 0008, Peng Zhang 0005, Huijun Ding, Dong Ni 0001 |
MICCAI (4) | 1 |
| 2019 | A novel deep residual network-based incomplete information competition strategy for four-players Mahjong games
Tianwei Yan 0001, Mingyuan Luo, Wei Huang 0013 |
Multim. Tools Appl. | 4 |
| 2019 | Arterial Spin Labeling Images Synthesis From sMRI Using Unbalanced Deep Discriminant LearningabstractAdequate medical images are often indispensable in contemporary deep learning-based medical imaging studies, although the acquisition of certain image modalities may be limited due to several issues including high costs and patients issues. However, thanks to recent advances in deep learning techniques, the above tough problem can be substantially alleviated by medical images synthesis, by which various modalities including T1/T2/DTI MRI images, PET images, cardiac ultrasound images, retinal images, and so on, have already been synthesized. Unfortunately, the arterial spin labeling (ASL) image, which is an important fMRI indicator in dementia diseases diagnosis nowadays, has never been comprehensively investigated for the synthesis purpose yet. In this paper, ASL images have been successfully synthesized from structural magnetic resonance images for the first time. Technically, a novel unbalanced deep discriminant learning-based model equipped with new ResNet sub-structures is proposed to realize the synthesis of ASL images from structural magnetic resonance images. The extensive experiments have been conducted. Comprehensive statistical analyses reveal that: 1) this newly introduced model is capable to synthesize ASL images that are similar towards real ones acquired by actual scanning; 2) synthesized ASL images obtained by the new model have demonstrated outstanding performance when undergoing rigorous tests of region-based and voxel-based corrections of partial volume effects, which are essential in ASL images processing; and 3) it is also promising that the diagnosis performance of dementia diseases can be significantly improved with the help of synthesized ASL images obtained by the new model, based on a multi-modal MRI dataset containing 355 demented patients in this paper. Wei Huang 0013, Mingyuan Luo, Xi Liu 0008, Peng Zhang 0005, Huijun Ding, Wufeng Xue, Dong Ni 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Pixel-wise partial volume effects correction on arterial spin labeling magnetic resonance images
Wei Huang 0013, Chuyu Wan, Huijun Ding, Peng Zhang 0005, Guang Chen 0004 |
Multim. Tools Appl. | 1 |
| 2018 | Image-based dementia disease diagnosis via deep low-resource pair-wise learning
Wei Huang 0013, Chuyu Wan, Huijun Ding, Guang Chen 0004 |
Multim. Tools Appl. | 1 |
| 2018 | Single-target localization in video sequences using offline deep-ranked metric learning and online learned models updating
Wei Huang 0013, Peng Zhang 0005, Guang Chen 0004, Huijun Ding |
Multim. Tools Appl. | 1 |
| 2018 | Going deeper with two-stream ConvNets for action recognition in video surveillance
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 4 |
| 2018 | A novel deep multi-channel residual networks-based metric learning method for moving human localization in video surveillance
Wei Huang 0013, Huijun Ding, Guang Chen 0004 |
Signal Process. | 1 |
| 2017 | Online object tracking based on CNN with spatial-temporal saliency guided sampling
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Kangli Chen, Mohan Kankanhalli |
Neurocomputing | 3 |
| 2016 | Medical media analytics via ranking and big learning: A multi-modality image-based disease severity prediction study
Wei Huang 0013, Shuru Zeng, Guang Chen 0004 |
Neurocomputing | 1 |
| 2016 | A new image-based immersive tool for dementia diagnosis using pairwise ranking and learning
Wei Huang 0013, Shuru Zeng, Jing Li 0027, Guang Chen 0004 |
Multim. Tools Appl. | 1 |
| 2016 | A novel dementia diagnosis strategy on arterial spin labeling magnetic resonance images via pixel-wise partial volume correction and ranking
Wei Huang 0013, Peng Zhang 0005, Minmin Shen |
Multim. Tools Appl. | 1 |
| 2016 | A novel disease severity prediction scheme via big pair-wise ranking and learning techniques using image-based personal clinical data
Wei Huang 0013 |
Signal Process. | 1 |
| 2015 | Region-based image retrieval based on medical media data using ranking and multi-view learningabstractIn this study, a novel region-based image retrieval approach via ranking and multi-view learning techniques is introduced for the first time based on medical multi-modality data. A surrogate ranking evaluation measure is derived, and direct optimization via gradient ascent is carried out based on the surrogate measure to realize ranking and learning. A database composed of 1000 real patients data is constructed and several popular pattern recognition methods are implemented for performance evaluation compared with ours. It is suggested that our new method is superior to others in this medical image retrieval utilization from the statistical point of view. Wei Huang 0013, Shuru Zeng, Guang Chen 0004 |
ACII | 1 |
| 2015 | Pan-Sharpening via Coupled Unitary Dictionary Learning
Shumiao Chen, Liang Xiao 0001, Zhihui Wei, Wei Huang 0013 |
ICIG (3) | 4 |
| 2015 | A New Pan-Sharpening Method With Deep Neural NetworksabstractA deep neural network (DNN)-based new pansharpening method for the remote sensing image fusion problem is proposed in this letter. Research on representation learning suggests that the DNN can effectively model complex relationships between variables via the composition of several levels of nonlinearity. Inspired by this observation, a modified sparse denoising autoencoder (MSDA) algorithm is proposed to train the relationship between high-resolution (HR) and low-resolution (LR) image patches, which can be represented by the DNN. The HR/LR image patches only sample from the HR/LR panchromatic (PAN) images at hand, respectively, without requiring other training images. By connecting a series of MSDAs, we obtain a stacked MSDA (S-MSDA), which can effectively pretrain the DNN. Moreover, in order to better train the DNN, the entire DNN is again trained by a back-propagation algorithm after pretraining. Finally, assuming that the relationship between HR/LR multispectral (MS) image patches is the same as that between HR/LR PAN image patches, the HR MS image will be reconstructed from the observed LR MS image using the trained DNN. Comparative experimental results with several quality assessment indexes show that the proposed method outperforms other pan-sharpening methods in terms of visual perception and numerical measures. Wei Huang 0013, Liang Xiao 0001, Zhihui Wei, Hongyi Liu 0001, Songze Tang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Empirical mode decomposition based blind audio watermarking
Zhaoyang Fu, Peng Zhang 0005, Wei Huang 0013, Liang Wang 0001, Sabu Emmanuel, Guang Chen 0004 |
Multim. Tools Appl. | 3 |
| 2015 | A novel marker-less lung tumor localization strategy on low-rank fluoroscopic images with similarity learning
Wei Huang 0013, Jing Li 0027, Peng Zhang 0005, Can Fang, Minmin Shen |
Multim. Tools Appl. | 1 |
| 2015 | Interactive tracking of insect posture
Minmin Shen, Chen Li 0022, Wei Huang 0013, Paul Szyszka, Kimiaki Shirahama, Marcin Grzegorzek, Dorit Merhof, Oliver Deussen |
Pattern Recognit. | 3 |
| 2015 | Multiple pedestrian tracking based on couple-states Markov chain with semantic topic learning for video surveillance
Peng Zhang 0005, Liang Wang 0001, Wei Huang 0013, Lei Xie 0001, Guang Chen 0004 |
Soft Comput. | 3 |
| 2014 | Interactive Framework for Insect Tracking with Active LearningabstractExtracting motion trajectories of insects is an important prerequisite in many behavioral studies. Despite great efforts to design efficient automatic tracking algorithms, tracking errors are unavoidable. In this paper, we propose general principles that help to minimize the human effort required for accurate multi-target tracking in the form of applications that can track the antennae and mouthparts of a honey bee based on a set of low frame rate videos. This interactive framework estimates which key frames will require user correction, i.e. those that are used for user correction, which are used for 1) incrementally learning an object classifier and 2) data association based tracking. To this framework we apply a standard classification algorithm (i.e. naive Bayesian classification) and an association optimization algorithm (i.e. Hungarian algorithm). The precision of tracking results by our framework on real-world video data is above 98%. Minmin Shen, Wei Huang 0013, Paul Szyszka, C. Giovanni Galizia, Dorit Merhof |
ICPR | 2 |
| 2014 | Spatial-spectral compressive sensing for hyperspectral images super-resolution over learned dictionaryabstractThis paper proposes a new hyperspectral images superresolution (HSI-SR) method based on compressive sensing (CS) theory, spatial sparsity and spectral similarity prior. First, according to sparsity and incoherence of CS theory, we propose a new dictionary learning method, ensuring that the learned dictionary not only has less dimensionality to speed up the sparse decomposition, but also satisfies sparsity well. Then, we introduce the spatial sparsity and spectral similarity regularizations into HSI-SR model, which can recover the spatial information effectively and preserve the spectral information well. The experimental results show the proposed method outperforms other well-known methods in terms of both objective measurements and visual evaluation. Wei Huang 0013, Zebin Wu 0001, Hongyi Liu 0001, Liang Xiao 0001, Zhihui Wei |
IGARSS | 1 |
| 2014 | Object Tracking using Reformative Transductive Learning with Sample Variational CorrespondenceabstractTracking-by-learning strategies have effectively solved many challenging problems for visual tracking. When labeled samples are limited, the learning performance can be improved by exploiting unlabeled ones. Thus, a key issue for semi-supervised learning is the label assignment of the unlabeled samples, which is the principal focus of transductive learning. Unfortunately, the optimization scheme employed by the transductive learning is hard to be applied to online tracking because of its large amount of computation for sample labeling. In this paper, a reformative transductive learning was proposed with the variational correspondence between the learning samples, which are utilized to build an effective matching cost function for more efficient label assignment during the learning of representative separators. By using a weighted accumulative average to update the coefficients via a fixed budget of support vectors, the proposed tracking has been demonstrated to outperform most of the state-of-art trackers. Tao Zhuo, Peng Zhang 0005, Yanning Zhang 0001, Wei Huang 0013, Hichem Sahli |
ACM Multimedia | 4 |
| 2014 | Building recognition in urban environments: A survey of state-of-the-art and future challenges
Jing Li 0027, Wei Huang 0013, Ling Shao 0001, Nigel M. Allinson |
Inf. Sci. | 2 |
| 2013 | A novel marker-less tumor tracking strategyonlow-rank fluoroscopic images for image-guided lung cancer radiotherapyabstractFluoroscopic images recording the real-time motion of lung tumor lesion play an important role on lung cancer radiotherapy, as these images help to facilitate the accurate delivery of radiation dose on target tumor lesion. Derivation of tumor position in conventional lung tumor tracking strategies is realized via either placing external surrogates on patients or implanting internal fiducial markers in patients. Inaccurate tumor tracking and patient safety problems are often inevitable for these strategies. In this study, a novel marker-less tumor tracking strategy is presented for image-guided lung cancer radiotherapy. A fluoroscopic image is first decomposed into low-rank and sparse components based on robust-PCA via a split Bregman method. Then, a series of techniques, including K-means clustering, morphological processing, connected component analysis, etc are employed on obtained low-rank fluoroscopic images for tumor tracking. Clinical data obtained from 45 patients is incorporated for experimental evaluation. Promising results are demonstrated from the introduced strategy. Wei Huang 0013, Jing Li 0027, Peng Zhang 0005 |
ICIP | 1 |
| 2013 | Non-rigid target tracking based on 'flow-cut' in pair-wise frames with online hough forestsabstractIn conventional online learning based tracking studies, fixed-shape appearance modeling is often incorporated for training samples generation, as it is simple and convenient to be applied. However, for more general non-rigid and articulated object, this strategy may regard some background areas as foreground, which is likely to deteriorate the learning process. Recently published works utilize more than one patches to represent non-rigid object with foreground object segmentation, but most of these segmentation for target representation are performed only in single frame manner. Since the motion information between the consecutive frames was not considered by these approaches, when the backgrounds are similar to the target, accurate segmentation is hard to be achieved. In this work, we propose a novel model for non-rigid object segmentation by incorporating consecutive gradients flow between pair-wise frames into a Gibbs energy function. With help from motion information, the irregular target areas can be segmented more accurately during precise boundary convergence. The proposed segmentation model is incorporated into a semi-supervised online tracking framework for training samples generation. We test the proposed tracking on challenging videos involving heavy intrinsic variations and occlusions. As a result, the experiments demonstrate a significant improvement in tracking accuracy and robustness in comparison with other state-of-art tracking works. Yanning Zhang 0001, Peng Zhang 0005, Wei Huang 0013, Hichem Sahli |
ACM Multimedia | 4 |
| 2011 | A Computer Assisted Method for Nuclear Cataract Grading From Slit-Lamp Images Using RankingabstractIn clinical diagnosis, a grade indicating the severity of nuclear cataract is often manually assigned by a trained ophthalmologist to a patient after comparing the lens' opacity severity in his/her slit-lamp images with a set of standard photos. This grading scheme is often subjective and time-consuming. In this paper, a novel computer-aided diagnosis method via ranking is proposed to facilitate nuclear cataract grading following conventional clinical decision-making process. The grade of nuclear cataract in a slit-lamp image is predicted using its neighboring labeled images in a ranked image list, which is achieved using a learned ranking function. This ranking function is learned via direct optimization on a newly proposed approximation to a ranking evaluation measure. Our proposed method has been evaluated by a large dataset composed of 1000 different cases, which are collected from an ongoing clinical population-based study. Both experimental results and comparison with several existing methods demonstrate the benefit of grading via ranking by our proposed method. Wei Huang 0013, Kap Luk Chan, Huiqi Li, Joo-Hwee Lim, Jiang Liu 0001, Tien Yin Wong |
IEEE Trans. Medical Imaging | 1 |
| 2009 | A Computer-Aided Diagnosis System of Nuclear Cataract via Ranking
Wei Huang 0013, Huiqi Li, Kap Luk Chan, Joo-Hwee Lim, Jiang Liu 0001, Tien Yin Wong |
MICCAI (1) | 1 |
| 2008 | Semi-supervised Nasopharyngeal Carcinoma Lesion Extraction from Magnetic Resonance Images Using Online Spectral Clustering with a Learned Metric
Wei Huang 0013, Kap Luk Chan, Jiayin Zhou, Vincent Chong |
MICCAI (1) | 1 |