Luwen Huangfu

dblp:119/3764 · also Vivian L. Huangfu · DBLP profile ↗
← Back
39ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0003-3926-7901ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DCE-MBR: Multi-behavior recommendation via hierarchical denoising and cascade enhancement
abstract
Graph neural networks (GNNs) have shown considerable promise in multi-behavior recommendation tasks, particularly for target behavior prediction (e.g., purchase conversion), by effectively integrating auxiliary behavioral signals (e.g., item browsing, cart addition). Recent advances in the field have substantially improved the modeling of hierarchical interactions and multi-behavior dependencies, effectively mitigating foundational challenges such as data sparsity. However, these state-of-the-art methods frequently overlook the semantic heterogeneity and inherent noise within auxiliary interactions, often resorting to uniform or simplistic denoising strategies that risk discarding valuable signals. To overcome this persistent limitation, the Denoising Cascade-Enhanced Multi-Behavior Recommendation (DCE-MBR) framework is introduced. DCE-MBR is designed to simultaneously suppress noise and preserve semantically informative interactions through a dual-stage architecture. First, a hierarchical graph denoising module dynamically removes noisy edges by applying behavior-specific thresholds across multiple levels of granularity, thereby preserving essential neighbor relations. Next, a cascade-enhanced module incrementally refines user preferences by propagating target behavior signals through auxiliary behavior paths, leading to improved feature representations. Comprehensive evaluations based on the Taobao as well as Tmall data collections show that DCE-MBR outperforms state-of-the-art baselines, achieving relative gains of 18.66% in Hit@10 and 15.42% in NDCG@10. These results confirm the model’s robustness against noisy interactions and its effectiveness in capturing intricate multi-behavior dependencies. The source code is publicly available at DCEMBR 1 .
Shuangdi Ma, Wei Zhou 0028, Jun Zeng 0003, Junhao Wen 0001, Luwen Huangfu
Expert Syst. Appl.5
2026 Multi-granularity preference enhancement with hierarchical feature extraction for session-based recommendations
abstract
Session-based recommendation predicts the next item a user will interact with based on their short-term session behavior, typically without long-term user profiles. Existing approaches often fail to capture the hierarchical nature of user preferences, leading to suboptimal personalization and limited recommendation accuracy. In this work, we argue that user preferences exhibit coarse-grained and fine-grained characteristics, and item features should be modeled accordingly across these two levels to capture users' preference signals more accurately. To this end, we propose a novel method, Multi-Granularity Preference Enhancement with Hierarchical Feature Extraction (MPEHFE), for session-based recommendation. MPEHFE explicitly captures semantic item relationships at each granularity and enhances fine-grained preference modeling through a differentiable architecture search mechanism. It also identifies interactions inconsistent with the user's general intent as noise, leveraging contrastive learning to reinforce the representation of coarse-grained preferences. Moreover, experiments on three real-world benchmark datasets demonstrate that MPEHFE consistently outperforms state-of-the-art baselines, achieving relative improvements of 3%-9% in P@20 and 11%-56% in MRR@20.
Yongjian Zhou, Wei Zhou 0028, Luwen Huangfu, Jun Zeng 0003, Tingyue He, Junhao Wen 0001
Neural Networks3
2026 Corner Case Detection and Generation for Autonomous Driving: An Overview
abstract
Safety concerns remain one of the most significant obstacles to the large-scale deployment and continued advancement of autonomous driving (AD) systems. A major underlying cause of many safety-related incidents in AD systems is suboptimal or erroneous decision-making when the vehicle encounters corner cases (CCs)—rare, unexpected, or extreme situations that fall outside typical operating conditions. Although recent advances in artificial intelligence have driven substantial progress in both autonomous driving and corner-case research, the field still lacks a coherent conceptual foundation and a systematic, widely accepted categorization of CCs. In this survey, we address this gap by offering a structured review of the existing corner-case literature along three key dimensions: understanding, detection, and generation. The key contribution is a three-level classification of corner cases—spanning data-level, model-level, and semantic-level CCs—that clarifies and disambiguates competing definitions and perspectives on AD corner cases. Building on this framework, we examine simulator-based methods for corner-case selection and data generation, and we identify open challenges, promising directions, and potential solutions to more effectively handle corner cases in autonomous driving systems.
Yunji Liang, Junteng Liu, Xiaokai Yan, Xiaolong Zheng 0001, Lei Tang 0002, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001
IEEE Trans. Intell. Transp. Syst.6
2026 Hard Sample Mining: A New Paradigm of Efficient and Robust Model Training
abstract
Over the past two decades, deep learning (DL) has achieved unprecedented breakthroughs across diverse application domains spanning computer vision (CV) to natural language processing (NLP). However, despite significant advances in computational resources and algorithmic frameworks, the training of deep neural networks continues to present formidable challenges due to persistent issues of training inefficiency and inherent data distribution biases. Recent years have witnessed the emergence of hard sample mining (HSM) as a promising paradigm to mitigate training inefficiencies and enhance model robustness through representative sample selection. Although HSM is reshaping contemporary AI research, its critical role in enabling efficient and robust model training has not yet been systematically explored. This article presents a comprehensive survey of HSM methodologies by: 1) establishing unified definitions of hard samples through rigorous sample complexity quantification criteria; 2) proposing a systematic taxonomy of HSM approaches with in-depth technical analysis; and 3) identifying pivotal research frontiers in this evolving field. This survey not only consolidates the foundations of HSM but also provides a roadmap for advancing efficient, robust, and generalizable deep learning models.
Lei Liu 0073, Yunji Liang, Xiaokai Yan, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001, Yanyong Zhang, Daniel Dajun Zeng
IEEE Trans. Neural Networks Learn. Syst.4
2025 Learning complementary visual information for few-shot food recognition by Regional Erasure and Reactivation
Yi Zhang 0113, Luwen Huangfu, Lili Balazs, Sheng Huang 0001
Expert Syst. Appl.3
2025 GNPSum: A code summarization enhancement framework based on Graph Node Position
Haogang Cheng, Luwen Huangfu, Chao Liu 0014, Meng Yan 0001, Yan Lei 0005
Inf. Softw. Technol.3
2025 Rethinking the sample relations for few-shot classification
Guowei Yin, Sheng Huang 0001, Luwen Huangfu, Yi Zhang 0113, Xiaohong Zhang 0002
Image Vis. Comput.3
2025 Enhanced prototype network with gated point recyclable feature mining for few-shot 3D point cloud classification
Sheng Huang 0001, Luwen Huangfu, Ma Rui, Bo Liu 0005
Knowl. Based Syst.3
2025 Learning Feature Exploration and Selection With Handcrafted Features for Few-Shot Learning
abstract
Interest in few-shot learning (FSL) has grown recently, but the value of feature learning, which bridges the gap between base and novel classes, remains largely understudied. The limited availability of labeled samples for each class poses a major challenge. To tackle this, we propose a simple yet effective approach called deep discriminative handcrafted feature regression (DDHFR) to explore intrinsic information and select improved discriminative features in few-shot data by mining knowledge from classical handcrafted features. To explore intrinsic information, we design several deep handcrafted feature regression (DHFR) modules and plugged them separately into different layers of the backbone to use feature engineering knowledge for feature learning optimization at different granularities. To achieve discriminative feature selection, we incorporate an auxiliary classifier (AC) into each DHFR module to enhance the acquisition of discriminative information. Furthermore, we employed self-distillation to boost ability of ACs ot be classified. Experimental results in three backbones on three datasets show that DDHFR can generally improve the performance of existing FSL methods. On average, it improves the recognition accuracy by 1.16% in two common few-shot settings.
Yi Zhang 0113, Sheng Huang 0001, Luwen Huangfu, Daniel Dajun Zeng
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Data Distribution Distilled Generative Model for Generalized Zero-Shot Recognition
abstract
In the realm of Zero-Shot Learning (ZSL), we address biases in Generalized Zero-Shot Learning (GZSL) models, which favor seen data. To counter this, we introduce an end-to-end generative GZSL framework called D3GZSL. This framework respects seen and synthesized unseen data as in-distribution and out-of-distribution data, respectively, for a more balanced model. D3GZSL comprises two core modules: in-distribution dual space distillation (ID2SD) and out-of-distribution batch distillation (O2DBD). ID2SD aligns teacher-student outcomes in embedding and label spaces, enhancing learning coherence. O2DBD introduces low-dimensional out-of-distribution representations per batch sample, capturing shared structures between seen and un seen categories. Our approach demonstrates its effectiveness across established GZSL benchmarks, seamlessly integrating into mainstream generative frameworks. Extensive experiments consistently showcase that D3GZSL elevates the performance of existing generative GZSL methods, under scoring its potential to refine zero-shot learning practices. The code is available at: https://github.com/PJBQ/D3GZSL.git
Mingjian Hong, Luwen Huangfu, Sheng Huang 0001
AAAI3
2024 SAM-MIL: A Spatial Contextual Aware Multiple Instance Learning Approach for Whole Slide Image Classification
abstract
Multiple Instance Learning (MIL) represents the predominant framework in Whole Slide Image (WSI) classification, covering aspects such as sub-typing, diagnosis, and beyond. Current MIL models predominantly rely on instance-level features derived from pretrained models such as ResNet. These models segment each WSI into independent patches and extract features from these local patches, leading to a significant loss of global spatial context and restricting the model's focus to merely local features. To address this issue, we propose a novel MIL framework, named SAM-MIL, that emphasizes spatial contextual awareness and explicitly incorporates spatial context by extracting comprehensive, image-level information. The Segment Anything Model (SAM) represents a pioneering visual segmentation foundational model that can capture segmentation features without the need for additional fine-tuning, rendering it an outstanding tool for extracting spatial context directly from raw WSIs. Our approach includes the design of group feature extraction based on spatial context and a SAM-Guided Group Masking strategy to mitigate class imbalance issues. We implement a dynamic mask ratio for different segmentation categories and supplement these with representative group features of categories. Moreover, SAM-MIL divides instances to generate additional pseudo-bags, thereby augmenting the training set, and introduces consistency of spatial context across pseudo-bags to further enhance the model's performance. Experimental results on the CAMELYON-16 and TCGA Lung Cancer datasets demonstrate that our proposed SAM-MIL model outperforms existing mainstream methods in WSIs classification. Our open-source implementation code is is available at https://github.com/FangHeng/SAM-MIL.
Sheng Huang 0001, Luwen Huangfu, Bo Liu 0005
ACM Multimedia4
2024 Category-Prompt Refined Feature Learning for Long-Tailed Multi-Label Image Classification
abstract
Real-world data consistently exhibits a long-tailed distribution, often spanning multiple categories. This complexity underscores the challenge of content comprehension, particularly in scenarios requiring Long-Tailed Multi-Label image Classification (LTMLC). In such contexts, imbalanced data distribution and multi-object recognition pose significant hurdles. To address this issue, we propose a novel and effective approach for LTMLC, termed Category-Prompt Refined Feature Learning (CPRFL), utilizing semantic correlations between different categories and decoupling category-specific visual representations for each category. Specifically, CPRFL initializes category-prompts from the pretrained CLIP's embeddings and decouples category-specific visual representations through interaction with visual features, thereby facilitating the establishment of semantic correlations between the head and tail classes. To mitigate the visual-semantic domain bias, we design a progressive Dual-Path Back-Propagation mechanism to refine the prompts by progressively incorporating context-related visual information into prompts. Simultaneously, the refinement process facilitates the progressive purification of the category-specific visual representations under the guidance of the refined prompts. Furthermore, taking into account the negative-positive sample imbalance, we adopt the Asymmetric Loss as our optimization objective to suppress negative samples across all classes and potentially enhance the head-to-tail recognition performance. We validate the effectiveness of our method on two LTMLC benchmarks and extensive experiments demonstrate the superiority of our work over baselines.The code is available at https://github.com/jiexuanyan/CPRFL.
Jiexuan Yan, Sheng Huang 0001, Nankun Mu, Luwen Huangfu, Bo Liu 0005
ACM Multimedia4
2024 Cross-Modality 3D Multiobject Tracking Under Adverse Weather via Adaptive Hard Sample Mining
abstract
3-D multiobject tracking (MOT) is an important task in numerous applications, including robotics and autonomous driving. Nevertheless, existing 3-D MOT solutions suffer from significant performance degradation under adverse weather conditions. Inspired by the fact that hard objects (e.g., missed detections or wrongly-associated objects) are more constructive to performance improvement, in this article, we leverage hard samples for robust 3-D MOT in adverse weather conditions. Specifically, we implement a cross-modality 3-D MOT framework to learn the 3-D region proposals from point clouds and RGB images, respectively. To minimize the risk of missed detection and wrong association, we introduce an adaptive hard sample mining scheme to align the 3-D region proposals provided by two modalities. We quantify the hard level by comparing the confidence values of the same object in the two branches and their distance in the embedding space. Meanwhile, we dynamically adjust the weights of hard samples during training to enhance the representation learning for robust 3-D MOT. Extensive experimental results showcase that our proposed solution effectively mitigates the missed detection and reduces wrong association with good generalization.
Lifeng Qiao, Peng Zhang 0139, Yunji Liang, Xiaokai Yan, Luwen Huangfu, Xiaolong Zheng 0001, Zhiwen Yu 0001
IEEE Internet Things J.5
2024 Code semantic enrichment for deep code search
Zhongyang Deng, Chao Liu 0014, Luwen Huangfu, Meng Yan 0001
J. Syst. Softw.4
2024 Query-oriented two-stage attention-based model for code search
Huanhuan Yang, Chao Liu 0014, Luwen Huangfu
J. Syst. Softw.4
2024 Learning Entangled Interactions of Complex Causality via Self-Paced Contrastive Learning
abstract
Learning causality from large-scale text corpora is an important task with numerous applications—for example, in finance, biology, medicine, and scientific discovery. Prior studies have focused mainly on simple causality, which only includes one cause-effect pair. However, causality is notoriously difficult to understand and analyze because of multiple cause spans and their entangled interactions. To detect complex causality, we propose a self-paced contrastive learning model, namely N2NCause, to learn entangled interactions between multiple spans. Specifically, N2NCause introduces data enhancement operations to convert implicit expressions into explicit expressions with the most rational causal connectives for the synthesis of positive samples and to invert the directed connection between a cause-effect pair for the synthesis of negative samples. To learn the semantic dependency and causal direction of positive and negative samples, self-paced contrastive learning is proposed to learn the entangled interactions among spans, including the interaction direction and interaction field. We evaluated the performance of N2NCause in three cause-effect detection tasks. The experimental results show that, with the least data annotation efforts, N2NCause demonstrates competitive performance in detecting simple cause-effect relations, and it is superior to existing solutions for the detection of complex causality.
Yunji Liang, Lei Liu 0073, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001, Daniel Dajun Zeng
ACM Trans. Knowl. Discov. Data3
2023 ASDFL: An adaptive super-pixel discriminative feature-selective learning for vehicle matching
abstract
Abstract There are a large number of cameras in modern transportation system that capture numerous vehicle images continuously. Therefore, automatic analysis of these vehicle images is helpful for traffic flow management, criminal investigations and vehicle inspections. Vehicle matching, which aims to determine whether two input images depict an identical vehicle, is one of the core tasks in vehicle analysis. Recent relevant studies have focused on local feature extraction instead of global extraction, since local details can provide crucial cues to distinguish between cars. However, these methods do not select local features; that is, they do not assign weights to local features. Therefore, in this research, we systematically study the vehicle matching task, and present a novel annotation‐free local‐based deep learning method called Adaptive super‐pixel discriminative feature‐selective learning (ASDFL) to address this issue. In ASDFL, vehicle images are segmented into clusters of super‐pixels of similar size by considering the location and colour similarities of pixels without using any component‐level annotation. These super‐pixels are deemed to be the virtual components of vehicles. Moreover, a convolutional neural network is used to extract the deep features of these virtual components. Thereafter, an instance‐specific mask generation module driven by the extracted global features is enhanced to produce a mask to select the most distinctive virtual components of each vehicle image pair in the feature space. Finally, the vehicle matching task is accomplished by classifying the selected virtual component features of each imaged vehicle pair. Extensive experiments on two popular vehicle identification benchmarks demonstrate that our method is 1.57% and 0.8% more accurate than the previous baselines in a vehicle matching task on the VeRi and VehicleID datasets, respectively, which demonstrates the effectiveness of our method.
Rong Qin 0001, Huanhuan Lv, Yi Zhang 0113, Luwen Huangfu, Sheng Huang 0001
Expert Syst. J. Knowl. Eng.4
2023 Identifying emotional causes of mental disorders from social media for effective intervention
abstract
Identifying the emotional causes of mental illnesses is key to effective intervention. Existing emotion-cause analysis approaches can effectively detect simple emotion-cause expressions where only one cause and one emotion exist. However, emotions may often result from multiple causes, implicitly or explicitly, with complex interactions among these causes. Moreover, the same causes may result in multiple emotions. How to model the complex interactions between multiple emotion spans and cause spans remains under-explored. To tackle this problem, a contrastive learning-based framework is presented to detect the complex emotion-cause pairs with the introduction of negative samples and positive samples. Additionally, we developed a large-scale emotion-cause dataset with complex emotion-cause instances based on subreddits associated with mental health. Our proposed approach was compared to prevailing CNN-based, LSTM-based, Transformer-based and GNN-based methods. Extensive experiments have been conducted and the quantifiable outcomes indicate that our proposed solution achieves competitive performance on simple emotion-cause pairs and significantly outperformed baseline methods in extracting complex emotion-cause pairs. Empirical studies further demonstrated that our proposed approach can be used to reveal the emotional causes of mental disorders for effective intervention.
Yunji Liang, Lei Liu 0073, Yapeng Ji, Luwen Huangfu, Daniel Dajun Zeng
Inf. Process. Manag.4
2023 MSTIL: Multi-cue Shape-aware Transferable Imbalance Learning for effective graphic API recommendation
abstract
Application Programming Interface (API) recommendation based on graphs is a valuable task in the fields of data visualization and software engineering. However, this task was previously undefined until a recently published paper coining the task as Plot2API and utilizing a deep learning-based method named SPGNN. Compared to general image classification methods, this dedicated approach uses semantic parsing to exploit deep features and yields better performance. However, its performance declines sharply in unbalanced datasets, thus limiting its generalizability. To address this issue, we propose a method named Multi-cue Shape and software engineering-aware Transferable Imbalance Learning (MSTIL), consisting of three major components: Cross-Language Shape-Aware Plot Transfer Learning (CLSAPTL), Cross-Language API Semantic Similarity-based Data Augmentation (CLASSDA), and Imbalance Plot2API Learning (IPL). Motivated by the hierarchical classification of the graphs, CLSAPTL guides the model to learn the graphs’ class hierarchy and thereby enabling the model to learn more transferable visual features. Given that a graph can be associated with multiple APIs and motivated by the fact that many APIs that exert similar functions in different languages have semantically similar names, CLASSDA leverages the samples of APIs with semantically similar names to assist in feature learning. Finally, inspired by the essence of softmax cross entropy loss, IPL alleviates the imbalances between positive and negative samples during training. We conduct our experiments on two public datasets. Extensive experimental results shows that MSTIL improves the performance of classic CNNs along with the state-of-the-art method, demonstrating its effectiveness. Specifically, MSTIL has an average relative mAP improvement of 12.94% across the models on all datasets.
Rong Qin 0001, Zeyu Wang 0001, Sheng Huang 0001, Luwen Huangfu
J. Syst. Softw.4
2023 Adaptively Weighted k-Tuple Metric Network for Kinship Verification
abstract
Facial image-based kinship verification is a rapidly growing field in computer vision and biometrics. The key to determining whether a pair of facial images has a kin relation is to train a model that can enlarge the margin between the faces that have no kin relation while reducing the distance between faces that have a kin relation. Most existing approaches primarily exploit duplet (i.e., two input samples without cross pair) or triplet (i.e., single negative pair for each positive pair with low-order cross pair) information, omitting discriminative features from multiple negative pairs. These approaches suffer from weak generalizability, resulting in unsatisfactory performance. Inspired by human visual systems that incorporate both low-order and high-order cross-pair information from local and global perspectives, we propose to leverage high-order cross-pair features and develop a novel end-to-end deep learning model called the adaptively weighted k -tuple metric network (AW k -TMN). Our main contributions are three-fold. First, a novel cross-pair metric learning loss based on k -tuplet loss is introduced. It naturally captures both the low-order and high-order discriminative features from multiple negative pairs. Second, an adaptively weighted scheme is formulated to better highlight hard negative examples among multiple negative pairs, leading to enhanced performance. Third, the model utilizes multiple levels of convolutional features and jointly optimizes feature and metric learning to further exploit the low-order and high-order representational power. Extensive experimental results on three popular kinship verification datasets demonstrate the effectiveness of our proposed AW k -TMN approach compared with several state-of-the-art approaches. The source codes and models are released.1.
Sheng Huang 0001, Jingkai Lin, Luwen Huangfu, Junlin Hu 0001, Daniel Dajun Zeng
IEEE Trans. Cybern.3
2023 An Escalated Eavesdropping Attack on Mobile Devices via Low-Resolution Vibration Signals
abstract
With the global prevalence of mobile devices, concerns about mobile devices regarding privacy breaches and data leakage are rising. Although sensor permissions are required for mobile applications to access outputs of built-in sensors, motion sensors (e.g., accelerometer and gyroscope) can be visited directly without permission requirement. Extant studies have shown that motion sensors may cause breaches of confidential information, such as passwords, digits, and voice-based commands, but whether it is possible to synthesize intelligible speech waveforms from low-resolution motion sensors has been understudied. In this article, we present an escalated side-channel attack of built-in speakers by synthesizing intelligible speech waveforms from low-resolution vibration signals. Opposite to traditional classification problems, we formulate this task as a generative problem and introduce an end-to-end synthesis framework dubbed asAccMyrinxto eavesdrop on the speaker via the low-resolution vibration signals. InAccMyrinx, we introduce the data alignment solution to provide the pair-wise voice-vibration sequences and present wavelet-based MelGAN (WMelGAN) with multi-scale time-frequency domain discriminators to generate intelligible acoustic waveforms. We conducted intensive experiments and demonstrated the feasibility of synthesizing the intelligible acoustic signals from low-resolution solid-borne vibration signals. Compared with existing synthesis solutions, our proposed solution outperforms the baselines in both subject and object metrics with the smoothed word error rate of 42.67% and the Mel-Cepstral distortion of 0.298. In addition, the quality of synthetic speeches could be impacted by several factors, including gender, speech rate, volume, and sampling frequency.
Yunji Liang, Yuchen Qin, Qi Li 0048, Xiaokai Yan, Luwen Huangfu, Sagar Samtani, Bin Guo 0001, Zhiwen Yu 0001
IEEE Trans. Dependable Secur. Comput.5
2023 Weakly Supervised Patch Label Inference Networks for Efficient Pavement Distress Detection and Recognition in the Wild
abstract
Automatic image-based pavement distress detection and recognition are vital for pavement maintenance and management. However, existing deep learning-based methods largely omit the specific characteristics of pavement images, such as high image resolution and low distress area ratio, and are not end-to-end trainable. In this paper, we present a series of simple yet effective end-to-end deep learning approaches named Weakly Supervised Patch Label Inference Networks (WSPLIN) for efficiently addressing these tasks under various application settings. WSPLIN transforms the fully supervised pavement image classification problem into a weakly supervised pavement patch classification problem for solutions. Specifically, WSPLIN first divides the pavement image under different scales into patches with different collection strategies and then employs a Patch Label Inference Network (PLIN) to infer the labels of these patches to fully exploit the resolution and scale information. Notably, we design a patch label sparsity constraint based on the prior knowledge of distress distribution and leverage the Comprehensive Decision Network (CDN) to guide the training of PLIN in a weakly supervised way. Therefore, the patch labels produced by PLIN provide interpretable intermediate information, such as the rough location and the type of distress. We evaluate our method on a large-scale bituminous pavement distress dataset named CQU-BPDD and the augmented Crack500 (Crack500-PDD) dataset, which is a newly constructed pavement distress detection dataset augmented from the Crack500. Extensive results demonstrate the superiority of our method over baselines in both performance and efficiency. The source codes of WSPLIN are released onhttps://github.com/DearCaat/wsplin.
Sheng Huang 0001, Guixin Huang, Luwen Huangfu, Dan Yang 0001
IEEE Trans. Intell. Transp. Syst.4
2023 DeepApp: characterizing dynamic user interests for mobile application recommendation
Yunji Liang, Lei Liu 0073, Luwen Huangfu, Zhu Wang 0001, Bin Guo 0001
World Wide Web (WWW)3
2022 Kernel Inversed Pyramidal Resizing Network for Efficient Pavement Distress Recognition
Rong Qin 0001, Luwen Huangfu, Devon Hood, James Ma, Sheng Huang 0001
ICONIP (6)2
2022 Boosting Multi-Label Image Classification with Complementary Parallel Self-Distillation
abstract
Multi-Label Image Classification (MLIC) appro-aches usually exploit label correlations to achieve good performance. However, emphasizing correlation like co-occurrence may overlook discriminative features and lead to model overfitting. In this study, we propose a generic framework named Parallel Self-Distillation (PSD) for boosting MLIC models. PSD decomposes the original MLIC task into several simpler MLIC sub-tasks via two elaborated complementary task decomposition strategies named Co-occurrence Graph Partition (CGP) and Dis-occurrence Graph Partition (DGP). Then, the MLIC models of fewer categories are trained with these sub-tasks in parallel for respectively learning the joint patterns and the category-specific patterns of labels. Finally, knowledge distillation is leveraged to learn a compact global ensemble of full categories with these learned patterns for reconciling the label correlation exploitation and model overfitting. Extensive results on MS-COCO and NUS-WIDE datasets demonstrate that our framework can be easily plugged into many MLIC approaches and improve performances of recent state-of-the-art approaches. The source code is released at https://github.com/Robbie-Xu/CPSD.
Jiazhi Xu, Sheng Huang 0001, Fengtao Zhou, Luwen Huangfu, Daniel Dajun Zeng, Bo Liu 0005
IJCAI4
2022 PicT: A Slim Weakly Supervised Vision Transformer for Pavement Distress Classification
abstract
Automatic pavement distress classification facilitates improving the efficiency of pavement maintenance and reducing the cost of labor and resources. A recently influential branch of this task divides the pavement image into patches and infers the patch labels for addressing these issues from the perspective of multi-instance learning. However, these methods neglect the correlation between patches and suffer from a low efficiency in the model optimization and inference. As a representative approach of vision Transformer, Swin Transformer is able to address both of these issues. It first provides a succinct and efficient framework for encoding the divided patches as visual tokens, then employs self-attention to model their relations. Built upon Swin Transformer, we present a novel vision Transformer named Pavement Image Classification Transformer (PicT) for pavement distress classification. In order to better exploit the discriminative information of pavement images at the patch level, the Patch Labeling Teacher is proposed to leverage a teacher model to dynamically generate pseudo labels of patches from image labels during each iteration, and guides the model to learn the discriminative features of patches via patch label inference in a weakly supervised manner. The broad classification head of Swin Transformer may dilute the discriminative features of distressed patches in the feature aggregation step due to the small distressed area ratio of the pavement image. To overcome this drawback, we present a Patch Refiner to cluster patches into different groups and only select the highest distress-risk group to yield a slim head for the final image classification. We evaluate our method on a large-scale bituminous pavement distress dataset named CQU-BPDD. Extensive results demonstrate the superiority of our method over baselines and also show that PicT outperforms the second-best performed model by a large margin of +2.4% in [email protected] on detection task, +3.9% in F1 on recognition task, and 1.8x throughput, while enjoying 7x faster training speed using the same computing resources. Our codes and models have been released on https://github.com/DearCaat/PicT.
Sheng Huang 0001, Xiaoxian Zhang, Luwen Huangfu
ACM Multimedia4
2022 Multi-label out-of-distribution detection via exploiting sparsity and co-occurrence of labels
Lei Wang 0062, Sheng Huang 0001, Luwen Huangfu, Bo Liu 0005, Xiaohong Zhang 0002
Image Vis. Comput.3
2022 An Iteratively Optimized Patch Label Inference Network for Automatic Pavement Distress Detection
abstract
We present a novel deep learning framework named the Iteratively Optimized Patch Label Inference Network (IOPLIN) for automatically detecting various pavement distresses that are not solely limited to specific ones, such as cracks and potholes. IOPLIN can be iteratively trained with only the image label via the Expectation-Maximization Inspired Patch Label Distillation (EMIPLD) strategy, and accomplish this task well by inferring the labels of patches from the pavement images. IOPLIN enjoys many desirable properties over the state-of-the-art single branch CNN models such as GoogLeNet and EfficientNet. It is able to handle images in different resolutions, and sufficiently utilize image information particularly for the high-resolution ones, since IOPLIN extracts the visual features from unrevised image patches instead of the resized entire image. Moreover, it can roughly localize the pavement distress without using any prior localization information in the training phase. In order to better evaluate the effectiveness of our method in practice, we construct a large-scale Bituminous Pavement Disease Detection dataset named CQU-BPDD consisting of 60,059 high-resolution pavement images, which are acquired from different areas at different times. Extensive results on this dataset demonstrate the superiority of IOPLIN over the state-of-the-art image classification approaches in automatic pavement distress detection. The source codes of IOPLIN are released onhttps://github.com/DearCaat/ioplin, and the CQU-BPDD dataset is able to be accessed onhttps://dearcaat.github.io/CQU-BPDD/.
Sheng Huang 0001, Qiming Zhao, Luwen Huangfu
IEEE Trans. Intell. Transp. Syst.5
2021 Weakly Supervised Patch Label Inference Network with Image Pyramid for Pavement Diseases Recognition in the Wild
abstract
Automatic pavement disease recognition is vital for pavement maintenance and management. In this paper, we present an end-to-end deep learning approach named Weakly Super-vised Patch Label Inference Network with Image Pyramid (WSPLIN-IP) for recognizing various types of pavement diseases that are not just limited to the specific ones, such as crack and pothole. WSPLIN-IP first divides the pavement image into patches with an image pyramid for fully exploiting the resolution and scale information. Then, a Patch Label Inference Network (PLIN) is employed for inferring the labels of these patches constrained with a patch label sparsity loss. Finally, the patch labels are fed into a Comprehensive Decision Network (CDN) for disease recognition. Since only the image label is available during whole training, the training of PLIN is conducted in a weakly supervised way under the guidance of CDN and the trained PLIN can provide the interpretable intermediate information. We evaluate our method on a large-scale Bituminous Pavement Disease Dataset named CQU-BPDD whose samples are acquired in the real world. Extensive results demonstrate the superiority of our method over baselines.
Guixin Huang, Sheng Huang 0001, Luwen Huangfu, Dan Yang 0001
ICASSP3
2021 Domain-oriented News Recommendation in Security Applications
abstract
The unprecedented growth of information on the Internet has brought about the problem of information overload. To alleviate this problem, news recommendation aims to select news articles for users according to their personal interests. In security applications such as intelligence collection and public opinion monitoring, it is of great importance to obtain valuable information quickly from massive news resources. Different from other application settings, users in security-related scenarios tend to browse news with a domain-oriented purpose. In contrast to the existing news recommendation methods which focus on general-purpose solutions, news recommendation in security applications needs domain-oriented solutions to incorporate users’ interests in a specific domain. To this end, in this paper, we propose the problem of domain-oriented news recommendation and develop a specific news recommendation model for security applications. Specifically, our proposed Domain-oriented News Recommendation (DNR) model extracts both general and specific preferences of the user, and performs matching between the user and the candidate news from the above two aspects to combine into the final result. We construct three security-related datasets using a large-scale real-world dataset and validate the effectiveness of our method.
Qingchao Kong, Luwen Huangfu
ISI3
2020 Robust Bidirectional Generative Network For Generalized Zero-Shot Learning
abstract
In this work, we propose a novel generative approach named Robust Bidirectional Generative Network (RBGN) based on Conditional Generative Adversarial Network (CGAN) for Generalized Zero-shot Learning (GZSL). RBGN employs the adversarial attack to train a more rigorous discriminator, thus enhancing the generalizability and robustness of the feature generator under minimax strategy. Moreover, RBGN decodes the generated visual features back to their semantic representations to further improve the representational ability of generated visual features and alleviate the hubness problem. The experimental results of GZSL on four datasets, i.e. CUB, SUN, AWA1, AWA2, demonstrate that our model achieves competitive performance compared to state-of-the-art approaches and owns better generalizability to the unseen classes over conventional generative GZSL models. Further robustness analysis also validates the strong robustness of our model to the different types of semantic disturbance.
Sheng Huang 0001, Luwen Huangfu, Feiyu Chen 0002, Yongxin Ge
ICME3
2020 Corner detection using the point-to-centroid distance technique
abstract
Corners, highly important local features of images and corner finding, play a crucial role in computer vision and image processing, such as object tracking and vehicle detection. Proposing effective and efficient corner detectors is the aim of corner detection. In this study, the authors first present a new measure of corner sharpness termed as the point‐to‐centroid distance (PCD) and then examine its behaviours, which display beneficial characteristics that help distinguish corners from non‐corners. Based on PCD behaviours, the authors propose a novel corner detector. Extensive experimental results demonstrate that the PCD technique is effective and simultaneously efficient for corner detection compared with six other contour‐based corner detectors in terms of two commonly used evaluation metrics – average repeatability and localisation error.
Shizheng Zhang, Luwen Huangfu, Zhifeng Zhang 0002, Sheng Huang 0001, Heng Wang 0004
IET Image Process.2
2020 Class-Prototype Discriminative Network for Generalized Zero-Shot Learning
abstract
We present a novel end-to-end deep metric learning model named Class-Prototype Discriminative Network (CPDN) for Generalized Zero-Shot Learning (GZSL). It consists of a generative network for producing the visual prototype of each class by feeding its semantic representation, and a metric network for measuring the similarities between the sample and the generated class-prototypes to accomplish the classification. In CPDN, a query sample intends to posses a higher similarity with its homogenous class-prototypes while the lower similarities with the inhomogenous ones, and the class-prototypes also intend to be distinguished with each other through the metric network. Moreover, a discriminative version of Relation Network (RN) named Discriminative Relation Network (DRN) is presented by incorporating the aforementioned idea into the conventional RN model for further achieving the complementation CDPN and RN in metric learning. Extensive experimental results on standard benchmarks demonstrate that our proposed approaches consistently outperform RN, and achieve the competitive performances compared with the state-of-the-arts in GZSL.
Sheng Huang 0001, Jingkai Lin, Luwen Huangfu
IEEE Signal Process. Lett.3
2018 Bootstrapping Polar-Opposite Emotion Dimensions from Online Reviews
Luwen Huangfu, Mihai Surdeanu
LREC1
2018 Improved hypergraph regularized Nonnegative Matrix Factorization with sparse representation
abstract
As a commonly used data representation technique, Nonnegative Matrix Factorization (NMF) has received extensive attentions in the pattern recognition and machine learning communities over decades, since its working mechanism is in accordance with the way how the human brain recognizes objects. Inspired by the remarkable successes of manifold learning, more and more researchers attempt to incorporate the manifold learning into NMF for finding a compact representation ,which uncovers the hidden semantics and respects the intrinsic geometric structure simultaneously. Graph regularized Nonnegative Matrix Factorization (GNMF) is one of the representative approaches in this category. The core of such approach is the graph, since a good graph can accurately reveal the relations of samples which benefits the data geometric structure depiction. In this paper, we leverage the sparse representation to construct a sparse hypergraph for better capturing the manifold structure of data, and then impose the sparse hypergraph as a regularization to the NMF framework to present a novel GNMF algorithm called Sparse Hypergraph regularized Nonnegative Matrix Factorization (SHNMF). Since the sparse hypergraph inherits the merits of both the sparse representation and the hypergraph model, SHNMF enjoys more robustness and can better exploit the high-order discriminant manifold information for data representation . We apply our work to address the image clustering issue for evaluation. The experimental results on five popular image databases show the promising performances of the proposed approach in comparison with the state-of-the-art NMF algorithms.
Sheng Huang 0001, Hongxing Wang 0001, Yongxin Ge, Luwen Huangfu, Xiaohong Zhang 0002, Dan Yang 0001
Pattern Recognit. Lett.4
2016 Towards Using Social Media to Identify Individuals at Risk for Preventable Chronic Illness
Dane Bell, Daniel Fried, Luwen Huangfu, Mihai Surdeanu, Stephen G. Kobourov
LREC3
2015 Class specific sparse representation for classification
Sheng Huang 0001, Yu Yang 0010, Dan Yang 0001, Luwen Huangfu, Xiaohong Zhang 0002
Signal Process.4
2013 OCC model-based emotion extraction from online reviews
abstract
Extracting emotions from online reviews is crucial to many security-related applications as well as applications in other domains. Traditional approaches to emotion extraction have mainly focused on mining the polarities of opinions or using annotated data to extract emotion types. Emotion theories, which identify the underlying cognitive structure and emotional dimensions that are key to generate emotions, have almost been totally ignored in previous work. To facilitate the automatic extraction of emotions from textual data, in this paper, we propose an emotion model based approach to emotion extraction from online reviews. Informed by the widely used OCC emotion model, we employ a statistical method to extract emotion words with their dimension values from texts, and implement OCC model to obtain emotions based on the emotion-dimension dictionary. We conduct an empirical study using security-related news reviews. The experimental results demonstrate the effectiveness of our proposed approach.
Luwen Huangfu, Wenji Mao, Daniel Dajun Zeng, Lei Wang 0062
ISI1
2012 Extracting opinion explanations from Chinese online reviews
abstract
Opinion mining has gained increasing attention and shown great practical value in recent years. Existing research on opinion mining mainly focuses on the extraction of lexicon orientation and opinion targets. The explanations of opinions, which are potentially valuable for many applications, are totally ignored. To address this specific research challenge, in this paper, we propose an approach to extract the explanation of reason and/or consequence behind an opinion via learning word pairs and using causal indicators from Chinese online reviews. We also improve our word pair based method by constructing clusters of word paris. Experiments on a Chinese business review corpus show that our method is feasible and effective.
Yuequn Li, Wenji Mao, Daniel Dajun Zeng, Luwen Huangfu
ISI4