EDBT 2026 Demo / reviewers in the wild / expert
Ruohong Huan
dblp:65/9619
· DBLP profile ↗
40ranked-venue papers
15as first author
31since 2021 · last 2026
0000-0003-2555-343XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRISM-GCN: Principled Representation Integrationon Graphs for Multi-Behavior RecommendationabstractPersonalized retrieval in modern multimedia systems must interpret heterogeneous user interactions (e.g., view, favorite, purchase) to infer user intent. Multi-Behavior Recommendation (MBR) leverages auxiliary behaviors to alleviate data sparsity, yet existing Graph Neural Network (GNN) methods face two major bottlenecks for retrieval: (i) learning representations that remain consistent across views and granularities, where naive alignment can over-homogenize heterogeneous views and degrade view-specific semantics (semantic collapse), and (ii) optimizing multiple behavior-specific objectives under inherent behavior imbalance. We propose PRISM-GCN (Principled Representation Integration via Structure-aware Multi-view learning), which decomposes each behavior into three complementary views—interaction, hypergraph, and semantic—to capture structural dependencies at different orders. The views are coordinated by Structure-Aware Contrastive Alignment (SCA), which promotes cross-view consistency according to structural hierarchy while preserving view-specific information. In addition, we mitigate behavior imbalance with a personalized multi-task reweighting strategy that adapts supervision strength to user-specific behavior relevance. Experiments on three real-world datasets show that PRISM-GCN consistently outperforms state-of-the-art baselines on top-K retrieval. Gaoquan Liu, Changqiang Xu, Ruohong Huan |
ICMR | 4 |
| 2026 | HCC-MBR: Escaping the quantity-over-quality trap via hierarchical and counterfactual calibration for multi-Behavior recommendation
Gaoquan Liu, Xilin Wen, Changqiang Xu, Ruohong Huan |
Expert Syst. Appl. | 5 |
| 2026 | Bridging accuracy and diversity in graph-based recommenders via adaptive granularity fusion
Jiamin Ma, Ruohong Huan |
Neurocomputing | 5 |
| 2026 | GCFL: Gray-Modality Conversion and Feature Learning for Visible-Infrared Person Re-Identification
Ruohong Huan, Mingzhen Wu, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2026 | A Dual Asynchronous-Synchronous Relation Graph Method for Sensor-Based Group Activity Recognition
Ai Bo 0001, Ruohong Huan, Peng Chen 0008, Ronghua Liang |
IEEE Internet Things J. | 2 |
| 2026 | FreTransLS: Frequency Transformer based large-scale group activity recognition model for sensor data
Ruohong Huan, Meijiao Cao, Yantong Zhou, Peng Chen 0008, Guodao Sun, Ronghua Liang |
Pervasive Mob. Comput. | 1 |
| 2025 | FLEX-YOLO: A Feedback-Driven Lightweight and Enhanced Exchange Network for UAV Aerial ImagesabstractTo address the challenges of densely distributed small objects, blurry features, and complex backgrounds in unmanned aerial vehicle (UAV) aerial images, as well as the inability of traditional algorithms to balance accuracy and computational cost, this paper proposes a feedback-driven lightweight enhanced exchange network, FLEX-YOLO. We design a Multi-Path Feedback Perception (MPFP) module. It balances spatial and channel feature representations and integrates global context with local receptive fields, significantly improving feature extraction precision and efficiency. Meanwhile, a Joint Feature Interaction Network (JFIN) is proposed, which optimizes cross-scale information fusion to address the limited interaction between non-adjacent layers in traditional feature fusion methods. We further optimize the feature concatenation method and design the lightweight GhostPlus module using the pre-activation concept. The model additionally reduces computational overhead by adopting the LAMP pruning algorithm, which removes redundant network parameters while maintaining performance. Experimental results show that FLEX-YOLO outperforms YOLOv11, achieving a 4.9% higher mAP (from 40.4% to 45.3%) on the VisDrone2019 dataset. It also reduces parameters by 42.9% (from 9.416M to 5.372M). On the DOTA dataset, FLEX-YOLO achieves a 1.7% higher accuracy, demonstrating its robustness. Jiamin Ma, Ruohong Huan |
IJCNN | 4 |
| 2025 | RFA: Regularized Feature Alignment Method for Cross-Subject Human Activity RecognitionabstractCross-subject activity recognition is challenging in the human activity recognition field. Previous studies have often assumed that training and test data follow the same distribution, which is impractical in real-world applications. Thus, models’ performance will significantly decline when applied to data collected from new unseen subjects because of the different physical conditions and human habits. To solve the above challenges, we proposed the regularized feature alignment (RFA) network. The RFA introduces a source domain selection mechanism (SDSM) based on calculating the Wasserstein distance between different subjects. Through SDSM, subjects with high similarity in the source domain can be retained, which implicitly compacts the feature subspace distribution. We implemented linear data augmentation on the retained subjects to mitigate the effects of the decline in the training set. In addition, the regularized dropout method was adopted to explicitly compact the feature subspace distributions. Finally, multi-level feature alignment is performed via maximum mean discrepancy regularization to precisely match the source and target domain. To demonstrate the effectiveness of RFA, comprehensive experiments were conducted on four public datasets under the iterative left-one-subject-out setting. The experimental results demonstrate that RFA outperformed the state-of-the-art methods in datasets with a large divergence between subjects and achieved performance comparable to the state-of-the-art methods in a subject-balanced dataset. Hao Fu 0027, Ruohong Huan, Mengjie Qu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2025 | MulDeF: A Model-Agnostic Debiasing Framework for Robust Multimodal Sentiment AnalysisabstractIn recent years, multimodal sentiment analysis (MSA) has gained prominence with the proliferation of social media. However, prior studies have often disregarded the possibility of spurious correlations between multimodal data and sentiment labels. Neglecting these factors often results in significant performance degradation, hampering the model's ability to generalize in out-of-distribution (OOD) scenarios. To gain a comprehensive understanding of multimodal knowledge and enhance the model's generalization across diverse distribution scenarios, we present the Multimodal Debiasing Framework (MulDeF). This model-agnostic framework addresses label bias through causal intervention and tackles multimodal biases using counterfactual reasoning. During the training phase, MulDeF rectifies multimodal representations through frontdoor adjustment in causal intervention, effectively eliminating label bias. In order to model conditional expectation calculations within the context of frontdoor adjustment, we introduce multimodal causal attention (MCA). In the inference phase, it employs counterfactual reasoning to eliminate multimodal biases. To further refine our debiasing strategy, we categorize multimodal biases into two distinct types: nonverbal bias and verbal bias. Nonverbal bias is addressed at the utterance level, involving the establishment of unimodal models for audio and visual modalities to estimate their biases concerning sentiment labels. Conversely, verbal bias mitigation occurs at the word level. Here, we mask “harmless” words to generate corresponding counterfactual texts, which are then assessed by the text model to identify word-level bias. Experimental results validate the effectiveness of MulDeF, showcasing its superior performance in OOD settings compared to state-of-the-art methods, while also achieving competitive results in independent and identically distributed (IID) settings. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Multim. | 1 |
| 2025 | Generate anomalies from normal: a partial pseudo-anomaly augmented approach for video anomaly detection
Yuanjie Dang, Jiangyun Chen, Peng Chen 0008, Nan Gao 0001, Ruohong Huan |
Vis. Comput. | 5 |
| 2024 | Task-Agnostic Self-Distillation for Few-Shot Action Recognition
Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Nan Gao 0001, Ruohong Huan, Xiaofei He 0001 |
IJCAI | 6 |
| 2024 | Towards a Deeper Insight Into Face Detection in Neonatal Wards
Yisheng Zhao, Huaiyu Zhu 0004, Qi Shu, Ruohong Huan, Shuohui Chen |
MICCAI (5) | 4 |
| 2024 | Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality AssessmentabstractCurrent state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation. These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality-and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we propose a semantic-Aware and quality-Aware Interaction Network (A2INet) for blind VQA. For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic-and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling. Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released at https://github.com/JianjunXiang/A2INet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan, Nan Gao 0001 |
ACM Multimedia | 5 |
| 2024 | Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionabstractTemporal relation modeling is one of the core aspects of few-shot action recognition. Most previous works mainly focus on temporal relation modeling based on coarse-level actions, without considering the atomic action details and fine-grained temporal information. This oversight represents a significant limitation in this task. Specifically, coarse-level temporal relation modeling can make the few-shot models overfit in high-discrepancy temporal context, and ignore the low-discrepancy but high-semantic relevance action details in the video. To address these issues, we propose a saliency-guided fine-grained temporal mask learning method that models the temporal atomic action relation for few-shot action recognition in a finer manner. First, to model the comprehensive temporal relations of video instances, we design a temporal mask learning architecture to automatically search for the best matching of each atomic action snippet. Next, to exploit the low-discrepancy atomic action features, we introduce a saliency-guided temporal mask module to adaptively locate and excavate the atomic action information. After that, the few-shot predictions can be obtained by feeding the embedded rich temporal-relation features to a common feature matcher. Extensive experimental results on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Yuanjie Dang, Peng Chen 0008, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 4 |
| 2024 | Focus on Subtle Actions: Semantic and Saliency Knowledge Co-Propagation Method for Weakly-Supervised Temporal Action Localization
Yuanjie Dang, Haoyu Shou, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Yilong Zhang 0001 |
PRCV (10) | 5 |
| 2024 | Credit Evaluation Model and Its Application in Healthcare Insurance Fraud DetectionabstractHealthcare insurance fraud has become a major problem worldwide in recent decades, resulting in significant financial losses for every affected country. Traditional fraud detection methods, however, often fall short as they primarily focus on analyzing data from the current period, thereby neglecting valuable historical information. In our study, we introduce a novel approach inspired by the financial concept of “credit” to detect fraudulent activities in various domains, such as healthcare insurance, credit card, and online retail transactions. Our approach aims to build a credit evaluation model (CEM) that can distinguish between fraudulent and normal activities by analyzing their historical records. We acknowledge that numerous fraud detection methods have been proposed, but they often struggle to detect edge cases, which limits their practical effectiveness. To address this challenge, our proposed CEM employs a time interval-aware long short-term memory (LSTM) algorithm to assist fraud detection. Furthermore, we propose an innovative approach that transforms traditional binary classification into a multi-classification problem, which improves the model’s ability to handle diverse fraudulent activities. We conducted experiments to evaluate the effectiveness of our proposed approach and model, comparing them against baseline algorithms and recently proposed methods. The results indicate that our approach outperforms the others, demonstrating its potential for practical use in detecting fraudulent activities across various domains. Zeyu Ding 0008, Ruohong Huan |
Int. J. Comput. Intell. Appl. | 3 |
| 2024 | TLCSFI: A Pose-Guided Person Re-Identification Method with Two-Level Channel-Spatial Feature IntegrationabstractPerson re-identification methods currently encounter challenges in feature learning, primarily due to difficulties in expressing the correlation between local features and integrating global and local features effectively. To address these issues, a pose-guided person re-identification method with Two-Level Channel–Spatial Feature Integration (TLCSFI) is proposed. In TLCSFI, a two-level integration mechanism is implemented. At the first level, TLCSFI integrates the spatial information from local features to generate fine-grained spatial features. At the second level, the fine-grained spatial feature and the coarse-grained channel feature are integrated together to complete channel–spatial feature integration. In the method, a Pose-based Spatial Feature Integration (PSFI) module is introduced to generate the pose union feature, which calculates intra-body affinity to guide the integration of spatial information among local pose feature maps. Then, a Channel and Spatial Union Feature Integration (CSUFI) module is proposed to efficiently integrate the channel information of the global feature and the spatial information of the pose union feature. Two individual networks are designed to extract channel and spatial information, respectively, in CSUFI, which are then weighted and integrated. Experiments are conducted on three publicly available datasets to evaluate TLCSFI, and the experimental results demonstrate its competitive performance. Ruohong Huan, Nan Gao 0001, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2024 | Learning Reliable Dense Pseudo-Labels for Point-Level Weakly-Supervised Action LocalizationabstractAbstract Point-level weakly-supervised temporal action localization aims to accurately recognize and localize action segments in untrimmed videos, using only point-level annotations during training. Current methods primarily focus on mining sparse pseudo-labels and generating dense pseudo-labels. However, due to the sparsity of point-level labels and the impact of scene information on action representations, the reliability of dense pseudo-label methods still remains an issue. In this paper, we propose a point-level weakly-supervised temporal action localization method based on local representation enhancement and global temporal optimization. This method comprises two modules that enhance the representation capacity of action features and improve the reliability of class activation sequence classification, thereby enhancing the reliability of dense pseudo-labels and strengthening the model’s capability for completeness learning. Specifically, we first generate representative features of actions using pseudo-label feature and calculate weights based on the feature similarity between representative features of actions and segments features to adjust class activation sequence. Additionally, we maintain the fixed-length queues for annotated segments and design a action contrastive learning framework between videos. The experimental results demonstrate that our modules indeed enhance the model’s capability for comprehensive learning, particularly achieving state-of-the-art results at high IoU thresholds. Yuanjie Dang, Guozhu Zheng, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Ronghua Liang |
Neural Process. Lett. | 5 |
| 2024 | TriSAT: Trimodal Representation Learning for Multimodal Sentiment AnalysisabstractTransformer-based multimodal sentiment analysis frameworks commonly facilitate cross-modal interactions between two modalities through the attention mechanism. However, such interactions prove inadequate when dealing with three or more modalities, leading to increased computational complexity and network redundancy. To address this challenge, this paper introduces a novel framework, Trimodal representations for Sentiment Analysis from Transformers (TriSAT), tailored for multimodal sentiment analysis. TriSAT incorporates a trimodal transformer featuring a module called Trimodal Multi-Head Attention (TMHA). TMHA considers language as the primary modality, combines information from language, video, and audio using a single computation, and analyzes sentiment from a trimodal perspective. This approach significantly reduces the computational complexity while delivering high performance. Moreover, we propose Attraction-Repulsion (AR) loss and Trimodal Supervised Contrastive (TSC) loss to further enhance sentiment analysis performance. We conduct experiments on three public datasets to evaluate TriSAT's performance, which consistently demonstrates its competitiveness compared to state-of-the-art approaches. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Heterogeneous Graph Network for Action DetectionabstractSpatio-temporal action detection is a fundamental task that detects persons and recognizes their actions from videos. It requires reasoning about the spatial-temporal interactions between persons and their surroundings. Recently, more modalities have been found by researchers, which puts higher demands on the reasoning capability of the method, yet a method capable of holistic reasoning is still lacking. To this end, we propose a heterogeneous graph network, which aims to reason the spatial-temporal interactions among different types of nodes (video entities) and edges (inter-entity relations). Concretely, it includes spatial and temporal graphs, which are alternately updated. The spatial graph contains nodes of person appearance, person pose, object appearance, and hand interaction, and the temporal graph has person nodes at different moments. For information aggregation, we propose a person-centric heterogeneous graph reasoning algorithm, which introduces heterogeneity into the graphs through node-type-specific projections and modulated edge-type-specific representations. We find that the introduction of heterogeneity enriches the model’s ability to understand multi-modality, which facilitates better parsing of complex semantic relations in videos and potentially leads to further mining of spatial-temporal interactions between entities in the future. Experimental results on four public datasets demonstrate the superiority of our method. Code will be available after acceptance. Yisheng Zhao, Huaiyu Zhu 0004, Ruohong Huan, Yaoqi Bao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | GMAEEG: A Self-Supervised Graph Masked Autoencoder for EEG Representation LearningabstractAnnotated electroencephalogram (EEG) data is the prerequisite for artificial intelligence-driven EEG autoanalysis. However, the scarcity of annotated data due to its high-cost and the resulted insufficient training limits the development of EEG autoanalysis. Generative self-supervised learning, represented by masked autoencoder, offers potential but struggles with non-Euclidean structures. To alleviate these challenges, this work proposes a self-supervised graph masked autoencoder for EEG representation learning, named GMAEEG. Concretely, a pretrained model is enriched with temporal and spatial representations through a masked signal reconstruction pretext task. A learnable dynamic adjacency matrix, initialized with prior knowledge, adapts to brain characteristics. Downstream tasks are achieved by finetuning pretrained parameters, with the adjacency matrix transferred based on task functional similarity. Experimental results demonstrate that with emotion recognition as the pretext task, GMAEEG reaches superior performance on various downstream tasks, including emotion, major depressive disorder, Parkinson's disease, and pain recognition. This study is the first to tailor the masked autoencoder specifically for EEG representation learning considering its non-Euclidean characteristics. Further, graph connection analysis based on GMAEEG may provide insights for future clinical studies. Zanhao Fu, Huaiyu Zhu 0004, Yisheng Zhao, Ruohong Huan, Yi Zhang 0126, Shuohui Chen |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | UniMF: A Unified Multimodal Framework for Multimodal Sentiment Analysis in Missing Modalities and Unaligned Multimodal SequencesabstractIn current multimodal sentiment analysis, aligned and complete multimodal sequences are often crucial. Obtaining complete multimodal data in the real world presents various challenges, and aligning multimodal sequences often requires a significant amount of effort. Unfortunately, most multimodal sentiment analysis methods fail when dealing with missing modalities or unaligned multimodal sequences. To tackle these two challenges simultaneously in a simple and lightweight manner, we present the Unified Multimodal Framework (UniMF). The primary components of UniMF comprise two distinct modules. The first module, Translation Module, translates missing modalities using information from existing modalities. The second module, Prediction Module, uses the attention mechanism to fuse the multimodal information and generate predictions. To enhance the translation performance of the Translation Module, we introduce the Multimodal Generation Mask (MGM) and utilize it to construct the Multimodal Generation Transformer (MGT). The MGT can generate the missing modality while focusing on information from existing modalities. Furthermore, we introduce the Multimodal Understanding Transformer (MUT) in the Prediction Module, which includes the Multimodal Understanding Mask (MUM) and a unique sequence,MultiModalSequence(MMSeq), representing a unified multimodality. To assess the performance of UniMF, we perform experiments on four multimodal sentiment datasets, and UniMF attains competitive or state-of-the-art outcomes with fewer learnable parameters. Furthermore, the experimental outcomes signify that UniMF, supported by MGT and MUT - two transformers utilizing special attention mechanisms, can efficiently manage both generating task of missing modalities and understanding task of unaligned multimodal sequences. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Multim. | 1 |
| 2024 | Discriminative Action Snippet Propagation Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to classify and localize actions in untrimmed videos with only video-level labels. Recent studies have attempted to obtain more accurate temporal boundaries by exploiting latent action instances in ambiguous snippets or propagating representative action features. However, empirically handcrafted ambiguous snippet extraction and the imprecise alignment of representative snippet propagation lead to challenges in modeling the completeness of actions for these methods. In this article, we propose a Discriminative Action Snippet Propagation Network (DASP-Net) to accurately discover ambiguous snippets in videos and propagate discriminative instance-level features throughout the video for improving action completeness. Specifically, we introduce a novel discriminative feature propagation module for capturing the global contextual attention and propagating the action concept across the whole video by perceiving the discriminative action snippets with instance information from the same video. Simultaneously, we incorporate denoised pseudo-labels as supervision, where we correct the controversial prediction based on the feature space distribution during training, thereby alleviating false detection caused by noise background features. Furthermore, we design an ambiguous feature mining module, which maximizes the feature affinity information of action and background in ambiguous snippets to generate more accurate latent action and background snippets and learns more precise action instance boundaries through contrastive learning of action and background snippets. Extensive experiments show that DASP-Net achieves state-of-the-art results on THUMOS14 and ActivityNet1.2 datasets. Yuanjie Dang, Chunxia Huang, Peng Chen 0008, Nan Gao 0001, Ronghua Liang, Ruohong Huan |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Spatial-angular Quality-aware Representation Learning for Blind Light Field Image Quality AssessmentabstractBlind light field image quality assessment (BLFIQA) remains a challenging task in deep learning due to the unique spatial-angular structure of light field images (LFIs) and the lack of large-scale labeled data for training. In this work, we propose a novel BLFIQA method using spatial-angular quality-aware representation learning in a self-supervised learning manner. Visual content and distortion type are important factors affecting the perceived quality of LFIs. In our observation, the band-pass transform maps of LFIs with the same distortion type exhibit similar Gaussian distributions. Thus, we learn spatial-angular quality-aware representations by minimizing the distance in the embedding space between the luminance map and the band-pass transform map of the same LFI. To implement spatial-angular quality-aware representations of LFI, we also build a large-scale unlabeled dataset containing 40k distorted LFIs with different distortion types and visual content. Further, we propose a fusion-separation-fusion network (FSFNet) to extract features for representing the intrinsic spatial-angular structure of the LFI. After pre-training on the unlabeled dataset using the proposed self-supervised learning, the FSFNet is employed for downstream BLFIQA tasks and achieves good performance. Experimental results show that our proposed method outperforms seventeen state-of-the-art models on the Win5-LID, NBU-LF1.0 and LFDD datasets, and achieves 3.78%, 6.61% and 4.06% SRCC improvements, respectively. The code and dataset will be publicly available in https://github.com/JianjunXiang/SSL_and_FSFNet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan |
ACM Multimedia | 5 |
| 2023 | Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionabstractFew-shot action recognition (FSAR) aims to classify unseen query actions into categories represented by a few labeled support videos. Most current FSAR methods adopt the frame-level matching mechanism that requires continuous actions to be represented by a fixed number of frame features. However, this could compromise the completeness of the contextual video information and make it difficult to handle video features of varying frame sampling speeds. In this paper, we propose a multi-speed global contextual subspace matching (MGCSM) method that generates global contextual action subspace representations from videos containing different numbers of frames to preserve contextual semantic information. Specifically, we propose to obtain the scale-agnostic information of embedding video features using a global contextual aggregation (GCA) module and then generate the discriminative action subspace representation with an action subspace generation (ASG) module. Furthermore, we introduce a multi-speed subspace matching (MSM) mechanism that generates a multi-speed classification score by integrating the similarities between query videos and support subspaces of varying sampling speeds. The proposed method is embedding-agnostic and can be combined with most mainstream embedding networks without model re-designs. Comprehensive and reproducible experiments on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Tianwei Yu, Peng Chen 0008, Yuanjie Dang, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 4 |
| 2023 | HPAN: A Hybrid Pose Attention Network for Person Re-Identification
Ruohong Huan, Tianya Chen, Ziwei Zhan, Peng Chen 0008, Ronghua Liang |
PRCV (12) | 1 |
| 2023 | DOMOPT: A Detection-Based Online Multi-Object Pedestrian Tracking Network for VideosabstractDue to the problem of low tracking accuracy and weak tracking stability of current multi-object pedestrian tracking algorithms in complex scenes for videos, a Detection-based Online Multi-Object Pedestrian Tracking (DOMOPT) network is proposed. First, a Multi-Level Feature Fusion (MLFF) pedestrian detection network is proposed based on the Center and Scale Prediction (CSP) algorithm. The pyramid convolutional neural network is used as the backbone to enhance the feature extraction capability for small objects. The shallow features and deep features at multiple levels are integrated to fully obtain the position and semantic information to further improve the detection performance for small objects. Then, on the basis of Joint Detection and Embedding (JDE) architecture, a Multi-Branch Pedestrian Appearance (MBPA) feature extraction network is proposed and added into the pedestrian detection network to extract the appearance feature vector corresponding to each pedestrian. The pedestrian appearance feature extraction is treated as a classification task jointly training with the pedestrian detection task, using the multi-task learning strategy. Experimental results show that the proposed network has better tracking accuracy and stability compared with state-of-the-art algorithms. Ruohong Huan, Shuaishuai Zheng, Chaojie Xie, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2023 | MLFFCSP: a new anti-occlusion pedestrian detection network with multi-level feature fusion for small targets
Ruohong Huan, Chaojie Xie, Ronghua Liang, Peng Chen 0008 |
Multim. Tools Appl. | 1 |
| 2022 | Privacy-Aware Fuzzy Range Query Processing Over Distributed Edge DevicesabstractRange query processing is a common edge computing and service in the Internet of things, which can extract user-interest information from distributed edge devices. How to design lightweight privacy-preserving range query processing methods remains a challenging task. Existing secure range query approaches suffer from both high communication cost and long response time, which makes them unsuitable for edge computing over resource-constrained edge devices. In this article, we propose two privacy-aware fuzzy query processing schemes based on fuzzy theory. Linguistic range variables, fuzzy overlap information, and its recovery mechanism are introduced. In addition, two distributed privacy-aware fuzzy range query processing algorithms are devised. Our approaches not only serve for privacy protection, but also aim to provide other optimal performances in terms of reliability, energy efficiency, and real-time response. Theoretical analysis and experimental evaluations based on real-world datasets validated our motivation. Yinglong Li, Weiru Liu, Hong Chen 0001, Hongbing Cheng, Tieming Chen, Ruohong Huan |
IEEE Trans. Fuzzy Syst. | 8 |
| 2021 | Video multimodal emotion recognition based on Bi-GRU and attention fusion
Ruohong Huan, Jia Shu, Shenglin Bao, Ronghua Liang, Peng Chen 0008, Kaikai Chi |
Multim. Tools Appl. | 1 |
| 2021 | A hybrid CNN and BLSTM network for human complex activity recognition with multi-feature fusion
Ruohong Huan, Ziwei Zhan, Luoqi Ge, Kaikai Chi, Peng Chen 0008, Ronghua Liang |
Multim. Tools Appl. | 1 |
| 2020 | SAR multi-target interactive motion recognition based on convolutional neural networksabstractSynthetic aperture radar (SAR) multi‐target interactive motion recognition classifies the type of interactive motion and generates descriptions of the interactive motions at the semantic level by considering the relevance of multi‐target motions. A method for SAR multi‐target interactive motion recognition is proposed, which includes moving target detection, target type recognition, interactive motion feature extraction, and multi‐target interactive motion type recognition. Wavelet thresholding denoising combined with a convolutional neural network (CNN) is proposed for target type recognition. The method performs wavelet thresholding denoising on SAR target images and then uses an eight‐layer CNN named EilNet to achieve target recognition. After target type recognition, a multi‐target interactive motion type recognition method is proposed. A motion feature matrix is constructed for recognition and a four‐layer CNN named FolNet is designed to perform interactive motion type recognition. A motion simulation dataset based on the MSTAR dataset is built, which includes four kinds of interactive motions by two moving targets. The experimental results show that the recognition performance of the authors’ Wavelet + EilNet method for target type recognition and FolNet for multi‐target interactive motion type recognition are both better than other methods. Thus, the proposed method is an effective method for SAR multi‐target interactive motion recognition. Ruohong Huan, Luoqi Ge, Chaojie Xie, Kaikai Chi, Keji Mao |
IET Image Process. | 1 |
| 2018 | Anti-occlusion particle filter object-tracking method based on feature fusionabstractA new anti‐occlusion particle filter object‐tracking method based on feature fusion is proposed in this study. Colour and local binary pattern features are extracted and additively fused with a deterministic coefficient, which is calculated based on the difference between the object features and the background. An integral cumulative histogram is proposed to reduce the computational cost of feature extraction. A new occlusion determination method is proposed, and corresponding tracking strategies are also put forward for various occlusion conditions; in the case of partial occlusion, block tracking is carried out, and in the case of serious occlusion, the least‐square method is used to predict the object position. Context Aware Vision using Image‐based Active Recognition (CAVIAR) and Video Image Retrieval and Analysis Tool (VIRAT) video libraries are used to validate the method. The experimental results show that the proposed method can describe an object effectively and improve tracking stability and robustness under the occlusion conditions. Ruohong Huan, Shenglin Bao |
IET Image Process. | 1 |
| 2016 | New structure for multi-aspect SAR image target recognition with multi-level joint consideration
Ruohong Huan, Yifan Tao |
Multim. Tools Appl. | 1 |
| 2013 | Particle state compression scheme for centralized memory-efficient particle filtersabstractIn this paper, particle state compression scheme is proposed together with its architecture for centralized implementation of particle filters. In the scheme, state values are processed in original bit-width, stored in a compressed way and recovered before sampling of next iteration. The advantage of the scheme is that particle states memory requirement can be greatly reduced while the trade-off is the deviations between original and recovered states introduced by the process. A case study in Nearly Constant Turn (NCT) scenario shows that while achieving the same level of filtering accuracy, proposed scheme can save up to 49.69% memory overhead for storing particle state values compared to traditional realizations. Qinglin Tian, Xiaolang Yan, Ruohong Huan |
ICASSP | 5 |
| 2012 | Hierarchical resampling architecture for distributed particle filtersabstractIn this paper, a hierarchical resampling (HR) architecture has been presented for distributed particle filters (PFs). The proposed architectures decomposes the resampling step into two hierarchies, of which the first one, called intermediate resampling, is conducted consecutively among processing elements (PEs) the moment new particles and their weights are generated by each PE, and the second one, named unitary resampling, is performed sequentially after the whole intermediate resampling procedure and shared by all PEs. Compared with traditional distributed architectures, the HR architecture eliminates the particle redistribution step, and has such advantages as short execution time, high memory efficiency and well scalability. Xiaolang Yan, Ruohong Huan |
ICASSP | 4 |
| 2012 | Weight sorting based scheme and architecture for distributed particle filtersabstractThis paper presents an efficient weight sorting based scheme and architecture for distributed particle filters (PFs). Instead of redistributing particles among processing elements (PEs) after the resampling step, the proposed scheme sorts the newly generated weights, along with the corresponding propagated particles, in reverse order of the current partial weight sums (PWSs) of PEs so that the final local weight sums in each PE are as close to each other as possible. Resampling is then performed locally in each PE without the redistribution procedure before the next iteration. The corresponding architecture uses two levels of multiplexers to realize the weight/state sorting operation, and features regular structure, low execution time, high memory efficiency and well scalability. Xiaolang Yan, Ruohong Huan |
ISCAS | 4 |
| 2011 | Behavioral modeling of direct sampling mixerabstractThis paper presents a behavioral model of the direct sampling mixer (DSM) using MATLAB SIMULINK environment. The proposed model is integrated into a SIMULINK block with tunable parameters, which allows fast and convenient time domain simulations. Nonidealities, including transconductance nonlinearity, jitter effects and on-resistance thermal noise were addressed. Among these nonidealities, jitter problem is most difficult problem. We focused our attention on jitter effects in the SC filters which, to the best of our knowledge, is the first attempt to address this issue. In handling this problem, we adopted a quantified probability distribution method. Xiaolang Yan, Ruohong Huan |
ISCAS | 4 |
| 2011 | A general communication performance evaluation model based on routing path decompositionabstractThe network-on-chip (NoC) architecture is a main factor affecting the system performance of complicated multi-processor systems-on-chips (MPSoCs). To evaluate the effects of the NoC architectures on communication efficiency, several kinds of techniques have been developed, including various simulators and analytical models. The simulators are accurate but time consuming, especially in large space explorations of diverse network configurations; in contrast, the analytical models are fast and flexible, providing alternative methods for performance evaluation. In this paper, we propose a general analytical model to estimate the communication performance for arbitrary NoCs with wormhole routing and virtual channel flow control. To resolve the inherent dependency of successive links occupied by one packet in wormhole routing, we propose the routing path decomposition approach to generating a series of ordered link categories. Then we use the traditional queuing system to derive the fine-grained transmission latency for each network component. According to our experiments, the proposed analytical model provides a good approximation of the average packet latency to the simulation results, and estimates the network throughput precisely under various NoC configurations and workloads. Also, the analytical model runs about 10 5 times faster than the cycle-accurate NoC simulator. Practical applications of the model including bottleneck detection and virtual channel allocation are also presented. Ai-lian Cheng, Xiaolang Yan, Ruohong Huan |
J. Zhejiang Univ. Sci. C | 4 |
| 2008 | SAR Target Recognition Based on MRF and Gabor Wavelet Feature ExtractionabstractThis paper presents a method for synthetic aperture radar (SAR) target recognition based on Markov random field (MRF) segmentation and Gabor wavelet transform feature extraction. ICM algorithm which based on Markov random field is used to segment a SAR target chip into target and background two fields and generate a two-value figure. Gabor wavelet transform is applied to extract feature vectors from the two-value figure. The method is verified by recognizing three-class targets in MSTAR database. The highest average probability of correct classification arrives at 93.11%, which indicates that the approach proposed in this paper is an effective method for SAR target recognition. Ruohong Huan, Ruliang Yang |
IGARSS (2) | 1 |