Ruimin Hu

dblp:97/1491 · DBLP profile ↗
← Back
314ranked-venue papers
2as first author
120since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 209 · 1 first-author · 53 since 2021Artificial intelligence and machine learning · 79 · 51 since 2021Databases, data management, data science and information retrieval · 20 · 7 since 2021Computer networks · 13 · 10 since 2021Systems, architecture and hardware · 11 · 2 since 2021Security and privacy · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021
YearPublicationVenuePosition
2026 Revisiting Attention in the Dark for Low-Light Person Re-Identiffcation
abstract
Person re-identification (Re-ID) under extremely low-light conditions suffers from severe image degradation, which significantly impairs the extraction of identity-discriminative features. Existing methods struggle to recover semantic information that is obscured under poor illumination. To better understand this problem, we conduct a comprehensive analysis of the semantic modeling behavior of Re-ID models in low-light settings. For the first time, we investigate the norm distributions of Query (Q), Key (K), and Value (V) vectors within the attention module and observe that, as illumination decreases, the norm of Query vectors in pedestrian regions drops significantly. This leads to dispersed attention and degraded feature representations. To address this issue, we propose a novel framework named Norm-Ratio Attention and Semantic Recovery Distillation Network (NRSRD), which consists of two key components: a Norm-Ratio Attention Module (NRA) and a Semantic Recovery Distillation Module (SRD). The former dynamically adjusts attention responses based on the ratio of K/Q vector norms, enhancing structural region perception while suppressing background interference. The latter transfers discriminative semantic knowledge from high-illumination auxiliary data to the low-light model, compensating for the semantic degradation caused by poor lighting. Extensive experiments on multiple publicly available low-light Re-ID benchmarks demonstrate the effectiveness and superiority of the proposed method.
Ruimin Hu, Dongliang Zhu 0001
AAAI2
2026 Deception Detection Meets Vision-Language Models
Dongliang Zhu 0001, Ruimin Hu, Mang Ye
Int. J. Comput. Vis.2
2026 Symmetrical bidirectional knowledge alignment for zero-shot sketch-based image retrieval
abstract
This paper studies the problem of zero-shot sketch-based image retrieval (ZS-SBIR), which aims to use sketches from unseen categories as queries to match the images of the same category. Due to the large cross-modality discrepancy, ZS-SBIR is still a challenging task and mimics realistic zero-shot scenarios. The key is to leverage transferable knowledge from the pre-trained model to improve generalizability. Existing researchers often utilize the simple fine-tuning training strategy or knowledge distillation from a teacher model with fixed parameters, lacking efficient bidirectional knowledge alignment between student and teacher models simultaneously for better generalization. In this paper, we propose a novel Symmetrical Bidirectional Knowledge Alignment for zero-shot sketch-based image retrieval (SBKA). The symmetrical bidirectional knowledge alignment learning framework is designed to effectively learn mutual rich discriminative information between teacher and student models to achieve the goal of knowledge alignment. Instead of the former one-to-one cross-modality matching in the testing stage, a one-to-many cluster cross-modality matching method is proposed to leverage the inherent relationship of intra-class images to reduce the adverse effects of the existing modality gap. Experiments on several representative ZS-SBIR datasets (Sketchy Ext dataset, TU-Berlin Ext dataset and QuickDraw Ext dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is publicly available at https://github.com/zermatt-luo/SBKA.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
Neural Networks5
2026 Multi-dimensional fine-grained modeling and assessment for non-stationary fatigue processes
Ruimin Hu, Xiaojie Zhu, Dongliang Zhu 0001
Pattern Recognit.2
2026 DeepFidelity: Perceptual Forgery Fidelity Assessment for Deepfake Detection
abstract
Deepfake detection refers to detecting artificially generated or edited faces in images or videos, which plays an essential role in visual information security. Despite promising progress in recent years, Deepfake detection remains a challenging problem due to the complexity and variability of face forgery techniques. Existing Deepfake detection methods are often devoted to extracting features by designing sophisticated networks but ignore the influence of perceptual quality of faces. Considering the complexity of the quality distribution of real and fake faces, we propose a deepfake detection framework called DeepFidelity, which mines the perceptual forgery fidelity of face images and introduces a quality-aware scoring mechanism to distinguish real and fake faces of different image qualities. Specifically, we improve the model’s ability to identify complex samples by mapping real and fake face data of different qualities to different scores to distinguish them in a more detailed way. In addition, we propose a network structure called Symmetric Spatial Attention Augmentation based vision Transformer (SSAAFormer), which uses the symmetry of face images to promote the network to model the geographic long-distance relationship at the shallow level and augment local features. Extensive experiments on multiple benchmark datasets demonstrate the superiority of the proposed method over state-of-the-art methods. The code is available athttps://github.com/shimmer-ghq/DeepFidelity.
Chunlei Peng, Huiqing Guo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 HPRNet: Human Parsing Reconstruction With Non-Local Multi-Scale Perception Network for Cloth-Changing Person Re-Identification
abstract
Cloth-changing Person Re-Identification (CC-ReID) is a challenging data modeling task that involves identifying specific pedestrians wearing different outfits. Existing methods primarily focus on altering clothing color and directly reconstructing appearance to extract features independent of the clothes. Real pedestrians differ in height, body shape, etc. Such methods are prone to losing the intrinsic information of the original sample (i.e., the person identity) owing to the absence of contextual phenomena (e.g., texture structure and local correlation), which decreases the recognition performance. To address this problem, we propose a framework called HPRNet, or ”Human Parsing Reconstruction with Non-Local Multi-Scale Perception Network,” which includes a non-local weighted multi-scale perception (NWMP) module and a parsing reconstruction exploration (PRE) module. In particular, the proposed NWMP module effectively captures the global receptive field of a sample and obtains a contextual correlation between non-neighboring pixels within the sample image. The PRE module was used to achieve a more accurate reconstruction of human body components with a clothing parsing model to better distinguish features related to or unrelated to clothes. Extensive experiments were conducted on CC-ReID public datasets (LTCC, PRCC, and CCVID) to demonstrate the effectiveness and competitiveness of the proposed method with state-of-the-art (SOTA) baselines for this complex modeling task.
Mingfu Xiong, Longlong Ge, Ruimin Hu, Khan Muhammad 0001, Sambit Bakshi, Javier Del Ser, Xiaokang Yang 0001, Bin Sheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and Evaluations
abstract
Due to the successful development of deep image generation technology, forgery detection plays a more important role in social and economic security. Racial bias has not been explored thoroughly in the deep forgery detection field. In the paper, we first contribute a dedicated dataset called the Fair Forgery Detection (FairFD) dataset, where we prove the racial bias of public state-of-the-art (SOTA) methods. Different from existing forgery detection datasets, the self-constructed FairFD dataset contains a balanced racial ratio and diverse forgery generation images with the largest-scale subjects. Additionally, we identify the problems with naive fairness metrics when benchmarking forgery detection models. To comprehensively evaluate fairness, we design novel metrics including Approach Averaged Metric and Utility Regularized Metric, which can avoid deceptive results. We also present an effective and robust post-processing technique, Bias Pruning with Fair Activations (BPFA), which improves fairness without requiring retraining or weight updates. Extensive experiments conducted with 12 representative forgery detection models demonstrate the value of the proposed dataset and the reasonability of the designed fairness metrics. By applying the BPFA to the existing fairest detector, we achieve a new SOTA. Furthermore, we conduct more in-depth analyses to offer more insights to inspire researchers in the community.
Decheng Liu, Zongqi Wang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
AAAI5
2025 MDFG: Multi-Dimensional Fine-Grained Modeling for Fatigue Detection
abstract
Fatigue is a critical factor contributing to accidents in industries such as safety monitoring and engineering construction. Fatigue exhibits dynamic complexity and non-stationary characteristics, so there are many intermediate states of short-term variation between alert and fatigue. Capturing and learning the signs of these intermediate states is essential for accurate fatigue assessment. However, current fatigue detection methods primarily rely on coarse-grained labels, typically spanning minutes to hours, and commonly treat alert and fatigue as two distinctly separate distributions, overlooking the expression of intermediate states and oversimplifying the rich distribution information of fatigue types and levels, thereby limiting detection effectiveness. To address these, this paper explores a refined representation of fatigue in terms of three dimensions: time, type, and level, and proposes a Multi-Dimensional Fine-Grained Modeling for Fatigue Detection (MDFG). This introduces the SmallLoss to extract trustworthy samples, utilizes clustering to identify diverse subtypes under alert and fatigued states, and establishes base class sets in each state. Subsequently, a complete base class set containing intermediate state bases is constructed using the base class synthesis method, which achieves the expression of intermediate fatigue states from absence to presence. Finally, fatigue levels are quantified based on the matching between samples and the complete base class set. Moreover, to cope with the complex variability of fatigue states, MDFG employs meta-learning for training. MDFG achieves an Average accuracy improvement of 10.0% and 12.1% on two real datasets compared to methods that do not consider fine-grained information. Extensive experiments demonstrate that the MDFG exhibits superior robustness and stability among current fatigue detection methods.
Xiaojie Zhu, Ruimin Hu, Dongliang Zhu 0001, Mang Ye
AAAI3
2025 TSCA: Enhancing Smartphone Security with a Touchpoints-Sensitive and Context-Aware Model for Touchscreen-Based Authentication
abstract
With the widespread use of smartphones, more and more private information is stored on them, which requires a more secure authentication method. Continuous authentication methods show potential for privacy protection due to their continuous and transparent nature, as opposed to the entry-point authentication methods like PIN codes and facial recognition. Users' touchscreen behavior, being both common and unique, has been widely studied in the field. However, previous methods based on manual features and deep learning have certain shortcomings when mining patterns from touchscreen behaviors that are beneficial for authentication. First, the operating habits of different users lead to different importance of touch points in a single touchscreen swipe. Second, extracting features from each touchscreen swipe independently ignores useful information from the context. To address this problem, we propose a novel Touchpoints-Sensitive and Context-Aware model called TSCA, which is able to adaptively assign weights to different touch points in a touchscreen swipe while also taking into account information from the context of the touchscreen swipe. We evaluate the TSCA model on two publicly available datasets and show that the TSCA is able to deliver significant performance gains compared to baseline methods while reducing model complexity.
Ruimin Hu
CSCWD2
2025 SE2E: Recognizing Emotion behind Societal Behavior
abstract
Emotion recognition, as a core technology in mental health monitoring, has long been constrained by the intrusive nature of data collection methods relying on physiological signals and behavioral cues. Although existing motion-based approaches enable non-intrusive data acquisition, they often overlook the societal dimensions inherent in human behavior. As a result, they often exhibit a significant performance drop in real-world scenarios compared to laboratory settings. In this study, we analyzed the spatial distribution of participants' spatiotemporal trajectories and their visited Points of Interest (POIs), and observed significant differences under varying emotional states. Building on this observation, we propose a novel emotion recognition framework, SE2E, which innovatively incorporates the semantic information of POIs into the emotion recognition task. Specifically, SE2E employs a category-aware semantic embedding mechanism combined with a masked prediction task to ensure that the POI embeddings capture both categorical semantics and contextual information. It then structurally represents individual societal event patterns through a personalized spatiotemporal flow. Finally, a temporal-region consistency attention module is employed to extract continuous representations of societal events, thereby enabling a robust mapping from societal behavior to emotional state. Extensive experimental results demonstrate that SE2E outperforms state-of-the-art methods across multiple benchmarks. To the best of our knowledge, this is the first study to leverage societal event for emotion recognition, offering a new technical direction, benchmark, and insight for future research in the field.
Wending Xiong, Ruimin Hu, Lingfei Ren, Dengshi Li
ACM Multimedia2
2025 Multi-layer Spatio-Temporal LSTM-Transformer Model for Individual Mobility Prediction
Ruimin Hu
PRCV (4)2
2025 Spatio-Temporal Multi-Granularity Gated Recurrent Transformer Model for Individual Mobility Prediction
Ruimin Hu
PRICAI2
2025 Efficient Extended Neighborhoods Dynamic Selection Re-Ranking for Person Re-Identification
abstract
Person re-identification (re-ID) is a challenging retrieval task that requires matching a person’s captured images across non-overlapping camera views, with re-ranking being a critical step for improving accuracy. The k-nearest neighbors relationship is commonly used to determine the rank results by selecting only k fixed pedestrian image values for distance calculations. This operation, however, generates additional distance errors due to changes in the appearance of pedestrians. This paper addresses the above issue by proposing a simple but effective Extended Neighborhood Dynamic Selection (ENDS), distance to optimize the performance of ReID reranking. The number of images selected is then distributed over an interval. A limit of upper and lower is placed on the selected number to ensure that it is neither too low nor too high. The automatic selection of adjacent images is achieved using this method. This distance is determined by combining ENDS distances with Jaccard distances. It is the core principle of this method that, instead of using fixed values, the choice of the number of images to be included in each neighborhood should be made automatically. It also allows the removal of images that are dissimilar in favour of those that are more representative. Experimental results demonstrate the novel method of ranking by using Market-1501 and DukeMTMC’s reID dataset. In this paper, we propose a method that increases the Market-1501 mAP/Rank1 by 29.4%/12.9% while DukeMTMC-reID reranking by 35%/20.7%.
Chao Wang 0084, Zhongyuan Wang 0001, Xiaochen Wang 0001, Ruimin Hu, Mithun Mukherjee 0001
SMC4
2025 Power on graph: Mining power relationship via user interaction correlation
Yilong Zang, Lingfei Ren, Junhang Wu, Yilin Xiao 0002, Ruimin Hu
Expert Syst. Appl.5
2025 Multimodal and multichannel speech separation using location-guided speech feature mapping network
Yulin Wu 0003, Xiaochen Wang 0001, Dengshi Li, Ruimin Hu
Neurocomputing4
2025 Toward trustworthy identity tracing via multi-attribute synergistic identification
Wenbin Feng, Decheng Liu, Ruimin Hu
Inf. Sci.3
2025 Devil in Shadow: Attacking NIR-VIS Heterogeneous Face Recognition via Adversarial Shadow
abstract
Near infrared-visible (NIR-VIS) heterogeneous face recognition aims to match face identities in cross-modality settings, which has achieved significant development recently. The work on adversarial attack and security issues of the heterogeneous face recognition task is still lacking. Existing adversarial face generation methods can’t deploy directly because of the inevitable large modality discrepancy. Besides, the ideal adversarial attacking generated images should maintain both high capabilities and low detectability. Considering the properties of near-infrared face images, our basic idea is to construct adversarial shadows for good stealthiness and high attack capability. In this paper, we propose a novel face adversarial shadow generation framework for NIR-VIS heterogeneous face recognition, which can synthesize fine-crafted lighting conditions containing strong identity attacking ability. Specifically, we design the variance consistency-based symmetric face attacking loss to improve the attacking generalization and the synthesized image quality. Extensive qualitative and quantitative experiments on the public large-scale NIR-VIS heterogeneous face dataset prove the proposed method achieves superior performance compared with the state-of-the-art methods. The source code is publicly available athttps://github.com/GEaMU/Devil-in-Shadow.
Decheng Liu, Rong Sheng, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Masked Text Adversarial Training for Cloth-Changing Person Re-Identification
Chengrui Hao, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.6
2025 Attention Consistency Refined Masked Frequency Forgery Representation for Generalizing Face Forgery Detection
abstract
Due to the successful development of deep image generation technology, visual data forgery detection would play a more important role in social and economic security. Existing forgery detection methods suffer from unsatisfactory generalization ability to determine the authenticity in the unseen domain. In this paper, we propose a novel Attention Consistency Refined masked frequency forgery representation model toward a generalizing face forgery detection algorithm (ACMF). Most forgery technologies always bring in high-frequency aware cues, which make it easy to distinguish source authenticity but difficult to generalize to unseen artifact types. The masked frequency forgery representation module is designed to explore robust forgery cues by randomly discarding high-frequency information. In addition, we find that the forgery saliency map inconsistency through the detection network could affect the generalizability. Thus, the forgery attention consistency is introduced to force detectors to focus on similar attention regions for better generalization ability. Experiment results on several public face forgery datasets (FaceForensic++, DFD, Celeb-DF, WDF and DFDC datasets) demonstrate the superior performance of the proposed method compared with the state-of-the-art methods. The source code and models are publicly available athttps://github.com/chenboluo/ACMF.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Improving Adversarial Robustness via Decoupled Visual Representation Masking
abstract
Deep neural networks are proven to be vulnerable to finely designed adversarial examples, and adversarial defense algorithms draw more and more attention nowadays. Pre-processing based defense is a major strategy, as well as learning robust feature representation, has been proven an effective way to boost generalization. However, existing defense works lack considering different depth-level visual features in the training process. In this paper, we first highlight two novel properties of robust features from the feature distribution perspective: 1) Diversity (robust features within the same class should maintain appropriate variety). 2) Discriminability (robust features from different classes should be sufficiently separated). We find that state-of-the-art defense methods aim to address both of these mentioned issues well. It motivates us to increase intra-class variance and decrease inter-class discrepancy simultaneously in adversarial training. Specifically, we propose a simple but effective defense based on decoupled visual representation masking. The designed Decoupled Visual Feature Masking (DFM) block can adaptively disentangle visual discriminative features and non-visual features with diverse mask strategies, while the suitable discarding information can disrupt adversarial noise to improve robustness. Our work provides a generic and easy-to-plugin block unit for any former adversarial training algorithm to achieve better protection integrally. Extensive experimental results prove that the proposed method can achieve superior performance compared with state-of-the-art defense approaches. The code is publicly available at https://github.com/chenboluo/Adversarial-defense.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Toward Fair Adversarial Defense via Class Encourage-Suppress Robust Learning
abstract
Deep Neural Networks with natural training are quite vulnerable to adversarial attacks, so it’s necessary to defend these attacks with effective defense methods like adversarial training. However, while defending against adversarial attacks, adversarially trained models’ robustness between classes shows severe unfairness. To mitigate the disparity, a lot of methods have been proposed, while many of them sacrifice the overall accuracy to leverage the worst-class accuracy. Inspired by previous works, we propose a new fairness methodology, and we name it Class Encourage-suppress Robust Learning (CRL). Based on the overall accuracy and the class-wise accuracies of the dataset, we introduce a new measurement named Class Diversity Ratio to adjust the weights of different classes in the loss function. Additionally, we propose a new learning strategy called the Competitor Encourage-suppress Strategy to mitigate the disparity between diverse classes, which is simple but effective. Experimental results on representative datasets show that our method outperforms state-of-the-art (SOTA) methods.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Semantic Token Transformer for Face Forgery Detection
abstract
In the era of digital media, the proliferation of forged images and videos poses a significant threat to societal stability. With the rapid advancement of deep learning, the generation of realistic fake images has become increasingly simple, presenting unprecedented challenges in discerning the authenticity of images. While some existing methods have shown promising results in forgery detection, they often underutilize facial semantic information. To address this issue, this paper introduces the Semantic Token Transformer for Face Forgery Detection. By incorporating facial semantic information with a transformer network, the input tokens of the transformer are transformed into tokens of varying shapes and sizes based on their importance, thereby enhancing the accuracy of the detector. To achieve this objective, we first employ an image processing stage to manipulate the image based on facial semantic information. Subsequently, we introduce a scoring network, guided by prior knowledge, which adaptively categorizes tokens into different clusters based on their importance and relevance to the results of the preprocessing stage. Finally, we merge the tokens within the clusters using an attention mechanism and input them into the detector for forgery detection. Through experiments conducted on multiple datasets and cross-dataset evaluations, we demonstrate that our approach outperforms state-of-the-art detection methods.
Chunlei Peng, Xiaoyi Luo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Detecting Deceptive Behavior via Learning Relation-Aware Visual Representations
abstract
With the rapid development and widespread adoption of digital media, deceptive behaviors have raised numerous ethical and security issues, making the research and advancement of deception detection technology particularly important. Most previous automated deception detection methods primarily focus on facial information in a visual context. However, from a psychological perspective, deceptive behavior extends beyond mere changes in facial expressions; it can also manifest through limb behaviors and subtle incoordination among body components. Motivated by this inconsistency, this paper attempts to model body behaviors and their relationships for deception detection. It is worth noting that some mainstream video understanding methods can roughly model head and limb information, but their holistic video input approach is easily affected by background interference. This limits their ability to focus on key body regions and subtle motion cues that reflect deception, thereby restricting detection performance. To address the above challenges, this paper proposes a Dynamic Learning Framework leveraging Body Part Relationship-Aware Modeling (DLF-BRAM). Within this framework, we segment and model the head and limb regions to reduce irrelevant background interference and enhance the accuracy of feature learning. The framework includes two main components: the Head-Limb Relationship-Aware Representation (HLRAR) module and the Dynamic Assessment Learning Strategy (DALS). The HLRAR module reveals the spatiotemporal relationship of the head, limbs, and their interactions, and learns deep feature representations for each cue, thereby highlighting the uniqueness of these cues. DALS evaluates the learning effectiveness of the three spatiotemporal relationships during training and dynamically adjusts their learning weights, preventing dominance by any single branch and promoting balanced learning. Extensive benchmark and ablation experiments demonstrate that our method outperforms most existing approaches, verifying its effectiveness.
Dongliang Zhu 0001, Ruimin Hu, Mang Ye
IEEE Trans. Inf. Forensics Secur.3
2025 PrivacyHFR: Visual Privacy Preserving for Heterogeneous Face Recognition
abstract
Face recognition has achieved remarkable progress and is widely deployed in real-world scenarios. Recently more and more attention has been given to individual privacy protection, due to unauthorized sensitive image leakage by malicious attackers. Multi-modality face images captured by diverse sensors, also called heterogeneous faces, bring in more challenges in face privacy protection while lacking related research. In this paper, we propose a novel visual Privacy preserving method for Heterogeneous Face Recognition (Privacy-HFR) to protect perceptual visual information and maintain essential identity information in multi-modality face analysis scenarios. Frequency domain analysis is a vital strategy to bridge the inevitable modality gap for heterogeneous face images. Meanwhile, recent theoretical insights also inspire us to design a suitable frequency component adjustment to balance human visual sensitivity and identity discriminative information. In addition, the ability to defend against recovery attacks has emerged as an essential criterion for privacy preserving face recognition. Noting that there seems to exist a dilemma that reducing accessible information by the attack model will affect the extracted identity information for recognition. It is because these two kinds of information are mutually blended in the frequency domain, which makes it a challenge to simultaneously maintain visual privacy and identity distinguishability. Thus, we provide a novel perspective to leverage the randomly optimal solutions and design the specific adversarial perturbations against the recovery attack. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ, Tufts Face and CUFSF datasets) prove that the proposed method outperforms existing privacy-preserving face recognition methods in terms of recognition accuracy and privacy protection capability. The code is available in https://github.com/xiyin11/Privacy-HFR.
Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.5
2025 SketchAging: Face Photo-Sketch Synthesis and Aging With Multi-Scale Feature Extraction
abstract
With the rapid development of Generative Adversarial Networks (GANs), facial sketch generation and age transformation have advanced considerably. These technologies show great potential in digital media, entertainment, and forensic applications, particularly in helping law enforcement reconstruct the appearance of long-term fugitives. However, current methodologies exhibit notable limitations: existing approaches typically specialize in either facial sketch generation or age progression independently, lacking an effective integration for cross-domain synthesis. Moreover, preserving identity information while ensuring high-quality image generation remains a challenge. This paper proposes Multi-Scale Feature Extraction Networks (MSFE), an image-to-image translation framework that enables continuous age transformation while maintaining the stylistic characteristics of sketch domains. The core of the MSFS network uses a Dual Conditional Normalization Attention (DCNA) architecture to extract sketch features and encode facial images into the latent space of a pre-trained StyleGAN based on the desired age change. Experimental results on public datasets demonstrate that our approach outperforms existing methods, achieving superior facial photo-sketch synthesis with enhanced realism, identity preservation, and age accuracy.
Chunlei Peng, Zhuang Tang, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.5
2025 Face Forgery Detection With CLIP-Enhanced Multi-Encoder Distillation
abstract
With the development of face forgery technology, fake faces are rampant, threatening the security and authenticity of many fields. Therefore, it is of great significance to study face forgery detection. At present, existing detection methods have deficiencies in the comprehensiveness of feature extraction and model adaptability, and it is difficult to accurately deal with complex and changeable forgery scenarios. However, the rise of multimodal models provides new insights for current forgery detection methods. At present, most methods use relatively simple text prompts to describe the difference between real and fake faces. However, these researchers ignore that the CLIP model itself does not have the relevant knowledge of forgery detection. Therefore, our paper proposes a face forgery detection method based on multi-encoder fusion and cross-modal knowledge distillation. On the one hand, the prior knowledge of the CLIP model and the forgery model is fused. On the other hand, through the alignment distillation, the student model can learn the visual abnormal patterns and semantic features of the forged samples captured by the teacher model. Specifically, our paper extracts the features of face photos by fusing the CLIP text encoder and the CLIP image encoder, and uses the dataset in the field of forgery detection to pretrain and fine-tune the Deepfake-V2-Model to enhance the detection ability, which are regarded as the teacher model. At the same time, the visual and language patterns of the teacher model are aligned with the visual patterns of the pretrained student model, and the aligned representations are refined to the student model. This not only combines the rich representation of the CLIP image encoder and the excellent generalization ability of text embedding, but also enables the original model to effectively acquire relevant knowledge for forgery detection. Experiments show that our method effectively improves the performance on face forgery detection.
Chunlei Peng, Tianzhe Yan, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.5
2025 Multi-time-scale with clockwork recurrent neural network modeling for sequential recommendation
Nana Huang, Ruimin Hu, Pengfei Jiao, Zhidong Zhao, Bin Yang 0034
J. Supercomput.3
2025 Masked Attribute Description Embedding for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID) aims to match persons who change clothes over long periods. The key challenge in CC-ReID is to extract cloth-irrelated features, such as face, hairstyle, body shape, and gait. Current research mainly focuses on modeling body shape using multi-modal biological features (such as silhouettes and sketches). However, it does not fully leverage the personal description information hidden in the original RGB image. Considering that there are certain attribute descriptions that remain unchanged after the changing of cloth, we propose a Masked Attribute Description Embedding (MADE) method that unifies personal visual appearance and attribute description for CC-ReID. Specifically, handling variable cloth-sensitive information, such as color and type, is challenging for effective modeling. To address this, we mask the clothes type and color information (upper body type, upper body color, lower body type, and lower body color) in the personal attribute description extracted through an attribute detection model. The masked attribute description is then connected and embedded into Transformer blocks at various levels, fusing it with the low-level to high-level features of the image. This approach compels the model to discard cloth information. Experiments are conducted on several CC-ReID benchmarks, including PRCC, LTCC, Celeb-reID-light, and LaST. Results demonstrate that MADE effectively utilizes attribute description, enhancing cloth-changing person re-identification performance, and compares favorably with state-of-the-art methods.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Multim.5
2025 Adaptive Clustering and Weighted Regularization Contrastive Learning Framework for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (ReID) has recently gained significant attention from researchers. ReID matches images of the same person from different camera views in various scenes without any labels. Existing clustering methods primarily rely on a fixed threshold (the maximum distance between sample points and clustering centroids) and overlook the importance of adjusting this threshold during continuous model optimization. This mismatch between clustering thresholds and inter- or intra-class spacing reduces clustering accuracy. To address this issue, this study proposes an Adaptive Clustering and Weighted Regularization Contrastive Learning (ACWRCL) framework for unsupervised person ReID. The ACWRCL framework comprises two main components: (1) the Clustering Threshold Adaptive Adjustment (CTAA) module, and (2) the Weighted Regularization Contrastive Learning (WRCL) module. The CTAA module dynamically adjusts the clustering threshold to align with model optimization, ensuring that the threshold remains within an appropriate range to prevent under- or over-robustness in the clustering model. The WRCL module uses the similarity ratio between the query sample and the clustering centroid relative to the overall similarity of all samples with the same labels as the query sample. This ratio is used as the weight in the loss function to penalize incorrect clustering and improve pseudo-label generation accuracy. Extensive experiments on public ReID datasets—Market-1501, MSMT17, Veri776, CUHK03, and PersonX—demonstrate the effectiveness of the proposed method.
Mingfu Xiong, Kaikang Hu, Zhongyuan Wang 0001, Ruimin Hu, Khan Muhammad 0001, Javier Del Ser, Xiaokang Yang 0001, Bin Sheng 0001
IEEE Trans. Multim.4
2025 Uniform Light Transformer for Person Re-identification under Complex Illumination
abstract
The quality of pedestrian image retrieval is affected by the difference in illumination between images. Previous studies have used one-to-one lighting transformers to convert images taken under different lighting conditions into the target lighting. However, can we use a single lighting transformer to convert input images with various lighting conditions to the target lighting? This question motivated us to investigate the discrepancy between images generated by a Unified Lighting Transformer and the ground truth images across different illumination scales. We discovered that the modeling capability of the Unified Lighting Transformer for low-frequency information decreases gradually with an increase in the number of illuminant variations. Therefore, based on this insight, we proposed a Discriminative Feature Spectrum Consistency and Low-Frequency Information Constrained method. This method employs two constraints to enhance the Unified Lighting Transformer’s modeling capability for low-frequency information. The first mechanism enforces the constraint at the feature level by comparing the spectrum information between real and fake discriminative features. The second approach constrains the differences in pedestrian recognition features caused by the differences in low-frequency information between real and virtual images composed of low-frequency information from fake images and high-frequency information from authentic images. Our experiments show that our method outperforms other approaches and performs best across all metrics.
Ruimin Hu, Dongliang Zhu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Optimal Illumination Distance Metrics for Person Re-Identification in Complex Lighting Conditions
abstract
Person re-identification is extensively applied in public security and surveillance. However, environmental factors like time and location often lead to varying lighting conditions in captured pedestrian images, significantly impacting identification accuracy. Current approaches mitigate this issue through lighting transformation techniques, aiming to normalize images to a standard lighting condition for consistent person re-identification results. Yet, these methods overlook the fact that different content may hold distinct identification values under diverse lighting conditions. To address this, we conducted an analysis on the identification distance between images of the same or different pedestrians under pre-defined lighting conditions. From this analysis, we introduce the concept of optimal lighting: a condition where the distance between image pairs is minimized compared to other lighting scenarios. We propose utilizing this optimal lighting distance in the image retrieval process for final ranking. Our study, validated on synthetic datasets Market-IA and Duke-IA, demonstrates that optimal lighting is independent of image texture information. Each image pair exhibits a unique optimal lighting, yet consistently shows a minimum distance value.
Chao Wang 0084, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001, Wen Zhou 0029
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Hidden Follower Detection: How Is the Gaze-Spacing Pattern Embodied in Frequency Domain?
abstract
Spatiotemporal social behavior analysis is a technique that studies the social behavior patterns of objects and estimates their risks based on their trajectories. In social public scenarios such as train stations, hidden following behavior has become one of the most challenging issues due to its probability of evolving into violent events, which is more than 25%. In recent years, research on hidden following detection (HFD) has focused on differences in time series between hidden followers and normal pedestrians under two temporal characteristics: gaze and spatial distance. However, the time-domain representation for time series is irreversible and usually causes the loss of critical information. In this paper, we deeply study the expression efficiency of time/frequency domain features of time series, by exploring the recovery mechanism of features to source time series, we establish a fidelity estimation method for feature expression and a selection model for frequency-domain features based on the signal-to-distortion ratio (SDR). Experimental results demonstrate the feature fidelity of time series and HFD performance are positively correlated, and the fidelity of frequency-domain features and HFD performance are significantly better than the time-domain features. On both real and simulated datasets, the accuracy of the proposed method is increased by 3%, and the gaze-only module is improved by 10%. Related research has explored new methods for optimal feature selection based on fidelity, new patterns for efficient feature expression of hidden following behavior, and the mechanism of multimodal collaborative identification.
Ruimin Hu, Suhui Li
AAAI2
2024 Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model
abstract
Adversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low transferability and high detectability. In this paper, we propose a unified framework Adv-Diffusion that can generate imperceptible adversarial identity perturbations in the latent space but not the raw pixel space, which utilizes strong inpainting capabilities of the latent diffusion model to generate realistic adversarial images. Specifically, we propose the identity-sensitive conditioned diffusion generative model to generate semantic perturbations in the surroundings. The designed adaptive strength-based adversarial perturbation algorithm can ensure both attack transferability and stealthiness. Extensive qualitative and quantitative experiments on the public FFHQ and CelebA-HQ datasets prove the proposed method achieves superior performance compared with the state-of-the-art methods without an extra generative model training process. The source code is available at https://github.com/kopper-xdu/Adv-Diffusion.
Decheng Liu, Xijun Wang 0005, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
AAAI5
2024 Mutuality Attribute Makes Better Video Anomaly Detection
abstract
Video anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep features to characterize attributes, ignore the mutuality among multiple deep features. Therefore, the constructed representation is limited to indirectly representing the anomaly from isolated attributes, and makes the network difficult to capture the high-level causes of anomaly. In this paper, we proposed a novel Mutuality Attribute-based Representation framework (MAR-VAD) for the VAD task, which absorbs the mutuality among deep features to characterize the mutuality attribute. Specifically, the mutuality attribute encapsulates high-level semantic information, such as the specific abnormal object or action, which mutually utilizes information from multiple deep features. In this way, the system is able to directly capture the high-level causes of anomaly, thus providing a more comprehensive perspective to accurately detect anomaly events. Following a process-transparent density estimation, we produce the final anomaly scores. Experiments show that MAR-VAD achieves state-of-the-art performance on ShanghaiTech and Avenue.
Xingshuo Han, Xiao Wang 0029, Kui Jiang, Wei Liu 0183, Ruimin Hu, Xuefeng Pan, Xin Xu 0007
ICASSP5
2024 Robust Heterophilic Graph Learning against Label Noise for Anomaly Detection
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Yilong Zang
IJCAI2
2024 Heterophilic Graph Invariant Learning for Out-of-Distribution of Fraud Detection
abstract
Graph-based fraud detection (GFD) has garnered increasing attention due to its effectiveness in identifying fraudsters within multimedia data such as online transactions, product reviews, or telephone voices. However, the prevalent in-distribution (ID) assumption significantly impedes the generalization of GFD approaches to out-of-distribution (OOD) scenarios, which is a pervasive challenge considering the dynamic nature of fraudulent activities. In this paper, we introduce the Heterophilic Graph Invariant Learning Framework (HGIF), a novel approach to bolster the OOD generalization of GFD. HGIF addresses two pivotal challenges: creating diverse virtual training environments and adapting to varying target distributions. Leveraging edge-aware augmentation, HGIF efficiently generates multiple virtual training environments characterized by generalized heterophily distributions, thereby facilitating robust generalization against fraud graphs with diverse heterophily degrees. Moreover, HGIF employs a shared dual-channel encoder with heterophilic graph contrastive learning, enabling the model to acquire stable high-pass and low-pass node representations during training. During the Test-time Training phase, the shared dual-channel encoder is flexibly fine-tuned to adapt to the test distribution through graph contrastive learning. Extensive experiments showcase HGIF's superior performance over existing methods in OOD generalization, setting a new benchmark for GFD in OOD scenarios.
Lingfei Ren, Ruimin Hu, Zheng Wang 0007, Yilin Xiao 0002, Dengshi Li, Junhang Wu, Yilong Zang, Jinzhang Hu
ACM Multimedia2
2024 Optimal Illumination Distance Metrics for Person Re-identification
Chao Wang 0084, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001, Wen Zhou 0029
PRICAI (4)3
2024 Sensor-Based Authentication on Smartphones via Integrating Auxiliary Information
abstract
With the widespread use of smartphones, more and more private information is stored on the phone, and the loss or theft of the phone can lead to data leakage, theft of property, and other problems. Traditional active authentication methods based on passwords, faces, fingerprints, etc. authenticate only once at login, which still has some hidden dangers. Authentication methods based on behavioral biometrics such as walking gait, touch screen, keystroke, etc. can provide implicit and continuous authentication services, and thus have been widely studied recently. Considering the limited computational resources of smart devices, we first introduce the state-space model with linear model complexity, i.e., Mamba, to extract deep feature representations from raw signal sequences in the sensor-based authentication task. Then a series of auxiliary information from different domains are proposed and fused to the deep features to enhance the consistent semantic information and ultimately improve the discriminative ability of the model. Our model A-Mamba, i.e., Auxiliary-Mamba, is experimented on two public datasets. The experimental results show that our model outperforms existing approaches.
Ruimin Hu
SMC2
2024 Acoustic scene classification: A comprehensive survey
Biyun Ding, Tao Zhang 0025, Chao Wang 0135, Ganjun Liu, Jinhua Liang, Ruimin Hu, Yulin Wu 0003, Difei Guo
Expert Syst. Appl.6
2024 Do not ignore heterogeneity and heterophily: Multi-network collaborative telecom fraud detection
Lingfei Ren, Yilong Zang, Ruimin Hu, Dengshi Li, Junhang Wu, Jinzhang Hu
Expert Syst. Appl.3
2024 Adaptive subband partition encoding scheme for multiple audio objects using CNN and residual dense blocks mixture network
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001
Expert Syst. Appl.2
2024 A GNN-based fraud detector with dual resistance to graph disassortativity and imbalance
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilong Zang
Inf. Sci.2
2024 Learning with noisy labels for robust fatigue detection
Ruimin Hu, Xiaojie Zhu, Dongliang Zhu 0001, Xiaochen Wang 0001
Knowl. Based Syst.2
2024 Domain generalized person reidentification based on skewness regularity of higher-order statistics
Mingfu Xiong, Ruimin Hu, Zhongyuan Wang 0001, Javier Del Ser, Khan Muhammad 0001, Zixiang Xiong
Knowl. Based Syst.3
2024 Improving fraud detection via imbalanced graph structure learning
Lingfei Ren, Ruimin Hu, Yang Liu 0200, Dengshi Li, Junhang Wu, Yilong Zang, Wenyi Hu
Mach. Learn.2
2024 Beyond the individual: An improved telecom fraud detection approach based on latent synergy graph learning
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Yilong Zang
Neural Networks2
2024 Watch You Under Low-Resolution and Low-Illumination: Face Enhancement via Bi-Factor Degradation Decoupling
abstract
Face enhancement aims to improve low-quality face images to a higher-quality level. However, in real-world nighttime scenes, complex degradation factors often affect these images, making it challenging to preserve important facial details. Existing image enhancement algorithms typically focus on independently conducting image super-resolution and brightness enhancement, assuming a fixed degradation level based on simulated training datasets. Nonetheless, real nighttime scenes involve complex degradation processes, where degradation factors dynamically and variably manifest. Therefore, achieving effective face enhancement in such scenarios is particularly daunting. This work analyzes and unveils the multiple factors of low resolution and low illumination during degradation. Based on this analysis, we propose a Bi-factor Degradation Decoupling network. Our method leverages a decoupling network to generate qualitative and quantitative features corresponding to each factor’s degradation degree in the low-quality environment. These features are then combined with robust facial feature constraints to recover the details of low-quality faces. Extensive experiments demonstrate that our method surpasses state-of-the-art approaches in both enhancement and face super-resolution.
Zheng Wang 0007, Zhenyu Shu, Ruimin Hu, Chia-Wen Lin
IEEE Trans. Circuits Syst. Video Technol.5
2024 MRLReID: Unconstrained Cross-Resolution Person Re-Identification With Multi-Task Resolution Learning
abstract
Cross-resolution person re-identification (ReID) is a challenging task that addresses the issue of matching individuals across different resolution conditions. Traditional person ReID methods often assume that images have sufficiently high resolution and overlook the practical scenarios involving low-resolution or blurry images. Existing cross-resolution ReID approaches either utilize image super-resolution techniques to improve the quality of low-resolution images or extract and learn resolution invariant features for person representation. Although multi-task learning has been applied in ReID to integrate auxiliary tasks including attribute recognition, image super-resolution, and so on, how to incorporate the vital resolution learning task into cross-resolution ReID has rarely explored before. Therefore, we propose a novel multi-task resolution learning based ReID network named MRLReID. Our approach treats ross-resolution person ReID as the primary task and the resolution estimation as an auxiliary task. Our network simultaneously learns the resolution information and person identity information of images, aiming to improve cross-resolution person ReID performance. Considering that existing similuated cross-resolution datasets are too simple to mimic unconstrained scenario, we further employ image degradation technique to simulate more realistic cross-resolution ReID datasets. We evaluate our method on two real-world cross-resolution datasets and two newly simulated cross-resolution datasets, and both intra-dataset and cross-dataset evaluations demonstrate the effectiveness and superiority of our method in cross-resolution person ReID. The codes and datasets are available at https://github.com/amateurbo/MRLReID.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Where Deepfakes Gaze at? Spatial-Temporal Gaze Inconsistency Analysis for Video Face Forgery Detection
abstract
With the continuous development of generative models on face generation, how to distinguish the real and fake face has become an important problem for security. Because of the continuous improvement on the detection accuracy by facial physiological signals, video face forgery detection based on facial physiological signal analysis has received more and more attention, which has become an important research branch in the field of face forgery detection. Currently, most of the research on forgery detection based on physiological signal analysis use biometric features such as blinking patterns, head swings, heart rate signals, and lip movements. However, there hasn’t been much exploration on the usage of gaze features in face forgery detection. Through the analysis of gaze directions in face videos, we have observed differences in the distribution of gaze direction pattern between the real and forged videos. Specifically, real videos tend to have more concentrated gaze distribution within a short period of time, while forged videos have more dispersed gaze distributions. In this paper, we present a novel Deepfake gaze analysis method named DFGaze, to explore spatial-temporal gaze inconsistency for video face forgery detection. Our method uses the gaze analysis model (GAM) to analyze the gaze features of face video frames, and then applies a spatial-temporal feature aggregator to realize authenticity classification based on gaze features. In order to better mine the authenticity clues in the videos, we further use the texture analysis model (TAM) and attribute analysis model (AAM) to improve the representation ability of spatial-temporal feature differences between real and forged faces. Extensive experiments show that our method can achieve state-of-the-art performance with the help of gaze analysis. The source code is available at https://github.com/ziminMIAO/DFGaze.
Chunlei Peng, Zimin Miao, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2024 From Multi-Source Virtual to Real: Effective Virtual Data Search for Vehicle Re-Identification
abstract
Without tedious and time-consuming labeling processes, virtual datasets have recently shown their superiority for vehicle re-identification (re-ID). Existing virtual to real vehicle re-ID methods employ only a single virtual dataset for model training, while datasets from different generative sources are not jointly exploited. Multiple source virtual datasets contain more data diversity that can boost model performance. We thus propose a multi-source virtual to real vehicle re-ID pipeline, where multiple source virtual datasets are used during training. However, the multi-source virtual dataset suffers from more data redundancy than the single virtual dataset, which can affect the training efficiency. Intuitively, it can be mitigated by virtual data search. Unlike a single virtual dataset, a performance gap exists between multiple source virtual datasets, indicating their different contributions to model learning. Accordingly, we propose to split the multi-source virtual dataset into the main training set and the auxiliary training set, and then design the sampling strategy separately. For the main training set, the Consistent Attribute Distribution-FEature distance Trade-off (CAD-FET) strategy is designed to search for representative data. For the auxiliary training set, a cluster-based sampling strategy is further proposed to search for the most diverse subset. Besides, a simple yet effective two-stage training strategy is proposed to utilize these subsets reasonably. Extensive virtual-to-real vehicle re-ID experiments show that our data sampling method can reduce the volume of the multi-source virtual dataset by around 77%/96% and boost the model performance when tested on the VeRi776/VehicleID.
Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Zhixiang Wang 0001, Ruimin Hu
IEEE Trans. Intell. Transp. Syst.5
2024 An Image Arbitrary-Scale Super-Resolution Network Using Frequency-domain Information
abstract
Image super-resolution (SR) is a technique to recover lost high-frequency information in low-resolution (LR) images. Since spatial-domain information has been widely exploited, there is a new trend to involve frequency-domain information in SR tasks. Besides, image SR is typically application-oriented and various computer vision tasks call for image arbitrary magnification. Therefore, in this article, we study image features in the frequency domain to design a novel image arbitrary-scale SR network. First, we statistically analyze LR-HR image pairs of several datasets under different scale factors and find that the high-frequency spectra of different images under different scale factors suffer from different degrees of degradation, but the valid low-frequency spectra tend to be retained within a certain distribution range. Then, based on this finding, we devise an adaptive scale-aware feature division mechanism using deep reinforcement learning, which can accurately and adaptively divide the frequency spectrum into the low-frequency part to be retained and the high-frequency one to be recovered. Finally, we design a scale-aware feature recovery module to capture and fuse multi-level features for reconstructing the high-frequency spectrum at arbitrary scale factors. Extensive experiments on public datasets show the superiority of our method compared with state-of-the-art methods.
Yinbo Yu, Zhongyuan Wang 0001, Ruimin Hu
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Inter-camera Identity Discrimination for Unsupervised Person Re-identification
abstract
Unsupervised person re-identification (Re-ID) has garnered significant attention because of its data-friendly nature, as it does not require labeled data. Existing approaches primarily address this challenge by employing feature-clustering techniques to generate pseudo-labels. In addition, camera-proxy-based methods have emerged because of their impressive ability to cluster sample identities. However, these methods often blur the distinctions between individuals within inter-camera views, which is crucial for effective person re-ID. To address this issue, this study introduces an inter-camera-identity-difference-based contrastive learning framework for unsupervised person Re-ID. The proposed framework comprises two key components: (1) a different sample cross-view close-range penalty module and (2) the same sample cross-view long-range constraint module. The former aims at penalizing excessive similarity among different subjects across inter-camera views, whereas the latter mitigates the challenge of excessive dissimilarity among the same subject across camera views. To validate the performance of our method, we conducted extensive experiments on three existing person Re-ID datasets (Market-1501, MSMT17, and PersonX). The results demonstrate the effectiveness of the proposed method, which shows a promising performance. The code is available at https://github.com/hooldylan/IIDCL .
Mingfu Xiong, Kaikang Hu, Zhihan Lyu, Zhongyuan Wang 0001, Ruimin Hu, Khan Muhammad 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Crowd-Level Abnormal Behavior Detection via Multi-Scale Motion Consistency Learning
abstract
Detecting abnormal crowd motion emerging from complex interactions of individuals is paramount to ensure the safety of crowds. Crowd-level abnormal behaviors (CABs), e.g., counter flow and crowd turbulence, are proven to be the crucial causes of many crowd disasters. In the recent decade, video anomaly detection (VAD) techniques have achieved remarkable success in detecting individual-level abnormal behaviors (e.g., sudden running, fighting and stealing), but research on VAD for CABs is rather limited. Unlike individual-level anomaly, CABs usually do not exhibit salient difference from the normal behaviors when observed locally, and the scale of CABs could vary from one scenario to another. In this paper, we present a systematic study to tackle the important problem of VAD for CABs with a novel crowd motion learning framework, multi-scale motion consistency network (MSMC-Net). MSMC-Net first captures the spatial and temporal crowd motion consistency information in a graph representation. Then, it simultaneously trains multiple feature graphs constructed at different scales to capture rich crowd patterns. An attention network is used to adaptively fuse the multi-scale features for better CAB detection. For the empirical study, we consider three large-scale crowd event datasets, UMN, Hajj and Love Parade. Experimental results show that MSMC-Net could substantially improve the state-of-the-art performance on all the datasets.
Linbo Luo 0001, Yuanjing Li, Haiyan Yin, Shangwei Xie, Ruimin Hu, Wentong Cai 0001
AAAI5
2023 Gaze Behavior Patterns for Early Drowsiness Detection
Hongfei Gao, Ruimin Hu
ICANN (1)2
2023 User and Interaction Both Matter: Social Relationship Mining Via Interaction Graph Propagating
abstract
Social relationship mining benefits many applications such as leadership analysis and advisor recommendation. Existing methods focus on mining user relationships only from the perspective of user-level. To our knowledge, from this perspective, representing the user interactions by edges is not sufficient for the complex information about interactions between users. In addition, mining users' relationship independently ignores the propagation of social interaction across networks. In this paper, we investigate social relationship mining from a new perspective of interaction-level. We propose an Interaction Graph Propagating(IGP) model which constructs an interaction graph. It not only captures the user interaction information as the union but also exploits the propagation between user interactions. In particular, we utilize the graph attention mechanism to distinguish the contributions of each neighbor union. Experimental results on several public datasets demonstrate that IGP achieves significant improvements over state-of-the-art methods.
Yilong Zang, Ruimin Hu, Zheng Wang 0007, Dengshi Li
ICC2
2023 Hidden Follower Detection via Refined Gaze and Walking State Estimation
abstract
Hidden following is following behavior with special intentions, and detecting hidden following behavior can prevent many criminal activities in advance. The previous method uses gaze and spacing behaviors to distinguish hidden followers from normal pedestrians. However, they express gaze behaviors in a coarse-grained way with binary values, making it difficult to accurately depict the gaze state of pedestrians. To this end, we propose the Refined Hidden Follower Detection (RHFD) model by choosing a suitable mapping function based on the principle that the closer the gaze direction is to someone, the more likely it is to gaze at someone, which converts the gaze direction into a continuous estimated gaze state representing the complex and variable gaze behavior of pedestrians. Simultaneously, we introduce variations in the magnitude and direction of pedestrian velocity to refine the representation of pedestrian walking states. Experimental results on the surveillance dataset show that RHFD outperforms state-of-the-art methods.
Yaxi Chen, Ruimin Hu, Danni Xu, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li
ICME2
2023 Multi-speaker Direction of Arrival Estimation Using Audio and Visual Modalities with Convolutional Neural Network
abstract
In reality, audible and visible sound sources are closely aligned, and they can help humans locate sources exactly. To exploit the complementarity between audio and visual data in multi-speaker direction of arrival (DoA) estimation, we propose a novel network consisting of 3D convolution neural networks (3D-CNNs) and 2D-CNNs mixture networks with residual dense blocks. It has two main advantages: 1) both input audio and visual features are low-level signal representation: the real and imaginary parts of STFT coefficients for the audio feature and pixel coordinates for the visual feature, which can allow the network to learn to extract the most informative high-level features. 2) 3D-CNNs with the residual dense block are used for audio and visual feature mapping along the time and frequency axis. The following 2D-CNNs are to ensemble the high-level features along the DoA axis. Experimental results demonstrate promising SSL performance.
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001
ICME2
2023 Perceptual Audio Object Coding Using Adaptive Subband Grouping with CNN and Residual Block
abstract
Spatial audio content is becoming increasingly popular and is regarded as a set of object signals with associated metadata. The object-based content representation is independent of loudspeaker layouts and provides high spatial resolution when reproduced on more loudspeakers. The audio quality of the traditional spatial audio object coding (SAOC) method has severe aliasing distortion, which impairs the immersive listening experience. In this study, we reduce aliasing distortion by perceptual adaptive subband grouping strategy and use the convolutional neural network (CNN) and residual block to build the side information compressing model. Both objective and subjective experiments on benchmark datasets with different bitrates show that the proposed method achieves favorable performance against state-of-the-art methods.
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001
ICME2
2023 Don't Ignore Alienation and Marginalization: Correlating Fraud Detection
abstract
The anonymity of online networks makes tackling fraud increasingly costly. Thanks to the superiority of graph representation learning, graph-based fraud detection has made significant progress in recent years. However, upgrading fraudulent strategies produces more advanced and difficult scams. One common strategy is synergistic camouflage —— combining multiple means to deceive others. Existing methods mostly investigate the differences between relations on individual frauds, that neglect the correlation among multi-relation fraudulent behaviors. In this paper, we design several statistics to validate the existence of synergistic camouflage of fraudsters by exploring the correlation among multi-relation interactions. From the perspective of multi-relation, we find two distinctive features of fraudulent behaviors, i.e., alienation and marginalization. Based on the finding, we propose COFRAUD, a correlation-aware fraud detection model, which innovatively incorporates synergistic camouflage into fraud detection. It captures the correlation among multi-relation fraudulent behaviors. Experimental results on two public datasets demonstrate that COFRAUD achieves significant improvements over state-of-the-art methods.
Yilong Zang, Ruimin Hu, Zheng Wang 0007, Danni Xu, Jia Wu 0001, Dengshi Li, Junhang Wu, Lingfei Ren
IJCAI2
2023 Collaborative Fraud Detection: How Collaboration Impacts Fraud Detection
abstract
Collaborative fraud has become increasingly serious in telecom and social networks, but is hard to detect by traditional fraud detection methods. In this paper, we find a significant positive correlation between the increase of collaborative fraud and the degraded detection performance of traditional techniques, implying that those fraudsters that are difficult to detect with traditional methods are often collaborative in their fraudulent behavior. As we know, multiple objects may contact a single target object over a period of time. We define multiple objects with the same contact target as generalized objects, and their social behaviors can be combined and processed as the social behaviors of one object. We propose Fraud Detection Model based on Second-order and Collaborative Relationship Mining (COFD), exploring new research avenues for collaborative fraud detection. Our code and data are released at https://github.com/CatScarf/COFD-MM https://github.com/CatScarf/COFD-MM.
Jinzhang Hu, Ruimin Hu, Zheng Wang 0007, Dengshi Li, Junhang Wu, Lingfei Ren, Yilong Zang
ACM Multimedia2
2023 Modality-agnostic Augmented Multi-Collaboration Representation for Semi-supervised Heterogenous Face Recognition
abstract
Heterogeneous face recognition (HFR) aims to match input face identity across different image modalities. Due to the existing large modality gap and the limited number of training data, HFR is still a challenging problem in biometrics and draws more and more attention. Existing researchers always extract modality invariant features or generate homogeneous images to decrease the modality gap, lacking abundant labeled data to avoid the overfitting problem. In this paper, we proposed a novel Modality-Agnostic Augmented Multi-Collaboration representation for Heterogeneous Face Recognition (MAMCO-HFR) in a semi-supervised manner. The modality-agnostic augmentation strategy is proposed to generate adversarial perturbations to map unlabeled faces into the modality-agnostic domain. The multi-collaboration feature constraint is designed to mine the inherent relationships between diverse layers for discriminative representation. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ and Tufts Face dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is available at https://github.com/xiyin11/Semi-HFR.
Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
ACM Multimedia5
2023 Cross-Illumination Video Anomaly Detection Benchmark
abstract
Video anomaly detection is a critical problem with widespread applications in domains such as security surveillance. Most existing methods focus on video anomaly detection tasks under uniform illumination conditions. However, in the real world, the situation is much more complicated. Video anomalies are widespread across periods and under different illumination conditions, which can lead to the detector model incorrectly reporting high anomaly scores. To address this challenge, we design a benchmark framework for the cross-illumination video anomaly detection task. The framework restores videos under different illumination scales to the same illumination scale. This reduces domain differences between uniformly illuminated training videos and differently illuminated test videos. Additionally, to demonstrate the illumination change problem and evaluate our model, we construct three large-scale datasets with a wide range of illumination variations. We experimentally validate our approach on three cross-illuminance video anomaly detection datasets. Experimental results show that our method outperforms existing methods regarding detection accuracy and is more robust.
Dongliang Zhu 0001, Ruimin Hu, Zheng Wang 0007
ACM Multimedia2
2023 Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly Detection
abstract
Unsupervised image Anomaly Detection (UAD) aims to learn robust and discriminative representations of normal samples. While separate solutions per class endow expensive computation and limited generalizability, this paper focuses on building a unified framework for multiple classes. Under such a challenging setting, popular reconstruction-based networks with continuous latent representation assumption always suffer from the "identical shortcut" issue, where both normal and abnormal samples can be well recovered and difficult to distinguish. To address this pivotal issue, we propose a hierarchical vector quantized prototype-oriented Transformer under a probabilistic framework. First, instead of learning the continuous representations, we preserve the typical normal patterns as discrete iconic prototypes, and confirm the importance of Vector Quantization in preventing the model from falling into the shortcut. The vector quantized iconic prototypes are integrated into the Transformer for reconstruction, such that the abnormal data point is flipped to a normal data point. Second, we investigate an exquisite hierarchical framework to relieve the codebook collapse issue and replenish frail normal patterns. Third, a prototype-oriented optimal transport method is proposed to better regulate the prototypes and hierarchically evaluate the abnormal score. By evaluating on MVTec-AD and VisA datasets, our model surpasses the state-of-the-art alternatives and possesses good interpretability. The code is available at https://github.com/RuiyingLu/HVQ-Trans.
Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Ruimin Hu
NeurIPS7
2023 Multi-scale modeling temporal hierarchical attention for sequential recommendation
Nana Huang, Ruimin Hu, Xiaochen Wang 0001
Inf. Sci.2
2023 Cross-platform sequential recommendation with sharing item-level relevance data
Nana Huang, Ruimin Hu, Xiaochen Wang 0001, Xinjian Huang
Inf. Sci.2
2023 Dynamic graph neural network-based fraud detectors against collaborative fraudsters
Lingfei Ren, Ruimin Hu, Dengshi Li, Yang Liu 0200, Junhang Wu, Yilong Zang, Wenyi Hu
Knowl. Based Syst.2
2023 Dual-focus: person search from Coarse-Grained Focus to Fine-Grained Focus
Wenyi Hu, Xiao Wang 0029, Zheng Wang 0007, Xin Xu 0007, Ruimin Hu
Multim. Syst.5
2023 Who is your friend: inferring cross-regional friendship from mobility profiles
Lingfei Ren, Ruimin Hu, Dengshi Li, Zheng Wang 0007, Junhang Wu, Wenyi Hu
Multim. Tools Appl.2
2023 Face enhancement and hallucination in the wild
Ruimin Hu, Zhongyuan Wang 0001
Neural Comput. Appl.2
2023 Beyond fixed time and space: next POI recommendation via multi-grained context and correlation
Ruimin Hu, Zheng Wang 0007
Neural Comput. Appl.2
2023 Single-channel Multi-speakers Speech Separation Based on Isolated Speech Segments
Shanfa Ke, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001
Neural Process. Lett.3
2023 Where Have You Gone: Category-aware Multigraph Embedding for Missing Point-of-Interest Identification
Junhang Wu, Ruimin Hu, Dengshi Li, Yilin Xiao 0002, Lingfei Ren, Wenyi Hu
Neural Process. Lett.2
2023 Multi-speaker DoA Estimation Using Audio and Visual Modality
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Shanfa Ke
Neural Process. Lett.2
2023 CycMuNet+: Cycle-Projected Mutual Learning for Spatial-Temporal Video Super-Resolution
abstract
Spatial-Temporal Video Super-Resolution (ST-VSR) aims to generate high-quality videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR by directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal Video Super-Resolution (T-VSR) but ignore the reciprocal relations among them. 1) T-VSR to S-VSR: temporal correlations help accurate spatial detail representation; 2) S-VSR to T-VSR: abundant spatial information contributes to the refinement of temporal prediction. To this end, we propose a one-stage based Cycle-projected Mutual learning network (CycMuNet) for ST-VSR, which makes full use of spatial-temporal correlations via the mutual learning between S-VSR and T-VSR. Specifically, we propose to exploit the mutual information among them via iterative up- and down projections, where spatial and temporal features are fully fused and distilled, helping high-quality video reconstruction. In addition, we also show interesting extensions for efficient network design (CycMuNet+), such as parameter sharing and dense connection on projection units and feedback mechanism in CycMuNet. Besides extensive experiments on benchmark datasets, we also compare our proposed CycMuNet (+) with S-VSR and T-VSR tasks, demonstrating that our method significantly outperforms the state-of-the-art methods.
Mengshun Hu, Kui Jiang, Zheng Wang 0007, Xiang Bai, Ruimin Hu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 A Dual Self-Attention mechanism for vehicle re-Identification
Wenqian Zhu, Zhongyuan Wang 0001, Xiaochen Wang 0001, Ruimin Hu, Huikai Liu, Chao Wang 0084, Dengshi Li
Pattern Recognit.4
2023 From Collective Attribute Association of Groups to Precise Attribute Association of Individuals
abstract
Obscured person re-identification (Re-ID) aims to match an obscured image with a complete image of the same person captured by other cameras. As a major challenge in person identification, occlusion severely affects the effectiveness of most traditional person Re-ID methods. To solve this problem, this study proposes a trajectory association method, which, as a pre-processing technique for person Re-ID, can narrow the search range and reduce the problem of degradation caused by mixing. We investigate the method of converting the fuzzy association between sets into the precise association between elements for M video objects and N phone objects (trajectory information) with fuzzy group association relationships at the crime scene. First, we decompose the M-N precise association problem and analyze the similarity of the video objects in the source point and on the trajectories. Then, we define high-similarity points, study their distribution characteristics in different trajectories, and find that there is a significant difference between the distribution of high-similarity points in correct and incorrect matching trajectories. We simplify the full-path association problem into a partial-path high-similarity point distribution difference problem, which effectively reduces the difficulty in accurate association relationship construction. The association experiments in simple and mixed scenarios as well as Re-ID experiments on the PRPW and Market1501 demonstrate the effectiveness of our method.
Yun Lan, Ruimin Hu, Xin Xu 0007, Dengshi Li, Chao Wang 0084, Xiaochen Wang 0001
IEEE Trans. Multim.2
2022 How to Face Unseen Defects? UDGAN for Improving Unseen Defects Recognition
Yaxi Chen, Ruimin Hu, Zheng Wang 0007
ICANN (1)2
2022 Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve Refinement
abstract
Deep learning networks with deeper layers become a trend for their good performance but lacks the potential for real-time mobile deployment. Another challenge for paired training networks is the limited generalization capacity caused by the sample bias. To overcome these two challenges, we propose a lightweight self-supervised low-light image enhancement method, that trains with low light images only. Specifically, our method consists of a low-resolution dense CNN network stream and a full-resolution guidance stream, responsible for image-to-curve transformation with refinement and spatial guidance fusion, respectively. Then, a new self-supervised loss function is introduced to measure the restored patch-based color deviations among color channels. Experimental results show that our method gives competitive performance to the full-supervised approaches.
Wanyu Wu, Wei Wang 0170, Kui Jiang, Xin Xu 0007, Ruimin Hu
ICASSP5
2022 Discriminative Region Transfer Network for Cross-Database Micro-Expression Recognition
abstract
Compared to conventional micro-expression analysis, cross-database micro-expression recognition (CDMER) considers a more practical scenario, where the training and testing samples come from different databases. Under this problem setting, most previous micro-expression recognition methods may suffer severe performance degradation due to dataset bias or domain shift. This makes CDMER a challenging but interesting problem. Most existing CDMER methods transfer global micro-expression features without considering the different contributions of facial regions to micro-expression. To tackle this problem, we first confirm the following argument: for samples of the same category from different datasets, their discriminative facial regions are similar and relevant. Therefore, we can transfer the knowledge about discriminative facial regions learned from the source domain to the target domain directly. Based on this argument, we design a novel deep domain adaption method called discriminative region transfer network (DRTN). The DRTN uses an adversarial-based adaptation method to align the source and the target distribution in the learned feature space. Then during transfer, we focus on the discriminative facial regions by the feat of attention mechanism. We conducted extensive experiments on CASME II and SMIC datasets and achieves very competitive results. These evaluations convincingly demonstrate the effectiveness of our proposed method.
Jianbang Li, Ruimin Hu, Mithun Mukherjee 0001
ICC2
2022 A Bi-directional Category-Aware Multi-task Learning Framework for Missing Check-in POI Identification
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilong Zang
ICSOC2
2022 IDGL: An Imbalanced Disassortative Graph Learning Framework for Fraud Detection
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilong Zang
ICSOC2
2022 $(\alpha, \ \beta)$-AWCS: $(\alpha, \ \beta)$-Attributed Weighted Community Search on Bipartite Graphs
abstract
Community search on bipartite graphs aims to find a community closely associated with the query vertex for personalized recommendation, fraud detection, and team formation. During community search, considering both the structural closeness and the homogeneity of attributes of nodes is the key to improving the quality of the output community. Traditional work uses the$(\alpha,\beta)$-core model to guarantee structural cohesion of the nodes (i.e., degree of each upper vertex is at least$\alpha$and degree of each lower vertex is at least$\beta$). However, it ignores the attributes of nodes, resulting in an average attribute similarity of only about 0.17 for the node pairs in the output community. In this paper, a framework for$(\alpha,\ \beta)$-Attributed Weighted Community Search ($(\alpha,\ \beta)$-AWCS) was proposed. It output a connected subgraph of$G$containing the query vertex, which satisfies both structurally cohesive (i.e., ($(\alpha,\ \beta){-}$-core) and keyword cohesiveness (i.e., its vertices share common keywords). The framework includes a pruning strategy to strip vertices that do not contain query attributes, thus effectively reducing the search space, and two algorithms improve the attributes cohesiveness of the output community. One of the exact algorithms first obtains a subgraph of attribute cohesion and subsequently keeps the structure cohesive. The other approximate algorithm has higher robustness, which iteratively removes the vertex with the lowest attribute score until the structural cohesion cannot be maintained. We have conducted experiments on real datasets of different sizes. Experiments show that both algorithms can improve the attribute cohesiveness metric by more than 25% compared to the traditional method. Meanwhile, structural cohesion was appropriate.
Dengshi Li, Xiaocong Liang, Ruimin Hu, Xiaochen Wang 0001
IJCNN3
2022 ITC: Influential-Truss Community Search
abstract
Community search is a method of finding a com-munity closely related to a query node. The latest influence community search considers both the structural cohesion of the community and the influence between nodes. It sets the influence threshold to constrain the output community. However, artificially setting the influence threshold makes the output community too large or too small, which leads to low accuracy of the output community. In order to avoid the low accuracy of community search caused by artificially setting influence thresh-old constraints, this paper studies the community search problem based on community influence score. In this paper, an influence-truss community (ITC) model is proposed for community search by combining structural cohesion and community influence score. This model aims to obtain a connected subgraph in a social network containing the query node, which satisfies structural cohesion and satisfies the subgraph's maximum community in-fluence score. In order to obtain ITC, an effective pruning method is proposed, which strips other nodes far away from the query node. Then, the ITCS algorithm is designed, which firstly imposes structural cohesion constraints on query nodes. Then, the search community's influence scores are iteratively calculated until the community has the highest community influence score under the condition of meeting the structural cohesion. Experiments on real-world networks of different scales show that the community search accuracy index of ITCS is improved by about 20% compared with the traditional method.
Dengshi Li, Ruimin Hu, Xiaocong Liang, Yilong Zang
IJCNN3
2022 User Alignment Across Social Networks Based On ego-Network Embedding
abstract
Cross-social network user alignment is to find users with the same identity in multiple social networks. It has important applications in natural and scientific fields, such as link prediction and personality recommendation, and has certain research value in the field of data mining. Most current approaches embed social networks in a low-dimensional vector space and then align users in the low-dimensional space. However, because the social network is extremely complex and large, it is easy to be affected by error propagation and noise of different neighbors in the process of network embedding. Therefore, to obtain better embedding, we first form the user's EGO network, then use the random walk to extract the user node sequence, then use the framework of the natural language model to learn the low-dimensional vector representation of the user, and finally train a matrix to map the two social networks into the same feature space for alignment. Our experiments on real-world data set Foursquare-Twitter and Livejournal-myspace show some improvement over several baseline results.
Yu Zhen, Ruimin Hu, Dengshi Li, Yilin Xiao 0002
IJCNN2
2022 Cross-Regional Friendship Inference via Category-Aware Multi-Bipartite Graph Embedding
abstract
This paper proposes a novel problem of cross-regional friendship inference to solve the geographically restricted friends recommendation. Traditional approaches rely on a fundamental assumption that friends tend to be co-location, which is unrealistic for inferring friendship across regions. By reviewing a large-scale Location-based Social Networks (LBSNs) dataset, we spot that cross-regional users are more likely to form a friendship when their mobility neighbors are of high similarity. To this end, we propose Category-Aware Multi-Bipartite Graph Embedding (CMGE for short) for cross-regional friendship inference. We first utilize multi-bipartite graph embedding to capture users’ Point of Interest (POI) neighbor similarity and activity category similarity simultaneously, then the contributions of each POI and category are learned by a category-aware heterogeneous graph attention network. Experiments on the real-world LBSNs datasets demonstrate that CMGE outperforms state-of-the-art baselines.
Linfei Ren, Ruimin Hu, Dengshi Li, Junhang Wu, Yilong Zang, Wenyi Hu
LCN2
2022 Gaze- and Spacing-flow Unveil Intentions: Hidden Follower Discovery
abstract
We raise a new and challenging multimedia application in video surveillance system, i.e., Hidden Follower Discovery (HFD). In contrast to the common abnormal behaviors that are occurring, hidden following is not an ongoing activity, but a preparatory action. Hidden following behavior does not have salient features, making it hard to be discovered. Fortunately, from a socio-cognitive perspective, we found and verified the phenomena that the gaze-flow pattern and the spacing-flow pattern between hidden and normal followers are different. To promote HFD research, we construct two pioneering datasets and devise an HFD baseline network based on the recognition of both gaze-flow and spacing-flow patterns from surveillance videos. Extensive experiments demonstrate their effectiveness.
Danni Xu, Ruimin Hu, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li, Wenjun Zeng 0001
ACM Multimedia2
2022 Cover: International Journal of Intelligent Systems, Volume 37 Issue 5 May 2022
abstract
Cover Caption: The cover image is based on the Research Article Efficient virtual data search for annotationfree vehicle reidentification by Zhijing Wan et al., https://doi.org/10.1002/int.22829.
Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki, Xiaolong Zhang 0002, Ruimin Hu
Int. J. Intell. Syst.6
2022 Efficient virtual data search for annotation-free vehicle reidentification
abstract
Vehicle reidentification (re-ID) is the task of retrieving the same vehicle across nonoverlapping cameras, which has made significant progress with the help of abundant manually annotated real images. To avoid the time-consuming and tedious labeling of real images, virtual data sets with large-scale synthetic images have recently been constructed to perform annotation-free model training. However, current methods fail to exploit the potential of virtual data search, that is, searching valuable and representative virtual subdata set for efficient training. This paper presents a novel data sampling strategy from both semantic and feature levels to perform an effective data search. The semantic level determines the sample number of each vehicle identity via the consistency constraint of attribute distribution for source domain and target domain; while the feature level searches valuable and representative samples of each vehicle identity. To our knowledge, we are among the first attempts to search effective virtual data to perform annotation-free vehicle re-ID. Extensive cross-domain experiments from virtual vehicle re-ID data sets to real vehicle re-ID data sets show that our data sampling strategy can significantly reduce the training data volume and even boost the re-ID performance.
Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki, Xiaolong Zhang 0002, Ruimin Hu
Int. J. Intell. Syst.6
2022 Face hallucination based on degradation analysis for robust manifold
Ruimin Hu, Zheng He 0001, Chao Liang 0001, Zhongyuan Wang 0001
Neurocomputing2
2022 Single low-light image brightening using learning-based intensity mapping
Ruimin Hu
Neurocomputing2
2022 Where have you been: Dual spatiotemporal-aware user mobility modeling for missing check-in POI identification
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilin Xiao 0002
Inf. Process. Manag.2
2022 Next-point-of-interest recommendation based on joint mining of regularity and randomness
Ruimin Hu, Zheng Wang 0007
Knowl. Based Syst.2
2022 Spatiotemporal two-stream LSTM network for unsupervised video summarization
Ruimin Hu, Zhongyuan Wang 0001, Zixiang Xiong
Multim. Tools Appl.2
2022 Multi-scale Interest Dynamic Hierarchical Transformer for sequential recommendation
Nana Huang, Ruimin Hu, Mingfu Xiong, Xiaoran Peng, Xiaodong Jia 0005, Lingkun Zhang
Neural Comput. Appl.2
2022 High Parameter Frequency Resolution Encoding Scheme for Spatial Audio Objects Using Stacked Sparse Autoencoder
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Chenhao Hu, Shanfa Ke
Neural Process. Lett.2
2022 COVID-19 contact tracking by group activity trajectory recovery over camera networks
Chao Wang 0084, Xiaochen Wang 0001, Zhongyuan Wang 0001, Wenqian Zhu, Ruimin Hu
Pattern Recognit.5
2022 Towards generalizable person re-identification with a bi-stream generative model
Xin Xu 0007, Wei Liu 0183, Zheng Wang 0007, Ruimin Hu
Pattern Recognit.4
2022 Rank-in-Rank Loss for Person Re-identification
abstract
Person re-identification (re-ID) is commonly investigated as a ranking problem. However, the performance of existing re-ID models drops dramatically, when they encounter extreme positive-negative class imbalance (e.g., very small ratio of positive and negative samples) during training. To alleviate this problem, this article designs a rank-in-rank loss to optimize the distribution of feature embeddings. Specifically, we propose a Differentiable Retrieval-Sort Loss (DRSL) to optimize the re-ID model by ranking each positive sample ahead of the negative samples according to the distance and sorting the positive samples according to the angle (e.g., similarity score). The key idea of the proposed DRSL lies in minimizing the distance between samples of the same category along with the angle between them. Considering that the ranking and sorting operations are non-differentiable and non-convex, the DRSL also performs the optimization of automatic derivation and backpropagation. In addition, the analysis of the proposed DRSL is provided to illustrate that the DRSL not only maintains the inter-class distance distribution but also preserves the intra-class similarity structure in terms of angle constraints. Extensive experimental results indicate that the proposed DRSL can improve the performance of the state-of-the-art re-ID models, thus demonstrating its effectiveness and superiority in the re-ID task.
Xin Xu 0007, Xin Yuan 0009, Zheng Wang 0007, Kai Zhang 0002, Ruimin Hu
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Variance Weight Distribution Network Based Noise Sample Learning for Robust Person Re-identification
Xiaoyi Long, Ruimin Hu
CGI2
2021 Spatial Audio Object Coding Based on Time-Frequency Shifting and Scheduling
abstract
Spatial audio object coding (SAOC) is an effective method to transmit multiple audio objects. It divides the full frequency band into 28 subbands and extracts spatial parameters for de-coding. In this way, objects can be encoded into a downmix signal with a few parameters. However, using the same parameters in one subband will cause frequency aliasing distortion, which severely impacts the listening experience. Existing studies to enhance SAOC cannot eliminate the aliasing distortion of all objects effectively. This paper describes a new structure to balance the bit-rate and decode quality based on time-frequency (TF) shifting and scheduling. In this structure, a TF shifting strategy (contains global shifting and local shifting) is proposed to reduce frequency aliasing distortion. Furthermore, a scheduling strategy is used to decide which part should be shifted according to the degree of aliasing. From the experiment results, the performance of the proposed method is better than SAOC and other enhanced methods.
Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Yulin Wu 0003
ICME2
2021 Efficient Multi-Step Audio Object Coding with Limited Residual Information
abstract
Spatial audio object coding (SAOC) is an effective method to transmit multiple audio objects. Audio systems can provide personalized services under this framework. However, this method causes frequency aliasing distortion, which severely impacts the listening experience. The multi-step SAOC (MS-SAOC) scheme was proposed to enhance the sound quality of each audio object by using residual information. Compared with SAOC, the bit-rate increases three times due to the residual data of multiple objects. In this paper, an efficient multi-step residual coding method is proposed to reduce the residual bit-rate of MS-SAOC. A two-level filter is designed to remove redundant residual information, and the limited residual information can efficiently compensate for frequency aliasing distortion. From experiment results, the residual bit-rate is half of MS-SAOC, and the sound quality is maintained at the Good-Excellent level.
Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Yulin Wu 0003, Wenke Liu
ICME2
2021 Person Retrieval in Physical World
abstract
Person re-identification (re-ID) gains plenty of achievements as a retrieval problem in constrained camera networks. However, most of the researches are concentrated on visual appearance, they still suffer from the complicated environments in unconstrained urban/campus surveillance scenario due to unreliable visual representations with extremely challenging problems as lack of training samples, amounts of irrelevant crowds, etc. Besides, most of the existing person re-ID datasets neglect the physical truth of realistic investigation application: 1) investigators search only few suspects among amounts of crowds. Moreover, he may not go through by every camera in the surveillance area and may appear in the same camera several times; and 2) the corresponding characteristic in multi-space of the same ID can be verified with each other. Therefore, we propose a person retrieval in physical world (PRPW) dataset with large-scale unconstrained surveillance scenario. It contains over 1.4 million bounding boxes, including 20 labeled IDs and numerous irrelevant crowds captured by 86 cameras. Furthermore, over 30,000 records of 20 mobile trajectories are collected in this dataset, and the 20 mobile trajectories are partially overlapped while passing by 86 cameras. Finally, based on two common senses and a verification experiment, we provide a proposal to tackle with PRPW task on the basis of trajectory association which utilizes global optimization to compensate for the errors caused by visual expression on local observation points. The comparison experiments with two typical unsupervised person re-ID methods are implemented on the constructed dataset.
Wenxin Huang, Ruimin Hu, Chao Liang 0001, Xian Zhong
ICME3
2021 Low Bitrates Audio Object Coding Using Convolutional Auto-Encoder and Densenet Mixture Model
abstract
The efficient transmission of the audio objects can be achieved by spatial audio object coding (SAOC) method that conveys a mono downmix signal together with side information parameters that enable object reconstruction in the decoder. To allow the transmission of audio objects at low bitrates, we present a new audio coding method with convolutional auto-encoder (CAE) and dense convolutional network (DenseNet) mixture model, optimizing the compression of side information parameters of audio objects. It has two main advantages: 1) Different from the linear transform methods, CAE can dig the nonlinear relationship of side information parameters and can effectively reduce the dimension of side information parameters; 2) DenseNet is adding in the decoder to make full use of low dimensional features of side information parameters, which improves the audio quality at low bitrate. Experiments show that our method outperforms base-line methods permitting bitrates as low as 1 kbps per object.
Yulin Wu 0003, Ruimin Hu, Chenhao Hu, Shanfa Ke, Xiaochen Wang 0001
ICME2
2021 Location Predicts You: Location Prediction via Bi-direction Speculation and Dual-level Association
abstract
Location prediction is of great importance in location-based applications for the construction of the smart city. To our knowledge, existing models for location prediction focus on the users' preference on POIs from the perspective of the human side. However, modeling users' interests from the historical trajectory is still limited by the data sparsity. Additionally, most of existing methods predict the next location according to the individual data independently. But the data sparsity makes it difficult to mine explicit mobility patterns or capture the casual behavior for each user. To address the issues above, we propose a novel Bi-direction Speculation and Dual-level Association method (BSDA), which considers both users' interests in POIs and POIs' appeal to users. Furthermore, we develop the cross-user and cross-POI association to alleviate the data sparsity by similar users and POIs to enrich the candidates. Experimental results on two public datasets demonstrate that BSDA achieves significant improvements over state-of-the-art methods.
Ruimin Hu, Zheng Wang 0007, Toshihiko Yamasaki
IJCAI2
2021 Information Reuse Attention in Convolutional Neural Networks for Facial Expression Recognition in the Wild
abstract
Unlike the constraint frontal face condition, faces in the wild have various unconstrained interference factors, such as pose variations, illumination variations and occlusion. Because of this, facial expressions recognition (FER) in the wild is a challenging task and existing methods fail to performant well. However, for occluded faces (containing occlusion caused by other objects and self-occlusion caused by head posture changes), the attention mechanism has the ability to focus on the non-occluded regions automatically. In this paper, we propose an Information Reuse Attention Module (IRAM) for Convolutional Neural Network (CNN) to extract attention-aware features from faces. Our module reduces decay information in the process of generating attention maps by reusing the information of the previous layer and not reducing the dimensionality. Sequentially, we adaptively refine the feature responses by fusing the attention maps with the feature map. The proposed method is evaluated with two in-the-wild facial expression datasets RAF-DB and FER2013 and also compared with other state-of-the-art methods.
Ruimin Hu
IJCNN2
2021 Multi-level Graph Attention Network based Unsupervised Network Alignment
abstract
Network alignment is the matching of two networks with corresponding nodes that belong to the same user or entity. The most common application is to analyze which accounts belong to the same user in two social networks. Most of existing techniques rely on matrix factorization so that they cannot be scaled to large-scale networks, are constrained by strict constraints, and cannot learn node embedding without a training set. In this paper, we propose an unsupervised network alignment model based on multi-level graph attention networks. The model uses multi-level graph attention network to learn the embedded representation of nodes, satisfying attribute and structure constraints of alignment. Augmented learning process is proposed to simulate attribute noise and structural noise to improve adaptability of the model. Extensive experiments on real datasets show that the proposed model performs better than the state-of-the-art network alignment model. We also demonstrate the robustness of the proposed model.
Yilin Xiao 0002, Ruimin Hu, Dengshi Li, Junhang Wu, Yu Zhen, Lingfei Ren
LCN2
2021 Trajectory is not Enough: Hidden Following Detection
abstract
In outdoor crimes such as robbery and kidnapping, suspects generally secretly follow their victims in public places and then look for opportunities to commit crimes. Video anomaly detection (VAD) has achieved fruitful results through deep neural networks (DNN). However, as an abnormal behavior without obvious abnormal physical features, hidden following is highly similar to ordinary walking and accompanying behaviors, so it is difficult to effectively detect hidden dangerous followers using video anomaly detection methods or traditional trajectory analysis methods. We propose "hidden follower'' detection (HFD) task and a HFD model based on gaze pattern extraction. It extracts gaze pattern features of pedestrians from gaze-interval-series and introduces a time series classification model to classify pedestrians with or without hidden following purposes. Based on this model, we propose a hidden follower detection framework (HFDF) to detect hidden followers from normal pedestrians, which utilizes the trajectories and gaze patterns extracted from videos. To cope with the lack of test data, we construct a dataset of 1200 pedestrians from the crowd simulation model to simulate scenes including hidden followers, and we also collected a surveillance video dataset including the hidden following behaviors. The experiments conducted on these two datasets show that HFDF can consistently outperform the state-of-the-art method by a notable margin in the HFD task on the commonly-used F1 benchmark.
Danni Xu, Ruimin Hu, Zixiang Xiong, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li
ACM Multimedia2
2021 Unsupervised Temporal Attention Summarization Model for User Created Videos
Ruimin Hu, Rui Sheng
MMM (1)2
2021 Stacked Sparse Autoencoder for Audio Object Coding
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Chenhao Hu
MMM (1)2
2021 Unsupervised Person Re-identification via Diversity and Salience Clustering
abstract
Obtaining large-scale annotations for person re-identification tasks is very difficult and hard to deploy in real scenarios. Existing unsupervised person re-identification methods focus on one-shot learning and domain adaptation, but they still rely on the labels of the source domain or a small number of sample labels. To solve this problem, we formulate a very simple but effective unsupervised person re-identification method: Diversity and Salience Clustering (DSC). The method generates a stable clustering feature space for unsupervised re-identification by considering the diversity and salience of pedestrian samples. As a pure unsupervised person re-identification model, our approach does not use any camera annotations, pedestrian labels and source domain data. Extensive experiments on two image datasets demonstrate that our proposed method outperforms the state-of-the-art unsupervised person re-identification methods.
Xiaoyi Long, Ruimin Hu
SMC2
2021 Intelligent cloud computing platform for three-dimensional sound reproduction
abstract
Summary Three‐dimensional (3D) audio reproduction techniques reproduce realistic sound sources and spatial perception for listeners. However, it is a rather difficult work to configure all the parameters in the previous 3D reproduction systems since ordinary users hardly possess professional acoustic knowledge. In this article, we developed a sound source distance reproduction model and design an online parameter computing framework for users to generate reproduction parameters based on a cloud computing platform. The proposed model reproduces sound direction and distance cues under specific conditions. Subjective experiments are executed to assess the spatial reproduction performance in a real environment and objective experiments are performed in simulated scenarios. Both the subjective and objective experiments show that the proposed method improves the spatial perception of sound events. Since the reproduction model depends on environment and loudspeaker array configuration, a cloud parameter generation framework is developed for users to generate acoustic parameters for the reproduction system.
Maosheng Zhang, Ruimin Hu, Jiang Lin, Xiaochen Wang 0001
Concurr. Comput. Pract. Exp.2
2021 Silicone mask face anti-spoofing detection based on visual saliency and facial motion
Guangcheng Wang, Zhongyuan Wang 0001, Kui Jiang, Baojin Huang, Zheng He 0001, Ruimin Hu
Neurocomputing6
2021 Audio object coding based on N-step residual compensating
Chenhao Hu, Xiaochen Wang 0001, Ruimin Hu, Yulin Wu 0003
Multim. Tools Appl.3
2021 Optimization of sound fields reproduction based Higher-Order Ambisonics (HOA) using the Generative Adversarial Network (GAN)
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu
Multim. Tools Appl.3
2021 Estimation of spherical harmonic coefficients in sound field recording using feed-forward neural networks
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu
Multim. Tools Appl.3
2021 Occluded suspect search via channel-guided mechanism
Wenxin Huang, Ruimin Hu, Xiao Wang 0029, Chao Liang 0001, Jun Chen 0001
Neural Comput. Appl.2
2021 Trajectory Association for Person Re-identification
Ruimin Hu, Wenxin Huang, Dengshi Li, Xiaochen Wang 0001, Chenhao Hu
Neural Process. Lett.2
2021 Rethinking data collection for person re-identification: active redundancy reduction
Xin Xu 0007, Xiaolong Zhang 0002, Weili Guan, Ruimin Hu
Pattern Recognit.5
2021 Intelligibility Enhancement Via Normal-to-Lombard Speech Conversion With Long Short-Term Memory Network and Bayesian Gaussian Mixture Model
abstract
Speech communications and interactions frequently occur in a variety of environments. Noise in the environment significantly degrades speech intelligibility when speaking and listening. Especially in the listening stage, even if the multimedia terminal outputs clean speech, it is still difficult for listeners to obtain information. Intelligibility enhancement (IENH) of speech is a technique for overcoming the environmental noise in the listening stage. It implements a perceptual enhancement of non-noisy speech. This study focuses on IENH via normal-to-Lombard speech conversion, inspired by a well known acoustic mechanism named the Lombard effect. Our method combines the long short-term memory (LSTM) network and Bayesian Gaussian mixture model (BGMM) to build a conversion architecture. Compared with baselines, it has three main advantages: 1) an LSTM network is used for spectral tilt mapping with fully considering short-term correlations and high-dimensional expression abilities; 2) the aperiodicity (AP) is mapped together with the fundamental frequency ($F_0$) by a BGMM, which considers their relevance constraints and the importance of APs; 3) the gender-dependent mapping is used for$F_0$and APs to consider distribution differences between genders. Experiments indicate that our method gets better performance in both objective and subjective tests.
Xiaochen Wang 0001, Ruimin Hu, Huyin Zhang, Shanfa Ke
IEEE Trans. Multim.3
2021 Exploring Image Enhancement for Salient Object Detection in Low Light Images
abstract
Low light images captured in a non-uniform illumination environment usually are degraded with the scene depth and the corresponding environment lights. This degradation results in severe object information loss in the degraded image modality, which makes the salient object detection more challenging due to low contrast property and artificial light influence. However, existing salient object detection models are developed based on the assumption that the images are captured under a sufficient brightness environment, which is impractical in real-world scenarios. In this work, we propose an image enhancement approach to facilitate the salient object detection in low light images. The proposed model directly embeds the physical lighting model into the deep neural network to describe the degradation of low light images, in which the environment light is treated as a point-wise variate and changes with local content. Moreover, a Non-Local-Block Layer is utilized to capture the difference of local content of an object against its local neighborhood favoring regions. To quantitative evaluation, we construct a low light Images dataset with pixel-level human-labeled ground-truth annotations and report promising results on four public datasets and our benchmark dataset.
Xin Xu 0007, Shiqin Wang, Zheng Wang 0007, Xiaolong Zhang 0002, Ruimin Hu
ACM Trans. Multim. Comput. Commun. Appl.5
2020 Social-IFD: Personalized Influential Friends Discovery Based on Semantics in LBSN
abstract
Social influence is a hot topic in social network research, and this paper focuses on how to search for the most influential friends for a target user. The key point is to measure the influence between different users, such as adjacent users and the non-adjacent users. However, the traditional method, called the IS model, can only calculate the influence strength between neighboring users based on the inner-product of the Influence vector and the Susceptibility vector. In this paper, the social-IFD algorithm is proposed to compute the influence between different users (not only neighboring users but also non-adjacent users) based on network structure and the semantic information of users in LBSN, which has promoted the development of the IS model. Furthermore, we propose social-IFD ++ algorithm based on dynamic program to reduce the complexity of the social-IFD algorithm. Experiment results on two real large-scale network show that the average precision of the proposed social-IFD algorithm is 30.5% higher than the average precision of the IS model. In addition, the CPU running time of the proposed social-IFD ++ algorithm is nearly ten times lower than that of the IS model. It indicates that the proposed two algorithms have superior performance.
Ruimin Hu, Dengshi Li
ICC2
2020 Learning To See Faces In The Dark
abstract
Recovering details from dark images has received increasing attention due to its potential in applications such as video surveillance. We propose the first approach to detect and enhance human faces in extreme low-light images. Our method consists of two stages: a novel Face Location Network (FLNet) to locate the face, followed by a Face Enhancement Network (FE-Net) that uses concatenated sub-modules to progressively recover the face from coarse to fine grained details. Specifically, our enhancement modules exploit the semantic priors of facial landmarks to facilitate face recovery. Extensive experiments show our method is quantitatively and qualitatively superior to the state-of-the-art in terms of enhancement quality and face recognition. We have also collected a real-world dataset to support relevant research. All code and data will be shared for reproducing our experiments.
Ruimin Hu
ICME2
2020 Speech Intelligibility Enhancement Using Non-Parallel Speaking Style Conversion With Stargan And Dynamic Range Compression
abstract
Speech intelligibility enhancement is a perceptual enhancement technique for clean speech reproduced in noisy environments. It is typically used in the listening stage of multimedia communications. In this study, we enhance speech intelligibility by speaking style conversion (SSC), which is a data-driven approach inspired by a vocal mechanism named Lombard effect. The proposed SSC method combines star generative adversarial network (StarGAN) based mapping and dynamic range compression (DRC). It has two main advantages: 1) different from gender-independent conversion in previous studies, StarGAN can separately learn speech features of different genders to provide a differential conversion among genders with a single model and non-parallel training data; 2) we design a multi-level enhancement strategy with the use of DRC in the StarGAN architecture, which improves the SSC performance in strong noise interference. Experiments show that our method outperforms baseline methods.
Ruimin Hu, Shanfa Ke, Xiaochen Wang 0001
ICME2
2020 Normal-To-Lombard Speech Conversion by LSTM Network and BGMM for Intelligibility Enhancement of Telephone Speech
abstract
Noise in the environment significantly decreases the speech intelligibility of telephone conversations. Despite clean speech output from the device, the listener is still hard to get information. This study focuses on intelligibility enhancement (IENH) of telephone speech in near-end background noise based on normal-to-Lombard speech conversion. The proposed approach uses long short-term memory (LSTM) and Bayesian Gaussian mixture model (BGMM) to build the speech mapping model. Compared with previous studies, we fully consider the short-term correlations of speech and implement feature mappings with higher dimensional features and more types of features. Evaluations indicate that the proposed approach has achieved better results in both objective and subjective evaluation.
Xiaochen Wang 0001, Ruimin Hu, Huyin Zhang, Shanfa Ke
ICME3
2020 Tell The Truth From The Front: Anti-Disguise Vehicle Re-Identification
abstract
Recent efforts have been increasingly made on vehicle reidentification (re-ID), which has huge contributions to intelligent transportation and criminal investigation. However, most existing methods heavily rely on the color and texture features of vehicles to discern their identities, which turn invalid under adversarial social security occasions where vehicles' color and style are always tampered or forged by crime suspects. In this paper, we propose a local feature preservation method to learn the structure-aware features from the position distribution of individual local regions within vehicle front window area, which appears more robust and discriminative upon disguise. We further develop a two-branch deep convolutional network framework to integrate the structure-aware features with vehicle model features for vehicle Re-ID. The experimental results on datasets VehicleID and Vehicle-1M show that our end-to-end framework achieves promising performance and outperforms the state-of-the-art methods proposed so far.
Wenqian Zhu, Ruimin Hu, Zhongyuan Wang 0001, Dengshi Li, Xiyue Gao
ICME2
2020 HRTF Representation with Convolutional Auto-encoder
Wei Chen 0143, Ruimin Hu, Xiaochen Wang 0001, Dengshi Li
MMM (1)2
2020 Perceptual Localization of Virtual Sound Source Based on Loudspeaker Triplet
Duanzheng Guan, Dengshi Li, Xuebei Cai, Xiaochen Wang 0001, Ruimin Hu
MMM (2)5
2020 Multi-step Coding Structure of Spatial Audio Object Coding
Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Tingzhao Wu, Dengshi Li
MMM (1)2
2020 HMM-Based Person Re-identification in Large-Scale Open Scenario
Ruimin Hu, Wenxin Huang, Xiaochen Wang 0001, Dengshi Li
MMM (1)2
2020 Loudspeaker triplet selection based on low distortion within head for multichannel conversion of smart 3D home theater
abstract
Summary In recent years, with the vigorous development of 3D film industry, the demand for 3D Smart Home Theater, based on Internet of Things (IoT), continues to grow. Theaters are populated with a large number of loudspeakers for more realistic 3D sound effects. However the number of loudspeakers in home is limited. Therefore, multichannel conversion is required to achieve theater 3D sound effects in home. Traditionally, the replaced loudspeaker signal of the original system is assigned to a “loudspeaker triplet” of the converted system. A large amount of subjective evaluations is necessary to judge the consistency of the replaced loudspeaker position with the “phantom source” positions reconstructed by loudspeaker triplets. In this study, after calculating the least‐squares errors of the reproduced sound field within a given region, we explore the constraint between the low distortion of the reproduced sound field and the loudspeaker triplet positions. Using this constraint rule, we present a new loudspeaker triplet selection criteria that can greatly reduce the number and time of subjective evaluations for selecting the optimal loudspeaker triplet. Simulation and subjective evaluation experiments indicate that the proposed selection method outperforms the traditional method, and that the proposed method can be successfully applied to multichannel conversion.
Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Weiping Tu
Concurr. Comput. Pract. Exp.2
2020 Three-dimensional sound reproduction in vehicle based on data mining technique
abstract
Summary Three‐dimensional (3D) audio reproduction technique is a hot topic since MPEG proposed a proposal for 3D audio standard in 2012. The present 3D reproduction method is not available in vehicle because the loudspeaker configuration does not meet the requirements of the methods. In this paper, we develop a regression model to reproduce sound pressure in vehicle using a data mining algorithm. A big sound pressure dataset is established based on acoustic theory. We analyze this dataset and propose a stepwise regression model training and testing on this dataset. Both the Mean Square Error and Root Mean Square Error are rather small, and all the residuals such as raw residuals, Pearson Residuals, Studentized residuals, and Standardize residuals are small as well. The experiments indicate that the proposed reproduction system accurately simulates the theoretical sound reproduction system.
Maosheng Zhang, Ruimin Hu
Concurr. Comput. Pract. Exp.2
2020 SIST: Online Scale-Adaptive Object tracking with Stepwise Insight
Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Jun Chen 0001, Ruimin Hu
Neurocomputing5
2020 Learning latent geometric consistency for 6D object pose estimation in heavily cluttered scenes
Qingnan Li, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001, Yu Chen 0021
J. Vis. Commun. Image Represent.2
2020 Single Channel multi-speaker speech Separation based on quantized ratio mask and residual network
Shanfa Ke, Ruimin Hu, Xiaochen Wang 0001, Tingzhao Wu, Zhongyuan Wang 0001
Multim. Tools Appl.2
2020 A mapping model of spectral tilt in normal-to-Lombard speech conversion for intelligibility enhancement
Ruimin Hu, Xiaochen Wang 0001
Multim. Tools Appl.2
2020 Modeling and Optimizing of the Multi-Layer Nearest Neighbor Network for Face Image Super-Resolution
abstract
In this paper, we propose a face super-resolution (FSR) method to handle the decreasing face recognition rate caused by low-quality images. To better model the input images, we build a nearest neighbor network (NNN) which consists of nodes and paths by introducing the second-layer nearest neighbors (SLNNs), where the paths of the network represent the distance between nodes. As the SLNN is trained in the high-resolution (HR) space and is exponentially supplementary to the traditional first-layer nearest neighbors (FLNNs), the neighbor inadequacy problem can be effectively solved by enriching the neighbor candidate set via NNN. Furthermore, we solve the NNN for the optimal weights of neighbors. Finally, we fuse the refined weights and neighbors for better reconstruction results. The effectiveness of this fusion strategy is validated by both quantitative and qualitative experimental results. The extensive experimental results on the public face datasets and real-world challenging low-resolution (LR) images demonstrate that the proposed method performs favorably against the state-of-the-art methods.
Liang Chen 0026, Jinshan Pan, Ruimin Hu, Zhen Han 0002, Chao Liang 0001, Yi Wu 0010
IEEE Trans. Circuits Syst. Video Technol.3
2020 Long-Term Background Redundancy Reduction for Earth Observatory Video Coding
abstract
Huge earth observatory video data (EOVD) and the limited transmission bandwidth from satellites to terrestrial devices pose serious challenges to compression efficiency for satellite video. In this work, we deeply explore long-term background redundancy caused by periodical satellite revisit, and make utilization of long-term background reference (LTBR) to design a high-efficiency coding method specific for EOVD. Firstly, we use data of Google Earth as prior background knowledge and construct LTBR from it. Then, two novel prediction methods are proposed to make full use of prior information of LTBR, including color and definition correction based inter prediction with LTBR (CDIP-LTBR), structure and texture constrained intra prediction with LTBR (STIP-LTBR). Lastly, the proposed two new prediction schemes are integrated into one unified coding framework along with HEVC to achieve improved coding performance, and an improved RDO method is designed for additional prediction modes selection problem in the framework to obtain a higher prediction efficiency. Extensive experiments on real-world EOVD show that the proposed coding scheme exhibit significant improvement over HEVC and H.264, in terms of BD-Rate, BD-PSNR and rate distortion comparison.
Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004, Shin'ichi Satoh 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Ensemble Super-Resolution With a Reference Dataset
abstract
By developing sophisticated image priors or designing deep(er) architectures, a variety of image super-resolution (SR) approaches have been proposed recently and achieved very promising performance. A natural question that arises is whether these methods can be reformulated into a unifying framework and whether this framework assists in SR reconstruction? In this paper, we present a simple but effective single image SR method based on ensemble learning, which can produce a better performance than that could be obtained from any of SR methods to be ensembled (or called component super-resolvers). Based on the assumption that better component super-resolver should have larger ensemble weight when performing SR reconstruction, we present a maximum a posteriori (MAP) estimation framework for the inference of optimal ensemble weights. Especially, we introduce a reference dataset, which is composed of high-resolution (HR) and low-resolution (LR) image pairs, to measure the SR abilities (prior knowledge) of different component super-resolvers. To obtain the optimal ensemble weights, we propose to incorporate the reconstruction constraint, which states that the degenerated HR estimation should be equal to the LR observation one, as well as the prior knowledge of ensemble weights into the MAP estimation framework. Moreover, the proposed optimization problem can be solved by an analytical solution. We study the performance of the proposed method by comparing with different competitive approaches, including four state-of-the-art nondeep learning-based methods, four latest deep learning-based methods, and one ensemble learning-based method, and prove its effectiveness and superiority on some general image datasets and face image datasets.
Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Suhua Tang, Ruimin Hu, Jiayi Ma 0001
IEEE Trans. Cybern.5
2020 Intraspectrum Discrimination and Interspectrum Correlation Analysis Deep Network for Multispectral Face Recognition
abstract
Multispectral images contain rich recognition information since the multispectral camera can reveal information that is not visible to the human eye or to the conventional RGB camera. Due to this characteristic of multispectral images, multispectral face recognition has attracted lots of research interest. Although some multispectral face recognition methods have been presented in the last decade, how to fully and effectively explore the intraspectrum discriminant information and the useful interspectrum correlation information in multispectral face images for recognition has not been well studied. To boost the performance of multispectral face recognition, we propose an intraspectrum discrimination and interspectrum correlation analysis deep network (IDICN) approach. Multiple spectra are divided into several spectrum-sets, with each containing a group of spectra within a small spectral range. The IDICN network contains a set of spectrum-set-specific deep convolutional neural networks attempting to extract spectrum-set-specific features, followed by a spectrum pooling layer, whose target is to select a group of spectra with favorable discriminative abilities adaptively. IDICN jointly learns the nonlinear representations of the selected spectra, such that the intraspectrum Fisher loss and the interspectrum discriminant correlation are minimized. Experiments on the well-known Hong Kong Polytechnic University, Carnegie Mellon University, and the University of Western Australia multispectral face datasets demonstrate the superior performance of the proposed approach over several state-of-the-art methods.
Fei Wu 0004, Xiaoyuan Jing, Xiwei Dong, Ruimin Hu, Dong Yue 0001, Lina Wang 0001, Yimu Ji 0001, Ruchuan Wang 0001, Guoliang Chen 0008
IEEE Trans. Cybern.4
2019 CISI-net: Explicit Latent Content Inference and Imitated Style Rendering for Image Inpainting
abstract
Convolutional neural networks (CNNs) have presented their potential in filling large missing areas with plausible contents. To address the blurriness issue commonly existing in the CNN-based inpainting, a typical approach is to conduct texture refinement on the initially completed images by replacing the neural patch in the predicted region using the closest one in the known region. However, such a processing might introduce undesired content change in the predicted region, especially when the desired content does not exist in the known region. To avoid generating such incorrect content, in this paper, we propose a content inference and style imitation network (CISI-net), which explicitly separate the image data into content code and style code. The content inference is realized by performing inference in the latent space to infer the content code of the corrupted images similar to the one from the original images. It can produce more detailed content than a similar inference procedure in the pixel domain, due to the dimensional distribution of content being lower than that of the entire image. On the other hand, the style code is used to represent the rendering of content, which will be consistent over the entire image. The style code is then integrated with the inferred content code to generate the complete image. Experiments on multiple datasets including structural and natural images demonstrate that our proposed approach out-performs the existing ones in terms of content accuracy as well as texture details.
Jing Xiao 0004, Qiegen Liu, Ruimin Hu
AAAI4
2019 Long Term Background Reference Based Satellite Video Coding
abstract
Video transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellite calls for higher coding efficiency. In this paper, we propose a high efficiency satellite video coding method based on long term background reference (LTBR) to eliminate redundancy caused by periodical revisit. Firstly, data of Google Earth is used to provide prior information for establishing LTBR. Then a novel intra prediction method guided by pixels' cluster information from LTBR is introduced. Experiments demonstrate that our method outperforms HEVC and H.264 , in terms of rate-distortion, BD-PSNR and BD-Rate performance.
Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004
ICASSP2
2019 Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge
abstract
The rapidly increasing surveillance video data has challenged the existing video coding standards. Even though knowledge based video coding scheme proposed for moving objects so far has achieved high efficiency, it does not take full advantages of local information and highly relies on the accuracy of pose parameter of the objects, thus leading to large prediction residuals. In this paper, a novel surveillance video coding utilizing 3D and 2D knowledge is proposed. On the one hand, we generate a knowledge based reference frame from 3D models of the objects and incorporate it into the block based coding framework to remove global redundancy while improve the robustness to pose errors. On the other hand, 2D knowledge in the form of visual appearances of the objects in the previously encoded frames is employed to rectify the knowledge based reference frame for local redundancy removal. Experimental results demonstrate the effectiveness of our proposed method against HEVC and the knowledge based coding method.
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001
ICASSP2
2019 Vehicle Pose Estimation Using Mask Matching
abstract
In this paper, we present a conceptually novel framework for vehicle pose estimation from given RGB images. Our approach extends Mask R-CNN by adding two branches for coarse viewpoint estimation and keypoint detection, in parallel with the existing branches for mask segmentation and 2D object detection in the training stage. Capitalizing on the estimated mask and the mask renderings from ShapeNet in the inference stage, we propose a mask optimization scheme to recover the vehicle poses from 2D-3D correspondences. Then, we enforce geometric constraint on these vehicle poses in a coarse-to-fine hybrid approach for robustness. Experimentally, our framework outperforms the state-of-the-art approaches on the very challenging PASCAL3D+ dataset.
Qingnan Li, Ruimin Hu, Yu Chen 0021
ICASSP2
2019 Cross-view Identical Part Area Alignment for Person Re-identification
abstract
Person re-identification aims to associate images captured by non-overlapping cameras. It is a challenging task because images are often in different conditions such as background clutter, illumination variation, viewpoint changes and different camera settings. Viewpoint changes and pose variations often cause body part self-occlusion and misalignment. To deal with the problem, local features from human body parts are extracted. However, with viewpoint changes, the body parts also rotate horizontally. It is inappropriate to extract feature from entire area of body parts directly because the visible surface of body parts would turn away if viewpoint changes. Comparing identical areas provides a new way to pay attention to the details of person images. In this paper, we propose a Rotation Invariant Network to find the identical areas in cross-view images to extract robust local features. Extensive experiment show the effectiveness of our method on public datasets including CUHK03, Market1501 and DukeMTMC.
Dongshu Xu, Jun Chen 0001, Chao Liang 0001, Zheng Wang 0007, Ruimin Hu
ICASSP5
2019 Kullback-Leibler Divergence Frequency Warping Scale for Acoustic Scene Classification Using Convolutional Neural Network
abstract
Most of current best performing Acoustic Scene Classification (ASC) systems utilize Mel scale spectrograms with Convolutional Neural Networks (CNNs). Mel scale is a common way to suit frequency warping of human ears, with strict decreasing frequency resolution on low to high frequency range. However, we find that significant frequency bins are located at mid to high frequency range for some acoustic scenes, such as travelling by bus, tram or train. In this paper, we show that a better frequency warping scale for ASC can be automatically learned from raw spectrograms, using Kullback-Leibler (KL) divergence scale. Our KL scale spectrograms with CNN method is evaluated on two public ASC datasets. The results show that we outperform the Mel scale method on both datasets. In addition, we also employ a Conditional Generative Adversarial Nets (Conditional-GAN) model for data augmentation, to prevent overfitting problem and allow further improvements on ASC.
Yuhong Yang 0001, Weiping Tu, Haojun Ai, Linjun Cai, Ruimin Hu
ICASSP6
2019 Multi-speakers Speech Separation Based on Modified Attractor Points Estimation and GMM Clustering
abstract
In this paper, a new attractor points estimation method for DANet algorithm used in single channel multi-speaker speech separation has been proposed. A prerequisite is that there must be separate segments of each source in the mixture. This condition is met in the actual situation because the source signal is not overlapping at any time. With this prerequisite, an isolated source segments extracted from the mixture is converted to the embedding space. With the embedding of isolated source segments, a more accurately attractor point for each source will be created, due to it does not contain components of other sources. In addition, a gaussian mixture model(GMM) clustering method instead of K-means clustering method were used at run time. The experiment demonstrated that the proposed method gets a better separation performance than state of the art method up to 1.04dB in SDR.
Shanfa Ke, Ruimin Hu, Tingzhao Wu, Xiaochen Wang 0001, Zhongyuan Wang 0001
ICME2
2019 Fast Incremental PageRank on Dynamic Networks
Zexing Zhan, Ruimin Hu, Xiyue Gao, Nian Huai
ICWE2
2019 Deep Structural Feature Learning: Re-Identification of simailar vehicles In Structure-Aware Map Space
abstract
Vehicle re-identification (re-ID) has received more attention in recent years as a significant work, making huge contribution to the intelligent video surveillance. The complex intra-class and inter-class variation of vehicle images bring huge challenges for vehicle re-ID, especially for the similar vehicle re-ID. In this paper we focus on an interesting and challenging problem, vehicle re-ID of the same/similar model. Previous works mainly focus on extracting global features using deep models, ignoring the individual loa-cal regions in vehicle front window, such as decorations and stickers attached to the windshield, that can be more discriminative for vehicle re-ID. Instead of directly embedding these regions to learn their features, we propose a Regional Structure-Aware model (RSA) to learn structure-aware cues with the position distribution of individual local regions in vehicle front window area, constructing a FW structural map space. In this map sapce, deep models are able to learn more robust and discriminative spatial structure-aware features to improve the performance for vehicle re-ID of the same/similar model. We evaluate our method on a large-scale vehicle re-ID dataset Vehicle-1M. The experimental results show that our method can achieve promising performance and outperforms several recent state-of-the-art approaches.
Wenqian Zhu, Ruimin Hu, Zhongyuan Wang 0001, Dengshi Li, Xiyue Gao
MMAsia2
2019 Action Recognition Using Visual Attention with Reinforcement Learning
Jun Chen 0001, Ruimin Hu, Zengmin Xu
MMM (2)3
2019 Spectral Tilt Estimation for Speech Intelligibility Enhancement Using RNN Based on All-Pole Model
Ruimin Hu, Xiaochen Wang 0001
MMM (2)2
2019 Three-dimensional sound reproduction in vehicle based on data mining technique
abstract
Summary Three‐dimensional (3D) audio reproduction technique is a hot topic since MPEG proposed a proposal for 3D audio standard in 2012. The present 3D reproduction method is not available in vehicle because the loudspeaker configuration does not meet the requirements of the methods. In this paper, we develop a regression model to reproduce sound pressure in vehicle using a data mining algorithm. A big sound pressure dataset is established based on acoustic theory. We analyze this dataset and propose a stepwise regression model training and testing on this dataset. Both the Mean Square Error and Root Mean Square Error are rather small, and all the residuals such as raw residuals, Pearson Residuals, Studentized residuals, and Standardize residuals are small as well. The experiments indicate that the proposed reproduction system accurately simulates the theoretical sound reproduction system.
Maosheng Zhang, Ruimin Hu
Concurr. Comput. Pract. Exp.2
2019 Person re-identification with multiple similarity probabilities using deep metric learning for efficient smart security applications
Mingfu Xiong, Dan Chen 0001, Jun Chen 0001, Jingying Chen 0001, Benyun Shi, Chao Liang 0001, Ruimin Hu
J. Parallel Distributed Comput.7
2019 Multisource surveillance video coding with synthetic reference frame
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001
J. Vis. Commun. Image Represent.2
2019 Multisource surveillance video data coding with hierarchical knowledge library
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Liang Xu 0010, Zhongyuan Wang 0001
Multim. Tools Appl.2
2019 A near-end listening enhancement system by RNN-based noise cancellation and speech modification
Ruimin Hu, Xiaochen Wang 0001
Multim. Tools Appl.2
2019 Audio object coding based on optimal parameter frequency resolution
Tingzhao Wu, Ruimin Hu, Xiaochen Wang 0001, Shanfa Ke
Multim. Tools Appl.2
2019 Multi-Correlation Filters With Triangle-Structure Constraints for Object Tracking
abstract
Correlation filters (CFs) have been extensively used in tracking tasks due to their high efficiency although most of them regard the tracked target as a whole and are minimally effective in handling partial occlusion. In this study, we incorporate a part-based strategy into the framework of CFs and propose a novel multipart correlation tracker with triangle-structure constraints. Specifically, we train multiple CFs for the global object and local parts, which are then jointly applied to obtain the correlation response of any candidate during tracking. The tracker is robust in handling partial occlusion because of the use of part-based representation. The remaining global representation can contribute reliable cues in cases wherein several local filters drift away in a specific scene. We further propose a triangle-structure model to measure the structural similarity of candidates. The model employs multiple triangles to determine the spatial relationship among parts and helps constrain the location of the target. Moreover, we introduce an effective part selection scheme based on energy and integrity, which is generally applicable to part-tracking models. Extensive experiments on two public benchmarks demonstrate the superiority of the proposed method over the state-of-the-art approaches.
Weijian Ruan, Jun Chen 0001, Yi Wu 0001, Jinqiao Wang, Chao Liang 0001, Ruimin Hu, Junjun Jiang
IEEE Trans. Multim.6
2019 Semisupervised Discriminant Multimanifold Analysis for Action Recognition
abstract
Although recent semisupervised approaches have proven their effectiveness when there are limited training data, they assume that the samples from different actions lie on a single data manifold in the feature space and try to uncover a common subspace for all samples. However, this assumption ignores the intraclass compactness and the interclass separability simultaneously. We believe that human actions should occupy multimanifold subspace and, therefore, model the samples of the same action as the same manifold and those of different actions as different manifolds. In order to obtain the optimum subspace projection matrix, the current approaches may be mathematically imprecise owe to the badly scaled matrix and improper convergence. To address these issues in unconstrained convex optimization, we introduce a nontrivial spectral projected gradient method and Karush-Kuhn-Tucker conditions without matrix inversion. Through maximizing the separability between different classes by using labeled data points and estimating the intrinsic geometric structure of the data distributions by exploring unlabeled data points, the proposed algorithm can learn global and local consistency and boost the recognition performance. Extensive experiments conducted on the realistic video data sets, including JHMDB, HMDB51, UCF50, and UCF101, have demonstrated that our algorithm outperforms the compared algorithms, including deep learning approach when there are only a few labeled samples.
Zengmin Xu, Ruimin Hu, Jun Chen 0001, Chen Chen 0001, Junjun Jiang, Jiaofen Li
IEEE Trans. Neural Networks Learn. Syst.2
2018 Video-Based Person Re-Identification via Self Paced Weighting
abstract
Person re-identification (re-id) is a fundamental technique to associate various person images, captured by differentsurveillance cameras, to the same person. Compared to the single image based person re-id methods, video-based personre-id has attracted widespread attentions because extra space-time information and more appearance cues that can beused to greatly improve the matching performance. However, most existing video-based person re-id methods equally treatall video frames, ignoring their quality discrepancy caused by object occlusion and motions, which is a common phenomenonin real surveillance scenario. Based on this finding, we propose a novel video-based person re-id method via self paced weighting (SPW). Firstly, we propose a self paced outlier detection method to evaluate the noise degree of video sub sequences. Thereafter, a weighted multi-pair distance metric learning approach is adopted to measure the distance of two person image sequences. Experimental results on two public datasets demonstrate the superiority of the proposed method over current state-of-the-art work.
Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Weijian Ruan, Ruimin Hu
AAAI6
2018 Edge-Aware Context Encoder for Image Inpainting
abstract
We present Edge-aware Context Encoder (E-CE): an image inpainting model which takes scene structure and context into account. Unlike previous CE which predicts the missing regions using context from entire image, E-CE learns to recover the texture according to edge structures, attempting to avoid context blending across boundaries. In our approach, edges are extracted from the masked image, and completed by a full-convolutional network. The completed edge map together with the original masked image are then input into the modified CE network to predict the missing region. The experiments demonstrate that E-CE can generate images with better shapes and structures than CE.
Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001
ICASSP2
2018 Faster Seam Carving for Video Retargeting
abstract
Video retargeting is to resize a video to a desired resolution or aspect ratio while preserving its salient content without visual distortion. The key to video retargeting is to reconcile spatio-temporal coherence of video frames, and most existing works use seam carving to achieve that by employing the dynamic programming to find optimal seams. However, these methods are too time-consuming due to high computational complexity of the dynamic programming. To this end, we propose a novel method which uses discontinuous and suboptimal seams for seam carving. Concretely, we obtain the discontinuous seams by allowing seams to move freely in homogeneous regions of the frame, which helps preserve the spatio-temporal coherence effectively. Then, the genetic algorithm is employed to find suboptimal seams, so as to reduce computational complexity. Finally, each frame can be retargeted to a new aspect ratio or size by repeatedly carving out seams. Compared to the-state-of-the-art methods, the proposed algorithm achieves comparable results at an average expense of only one third of their running time.
Ruimin Hu, Chao Liang 0001, Chunxia Xiao, Weijian Ruan
ICIP2
2018 Individualization of Head Related Transfer Functions Based on Radial Basis Function Neural Network
abstract
Head Related Transfer Functions (HRTFs) contain sound localization cues and are commonly used in 3D audio reproduction. Due to HRTFs are closely related to anthropometric parameters (head, pinna, torso), which means HRTFs vary with each individual, how to obtain a set of suitable HRTFs for each individual remains to be solved. In this paper, we investigated the complex relationship between HRTFs and anthropometric parameters through Radial Basis Function neural network (RBF), and proposed a method of generating individualized HRTFs with listener's anthropometric parameters. Objective experiments show that the estimated HRTFs have good consistency with measured ones, and the spectral distortion values have an average reduction of 0.59 dB compared with other methods. Subjective listening tests show that using estimated HRTFs enable accurate auditory localization.
Lian Meng, Xiaochen Wang 0001, Wei Chen 0143, Chunling Ai, Ruimin Hu
ICME5
2018 VCF: Velocity Correlation Filter, Towards Space-Borne Satellite Video Tracking
abstract
Tracking a moving target of interests from a space-borne satellite video is really a difficulty, since the target usually occupies only a few pixels in each frame of satellite video. Even it is a long train. Most state-of-the-art tracking algorithms mainly rely on luminance or color features, failing to handle this video tracking problem due to the extremely inadequate quality of targets features. To overcome this difficulty, we propose a velocity correlation filter (VCF), employing velocity feature and inertia mechanism to construct a kernel correlation filter for satellite video targets tracking. The velocity feature has a high discriminative ability to detect moving targets in satellite videos and the inertia mechanism can prevent model drift adaptively. Experimental results on three real satellite video datasets show our proposed approach outperforms state-of-the-art tracking methods with a more than 100 frames per second.
Jia Shao, Bo Du 0001, Chen Wu 0003, Jia Wu 0001, Ruimin Hu, Xuelong Li 0001
ICME5
2018 TLR: Transfer Latent Representation for Unsupervised Domain Adaptation
abstract
Domain adaptation refers to the process of learning prediction models in a target domain by making use of data from a source domain. Many classic methods solve the domain adaptation problem by establishing a common latent space, which may cause the loss of many important properties across both domains. In this manuscript, we develop a novel method, transfer latent representation (TLR), to learn a better latent space. Specifically, we design an objective function based on a simple linear autoencoder to derive the latent representations of both domains. The encoder in the autoencoder aims to project the data of both domains into a robust latent space. Besides, the decoder imposes an additional constraint to reconstruct the original data, which can preserve the common properties of both domains and reduce the noise that causes domain shift. Experiments on cross-domain tasks demonstrate the advantages of TLR over competing methods.
Bo Du 0001, Jia Wu 0001, Lefei Zhang, Ruimin Hu, Xuelong Li 0001
ICME5
2018 A Novel Frontal Facial Synthesis Algorithm Based on Individual Residual Face
Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001
MMM (2)2
2018 Reinforcing Pedestrian Parsing on Small Scale Dataset
Jun Chen 0001, Junjun Jiang, Ruimin Hu
MMM (1)4
2018 Virtual Background Reference Frame Based Satellite Video Coding
abstract
Video transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellites calls for higher coding efficiency. In this paper, we propose a high efficiency satellite video compression method to eliminate long-term redundancy among multiple periodically revisited videos, based on virtual background reference frame (VBRF) obtained from Google Earth data. First, we make full use of Google Earth data to create VBRF for representing the constant ground background. Then, we encode all I frames by referring to VBRF and performing interprediction. Experiments demonstrate that our method outperforms H.264 and HEVC, in terms of rate distortion, bjøntegaard delta peak signal-to-noise rate (BD-PSNR), and BD-Rate performance.
Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004
IEEE Signal Process. Lett.2
2018 Person Reidentification via Discrepancy Matrix and Matrix Metric
abstract
Person reidentification (re-id), as an important task in video surveillance and forensics applications, has been widely studied. Previous research efforts toward solving the person re-id problem have primarily focused on constructing robust vector description by exploiting appearance's characteristic, or learning discriminative distance metric by labeled vectors. Based on the cognition and identification process of human, we propose a new pattern, which transforms the feature description from characteristic vector to discrepancy matrix. In particular, in order to well identify a person, it converts the distance metric from vector metric to matrix metric, which consists of the intradiscrepancy projection and interdiscrepancy projection parts. We introduce a consistent term and a discriminative term to form the objective function. To solve it efficiently, we utilize a simple gradient-descent method under the alternating optimization process with respect to the two projections. Experimental results on public datasets demonstrate the effectiveness of the proposed pattern as compared with the state-of-the-art approaches.
Zheng Wang 0007, Ruimin Hu, Chen Chen 0001, Yi Yu 0001, Junjun Jiang, Chao Liang 0001, Shin'ichi Satoh 0001
IEEE Trans. Cybern.2
2017 Multi-Kernel Low-Rank Dictionary Pair Learning for Multiple Features Based Image Classification
abstract
Dictionary learning (DL) is an effective feature learning technique, and has led to interesting results in many classification tasks. Recently, by combining DL with multiple kernel learning (which is a crucial and effective technique for combining different feature representation information), a few multi-kernel DL methods have been presented to solve the multiple feature representations based classification problem. However, how to improve the representation capability and discriminability of multi-kernel dictionary has not been well studied. In this paper, we propose a novel multi-kernel DL approach, named multi-kernel low-rank dictionary pair learning (MKLDPL). Specifically, MKLDPL jointly learns a kernel synthesis dictionary and a kernel analysis dictionary by exploiting the class label information. The learned synthesis and analysis dictionaries work together to implement the coding and reconstruction of samples in the kernel space. To enhance the discriminability of the learned multi-kernel dictionaries, MKLDPL imposes the low-rank regularization on the analysis dictionary, which can make samples from the same class have similar representations. We apply MKLDPL for multiple features based image classification task. Experimental results demonstrate the effectiveness of the proposed approach.
Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Di Wu 0014, Li Cheng 0006, Ruimin Hu
AAAI7
2017 Cruise UAV Video Compression Based on Long-Term Wide-Range Background
abstract
With the rapid development of Unmanned Aerial Vehicle (UAV), the compression of video data captured by UAV has become a growing critical issue. However, most advanced coding schemes, like H.264 and HEVC, are oriented for common videos and thus cannot afford ideal coding efficiency when applied to UAV platform. Considering the characteristics of UAV video, much more improvement could be imposed onto current coding schemes to make full use of UAV's sensor information. In this paper, we exploit long term redundancy existing in the video data captured by cruise UAV. Firstly, we establish a long-term wide-range background set for reference. Then we separate each frame into new-area part and overlapped part. Lastly, we use GPS information of each frame to get reference from background set and compress two parts individually. In the experiments, by comparing to standard HEVC, our method has given more than 20% reduction in bitrate and meanwhile more than 4% gain in PSNR.
Xu Wang 0015, Jing Xiao 0004, Ruimin Hu, Zhongyuan Wang 0001
DCC3
2017 Taichi distance for person re-identification
abstract
Metric learning is an important issue in person re-identification, and Mahalanobis-distance based metric learning methods prevail in this field. All of these approaches can be considered as equivalently projecting all samples to a new metric space and calculating the Euclidean distance there. However, the performance of distinguishing similar samples from dissimilar ones via absolute distance is limited. In this paper, we suggest using relative distance instead. We adopt a bi-target perspective. The core idea is to construct a virtual opposite target for each original target. Then, the similarity between a sample and the others is judged by using both the original and opposite targets of the sample. In this way, we propose a bi-target metric method, named TAICHI distance. Considering simplicity and efficiency, we follow the KISSME metric in this paper. Extensive evaluations on challenging datasets confirm the effectiveness of the proposed method.
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Chao Liang 0001, Chen Chen 0001
ICASSP2
2017 A joint learning based Face Super Resolution approach via contextual topological structure
abstract
Face Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch based methods are based on the consistency assumption that the neighbors in HR/LR space form similar local geometry. But when LR images are with low quality, the LR space is seriously contaminated that even two distinct patches look similar, which means that the consistency assumption is not well held anymore. To this end, in this paper we introduce the contextual topological structure of target patch to improve the consistency. The contextual topological structure consists of the target patch as well as its adjacent patches, we explore the relationship between them based on statistical probability and apply the relationship for joint learning progress of mapping from LR to HR. By incorporating the contextual topological structure, the robustness to noise of approach is increased as well as the LR/HR consistency. The effectiveness of proposed method is verified both quantitatively and qualitatively.
Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Qing Li 0001
ICASSP2
2017 Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction
abstract
NHK has proposed a famous 3D audio system: 22.2 multi-channel system, but its loudspeakers are too many and are troublesome to put in home. Ando and Wang has proposed two simplification methods to reduce its channel number, but only 3D sound field at the central listening point can be recovered well by NHK 22.2 system and its simplified systems, the listening experience at a non central listening point is worse than that at the central listening point. In real life, listeners may stay at arbitrary listening point: central or non central point. Conventional pressure matching and particle matching method could be used for non central zone sound field reproduction, but they have some theoretical shortcomings. To address these problems, this paper propose a universal non central listening point sound field reproduction method by matching sound physical property between a non central listening point and the central listening point. Subjective and objective experiments show the effectiveness of the proposed method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP2
2017 Transferring clothing parsing from fashion dataset to surveillance
abstract
In this paper we address the problem of automatic clothing parsing in surveillance video with the information from user-generated tags such as “jeans” and “T-shirt”. Although clothing parsing has achieved great success in fashion clothing, it is quite challenging to parse clothing in practical surveillance conditions due to complicated environmental interferences, such as illumination change, scale zooming, viewpoint variation and etc. Our method is developed to capture the clothing information from the fashion field and apply it to surveillance domain by weakly-supervised transfer learning. Most of attribute labels in surveillance images convey strong location information, which can be considered as weak labels to deal with the transfer method. Both quantitative and qualitative experiments conducted on practical surveillance datasets have shown the effectiveness of the proposed method.
Jun Chen 0001, Chao Liang 0001, Wenhua Fang, Xiaoyuan Jing, Ruimin Hu
ICASSP6
2017 Action recognition with gradient boundary convolutional network
abstract
Deep learning features for video action recognition are usually learned from RGB/gray images, image gradients, and optical flows. The single modality of the input data can describe one characteristic of the human action such as appearance structure or motion information. In this paper, we propose a high efficient gradient boundary convolutional network (ConvNet) to simultaneously learn spatio-temporal feature from the single modality data of gradient boundaries. The gradient boundaries represent both local spacial structure and motion information of action video. The gradient boundaries also have less background noise compared to RGB/gray images and image gradients. Extensive experiments are conducted on two popular and challenging action benchmarks, the UCF101 and the HMDB51 action datasets. The proposed deep gradient boundary feature achieves competitive performances on both benchmarks.
Jun Chen 0001, Chen Chen 0001, Ruimin Hu
ICIP4
2017 Multi-feature fusion based background subtraction for video sequences with strong background changes
abstract
Current background subtraction algorithms are sensitive to sudden changes. In this paper, we propose a multi-feature fusion scheme to background subtraction for video sequences with strong background changes. We reconstruct the whole videos frame by frame by fusing several video features. In this fusing step, we design an energy function based on enforcing every features with an equal weight. By comparing reconstruction videos with the original videos, pixels with small differences are classified as background pixels. Thus, we can identify background areas in advance and then we construct a contour-based mask combining mechanism. Experimental results conducted on the OTCBVS, BMC 2012 and PETS 2001 datasets show that our method improves the performance of the Zivkovic's GMM and SubSENSE for video sequences with strong background changes.
Zhenkun Huang, Ruimin Hu, Thierry Bouwmans
ICIP2
2017 Object tracking via online trajectory optimization with multi-feature fusion
abstract
The goal of object tracking is to estimate the state and trajectory of an interested target in a video sequence, thus both spatial and temporal information are of critical importance for tracking. However, most existing trackers usually determine targets just by the judgement like confidence from a single frame, which tend to treat tracking as a static detecting problem while neglecting the spatial-temporal relationship. In this paper, we propose a novel tracking method of online trajectory optimization with multi-feature fusion (TOFF). Considering the trajectory continuity, the accurate targets are determined by estimating the optimal short trajectories over all the video fragments, which can be obtained by the observation models that are iteratively updated based on the selected most reliable proposals. By employing the structured samples instead of binary-labeled samples, we construct a structured output model with representing descriptor of multi-feature fusion as the basic tracker. Extensive experiments on various challenging image sequences demonstrate the superiority of our method to several state-of-the-art methods.
Weijian Ruan, Jun Chen 0001, Chao Liang 0001, Yi Wu 0001, Ruimin Hu
ICME5
2017 Low-resolution pedestrian detection via a novel resolution-score discriminative surface
abstract
Pedestrian detection, as an important task in video surveillance and forensics applications, has been widely studied. However, its performance is unsatisfactory especially in the low resolution conditions. In realistic scenarios, the size of pedestrians in the images is often small, and detection can be challenging. To solve this problem, this paper proposes a novel resolution-score discriminative surface method to investigate the variation behaviors of detection scores under different pedestrian and non-pedestrian image resolutions. The discriminative surface consists of a series of positive and negative resolution-score lines, and each of them is a connected line to depict the variation relationship between pedestrian's detection scores under various image resolutions. On this basis, the resolution-score discriminative surface can classify a resolution-score line as a pedestrian or not according to whether it lies in the positive or the negative region. Experimental results on two public datasets and one campus surveillance dataset demonstrate the effectiveness of the proposed method.
Xiao Wang 0029, Jun Chen 0001, Chao Liang 0001, Chen Chen 0001, Zheng Wang 0007, Ruimin Hu
ICME6
2017 On Gleaning Knowledge from Multiple Domains for Active Learning
abstract
How can a doctor diagnose new diseases with little historical knowledge, which are emerging over time? Active learning is a promising way to address the problem by querying the most informative samples. Since the diagnosed cases for new disease are very limited, gleaning knowledge from other domains (classical prescriptions) to prevent the bias of active leaning would be vital for accurate diagnosis. In this paper, a framework that attempts to glean knowledge from multiple domains for active learning by querying the most uncertain and representative samples from the target domain and calculating the importance weights for re-weighting the source data in a single unified formulation is proposed. The weights are optimized by both a supervised classifier and distribution matching between the source domain and target domain with maximum mean discrepancy. Besides, a multiple domains active learning method is designed based on the proposed framework as an example. The proposed method is verified with newsgroups and handwritten digits data recognition tasks, where it outperforms the state-of-the-art methods.
Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Ruimin Hu, Dacheng Tao
IJCAI5
2017 Structural superpixel descriptor for visual tracking
abstract
Object representation is a major component in object tracking, however, most conventional patch-based methods just simply decompose the object into patches with grid or stochastic rectangles. This kind of decomposition ignores the intrinsic structure of object, leading to low discriminative power and weak representation effectiveness when similar objects appear or under background clutters. In this paper, we propose an effective object descriptor based on a hierarchical representation with superpixels for visual tracking, called Structural Superpixel Descriptor (SSD). The proposed SSD not only exploits the superpixels to capture the structural information of object, but also preserves the spatial layout structure among the superpixels inside each target candidate. Moreover, we propose an adaptive patch weighting method based on spatial constraint to alleviate various adverse impacts of background information, making the tracker more robust against background noises. We show that the proposed SSD makes full use of the intrinsic structure inside target candidates. Extensive experiments conducted on various challenging sequences demonstrate that the proposed tracker performs well against state-of-the-art algorithms.
Ruimin Hu, Chao Liang 0001, Weijian Ruan, Bo Luo
IJCNN2
2017 Statistical Inference of Gaussian-Laplace Distribution for Person Verification
abstract
Metric learning is an important issue in the person verification problem, which is to identify whether a pair of face or human body images is about the same person. Due to low running cost, the non-iterative statistical inference methods for metric learning show their efficiency and effectiveness to large scale datasets and on-line updating person verification applications. The KISSME method is a typical one that constructs the metric based on two assumptions that both of the discrepancy spaces of negative pairs and positive pairs should be Gaussian structures. However, we find that, in fact, the distribution of discrepancies of positive pairs might tend to the Laplace distribution rather than the Gaussian distribution. Based on this finding, we propose a metric learning method by exploiting Gaussian-Laplace distribution statistical inference, where the Gaussian distribution of negative discrepancies and the Laplace distribution of positive discrepancies are considered together. Experiments conducted on two human body datasets (VIPeR and Market-1501) and one face dataset (LFW) show its superiority in terms of effectiveness and efficiency as compared with the state-of-the-art approaches, no matter the appearance description is handcrafted or deep learned.
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Junjun Jiang, Jiayi Ma 0001, Shin'ichi Satoh 0001
ACM Multimedia2
2017 The Perceptual Lossless Quantization of Spatial Parameter for 3D Audio Signals
Xiaochen Wang 0001, Ruimin Hu, Dengshi Li
MMM (2)4
2017 3D Sound Field Reproduction at Non Central Point for NHK 22.2 System
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
MMM (1)2
2017 Efficient Pedestrian Detection in the Low Resolution via Sparse Representation with Sparse Support Regression
Wenhua Fang, Jun Chen 0001, Ruimin Hu
PAKDD (2)3
2017 Action recognition by saliency-based dense sampling
Zengmin Xu, Ruimin Hu, Jun Chen 0001, Chen Chen 0001, Qingquan Sun
Neurocomputing2
2017 Comprehensive Association Rules Mining of Health Examination Data with an Extended FP-Growth Method
Bowei Wang, Dan Chen 0001, Benyun Shi, Yifu Duan, Jingying Chen 0001, Ruimin Hu
Mob. Networks Appl.7
2017 Face super resolution based on parent patch prior for VLQ scenarios
Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Qing Li 0001, Zheng Lu 0002
Multim. Tools Appl.2
2017 A novel face super resolution approach for noisy images using contour feature and standard deviation prior
Liang Chen 0026, Ruimin Hu, Chao Liang 0001, Qing Li 0001, Zhen Han 0002
Multim. Tools Appl.2
2017 HRM graph constrained dictionary learning for face image super-resolution
Kebin Huang, Ruimin Hu, Junjun Jiang, Zhen Han 0002
Multim. Tools Appl.2
2017 A sensitive object-oriented approach to big surveillance data compression for social security applications in smart cities
abstract
Summary Surveillance has become a fairly common practice with the global boom in “smart cities”. How to efficiently store and manage the vast quantities of surveillance data is a persistent challenge in terms of analyzing social security problems. Developing data compression technology under the analytic requirements of surveillance data is the key to solving the storage problem. Criminal investigation demands the quality preservation of sensitive objects, typically pedestrians, human faces, vehicles, and license plates; however, the analytical value of surveillance data is rapidly lost as the compression ratio increases. In this paper, we propose a sensitive object‐oriented regions of interest‐based coding strategy for preserving the analytical value of surveillance data. In the proposed method, instead of generating a saliency map based on human visual perception, we consider saliency as a set of characteristics important for object detection and recognition. By making this modification, almost all sensitive objects necessary in a criminal investigation are assigned high saliency value rather than only one or two salient regions. Motions in the temporal domain are integrated to place emphasis on moving objects, namely moving sensitive objects, which then gain the highest saliency. Finally, a saliency‐based rate control algorithm embedded in High Efficiency Video Coding is used to maintain the quality of sensitive objects in the encoded video under a fixed bitrate. Experiments were conducted on two analytical indexes: Feature similarity and object detection accuracy. The results showed that by achieving the same feature similarity and object detection accuracy, our method can save 20% and 40% bitrate over High Efficiency Video Coding, respectively, for the storage of big surveillance data. Copyright © 2016 John Wiley & Sons, Ltd.
Jing Xiao 0004, Zhongyuan Wang 0001, Yu Chen 0021, Jun Xiao 0004, Gen Zhan, Ruimin Hu
Softw. Pract. Exp.7
2017 Super-Resolution Person Re-Identification With Semi-Coupled Low-Rank Discriminant Dictionary Learning
abstract
Person re-identification has been widely studied due to its importance in surveillance and forensics applications. In practice, gallery images are high resolution (HR), while probe images are usually low resolution (LR) in the identification scenarios with large variation of illumination, weather, or quality of cameras. Person re-identification in this kind of scenarios, which we call super-resolution (SR) person re-identification, has not been well studied. In this paper, we propose a semi-coupled low-rank discriminant dictionary learning (SLD2L) approach for SR person re-identification task. With the HR and LR dictionary pair and mapping matrices learned from the features of HR and LR training images, SLD2L can convert the features of the LR probe images into HR features. To ensure that the converted features have favorable discriminative capability and the learned dictionaries can well characterize intrinsic feature spaces of the HR and LR images, we design a discriminant term and a low-rank regularization term for SLD2L. Moreover, considering that low resolution results in different degrees of loss for different types of visual appearance features, we propose a multi-view SLD2L (MVSLD2L) approach, which can learn the type-specific dictionary pair and mappings for each type of feature. Experimental results on multiple publicly available data sets demonstrate the effectiveness of our proposed approaches for the SR person re-identification task.
Xiaoyuan Jing, Xiaoke Zhu, Fei Wu 0004, Ruimin Hu, Xinge You, Yunhong Wang 0001, Jing-Yu Yang 0001
IEEE Trans. Image Process.4
2017 PLTD: Patch-Based Low-Rank Tensor Decomposition for Hyperspectral Images
abstract
Recent years has witnessed growing interest in hyperspectral image (HSI) processing. In practice, however, HSIs always suffer from huge data size and mass of redundant information, which hinder their application in many cases. HSI compression is a straightforward way of relieving these problems. However, most of the conventional image encoding algorithms mainly focus on the spatial dimensions, and they need not consider the redundancy in the spectral dimension. In this paper, we propose a novel HSI compression and reconstruction algorithm via patch-based low-rank tensor decomposition (PLTD). Instead of processing the HSI separately by spectral channel or by pixel, we represent each local patch of the HSI as a third-order tensor. Then, the similar tensor patches are grouped by clustering to form a fourth-order tensor per cluster. Since the grouped tensor is assumed to be redundant, each cluster can be approximately decomposed to a coefficient tensor and three dictionary matrices, which leads to a low-rank tensor representation of both the spatial and spectral modes. The reconstructed HSI can then be simply obtained by the product of the coefficient tensor and dictionary matrices per cluster. In this way, the proposed PLTD algorithm simultaneously removes the redundancy in both the spatial and spectral domains in a unified framework. The extensive experimental results on various public HSI datasets demonstrate that the proposed method outperforms the traditional image compression approaches and other tensor-based methods.
Bo Du 0001, Mengfei Zhang, Lefei Zhang, Ruimin Hu, Dacheng Tao
IEEE Trans. Multim.4
2017 SRLSP: A Face Image Super-Resolution Algorithm Using Smooth Regression With Local Structure Prior
abstract
The performance of traditional face recognition systems is sharply reduced when encountered with a low-resolution (LR) probe face image. To obtain much more detailed facial features, some face super-resolution (SR) methods have been proposed in the past decade. The basic idea of a face image SR is to generate a high-resolution (HR) face image from an LR one with the help of a set of training examples. It aims at transcending the limitations of optical imaging systems. In this paper, we regard face image SR as an image interpolation problem for domain-specific images. A missing intensity interpolation method based on smooth regression with a local structure prior (LSP), named SRLSP for short, is presented. In order to interpolate the missing intensities in a target HR image, we assume that face image patches at the same position share similar local structures, and use smooth regression to learn the relationship between LR pixels and missing HR pixels of one position patch. Performance comparison with the state-of-the-art SR algorithms on two public face databases and some real-world images shows the effectiveness of the proposed method for a face image SR in general. In addition, we conduct a face recognition experiment on the extended Yale-B face database based on the super-resolved HR faces. Experimental results clearly validate the advantages of our proposed SR method over the state-of-the-art SR methods in face recognition application.
Junjun Jiang, Chen Chen 0001, Jiayi Ma 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu
IEEE Trans. Multim.6
2016 Person Re-Identification via Multiple Coarse-to-Fine Deep Metrics
abstract
Person re-identification, aiming to identify images of the same person from various cameras views in different places, has attracted a lot of research interests in the field of artificial intelligence and multimedia. As one of its popular research directions, the metric learning method plays an important role for seeking a proper metric space to generate accurate feature comparison. However, the existing metric learning methods mainly aim to learn an optimal distance metric function through a single metric, making them difficult to consider multiple similar relationships between the samples. To solve this problem, this paper proposes a coarse-to-fine deep metric learning method equipped with multiple different Stacked Auto-Encoder (SAE) networks and classification networks. In the perspective of the human's visual mechanism, the multiple different levels of deep neural networks simulate the information processing of the brain's visual system, which employs different patterns to recognize the character of objects. In addition, a weighted assignment mechanism is presented to handle the different measure manners for final recognition accuracy. The experimental results conducted on two public datasets, i.e., VIPeR and CUHK have shown the prospective performance of the proposed method.
Mingfu Xiong, Jun Chen 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu, Chao Liang 0001, Daming Shi 0001
ECAI5
2016 Multiple instance discriminative dictionary learning for action recognition
abstract
Action recognition from video is a prominent research area in computer vision, with far-reaching applications. Current state-of-the-art action recognition methods is Fisher Vector (FV) coding model based on spatio-temporal local features. Though high dimensional local features have more representative, the high dimensions are challenge for the dictionary learning of FV model. This paper proposes a Multiple Instance Discriminative Dictionary Learning (MIDDL) method for action recognition. We introduce cross-validation method in multiple instance learning procedure, which prevents training from prematurely locking onto erroneous initial instances. In order to balance the positive instance number between positive bags, only the top ranked instances are labeled as positive in the step of iterative training classifiers. Taking these classifiers as discriminative visual words, we get the video global representation based on classifier response. The experimental results demonstrate the effectiveness of applying the learned discriminative classifiers as visual word on challenging action data sets, i.e. UCF50 and HMDB51.
Jun Chen 0001, Zengmin Xu, Ruimin Hu
ICASSP5
2016 Regularizing Deep Convolutional Neural Networks with a Structured Decorrelation Constraint
abstract
Deep convolutional networks have achieved successful performance in data mining field. However, training large networks still remains a challenge, as the training data may be insufficient and the model can easily get overfitted. Hence the training process is usually combined with a model regularization. Typical regularizers include weight decay, Dropout, etc. In this paper, we propose a novel regularizer, named Structured Decorrelation Constraint (SDC), which is applied to the activations of the hidden layers to prevent overfitting and achieve better generalization. SDC impels the network to learn structured representations by grouping the hidden units and encouraging the units within the same group to have strong connections during the training procedure. Meanwhile, it forces the units in different groups to learn non-redundant representations by minimizing the cross-covariance between them. Compared with Dropout, SDC reduces the co-adaptions between the hidden units in an explicit way. Besides, we propose a novel approach called Reg-Conv that can help SDC to regularize the complex convolutional layers. Experiments on extensive datasets show that SDC significantly reduces overfitting and yields very meaningful improvements on classification performance (on CIFAR-10 6.22% accuracy promotion and on CIFAR-100 9.63% promotion).
Wei Xiong 0008, Bo Du 0001, Lefei Zhang, Ruimin Hu, Dacheng Tao
ICDM4
2016 Multichannel reduction based on sound field within two ears
abstract
People hope to use a small number of loudspeakers to get the experience of the film 3D sound at home. Considering that people use two ears to listen, this paper provides a method which reproduce the sound field within the region of two ears. We develop the fundamental performance limits for the truncated spherical harmonic function expansions of the sound field within the region of ears. Based on this, the low distortion of reproduced sound field within two ears is maintained in the processing of reducing loudspeakers from Q to Q-1. The 22.2 multichannel sound system without two low-frequency effect channels can be simplified to 6 channels automatically and the total of loudspeaker arrangements is ten. The subjective evaluation of the proposed method is better than that of the previous multichannel reduction method with the decrease of the number of loudspeakers.
Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Guo Wu, Weiping Tu
ICME2
2016 Boosted local classifiers for visual tracking
abstract
Most existing discriminative tracking methods model a target object as a whole and train a tracker based on holistic templates, which cannot effectively deal with partial occlusions. Instead, in this paper, by treating the target as a collection of local patches, we propose a novel tracking approach based on boosted local classifiers. Initially, a set of local patches are sampled to train a set of local classifiers, and the weight of each classifier is given based on the estimated error. In addition, the positive examples and negative examples are sampled for model update with two constraints during the tracking process, which helps obtain more negatives for updating the appearance model and improve the updating efficiency. With updating the weights of local classifiers based on the temporal stability, the tracker can effectively handle partial occlusions. Extensive experiments on various challenging image sequences demonstrate the superiority to several state-of-the-art methods.
Weijian Ruan, Jun Chen 0001, Jinqiao Wang, Bo Luo, Ruimin Hu
ICME6
2016 Distance learning by treating negative samples differently and exploiting impostors with symmetric triplet constraint for person re-identification
abstract
Distance learning (DL) is an effective technique for person reidentification (PR-ID). DL based methods learn the distance metric by exploiting the discriminative information contained in samples. In PR-ID, different types of negative samples own different amounts of discriminative information, and impostor samples usually own more than other well separable negative samples (WSN-samples). Therefore, how to make full use of the different discriminative information conveyed by all negative samples in the DL process is a critical issue to be investigated. In this paper, we propose a novel DL approach for PR-ID. Specifically, for each target sample, we divide its negative samples into impostors and WSN-samples. Then we learn the distance metric by utilizing impostors and WSN-samples differently. For impostors, we design a symmetric triplet constraint, which requires the impostor to be far away from both samples of its corresponding positive sample pair simultaneously; for WSN-samples, we require them to keep their favorable separability. Experimental results on three benchmark datasets demonstrate the effectiveness and efficiency of our approach.
Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Wei-Shi Zheng 0001, Ruimin Hu, Chunxia Xiao, Chao Liang 0001
ICME5
2016 Scale-Adaptive Low-Resolution Person Re-Identification via Learning a Discriminating Surface
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Junjun Jiang, Chao Liang 0001, Jinqiao Wang
IJCAI2
2016 Face Super Resolution for VLQ facial images via parent patch matching
abstract
Face Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch based methods are based on the consistency assumption that the neighbors in HR/LR space form similar local geometry. But when LR images are Very Low Quality(VLQ), the LR space is seriously contaminated that even two distinct patches look similar, which means that the consistency assumption is not well held anymore. To this end, in this paper we use the target patch as well as the surrounding pixels, which we called parent patch, to represent the target patch. By incorporating the peripheral information, the parent patch is much more robust to noise in the LR and HR consistency learning. The effectiveness of proposed method is verified both quantitatively and qualitatively.
Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Qing Li 0001, Zheng Lu 0002
IJCNN2
2016 Spatiotemporal saliency based on location prior model
abstract
Saliency detection for images and videos becomes increasingly popular due to its wide applicability. Enormous research efforts have been focused on saliency detection, but it still has some issues in maintaining spatiotemporal consistency of videos and uniformly highlighting entire objects. To address these issues, this paper proposes a superpixel-level spatiotemporal saliency model for saliency detection in videos. To detect salient object, we extract multiple spatiotemporal features combined with intra-consistency motion information preliminarily. Meanwhile, considering inter-consistency of foreground in videos, a set of foreground locations are obtained from previous frames. Then, we introduce foreground-background and local foreground contrast saliency cues of those features using the location prior information of foreground. These two improved contrast saliency cues uniformly highlight the entire object and suppress the background effectively. Finally, we use an interactively dynamic fusion method to integrate the output spatial and temporal saliency maps. The proposed approach is validated on challenging sets of video sequences. Subjective observations and objective evaluations demonstrate that the proposed model achieves a better performance on saliency detection compared with the state-of-the-art spatiotemporal saliency methods.
Liuyi Hu, Zhongyuan Wang 0001, Mang Ye, Jing Xiao 0004, Ruimin Hu
IJCNN5
2016 Face Image Super-Resolution Through Improved Neighbor Embedding
Kebin Huang, Ruimin Hu, Junjun Jiang, Zhen Han 0002
MMM (1)2
2016 Camera Network Based Person Re-identification by Leveraging Spatial-Temporal Constraint and Multiple Cameras Relations
Wenxin Huang, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Xian Zhong, Chunjie Zhang 0001
MMM (1)2
2016 Analysis and Comparison of Inter-Channel Level Difference and Interaural Level Difference
Tingzhao Wu, Ruimin Hu, Xiaochen Wang 0001, Shanfa Ke
MMM (1)2
2016 Global Contrast Based Salient Region Boundary Sampling for Action Recognition
Zengmin Xu, Ruimin Hu, Jun Chen 0001
MMM (1)2
2016 Level Ratio Based Inter and Intra Channel Prediction with Application to Stereo Audio Frame Loss Concealment
Yuhong Yang 0001, Yanye Wang, Ruimin Hu, Hongjiang Yu, Song Wang 0011
MMM (1)3
2016 Adaptive Multichannel Reduction Using Convex Polyhedral Loudspeaker Array
Lingkun Zhang, Ruimin Hu, Dengshi Li, Xiaochen Wang 0001, Weiping Tu
MMM (1)2
2016 Geometrically Based Linear Iterative Clustering for Quantitative Feature Correspondence
abstract
Abstract A major challenge in feature matching is the lack of objective criteria to determine corresponding points. Recent methods find match candidates first by exploring the proximity in descriptor space, and then rely on a ratio‐test strategy to determine final correspondences. However, these measurements are heuristic and subjectively excludes massive true positive correspondences that should be matched. In this paper, we propose a novel feature matching algorithm for image collections, which is capable of providing quantitative depiction to the plausibility of feature matches. We achieve this by exploring the epipolar consistency between feature points and their potential correspondences, and reformulate feature matching as an optimization problem in which the overall geometric inconsistency across the entire image set ought to be minimized. We derive the solution of the optimization problem in a simple linear iterative manner, where a k‐means‐type approach is designed to automatically generate consistent feature clusters. Experiments show that our method produces precise correspondences on a variety of image sets and retrieves many matches that are subjectively rejected by recent methods. We also demonstrate the usefulness of the framework in structure from motion task for denser point cloud reconstruction.
Qingan Yan, Long Yang 0001, Chao Liang 0001, Huajun Liu, Ruimin Hu, Chunxia Xiao
Comput. Graph. Forum5
2016 Filling Kinect depth holes via position-guided matrix completion
Zhongyuan Wang 0001, Shizheng Wang, Jing Xiao 0004, Ruimin Hu
Neurocomputing6
2016 Noise robust position-patch based face super-resolution via Tikhonov regularized neighbor representation
Junjun Jiang, Chen Chen 0001, Kebin Huang, Zhihua Cai, Ruimin Hu
Inf. Sci.5
2016 Heteroskedasticity tuned mixed-norm sparse regularization for face hallucination
Zhongyuan Wang 0001, Ruimin Hu, Junjun Jiang, Zhen Han 0002
Multim. Tools Appl.2
2016 Multi-view low-rank dictionary learning for image classification
Fei Wu 0004, Xiaoyuan Jing, Xinge You, Dong Yue 0001, Ruimin Hu, Jing-Yu Yang 0001
Pattern Recognit.5
2016 CDMMA: Coupled discriminant multi-manifold analysis for matching low-resolution face images
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhihua Cai
Signal Process.2
2016 Facial Image Hallucination Through Coupled-Layer Neighbor Embedding
abstract
As the facial image captured by a low-cost camera is typically very low resolution (LR), blurring, and noisy, traditional neighbor-embedding-based facial image hallucination methods from one single manifold (i.e., the LR image manifold) fail to reliably estimate the intention geometrical structure, consequently leading to a bias to the image reconstruction result. In this paper, we introduce the notion of neighbor embedding (NE) from the LR and the high-resolution (HR) image manifolds simultaneously and propose a novel NE model, termed the coupled-layer NE (CLNE), for facial image hallucination. CLNE differs substantially from other NE models in that it has two layers: the LR and the HR layers. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights of the LR patches; the HR layer is the intrinsic geometry that can geometrically constrain the reconstruction weights. With this coupled-constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust NE through iteratively updating the LR patch reconstruction weights and the estimated HR patch. The experimental results in simulation and real conditions confirm that the proposed method outperforms the related state-of-the-art methods in both quantitative and visual comparisons.
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002, Jiayi Ma 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 Multi-Label Dictionary Learning for Image Annotation
abstract
Image annotation has attracted a lot of research interest, and multi-label learning is an effective technique for image annotation. How to effectively exploit the underlying correlation among labels is a crucial task for multi-label learning. Most existing multi-label learning methods exploit the label correlation only in the output label space, leaving the connection between the label and the features of images untouched. Although, recently some methods attempt toward exploiting the label correlation in the input feature space by using the label information, they cannot effectively conduct the learning process in both the spaces simultaneously, and there still exists much room for improvement. In this paper, we propose a novel multi-label learning approach, named multi-label dictionary learning (MLDL) with label consistency regularization and partial-identical label embedding MLDL, which conducts MLDL and partial-identical label embedding simultaneously. In the input feature space, we incorporate the dictionary learning technique into multi-label learning and design the label consistency regularization term to learn the better representation of features. In the output label space, we design the partial-identical label embedding, in which the samples with exactly same label set can cluster together, and the samples with partial-identical label sets can collaboratively represent each other. Experimental results on the three widely used image datasets, including Corel 5K, IAPR TC12, and ESP Game, demonstrate the effectiveness of the proposed approach.
Xiaoyuan Jing, Fei Wu 0004, Zhiqiang Li 0003, Ruimin Hu, David Zhang 0001
IEEE Trans. Image Process.4
2016 Zero-Shot Person Re-identification via Cross-View Consistency
abstract
Person re-identification, aiming to identify images of the same person from various cameras configured in different places, has attracted much attention in the multimedia retrieval community. In this problem, choosing a proper distance metric is a crucial aspect, and many classic methods utilize a uniform learnt metric. However, their performance is limited due to ignoring the zero-shot and fine-grained characteristics presented in real person re-identification applications. In this paper, we investigate two consistencies across two cameras, which are cross-view support consistency and cross-view projection consistency. The philosophy behind it is that, in spite of visual changes in two images of the same person under two camera views, the support sets in their respective views are highly consistent, and after being projected to the same view, their context sets are also highly consistent. Based on the above phenomena, we propose a data-driven distance metric (DDDM) method, re-exploiting the training data to adjust the metric for each query-gallery pair. Experiments conducted on three public data sets have validated the effectiveness of the proposed method, with a significant improvement over three baseline metric learning methods. In particular, on the public VIPeR dataset, the proposed method achieves an accuracy rate of 42.09% at rank-1, which outperforms the state-of-the-art methods by 4.29%.
Zheng Wang 0007, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Junjun Jiang, Mang Ye, Jun Chen 0001, Qingming Leng
IEEE Trans. Multim.2
2016 Knowledge-Based Coding of Objects for Multisource Surveillance Video Data
abstract
Global object redundancy (GOR), as opposed to local spatial/temporal redundancies in a single video clip, is a new form of redundancy common in multisource surveillance video data (MSVD). GOR is induced by the repetition of foreground objects across multiple cameras, and becomes influential as the number of objects increases. Eliminating GOR considerably improves MSVD coding efficiency. In an effort to accomplish this, this study first proposes a knowledge-based representation of objects based on careful analysis of GOR composition. The representation contains a constant part and a variational part: the former is used to represent the common knowledge shared by an object across multiple cameras, while the latter is used to represent local variations on the object's surfaces. Based on the proposed representation, a knowledge-based coding (KBC) method is then proposed in which each foreground object is encoded with a hybrid prediction scheme, where the constant part of the object is generated via global prediction from a model library and the variational part is predicted via local reference frames with pose-based, short-term prediction. Experimental results showed that the KBC method saves more than 39% bits on average for encoding foreground objects in high-resolution video clips (compared to 16% for the entire videos). Applying the proposed coding method to surveillance videos in large spatial and temporal scale allows storage savings at the PB level.
Jing Xiao 0004, Ruimin Hu, Yu Chen 0021, Zhongyuan Wang 0001, Zixiang Xiong
IEEE Trans. Multim.2
2016 Person Reidentification via Ranking Aggregation of Similarity Pulling and Dissimilarity Pushing
abstract
Person reidentification is a key technique to match different persons observed in nonoverlapping camera views. Many researchers treat it as a special object-retrieval problem, where ranking optimization plays an important role. Existing ranking optimization methods mainly utilize the similarity relationship between the probe and gallery images to optimize the original ranking list, but seldom consider the important dissimilarity relationship. In this paper, we propose to use both similarity and dissimilarity cues in a ranking optimization framework for person reidentification. Its core idea is that the true match should not only be similar to those strongly similar galleries of the probe, but also be dissimilar to those strongly dissimilar galleries of the probe. Furthermore, motivated by the philosophy of multiview verification, a ranking aggregation algorithm is proposed to enhance the detection of similarity and dissimilarity based on the following assumption: the true match should be similar to the probe in different baseline methods. In other words, if a gallery blue image is strongly similar to the probe in one method, while simultaneously strongly dissimilar to the probe in another method, it will probably be a wrong match of the probe. Extensive experiments conducted on public benchmark datasets and comparisons with different baseline methods have shown the great superiority of the proposed ranking optimization method.
Mang Ye, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Qingming Leng, Chunxia Xiao, Jun Chen 0001, Ruimin Hu
IEEE Trans. Multim.8
2015 Super-resolution Person re-identification with semi-coupled low-rank discriminant dictionary learning
abstract
Person re-identification has been widely studied due to its importance in surveillance and forensics applications. In practice, gallery images are high-resolution (HR) while probe images are usually low-resolution (LR) in the identification scenarios with large variation of illumination, weather or quality of cameras. Person re-identification in this kind of scenarios, which we call super-resolution (SR) person re-identification, has not been well studied. In this paper, we propose a semi-coupled low-rank discriminant dictionary learning (SLD2L) approach for SR person re-identification. For the given training image set which consists of HR gallery and LR probe images, we aim to convert the features of LR images into discriminating HR features. Specifically, our approach learns a pair of HR and LR dictionaries and a mapping from the features of HR gallery images and LR probe images. To ensure that the converted features using the learned dictionaries and mapping have favorable discriminative capability, we design a discriminant term which requires the converted HR features of LR probe images should be close to the features of HR gallery images from the same person, but far away from the features of HR gallery images from different persons. In addition, we apply low-rank regularization in dictionary learning procedure such that the learned dictionaries can well characterize intrinsic feature space of HR and LR images. Experimental results on public datasets demonstrate the effectiveness of SLD2L.
Xiaoyuan Jing, Xiaoke Zhu, Fei Wu 0004, Xinge You, Qinglong Liu, Dong Yue 0001, Ruimin Hu, Baowen Xu
CVPR7
2015 Joint Weighted Sparse Representation Based Median Filter for Depth Video Coding
abstract
In order to promote the development of auto-stereoscopic display, MPEG has proposed multi-view plus depth (MVD) format. The depth video is encoded and transmitted with color video to synthesize virtual views at the receiver side. The existing video coding standards such as H.264/AVC introduces coding artifacts along the depth boundaries, which may seriously affects the synthesized view quality and coding efficiency. Many in-loop depth filters such as joint depth filter have been proposed to remove the artifacts in compressed depth video. However, their performance is unstable and affected by the outliers due to the weighted summation. In this paper, based on the sparse prior characteristic in local region of depth map, we propose a joint weighted sparse representation based median filter to select the most relevant neighboring depth pixel as the output during the filter process. Experimental results show the proposed method is more effective in improving the depth video coding efficiency.
Ruimin Hu, Yu Chen 0021, Jing Xiao 0004, Ruolin Ruan
DCC2
2015 Global Coding of Multi-source Surveillance Video Data
abstract
In this paper, we exploit a new type of data redundancy in the multisource surveillance video to reduce the huge gap between the growth rate of the data and the video compression rate. Global redundancy caused by correlated appearances of moving objects in multiple videos consists of model similarity, spatial correlation and temporal consistency. Therefore, we propose a global coding scheme of moving objects to eliminate the global redundancy: a model based object reconstruction is initially employed to reconstruct the objects in the video, then a pose-based residual error prediction is developed to compensate the difference between the real video appearance and the initial reconstruction from model. The experiment with two simulated surveillance videos has proved that the proposed coding scheme can achieve better coding performance than the main profile of HEVC and surveillance profile of IEEE 1857-2013.
Jing Xiao 0004, Yu Chen 0021, Ruimin Hu
DCC5
2015 A Block-Based Background Model for Surveillance Video Coding
abstract
Background model can help to improve the compression efficiency for surveillance video coding, but the existing frame-based background model is inefficient in some situations, for example, when a region of background changes frequently or periodically. In this paper, a block-based background model is proposed to solve this problem. We save the background blocks recognized from each reconstructed frame into a buffer, thus the background blocks are collected gradually. At the same time, we compose a new background frame for each frame to be encoded based on the background blocks currently available in the buffer. Compared with the pre-built background frame, the instantly composed background frame often predicts more accurately because of the accumulated information about background. Experimental results show that the proposed model achieves better rate-distortion performance over the existing frame-based model in most cases, while keeping almost the same computation complexity.
Liming Yin, Ruimin Hu, Jing Xiao 0004
DCC2
2015 Face hallucination via Cauchy regularized sparse representation
abstract
In dictionary-learning-based face hallucination, the testing image is represented as a linear combination of the training samples, and how to obtain the optimal coefficients is the primary issue. Sparse representation (SR) has ever been widely used in face hallucination, however, due to the fact that SR overemphasizes the sparsity, the obtained linear combination coefficients turn out far aggressively sparse, then leading to unsatisfactory hallucinated results. In this paper, we present a moderately sparse prior model for face hallucination problem with the L1 norm penalty in classic SR replaced by a Cauchy penalty term. An iterative optimization is further presented to solve the minimization of Cauchy regularized objective function. The experimental results on public face database demonstrate that our method is much more effective than state-of-the-art methods.
Shenming Qu, Ruimin Hu, Zhongyuan Wang 0001, Junjun Jiang
ICASSP2
2015 A down-mixing method for 22.2 multichannel system reproduction
abstract
This paper proposes a general multichannel system reproduction method. Firstly, relative to original multichannel system, a general global model is build up by guaranteeing sound pressure and the direction of particle velocity at the receiving point constant, and making the square error of particle velocity magnitude at the receiving point as little as possible. Then the model is equivalent to a least squares problems with non-negative constraints, it can be worked out by existing mature algorithms, and the global optimal solution of simplifying multichannel system are obtained. The proposed method can be used to simplify 22.2 multichannel system to 10.2 and 8.2 multichannel system, objective and subjective experimental results demonstrate that it performs better than traditional method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP2
2015 R2FP: Rich and Robust Feature Pooling for Mining Visual Data
abstract
The human visual system proves smart in extracting both global and local features. Can we design a similar way for unsupervised feature learning? In this paper, we propose anovel pooling method within an unsupervised feature learningframework, named Rich and Robust Feature Pooling (R2FP), to better explore rich and robust representation from sparsefeature maps of the input data. Both local and global poolingstrategies are further considered to instantiate such a methodand intensively studied. The former selects the most conductivefeatures in the sub-region and summarizes the joint distributionof the selected features, while the latter is utilized to extractmultiple resolutions of features and fuse the features witha feature balancing kernel for rich representation. Extensiveexperiments on several image recognition tasks demonstratethe superiority of the proposed techniques.
Wei Xiong 0008, Bo Du 0001, Lefei Zhang, Ruimin Hu, Wei Bian 0003, Jialie Shen 0001, Dacheng Tao
ICDM4
2015 Exploiting effects of parts in fine-grained categorization of vehicles
abstract
Fine-grained categorization has become a hot topic in computer vision. Based on the theory that part information is crucial for fine-grained categorization, we proposed a part-based categorization method for vehicles, consisting vehicle parts localization, part-based vehicle representation and classification. There were three contributions we made in this work: 1) we analyzed discriminative powers of parts for fine-grained categorization; 2) we proposed a frame of how to integrate discriminative powers of parts into categorization, and proved that it can achieve better performance than treating every part equally; 3) we provided an annotated dataset with parts for vehicle categorization.
Ruimin Hu, Jun Xiao 0004, Jing Xiao 0004, Jun Chen 0001
ICIP2
2015 How much bandwidth does surveillance system require?
abstract
One of the main challenges in surveillance systems lies in the massive amount of video involved in providing potential key content with sufficient resolution. This paper shows that there exists a sweet spot, which we term critical video quality that can be used to reduce bitrate of video transmission without significantly affecting the accuracy of the surveillance tasks. We present a new city surveillance dataset which was divided into three types of scenarios, and we analyze subjective data collected via human subjective testing for object identification. These data are then used to create objective measurements (models) to drive video compression ratio based on the detection probability. The main idea is to find out the lowest bitrate of video transmission while maximizes the probability of detecting objects which are carried or abandoned. Experiment results shown that our generalized models can predict acceptable video quality for object identification in rational ways.
Zengmin Xu, Ruimin Hu, Jun Chen 0001
ICIP2
2015 Locally regularized Anchored Neighborhood Regression for fast Super-Resolution
abstract
The goal of learning-based image Super-Resolution (SR) is to generate a plausible and visually pleasing High-Resolution (HR) image from a given Low-Resolution (LR) input. The problem is dramatically under-constrained, which relies on examples or some strong image priors to better reconstruct the missing HR image details. This paper addresses the problem of learning the mapping functions (i.e. projection matrices) between the LR and HR images based on a dictionary of LR and HR examples. One recently proposed method, Anchored Neighborhood Regression (ANR) [1], provides state-of-the-art quality performance and is very fast. In this paper, we propose an improved variant of ANR, namely Locally regularized Anchored Neighborhood Regression (LANR), which utilizes the locality-constrained regression in place of the ridge regression in ANR. LANR assigns different freedom for each neighbor dictionary atom according to its correlation to the input LR patch, thus the learned projection matrices are much more flexible. Experimental results demonstrate that the proposed algorithm performs efficiently and effectively over state-of-the-art methods, e.g., 0.1–0.4 dB in term of PSNR better than ANR.
Junjun Jiang, Jican Fu, Tao Lu 0001, Ruimin Hu, Zhongyuan Wang 0001
ICME4
2015 Spatial perception reproduction of sound events based on sound property coincidences
abstract
Sound pressure and particle velocity are used to reproduce sound signals in multichannel systems. The two sound properties were estimated step by step and particle velocity was scaled due to ill-conditioned equations in Ando's study. We explore a new system of equations to maintain both sound pressure and particle velocity. The weight equations are solved in a non-traditional way to figure out exact solutions. Based on the proposed method, the perception of the direction of a sound event and the distance to the listening point are both reproduced correctly in a three-dimension reproduction system. The comparison between the proposed method and Ando's method is outlined and the proposed method is more flexible and useful. Objective evaluation shows the wavefront in the proposed method is more accurate than Ando's method and subjective evaluation confirms that the proposed method improves the spatial perception of sound events.
Maosheng Zhang, Ruimin Hu, Xiaochen Wang 0001, Dengshi Li
ICME2
2015 Unequal error protection for S3AC coding based on expanding window fountain codes
abstract
This paper presents a coding scheme to improve the performance of three-dimensional (3D) audio. The scheme is designed by the idea of joint source channel coding (JSCC) and is implemented by Expanding Window Fountain (EWF) codes for the 3D audio bitstreams after source coding of spatial squeeze surround audio coding (S3AC). EWF is one of unequal error protection (UEP) LT codes and the proposed scheme is achieved by that method for the case of two levels. Different from other transmissions with equal error protection (EEP) for each part of bitstreams, when transmitting the two parts of the bitstreams with downmixed mono signals and spatial side information after S3AC coding, this approach provides more protection to the part of spatial side information and comparatively less protection to the downmixed mono signals, and results in the improvement of the performance of the reproduced 3D audio especially on the aspect of the spatial perception. Objective simulation experiment has shown the proposed UEP scheme achieves a better performance than the EEP scheme on the aspect of spatial perception, for the bits error rates (BER) of spatial parameters can decrease dramatically to a low value of about 10-4, but BERs of downmixed mono signals and the case of EEP just decrease slightly to about 10-3.
Liuyue Su, Ruimin Hu, Xiaochen Wang 0001
ISCC4
2015 A Unsupervised Person Re-identification Method Using Model Based Representation and Ranking
abstract
As a core technique supporting the multi-camera tracking task, person re-identification attracts increasing research interests in both academic and industrial communities. Its aim is to match individuals across a group of spatially non-overlapping surveillance cameras, which are usually interfered by various imaging conditions and object motions. Current methods mainly focus on robust feature representation and accurate distance measure, where intensive computations and expensive training samples prohibit their practical applications. To address the above problems, this paper proposes a new unsupervised person re-identification method featured by its competitive accuracy and high efficiency. Both merits stem from model based person image representation and ranking, with which, merely 4-dimension pixel-level features can achieve over 20% matching rate at Rank 1 on the challenging VIPeR dataset.
Chao Liang 0001, Bingyue Huang, Ruimin Hu, Chunjie Zhang 0001, Xiaoyuan Jing, Jing Xiao 0004
ACM Multimedia3
2015 Multi-Level Fusion for Person Re-identification with Incomplete Marks
abstract
Most video surveillance suspect investigation systems rely on the videos taken in different camera views. Actually, besides the videos, in the investigation process, investigators also manually label some marks, which, albeit incomplete, can be quite accurate and helpful in identifying persons. This paper studies the problem of Person Re-identification with Incomplete Marks (PRIM), aiming at ranking the persons in the gallery according to both the videos and incomplete marks. This problem is solved by a multi-step fusion algorithm, which consists of three key steps: (i) The early fusing step exploits both visual features and marked attributes to predict a complete and precise attribute vector. (ii) Based on the statistical attribute d ominance and saliency phenomena, a dominance-saliency matching model is suggested for measuring the distance between attribute vectors. (iii) The gallery is ranked separately by using visual features and attribute vectors, and the overall ranking list is the result of a late fusion. Experiments conducted on VIPeR dataset have validated the effectiveness of the proposed method in all the three key steps. The results also show that through introducing marks, the retrieval accuracy is significantly improved.
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Chao Liang 0001, Wenxin Huang
ACM Multimedia2
2015 Object Detection in Low-Resolution Image via Sparse Representation
Wenhua Fang, Jun Chen 0001, Chao Liang 0001, Xiao Wang 0029, Yuanyuan Nan, Ruimin Hu
MMM (1)6
2015 Azimuthal Perceptual Resolution Model Based Adaptive 3D Spatial Parameter Coding
Ruimin Hu, Yuhong Yang 0001, Weiping Tu, Tingzhao Wu
MMM (1)2
2015 Coupled Discriminant Multi-Manifold Analysis with Application to Low-Resolution Face Recognition
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Liang Chen 0026, Jun Chen 0001
MMM (1)2
2015 Person Re-identification Using Data-Driven Metric Adaptation
Zheng Wang 0007, Ruimin Hu, Chao Liang 0001, Junjun Jiang, Kaimin Sun, Qingming Leng, Bingyue Huang
MMM (2)2
2015 Signal-Aware Parametric Quality Model for Audio and Speech over IP Networks
Songbo Xie, Yuhong Yang 0001, Ruimin Hu, Yanye Wang, Hongjiang Yu, ShaoLong Dong
MMM (1)3
2015 Person re-identification with content and context re-ranking
Qingming Leng, Ruimin Hu, Chao Liang 0001, Jun Chen 0001
Multim. Tools Appl.2
2015 3D hybrid just noticeable distortion modeling for depth image-based rendering
Ruimin Hu, Zhongyuan Wang 0001, Shizheng Wang
Multim. Tools Appl.2
2014 Joint speech/audio coding based scalable perceptual audio coding
abstract
With the technical evolution of global mobile communications, various heterogeneous communication environments, frequently fluctuant bandwidth and multiform signals put new challenges to coding technology of multimedia signals. Scalable Audio Coding (SAC) can provide smooth transition between different coding qualities, which is an optimal choice for coding audio signals of different types and can produce more reliable and consistent service quality in multimedia communications. A scalable audio coding system based on joint speech/audio coding method and an auditory perceptual importance model based on bit-plane are proposed here. Both the audio content and the network bandwidth fluctuation will be considered in the system to obtain stable service qualities in mobile multimedia services. Experimental results indicate that with the same bit rates the subjective quality of proposed method is slightly better than G.729.1 and the SNR is improved by 0.3dB.
Ruimin Hu, Yuhong Yang 0001
ICIS2
2014 Uncorrelated Multi-View Discrimination Dictionary Learning for Recognition
abstract
Dictionary learning (DL) has now become an important feature learning technique that owns state-of-the-art recognition performance. Due to sparse characteristic of data in real-world applications, DL uses a set of learned dictionary bases to represent the linear decomposition of a data point. Fisher discrimination DL (FDDL) is a representative supervised DL method, which constructs a structured dictionary whose atoms correspond to the class labels. Recent years have witnessed a growing interest in multi-view (more than two views) feature learning techniques. Although some multi-view (or multi-modal) DL methods have been presented, there still exists much room for improvement. How to enhance the total discriminability of dictionaries and reduce their redundancy is a crucial research topic. To boost the performance of multi-view DL technique, we propose an uncorrelated multi-view discrimination DL (UMDDL) approach for recognition. By making dictionary atoms correspond to the class labels such that the obtained reconstruction error is discriminative, UMDDL aims to jointly learn multiple dictionaries with totally favorable discriminative power. Furthermore, we design the uncorrelated constraint for multi-view DL, so as to reduce the redundancy among dictionaries learned from different views. Experiments on several public datasets demonstrate the effectiveness of the proposed approach.
Xiaoyuan Jing, Ruimin Hu, Fei Wu 0004, Xilin Chen 0001, Qian Liu 0010, Yong-Fang Yao
AAAI2
2014 Intra-View and Inter-View Supervised Correlation Analysis for Multi-View Feature Learning
abstract
Multi-view feature learning is an attractive research topic with great practical success. Canonical correlation analysis (CCA) has become an important technique in multi-view learning, since it can fully utilize the inter-view correlation. In this paper, we mainly study the CCA based multi-view supervised feature learning technique where the labels of training samples are known. Several supervised CCA based multi-view methods have been presented, which focus on investigating the supervised correlation across different views. However, they take no account of the intra-view correlation between samples. Researchers have also introduced the discriminant analysis technique into multi-view feature learning, such as multi-view discriminant analysis (MvDA). But they ignore the canonical correlation within each view and between all views. In this paper, we propose a novel multi-view feature learning approach based on intra-view and inter-view supervised correlation analysis (I2SCA), which can explore the useful correlation information of samples within each view and between all views. The objective function of I2SCA is designed to simultaneously extract the discriminatingly correlated features from both inter-view and intra-view. It can obtain an analytical solution without iterative calculation. And we provide a kernelized extension of I2SCA to tackle the linearly inseparable problem in the original feature space. Four widely-used datasets are employed as test data. Experimental results demonstrate that our proposed approaches outperform several representative multi-view supervised feature learning methods.
Xiaoyuan Jing, Ruimin Hu, Yang-Ping Zhu, Chao Liang 0001, Jing-Yu Yang 0001
AAAI2
2014 A spatial priority based scalable audio coding
abstract
A spatial priority scheme for scalable audio coding is presented in this paper. To improve the coding quality of important sounds with high attention, especially the moving sound, spatial information is introduced to assign the priorities of frequency subbands. Spatial cues and distance features are extracted in frequency subbands to represent the sound with fast changing direction and distance. Coding priorities are assigned to different frequency subbands according to the energy and spatial information. With trivial added side information and complexity, experimental results show that the perceptual quality is improved especially for the sound with high attention, especially the moving sound in scalable audio coding.
Ruimin Hu, Yuhong Yang 0001
ICASSP2
2014 Auditory attention based mobile audio quality assessment
abstract
Mobile audio services are growing with rising popularity of smart mobile devices using WiFi or cellular networks. A major issue facing mobile audio quality assessment is occasional background noises due to the prospect of sound recording at anytime and anywhere with smart mobile devices. Psychological study reveals that people pay selective attention to their interested sound in complex auditory input. In this paper, we model the mobile audio objective quality assessment based on auditory attention mechanism, with attention based horizontal azimuth parameters and timbre distortion parameters as additional Model Output Variables (MOVs). The results show that the prediction accuracy can be obtained by using such a method.
Yuhong Yang 0001, Hongjiang Yu, Ruimin Hu, Song Wang 0011, Qing Zhai, Songbo Xie
ICASSP3
2014 Gabor-based patch covariance matrix for face sketch synthesis
abstract
In this paper, we propose a novel face sketch/photo synthesis method by utilizing Gabor-based Patch Covariance Matrix (GPCM) as face descriptor, a.k.a. symmetric positive definite matrix, which lie on a Riemannian manifold. In particular, both pixel locations and Gabor coefficients of one patch are employed to form the covariance matrix. In this way, the sketch/photo can be then transformed from the pixel space to the Riemannian manifold space. With the aid of the recently introduced Stein kernel theory, we advance to perform Regularized Least Square Representation (RLSR) in Stein space. Based on the assumption that the Stein divergence manifold of photo/sketch patch and the sketch/photo share the same topology, a new sketch/photo patch of the same position can be synthesized by keeping the weights and replacing the photo/sketch training image patches with the corresponding sketch/photo ones. Experimental results demonstrate the superiority of the proposed method.
Ruimin Hu, Junjun Jiang, Zhen Han 0002
ICIP2
2014 Robust tracking via saliency-based appearance model
abstract
We propose a novel local-based saliency measure (LBSM) method for object tracking problem. In LBSM method, salient patches are defined as the patches having great local changes. Then we apply the saliency information derived from LBSM to appearance model by giving weights to patches according to their saliency levels. The patches with higher saliency levels are given larger weights. As a result, the appearance model is improved owing to the use of saliency information. Extensive experiments conducted on various challenging sequences demonstrate the effectiveness of LB-SM in tracking procedure, and our saliency-based tracker performs well against state-of-the-art algorithms.
Bo Luo, Ruimin Hu, Chao Liang 0001, Chunjie Zhang 0001
ICIP2
2014 Pedestrian detection from salient regions
abstract
Classic algorithms of pedestrian detection usually locate the latent position via sliding window techniques, which resize the matching window and/or original images at different scales and scan the image. However, this method has two main drawbacks. First, resizing at a fix rate cannot search through the whole scale space, resulting in the failure of accurate object location. Second, resizing and scanning at various scales is usually time-consuming, which is improper for practical applications. To conquer the above difficulties, a novel pedestrian detection method with salient information is proposed. In this paper, the salient detection model and the traditional covariance matrix descriptor are combined in a Bayesian framework to detect pedestrians in the still image. Finally, the efficiency of our approach compared with state-of-the-art results is demonstrated on the public INRIA dataset.
Xiao Wang 0029, Jun Chen 0001, Wenhua Fang, Chao Liang 0001, Chunjie Zhang 0001, Ruimin Hu
ICIP6
2014 Face hallucination via re-identified K-nearest neighbors embedding
abstract
Based on locally linear embedding (LLE) manifold learning theory, which assumes that the low-resolution (LR) manifold and high-resolution (HR) manifold spaces share the same local geometry structure, neighbor embedding based super-resolution(SR) methods search K-nearest neighbors(K-NN) of LR patch, then use the counterpart HR patches to estimate HR patch. The primary issue of these methods is how to search the optimal K-NN. However, due to the “one-to-many” mapping between the LR image and HR ones in practice, the neighborhood relationship of the LR patch in LR space is very different with its HR counterpart's. In this paper, we explore a novel and effective re-identified K-NN(RIKNN) method to search neighbors of LR patch by taking into consideration the neighbor information in the HR space. It searches K-NN of LR patch in the LR space and then refines the searching results by re-identifying in the HR space, thus giving rise to accurate K-NN and improvement performance. Experimental results with application to face hallucination demonstrate that our method outperforms state of the art in terms of subjective and objective results and computational complexity.
Shenming Qu, Ruimin Hu, Junjun Jiang, Zhongyuan Wang 0001, Jun Chen 0001
ICME2
2014 A 3D audio coding technique based on extracting the distance parameter
abstract
This paper presents a compression technique to improve the quality of three-dimensional (3D) audio produced by multiple loudspeaker channels or by headphone. The approach is based on extracting the side information of spatial sound sources within the three-dimensional space when capturing the sound sources. Different from other compression technique, the distances of sound sources are included in the side information. The separated signals of different sound sources are downmixed into one mono or stereo audio signal with the side information. The resulting downmixed signal is then compressed with traditional audio coder, resulting in a better perceptual quality of 3D audio by adding the distance parameter in the side information, and maintaining a low bit rates comparable with directional audio coding (DirAC).
Ruimin Hu, Liuyue Su, Weiping Tu, Xiaochen Wang 0001, Yuhong Yang 0001, Shi Dong 0004, Song Wang 0011, Maosheng Zhang, Furong Lei, Shiqing Li
ICME2
2014 Cloud Model-Based Dynamic Texture Synthesis for Video Coding
abstract
This paper presents a novel cloud model which is designed for the inter prediction coding using virtual frame technique. The virtual frame obtained by typical dynamic texture synthesis methods can have a better prediction result in some regions of the encoding frame than the common reference frames, because the non-linear motion and global illumination change between frames is taken into account. However, there are still many limiting factors for current dynamic texture models, which make the video scenes in virtual frames can't better reflect the motion and changing trend of the real scenes. The prediction performance of virtual frame is to some degree reduced or covered up. The proposed method firstly combined the powerful computing capabilities of cloud platform and the extrapolation process of dynamic texture synthesis for video codec. Results show that this method can be used as a significant complement to the current virtual frame technique.
Ruimin Hu, Zhongyuan Wang 0001
ICPR2
2014 Efficient learning based face hallucination approach via facial standard deviation prior
abstract
Most state-of-the-art face hallucination approaches suffer from complicated learning patterns and highly intensive computation, which will lead to low efficiency and considerable computing resources. Therefore, how to restore real face image quickly and efficiently is still an important issue in this field. To solve or partially solve the problem, this paper proposed a novel facial standard deviation prior based approach which can provide superior results with high efficiency for real face images. The high frequency information of test image will be enhanced via a facial specific sharpening operator which is obtained through the learning of standard deviation correspondence of training set. Experiments in simulation and real world images verified the effectiveness of proposed approach, and the distinct advantage on runtime and resource requirement of proposed approach.
Liang Chen 0026, Ruimin Hu, Junjun Jiang, Zhen Han 0002
ISCAS2
2014 The Perceptual Characteristics of 3D Orientation
Ruimin Hu, Weiping Tu, Xiaochen Wang 0001
MMM (2)3
2014 Noise robust face hallucination employing Gaussian-Laplacian mixture model
Zhongyuan Wang 0001, Zhen Han 0002, Ruimin Hu, Junjun Jiang
Neurocomputing3
2014 Generating algorithm for integer DST radixes in video coding
Zhongyuan Wang 0001, Ruimin Hu
J. Vis. Commun. Image Represent.3
2014 Efficient single image super-resolution via graph-constrained least squares regression
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001
Multim. Tools Appl.2
2014 Face image super-resolution through locality-induced support regression
Junjun Jiang, Ruimin Hu, Chao Liang 0001, Zhen Han 0002, Chunjie Zhang 0001
Signal Process.2
2014 Fast Synopsis for Moving Objects Using Compressed Video
abstract
With the increasing volume of video data, how to analyze and browse video in a fast and effective way has become an urgent problem in applications. This letter proposes a novel video synopsis method in compressed domain for browsing video captured by static cameras. Synopsis video is a video abstraction, which displays moving objects from different periods simultaneously on the primary background contents of original video. To overcome the low efficiency of traditional video synopsis for compressed video, our method presents a new graph cut algorithm to extract objects tubes and meanwhile gives a fast solution to minimize energy function in compressed domain. Experimental results in H.264 video have demonstrated the high-efficiency of this new video synopsis scheme for massive video browsing.
Ruimin Hu, Zhongyuan Wang 0001, Shizheng Wang
IEEE Signal Process. Lett.2
2014 Comments on "Algorithmic Aspects of Hardware/Software Partitioning: 1D Search Algorithms"
abstract
In this paper, the work inis analyzed. An error in its theoretical description part is pointed out and illustrated by a simple example. A modification suggestion is proposed to make the theoretical description of the workmore deliberate and thus being used appropriately.
Hao-Jun Quan, Tao Zhang 0025, Qiang Liu 0011, Jichang Guo, Xiaochen Wang 0001, Ruimin Hu
IEEE Trans. Computers6
2014 Camera Compensation Using a Feature Projection Matrix for Person Reidentification
abstract
Matching individuals within a group of spatially nonoverlapping surveillance cameras, also known as person reidentification, has recently attracted a lot of research interest. Current methods mainly focus on feature representation or distance measure, which directly compare person images captured by different cameras. However, it is still a problem because of various surveillance conditions; for example, view switching, lighting variations, and image scaling. Although the brightness transfer function was proposed to address the problem of illumination variation, it could not handle view and scale changes among various cameras. In this paper, we propose a new approach to compensate for the inconsistency of feature distributions of person images captured by different cameras. More precisely, a feature projection matrix (FPM) is learned to project image features of one camera to the feature space of another camera, from which the latent device difference can be effectively eliminated for the person reidentification task. In particular, we formulate the FPM learning as a smooth unconstrained convex optimization problem and use a simple gradient descent algorithm with stochastic samples to accelerate the solving process. Extensive comparative experiments conducted on three standard datasets have shown the promising prospect of the proposed method.
Ruimin Hu, Chao Liang 0001, Chunjie Zhang 0001, Qingming Leng
IEEE Trans. Circuits Syst. Video Technol.2
2014 Face Hallucination Via Weighted Adaptive Sparse Regularization
abstract
Sparse representation-based face hallucination approaches proposed so far use fixed ℓ1norm penalty to capture the sparse nature of face images, and thus hardly adapt readily to the statistical variability of underlying images. Additionally, they ignore the influence of spatial distances between the test image and training basis images on optimal reconstruction coefficients. Consequently, they cannot offer a satisfactory performance in practical face hallucination applications. In this paper, we propose a weighted adaptive sparse regularization (WASR) method to promote accuracy, stability and robustness for face hallucination reconstruction, in which a distance-inducing weighted ℓqnorm penalty is imposed on the solution. With the adjustment to shrinkage parameter q , the weighted ℓqpenalty function enables elastic description ability in the sparse domain, leading to more conservative sparsity in an ascending order of q . In particular, WASR with an optimal q > 1 can reasonably represent the less sparse nature of noisy images and thus remarkably boosts noise robust performance in face hallucination. Various experimental results on standard face database as well as real-world images show that our proposed method outperforms state-of-the-art methods in terms of both objective metrics and visual quality.
Zhongyuan Wang 0001, Ruimin Hu, Shizheng Wang, Junjun Jiang
IEEE Trans. Circuits Syst. Video Technol.2
2014 Face Super-Resolution via Multilayer Locality-Constrained Iterative Neighbor Embedding and Intermediate Dictionary Learning
abstract
Based on the assumption that low-resolution (LR) and high-resolution (HR) manifolds are locally isometric, the neighbor embedding super-resolution algorithms try to preserve the geometry (reconstruction weights) of the LR space for the reconstructed HR space, but neglect the geometry of the original HR space. Due to the degradation process of the LR image (e.g., noisy, blurred, and down-sampled), the neighborhood relationship of the LR space cannot reflect the truth. To this end, this paper proposes a coarse-to-fine face super-resolution approach via a multilayer locality-constrained iterative neighbor embedding technique, which intends to represent the input LR patch while preserving the geometry of original HR space. In particular, we iteratively update the LR patch representation and the estimated HR patch, and meanwhile an intermediate dictionary learning scheme is employed to bridge the LR manifold and original HR manifold. The proposed method can faithfully capture the intrinsic image degradation shift and enhance the consistency between the reconstructed HR manifold and the original HR manifold. Experiments with application to face super-resolution on the CAS-PEAL-R1 database and real-world images demonstrate the power of the proposed algorithm.
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002
IEEE Trans. Image Process.2
2014 Noise Robust Face Hallucination via Locality-Constrained Representation
abstract
Recently, position-patch based approaches have been proposed to replace the probabilistic graph-based or manifold learning-based models for face hallucination. In order to obtain the optimal weights of face hallucination, these approaches represent one image patch through other patches at the same position of training faces by employing least square estimation or sparse coding. However, they cannot provide unbiased approximations or satisfy rational priors, thus the obtained representation is not satisfactory. In this paper, we propose a simpler yet more effective scheme called Locality-constrained Representation (LcR). Compared with Least Square Representation (LSR) and Sparse Representation (SR), our scheme incorporates a locality constraint into the least square inversion problem to maintain locality and sparsity simultaneously. Our scheme is capable of capturing the non-linear manifold structure of image patch samples while exploiting the sparse property of the redundant data representation. Moreover, when the locality constraint is satisfied, face hallucination is robust to noise, a property that is desirable for video surveillance applications. A statistical analysis of the properties of LcR is given together with experimental results on some public face databases and surveillance images to show the superiority of our proposed scheme over state-of-the-art face hallucination approaches.
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002
IEEE Trans. Multim.2
2013 LBP-Guided Depth Image Filter
abstract
The multi-view video plus depth (MVD) format has been put forward for the call for proposals in free view video (FVV) and 3DTV. Since representing the 3D scene geometry, depth maps are used for synthesizing virtual views. However, compression artifacts of the depth images always lead to geometry distortions in synthesized views. By exploiting LBP features of the corresponding color samples, we propose a novel local binary pattern (LBP) guided depth filter which enables the local neighborhood samples those are in the same object of the current pixel to be filtering input. In recognition of its ability for describing the object edges, the LBP operator is used to calculate the weighted values of the local depth pixels for the depth-map filter. Furthermore, the filter is incorporated into the framework of H.264/MVC as an in-loop filter. The experimental results demonstrate that the proposed approach offers 0.45dB and 0.66dB average PSNR gains in terms of video rendering quality and depth coding efficiency, as well as significant subjective improvement in rendering views.
Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002
DCC2
2013 An expanded Mid/Side coding for 3D audio signal compression
abstract
Three dimensional (3D) audio technologies are booming with the success of 3D video technology. The sharply increased audio channels make its huge data unacceptable for transmitting bandwidth and storage media. This paper investigates the conventional Mid/Side (M/S) coding method, and expands it to a Three-channel Dependent M/S coding (3D-M/S) method. 3D-M/S perform sum and difference coding based on three channels instead of conventional two channels, and corresponding transform matrixes are presented. Furthermore, a framework is proposed to enable 3D-M/S compress any number of audio channels. Experiment shows proposed method obtains 25.4% objective quality improvement comparing with independent channel coding, and only increases 11.3% complexity comparing with the 29.9% of PCA method.
Shi Dong 0004, Ruimin Hu, Xiaochen Wang 0001, Weiping Tu
ICASSP2
2013 Manifold regularized sparse support regression for single image super-resolution
abstract
In this paper, we present a novel single image super-resolution method. To simultaneously improve the resolution and perceptual image quality, we bring forward a practical solution combining manifold regularization and sparse support regression. The main contribution of this paper is twofold. Firstly, a mapping function from low resolution (LR) patches to high-resolution (HR) patches will be learned by a local regression algorithm called sparse support regression, which can be constructed from the support bases of the LR-HR dictionary. Secondly, we propose to preserve the geometrical structure of the image patch dictionary, which is critical for reducing the artifacts and obtaining better visual quality. Experimental results demonstrate that the proposed method produces high quality results both quantitatively and perceptually.
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002, Shi Dong 0004
ICASSP2
2013 Face hallucination via weighted sparse representation
abstract
By incorporating the priors of image positions, position-patch based face hallucination methods can produce high-quality results and save computation time. These methods represent the test image patch as a linear combination of the same position patches in a training dictionary, and the key issue is how to obtain the optimal coefficients. Due to stability and accuracy issues, methods based on least square estimation or sparse representation (SR) proposed so far are not satisfactory. In this paper, we improve existing SR methods by exploiting similarity between the test and training patches. In particular, we impose a similarity constraint (in terms of the distance between the test patch and bases in the dictionary) on the ℓ1minimization regularization term and obtain the coefficients by solving a weighted SR problem. We also provide a new prospective on weighted SR and investigate its robustness to illumination variations. Experiments on commonly used database demonstrate that our method outperforms state of the art.
Zhongyuan Wang 0001, Junjun Jiang, Zixiang Xiong, Ruimin Hu
ICASSP4
2013 A joint learning based face hallucination approach for low quality face image
abstract
This paper describes a novel method for single-image super-resolution (SR) based on a neighbor embedding technique which uses coupled feature spaces under surveillance scenarios. For surveillance face images, traditional neighbor embedding SR approaches could not offer counterintuitive results because consistency between high resolution images and low resolution images is destroyed by serious noise which caused by environmental impact factors and large distance between the camera and objects. In order to reinforce the consistency, we extend the learning space from single to a coupled feature space that combine image intensity feature and contour model. The contour model describes facial contour information as images generated from original low resolution ones. Simulation experiments show that this proposed approach could provide competitive results in simulation experiments in subjective and objective quality. Even in surveillance scenario the proposed method outperforms the traditional methods.
Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Junjun Jiang
ICIP2
2013 Kinect depth map based enhancement for low light surveillance image
abstract
High noise level from darkness and low dynamic range are two characteristics of low light surveillance image that severely degrade the visual quality. Traditional low light image enhancement methods merely use the 2D cues without the depth information of the scene. Recently, the depth based image enhancement methods are proposed to enhance the depth perception of the image. However, these depth based methods are focus on the normal light image and only enhance the local depth perception. In this paper, based on the characteristics that the depth map captured by Kinect is less affected by low light condition than color image, we propose a Kinect depth based enhancement algorithm to enlarge the dynamic range and meanwhile to enhance the depth perception for the low light surveillance image. In our algorithm, firstly, the depth level similarity is incorporated into the non-local means denoising to remove the noises while better preserve object edges. Then, the depth aware contrast stretching is performed to enlarge the dynamic range and meanwhile to enhance both globe and local depth perception for low light surveillance image. Experimental results on low light surveillance images show that our proposed algorithm achieves better perceptual quality than previous work.
Ruimin Hu, Zhongyuan Wang 0001, Mang Duan
ICIP2
2013 Coupled-layer neighbor embedding for surveillance face hallucination
abstract
As the face image captured by a surveillance camera is typically very low-resolution (LR), blurred and noisy, traditional neighbor embedding method considers only one manifold (the LR image manifold) and fails very often to reliably estimate the intention geometrical structure. In this paper, we introduce the notion of neighbor embedding from the LR image manifold and the high-resolution (HR) one simultaneously and propose a novel neighbor embedding model, termed the coupled-layer neighbor embedding (CLNE), for surveillance face hallucination. CLNE differs substantially from other neighbor embedding models in that the former has two layers: the LR layer and the the HR layer. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights; the HR layer in this model is a set of HR training patches that guide the K-nearest neighbor (K-NN) searching and geometrically constrain the reconstruction weights. By this coupled constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust neighbor embedding through the significant degradation process. Indeed, the experimental results confirm that our method outperforms the related state-of-the-art methods by having better objective values as well as better visual results.
Junjun Jiang, Ruimin Hu, Liang Chen 0026, Zhen Han 0002, Tao Lu 0001, Jun Chen 0001
ICIP2
2013 Similarity preserving analysis based on sparse representation for image feature extraction and classification
abstract
Sparse representation has been a very active research area in recent years. Similarity analysis is an attractive research topic in the field of pattern recognition. In this paper, we take advantage of sparse representation in similarity analysis, and propose a novel unsupervised feature extraction approach, named similarity preserving analysis based on sparse representation (SPASR). SPASR projects samples from a high-dimensional space into a low-dimensional subspace, where the sparse reconstructive similarity relations among samples and the similarities of original samples and sparsely reconstructed samples are preserved. Experiments on the AR face database and COIL-20 object database demonstrate that the proposed SPASR approach outperforms several representative unsupervised subspace learning methods.
Qian Liu 0010, Xiaoyuan Jing, Ruimin Hu, Yong-Fang Yao, Jing-Yu Yang 0001
ICIP3
2013 Locality-constraint iterative neighbor embedding for face hallucination
abstract
Based on the assumption that low-resolution (LR) and high-resolution (HR) patch manifolds are locally isometric, the neighbor embedding based super-resolution algorithms try to preserve the local geometry of the patch manifold for the reconstructed HR patch manifold. However, due to “one-to-many” mappings between LR and HR images, the neighborhood relationship of the LR patch manifold can't reflect the inherent data structure. In this paper, we explore the data structure by both considering the LR patch and HR patch manifolds instead of only considering one manifold (LR patch manifold). By incorporating the position prior of face and local geometry of HR patch manifold, we propose an improved neighbor embedding method to face hallucination, namely locality-constraint iterative neighbor embedding (LINE), in which we iteratively update the K-nearest neighbors (K-NN) and reconstruction weights based on the result (the hallucinated HR patch) from previous iteration, giving rise to improved performance compared with traditional neighbor embedding algorithms. Experimental results with application to face hallucination on simulated LR face images and real world ones demonstrate the effectiveness of the proposed method.
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Tao Lu 0001, Jun Chen 0001
ICME2
2013 Bidirectional ranking for person re-identification
abstract
This paper proposes a simple but efficient bidirectional ranking method to improve person re-identification results across non-overlapping cameras. Previous methods treat person reidentification as a special object retrieval problem, and compute the final rank result purely based on a unidirectional matching between the probe and all gallery images. However, the expected person image may be excluded from the probe's ??-nearest neighbor due to appearance changes caused by variations in illuminations, poses, viewpoints and occlusion. To solve the above problem, our method queries every gallery image in a new gallery composed of the original probe image and other gallery images, and revises the initial query result in accordance with both content and context similarities between bidirectional ranking lists. A latent assumption of our method is that images of the same person should not only have similar visual content, known as content similarity, but also possess similar k-nearest neighbors, known as context similarity. Extensive experiments conducted on a series of standard data sets have validated the effectiveness of our proposed method with an average improvement of 5-10% over original baseline methods.
Qingming Leng, Ruimin Hu, Chao Liang 0001, Jun Chen 0001
ICME2
2013 Camera compensation using feature projection matrix for person re-identification
abstract
Matching individuals across a group of spatially non-overlapping surveillance cameras, also known as person re-identification, has recently attracted a lot of research interests. Current methods mainly focus on feature extraction or metric learning, which directly compare person images captured by different cameras, but seldom consider device differences caused by various surveillance conditions, e.g. view switching, scale zooming and illumination variation. Although brightness transfer function was proposed to address the problem of illumination variation, it could not handle view and scale changes among various cameras. In this paper, we propose an effective data-driven method to conquer device differences in the practical surveillance camera network. More precisely, with the help of a set of labelled pair-wise person images captured by two disjoint cameras, a feature projection matrix can be learned to project the person images of one camera to the feature space of the other camera, and thus images from these two different cameras can be accurately compared in a common feature space. Extensive comparative experiments conducted on three standard datasets have shown the promising prospect of our proposed methods.
Ruimin Hu, Chao Liang 0001, Chunjie Zhang 0001, Qingming Leng
ICME2
2013 Sound intensity and particle velocity based three-dimensional panning methods by five loudspeakers
abstract
In this paper, we present two new 3D panning methods. One method guarantees that time-average sound intensity and sound pressure of a virtual sound source at the receiving point are the same as time-averaged sound intensity and sound pressure of five loudspeakers at the receiving point. Another method maintains the direction of particle velocity and sound pressure. These methods can realize using five loudspeakers to replace a virtual sound source. These approaches relies on the assumption that the virtual sound source and five loudspeakers are on the same sphere and the virtual sound source needs to be in the area of a spherical pentagon that consists of the five loudspeakers. In the situation of five loudspeakers replacing a virtual source, these new methods do not need to undertake loudspeakers grouping. Compared with traditional 3D panning methods, these new methods are more convenient.
Song Wang 0011, Ruimin Hu, Yuhong Yang 0001
ICME2
2013 Face hallucination based on stepwise sparse reconstruction
abstract
Face hallucination methods based on low-resolution (LR) and high-resolution (HR) dictionary pair scheme infer HR patches by directly reusing coding coefficients trained by LR patches over LR dictionary. This scheme implies that LR and HR patch manifolds share highly similar local geometric structure. However, latest preliminary studies argue that the manifold assumption does not hold well such that face hallucination performance inevitably suffers from inconsistency of coding coefficients between LR and HR patches. In this paper, we are the first to observe that coding coefficients of LR patches are more relevant to latent those of HR patches under conditions of involving small magnifying factor. On the basis of this finding, we suggest a stepwise reconstruction scheme to minimize inconsistency risk in solution space. In particular, this scheme divides face hallucination process into multiple cascaded incremental training-synthesis steps, in which each individual step allows smaller magnifying factor as well as the corresponding intermediate resolution (IR) dictionary rather than merely LR and HR dictionary based learning. Moreover, in order to keep sparse representation (SR) sufficiently sparse while favoring its locality, we introduce a weighted ℓ1/ℓ2mixed norms minimization SR method and formulate a unified framework together with stepwise scheme. Experiments on commonly used face database demonstrate that our framework achieves state-of-the-art results.
Zhongyuan Wang 0001, Shizheng Wang, Ruimin Hu
ICME4
2013 Support-driven sparse coding for face hallucination
abstract
By incorporating the prior of positions, position patch based face hallucination methods can produce high-quality results and save computation time. Given a low-resolution face image, the key issue of these methods is how to encode the input low-resolution patch. However, due to stability and accuracy issues, the coding approaches proposed so far are not satisfactory. In this paper, we present a novel sparse coding method via exploiting the support information on the coding coefficients. In particular, the support information is characterized by the locality of the image patch manifold, which has been shown to be critical in data representation and analysis. According to the distances between the input patch and bases in the dictionary, we first assign different weights to the coding coefficients and then obtain the coding coefficients by solving a weighted sparse problem. Our proposed method exploits the non-linear manifold structure of patch samples and the sparse property of the redundant data, leading to stable and accurate representation. Experiments on commonly used databases demonstrate that our method outperforms state of the art.
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zixiang Xiong, Zhen Han 0002
ISCAS2
2013 Robust super-resolution for face images via principle component sparse representation and least squares regression
abstract
Face image super-resolution (SR) reconstruction is the problem of inducing a high-resolution (HR) face image from a low-resolution (LR) one. Traditional face SR methods are either sensitive to noise, i.e., local patch based technologies, or lacking facial details, i.e., global face reconstruction, thus could not achieve a satisfying result. In order to overcome these problems, we propose in this paper a novel face SR method. Taking full advantages of Principle Component analysis and Sparse Representation (PCSR), it aims to obtain an accurate and noise robust representation, transforming the image patch to the principle component sparse feature space (PC-SFS). Moreover, in PC-SFS, we try to learn a mapping function between the LR image patches and HR ones through Least Squares Regression. Given a LR patch, we first transform it to the LR PC-SFS by PCSR to obtain the robust and accurate representation, and then project the representation to the HR PC-SFS thus get the target HR patch. Experiments on the frontal faces SR in noise conditions demonstrate our method outperforms state of the art.
Tao Lu 0001, Ruimin Hu, Zhen Han 0002, Junjun Jiang
ISCAS2
2013 Color image guided locality regularized representation for Kinect depth holes filling
abstract
The emergence of Microsoft Kinect has attracted the attention not only from consumers but also from researchers in the field of computer vision. It facilitates the possibility to capture the depth map of the scene in real time and with low cost. Nonetheless, due to the limitations of structured light measurements used by Kinect, the captured depth map suffers random depth missing in the occlusion or smooth regions, which affects the accuracy of many Kinect based applications. In order to fill in the holes existing in Kinect depth map, some approaches that adopted color image guided in-painting or joint bilateral filter have been proposed to represent the missing depth pixel by available depth pixels. However, they are not able to obtain the optimal weights, thus the obtained missing depth values are not best. In this paper, we propose a color image guided locality regularized representation (CGLRR) to reconstruct the missing depth pixels by comprehensively determining the optimal weights of the available depth pixels from collocated patches in color image. Experimental results demonstrate that the proposed algorithm can better fill in the holes of depth map both in smooth and edge region than previous works.
Ruimin Hu, Zhongyuan Wang 0001, Mang Duan
VCIP2
2013 From local representation to global face hallucination: A novel super-resolution method by nonnegative feature transformation
abstract
Most of global face hallucination methods treat the face as a whole, ignoring the fact that the face is composed by part-based organs. Therefore, the results obtained by these methods always lack of detailed information. Nonnegative matrix factorization (NMF) based face hallucination method is properly used to enhance the detailed information. Usually, NMF basis is only learnt from high-resolution (HR) samples, leading to over-smooth output and lack of high frequency details. In order to solve this problem, we propose a simple but novel face hallucination method using nonnegative feature transformation by two-step framework. In particular, we learn the NMF basis from low-resolution (LR) and HR samples separately, and then transform the local representation feature of input into the global representation subspaces, keeping the weights into the HR samples space for output. Furthermore, the maximum a posteriori (MAP) method is used to estimate a better output. Experiments show that the hallucinated face of the proposed method is not only more high-frequency details, but also has better performance than many state-of-art algorithms.
Tao Lu 0001, Ruimin Hu, Zhen Han 0002, Junjun Jiang, Yanduo Zhang
VCIP2
2013 Surveillance video synopsis in the compressed domain for fast video browsing
Shizheng Wang, Zhongyuan Wang 0001, Ruimin Hu
J. Vis. Commun. Image Represent.3
2013 Background Subtraction With Video Coding
abstract
The classic Gaussian mixture model is based on the statistical information of every pixel; it is not robust to light changes. Before analysing every pixel in videos, it must be decoded to raw videos. In this letter, the method combining video coding and the Gaussian mixture model together is proposed. We use intra mode and motion vectors to find the foreground macroblock, then add one overhead flag in the compressed video to indicate it. In the decoder, we just decode possible foreground areas and detect moving objects in these areas. In our experiments, we test this method on two datasets, both of them with unique, dynamic, illumination conditions. Results show that the proposed method is effective to detect moving objects and easily assemble to current automated video surveillance systems.
Zhenkun Huang, Ruimin Hu, Zhongyuan Wang 0001
IEEE Signal Process. Lett.2
2012 A super-resolution method for low-quality face image through RBF-PLS regression and neighbor embedding
abstract
In this paper, a new two-step method is proposed to infer a high-quality and high-resolution (HR) face image from a low-quality and low-resolution (LR) observation based on training samples in the database. First, a global face image is reconstructed based on the non-linear relationship between LR and HR face images, which is established according to radial basis function and partial least squares (RBF-PLS) regression. Based on the reconstructed global face patches manifold (formed by the image patches at the same position of all global face images), whose local geometry is more consistent with that of original HR face patches manifold than noisy LR one is, the Neighbor Embedding is applied to induce the target HR face image by preserving the similar local geometry between global face patches manifold and the original HR face patches manifold. A comparison of some state-of-the-art methods shows the superiority of our method, and experiments also demonstrate the effectiveness both under simulation and real conditions.
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang
ICASSP2
2012 Graph discriminant analysis on multi-manifold (GDAMM): A novel super-resolution method for face recognition
abstract
How to efficiently recognize low-resolution (LR) probe images of one face recognition system, in which high-resolution (HR) gallery of faces is enrolled, is still an open problem. In this paper, we develop a novel super-resolution method, namely Graph Discriminant Analysis on Multi-Manifold (GDAMM), to super-resolved the HR version of a LR probe image and then perform matching at the resolution of the HR gallery. Unlike classical super-resolution approaches considering only the data fidelity, GDAMM takes the advantages of both manifold learning and discriminant analysis to integrate the data constraint and discriminant constraint, seeking the mapping between LR images and HR ones. In the reconstructed HR image space, faces of one person in the same manifold are close and those in different manifolds are far apart. Experiments on Extended Yale-B database and AR face database demonstrate that the learned discriminant information is essential for improving recognition accuracy. Through the contrastive experiment, the results (recognition rates) indicate that the proposed GDAMM method can greatly surpass classical super-resolution approaches, even outperforming the ideal case of having probe images of HR gallery by a big margin (nearly 9% on Extended Yale-B database and 8% on AR face database).
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Kebin Huang, Tao Lu 0001
ICIP2
2012 Enhanced Principal Component Using Polar Coordinate PCA for Stereo Audio Coding
abstract
High efficiency audio compression is the basic technology in audio involved multimedia application. Down mixing and parametric coding are efficient coding scheme with widely applications in some up to date audio codecs such as PS in EAAC+ and MPEG-Surround, and PCA stereo coding followed this idea to map two channels to one channel with maximum energy and parameterize the secondary channel. This paper investigates the conventional PCA method performance under general stereo model with multiple sound sources and different directions, and then proposes a Polar Coordinate based PCA (PC-PCA) stereo coding method. It has been proved that when multiple sound sources exist with different directions, proposed method is better than the conventional PCA method in certain conditions. A stereo codec based on PC-PCA has also been proposed to validate the performance improvement of proposed method.
Shi Dong 0004, Ruimin Hu, Weiping Tu, Junjun Jiang, Song Wang 0011
ICME2
2012 Efficient Single Image Super-Resolution via Graph Embedding
abstract
We explore in this paper efficient algorithmic solutions to single image super-resolution (SR). We propose the GESR, namely Graph Embedding Super-Resolution, to super-resolve a high-resolution (HR) image from a single low-resolution (LR) observation. The basic idea of GESR is to learn a projection matrix mapping the LR image patch to the HR image patch space while preserving the intrinsic geometrical structure of original HR image patch manifold. While GESR resembles other manifold learning-based SR methods in persevering the local geometric structure of HR and LR image patch manifold, the innovation of GESR lies in that it preserves the intrinsic geometrical structure of original HR image patch manifold rather than LR image patch manifold, which may be contaminated because of image degeneration (e.g., blurring, down-sampling and noise). Experiments on benchmark test images show that GESR can achieve very competitive performance as Neighbor Embedding based SR (NESR) and Sparse representation based SR (SSR). Beyond subjective and objective evaluation, all experiments show that GESR is much faster than both NESR and SSR.
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Kebin Huang, Tao Lu 0001
ICME2
2012 Position-Patch Based Face Hallucination via Locality-Constrained Representation
abstract
Instead of using probabilistic graph based or manifold learning based models, some approaches based on position-patch have been proposed for face hallucination recently. In order to obtain the optimal weights for face hallucination, they represent image patches through those patches at the same position of training face images by employing least square estimation or convex optimization. However, they can hope neither to provide unbiased solutions nor to satisfy locality conditions, thus the obtained patch representation is not the best. In this paper, a simpler but more effective representation scheme- Locality-constrained Representation (LcR) has been developed, compared with the Least Square Representation (LSR) and Sparse Representation (SR). It imposes a locality constraint onto the least square inversion problem to reach sparsity and locality simultaneously. Experimental results demonstrate the superiority of the proposed method over some state-of-the-art face hallucination approaches.
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang
ICME2
2012 Improvements of dynamic texture synthesis for video coding
Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002
ICPR2
2012 Face hallucination via K-selection mean constrained sparse representation
Kebin Huang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Junjun Jiang
ICPR2
2012 Surveillance face hallucination via variable selection and manifold learning
abstract
In this paper, we propose a new two-step face hallucination method to induce a high-resolution (HR) face image from a low-resolution (LR) observation. Especially for low-quality surveillance face image, an RBF-PLS based variable selection method is presented for the reconstruction of global face image. Further more, in order to compensate for the reconstruction errors, which are lost high frequency detailed face features, the Neighbor Embedding (NE) based residue face hallucination algorithm is used. Compared with current methods, the proposed RBF-PLS based method can generate a global face more similar to the original face and less sensitive to noise, moreover, the NE algorithm can reduce the reconstruction errors caused by misalignment on the basis of a carefully designed search strategy. Experiments show the superiority of the proposed method compared with some state-of-the-art approaches and the efficacy both in simulation and real surveillance condition.
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang
ISCAS2
2012 Face image super-resolution via nearest feature line
abstract
In this paper, we propose a manifold learning based algorithm using 'Nearest Feature Line - NFL' to hallucinate high-resolution face image. According to the fact that existing NFL can effectively characterize the geometrical proportions to the face samples, we propose using NFL metric to define the neighborhood relations between face samples. Our algorithm can solve the problem that traditional method cannot effectively reveal the similar local geometry between high-resolution and low-resolution face manifolds under the condition that the training sample size is small. Moreover, in order to enhance the representation capacity of available face samples and reduce the computational complexity, we select neighborhood samples for each input LR image. Experimental results demonstrate that our algorithm can generates clearer local feature details, and the PSNR is 1.4 dB higher than that of the best manifold learning based method reported so far.
Zhen Han 0002, Junjun Jiang, Ruimin Hu, Tao Lu 0001, Kebin Huang
ACM Multimedia3
2012 Intracoding and Refresh With Compression-Oriented Video Epitomic Priors
abstract
In video compression, intracoding plays an important role in terms of coding efficiency and error resilience and has been an attractive research topic since the standardization of H.264/AVC. In this paper, we propose a high-performance intracoding scheme with the help of epitomic priors. Different from intracoding in H.264/AVC and other video standards, we construct image epitomes as coding priors and use them to generate predictions of intrablocks at the encoder. In addition, we losslessly code and transmit the image epitomes to the decoder. We perform compression-oriented video epitomic analysis and search for the best epitomic priors by using the expectation maximization algorithm. The resulting image epitomes for a video sequence can be viewed as the base layer in spatially scalable video coding. Experiments show that our proposed intracoding scheme improves the state of the art by an average of 0.53 dB in PSNR. Simulations under a packet loss environment also demonstrate that intrarefresh with epitomic priors outperforms random intrarefresh by up to 2 dB, leading to better subjective quality.
Qijun Wang, Ruimin Hu, Zhongyuan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2011 Camera Model Identification Based on the Characteristic of CFA and Interpolation
Guanshuo Xu, Ruimin Hu
IWDW3
2011 Audio steganalysis of spread spectrum information hiding based on statistical moment and distance metric
Ruimin Hu, Haojun Ai
Multim. Tools Appl.2
2010 A Novel Frame Error Concealment Algorithm Based on Dynamic Texture Synthesis
abstract
Dynamic textures are sequences of frames of moving scenes that exhibit certain stationary properties in time, which have great significance in applications such as video production, virtual simulation and virtual walkthroughs. This paper presents an algorithm for dynamic texture extrapolation using for H.264 decoding system. The synthesized frames can be used by the decoder for whole frames loss error concealment. The simulation results show that the proposed whole frames loss error concealment algorithm achieves significant improvement over the motion vector extrapolation method.
Ruimin Hu, Dan Mao, Zhongyuan Wang 0001
DCC2
2010 Spatially Scalable Video Coding Based on Hybrid Epitomic Resizing
abstract
Scalable video coding (SVC) is considered as a potentially promising solution to enable the adaptability of video to heterogonous networks and various devices. In spatially scalable video encoder, how to resize the captured high-resolution video to get low-resolution video has great effect on the quality of experience (QoE) in the clients receiving low-resolution video. In this paper, we propose a new resizing algorithm called hybrid epitomic resizing (HER), which can make the resized image preserve the same ‘physical’ resolution with original image by the way of utilizing texture similarity inside image and highlight regions of interest while avoiding potential artifacts. For hybrid epitomic resizing, we also design two new inter-layer prediction methods to eliminate the redundancy between adjacent spatial layers instead of conventional inter-layer prediction. Experimental results show that HER can get resized images with perceptually much better quality and the performance of new inter-layer prediction are comparable to that of conventional inter-layer prediction in H.264 SVC.
Qijun Wang, Ruimin Hu, Zhongyuan Wang 0001
DCC2
2010 Spatial audio cues based surveillance audio attention model
abstract
In this paper, we propose a bottom-up audio attention model based on spatial audio cues for unsupervised event detecting in stereo audio surveillance. Firstly, the spatial audio parameter Interaural Level Difference (ILD) is extracted to calculate and represent the attention events, which are caused by rapid moving sound source. Then an environment adaptive normalization is used to assess the normalized attention level. Experimental results demonstrate that our proposed audio attention model is effective for audio surveillance event detection.
Bo Hang, Ruimin Hu
ICASSP2
2010 A face super-resolution approach using shape semantic mode regularization
abstract
In actual imaging environment, a variety of factors have an impact on the quality of images, which leads to pixel distortion and aliasing. The traditional face super-resolution algorithm only uses the difference of image pixel values as similarity criterion, which degrades similarity and identification of reconstructed facial images. Image semantic information with human understanding, especially structural information, is robust to the degraded pixel values. In this paper, we propose a face super-resolution approach using shape semantic model. This method describes the facial shape as a series of fiducial points on facial image. And shape semantic information of input image is obtained manually. Then a shape semantic regularization is added to the original objective function. The steepest descent method is used to obtain the unified coefficient. Experimental results demonstrate that the proposed method outperforms the traditional schemes significantly both in subjective and objective quality.
Chengdong Lan, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001
ICIP2
2010 A novel method for generation of motion saliency
abstract
Motion saliency is the key component for the video saliency model, and attracts great research interest. However, there is few universal predictor of motion saliency. In this paper, a novel method for generation of motion saliency is proposed, in which motion saliency map is obtained through the multi reference frames, and enhanced by spatial saliency information. The proposal can obtain more detailed information about motion feature and extract the salient object more integrally. The experiment results shows that our proposal can achieve 95.4% of the ROC area of a human based control in motion channel, whereas the classical Itti's model achieves 78.5%.
Ruimin Hu, Zhenkun Huang, Yin Su
ICIP2
2010 Temporal color Just Noticeable Distortion model and its application for video coding
abstract
Just Noticeable Distortion (JND), which is utilized to reduce the bit rate without introducing noticeable visual distortion, plays an important role in perceptual image and video processing. For the temporal color JND, it takes into account not only the spatial and luminance HVS properties, but also the temporal and chroma HVS properties. In this paper, we first develop a spatio-temporal model estimating JND for color video by completely incorporating the color CSF, the frequency property of DCT coefficient, the contrast masking effect and the motion property. Then we incorporate the JND model into video encoding system via the residue filtering process. This method can work with any prevalent video coding standards. To demonstrate the effectiveness of the JND model, we has implemented it into the H.264/AVC reference software JM12.4 and the experimental results show that the bit rate can be reduced by average 18.20%, which reflects that our JND model is able to exploit the HVS bounds more aggressively without introducing noticeable visual distortions.
Ruimin Hu, Zhongyuan Wang 0001
ICME2
2010 Video coding using dynamic texture synthesis
abstract
In video coding system, because of the delay between current picture and the reference pictures, the temporal prediction is not good for the video sequences with nonlinear motion and global illumination change between frames. Dynamic textures are sequences of frames of moving scenes that exhibit certain stationary properties in time, which have great significance in applications such as video production and virtual simulation. This paper presents a new algorithm for dynamic texture extrapolation using for H.264 encoding and decoding system. The synthesized frames are used by the encoder for virtual reference frames choice in inter prediction and decoder for whole frames loss error concealment. The simulation results show that the proposed virtual reference frames choice algorithm improves encoding efficiency compared with the H.264 standard, and the proposed whole frames loss error concealment algorithm achieves significant improvement over the motion vector extrapolation method.
Ruimin Hu, Dan Mao, Zhongyuan Wang 0001
ICME2
2010 A bottom-up audio attention model for surveillance
abstract
This paper proposes a bottom-up audio attention model based on spatial audio cues and sub band energy change for unsupervised event detection in stereo audio surveillance. Firstly, the spatial audio parameter Interaural Level Difference (ILD) is extracted to calculate and represent the attention events, which are caused by rapid moving sound source. Then the sub band energy change is computed to present the salient energy distribution change in frequency domain. At last, an environment adaptive normalization is used to assess the normalized attention level. Experimental results demonstrate that the proposed audio attention model is effective for audio surveillance event detection.
Ruimin Hu, Bo Hang, Shi Dong 0004
ICME1
2010 Intra coding and refresh based on video epitomic analysis
abstract
In video coding, intra coding plays an important role in both coding efficiency and error resilience. In this paper, a new intra coding method based on video epitomic analysis is proposed. Image epitome suitable for coding is extracted through video epitomic analysis, and is used to generate the prediction for each block. The mapping between the original image and image epitome as well as image epitome should be coded and transmitted to the decoder side. Due to the independence of adjacent reconstructed blocks, the proposed intra coding method can also be applied to improve error resilience of intra refresh. Experimental results show that, the proposed intra coding method can significantly improve the performance of intra coding by 0.7dB in terms of PSNR. The simulation results under packet loss environment show that the proposed intra refresh can out-perform random intra refresh (RIR) by up to 1dB, and better subjective quality can also be obtained.
Qijun Wang, Ruimin Hu, Zhongyuan Wang 0001, Bo Hang
ICME2
2010 Global Face Super Resolution and Contour Region Constraints
Chengdong Lan, Ruimin Hu, Tao Lu 0001, Ding Luo, Zhen Han 0002
ISNN (2)2
2010 Face hallucination with shape parameters projection constraint
abstract
In real surveillance scenarios, a variety of factors have an impact on the quality of images, which leads to pixel distortion and aliasing. Traditional face super-resolution algorithms only use the difference of image pixel values as similarity criterion, which degrades similarity and identification of reconstructed facial images. Image semantic information with human understanding, especially structural data of shapes, is robust to the degraded images. In this paper, we propose a face hallucination with shape parameters projection constraint. This method uses a parameter model to represent face shapes, and shape information of input image is introduced to improving the quality of reconstructed image. The shape model regularization is first added to original objective function. Then shape parameters are projected into the domain of image parameters by a linear regression model. Finally, the gradient descent method is used to obtain the unified parameters. Experimental results demonstrate the proposed method outperforms the traditional schemes significantly both in subjective and objective quality.
Chengdong Lan, Ruimin Hu, Kebin Huang, Zhen Han 0002
ACM Multimedia2
2010 Inter prediction based on spatio-temporal adaptive localized learning model
abstract
Inter prediction based on block matching motion estimation is important for video coding. But this method suffers from the additional overhead in data rate representing the motion information that needs to be transmitted to the decoder. To solve this problem, we present an improved implicit motion information inter prediction algorithm for P slice in H.264/AVC based on the spatio-temporal adaptive localized learning (STALL) model. According to 4 × 4 block transform structure in H.264/AVC, we first adaptively choose nine spatial neighbors and nine temporal neighbors, and a localized 3D casual cube is designed as training window. By using these information, the model parameters could be adaptively computed based on the Least Square Prediction (LSP) method. Finally, we add a new inter prediction mode into H.264/AVC standard for P slice. The experimental results show that our algorithm improves encoding efficiency compared with H.264/AVC standard, with relatively increases in complexity.
Ruimin Hu, Zhongyuan Wang 0001
PCS2
2010 Spatial parameters for audio coding: MDCT domain analysis and synthesis
Shuixian Chen, Naixue Xiong, Jong Hyuk Park 0001, Min Chen 0003, Ruimin Hu
Multim. Tools Appl.5
2009 Estimating spatial cues for audio coding in MDCT domain
abstract
Although widely used otherwise, MDCT is excluded in the current scheme for spatial cues representation, due to its lacking of phase information and energy conservation. But combining MDCT with MDST overcomes the difficulties. Moreover, MDST spectra can be built perfectly from neighboring MDCT spectra. The MDCT-MDST conversion, in matrix form, is approximating to a banded sparse matrix. When applied to spatial audio coding using MDCT based core coders, this method avoids separate transforming for cues representation and saves significant computation. Listening tests also show that it has same audio quality as other complex transform based methods.
Shuixian Chen, Ruimin Hu
ICME2
2009 Improved Object Tracking Algorithm Based on New HSV Color Probability Model
Gang Tian, Ruimin Hu, Zhongyuan Wang 0001, Youming Fu
ISNN (2)2
2009 Camera-Model Identification Using Markovian Transition Probability Matrix
Guanshuo Xu, Yun Q. Shi 0001, Ruimin Hu, Wei Su 0001
IWDW4
2008 Speech technology in real world environment: early results from a long term study
abstract
Existing knowledge on how people use speech-based technologies in realistic settings is limited. We are conducting a longitudinal field study, spanning six months, to investigate how users with no physical impairments and users with upper body physical impairments use speech technologies when interacting with computers in their home environment. Digital data logs, time diaries, and interviews are being used to record the types of applications used, frequency of use of each application, and difficulties experienced as well as subjective data regarding the usage experience. While confirming many expectations, initial results have provided several unexpected insights including a preference to use speech for navigation instead of dictation tasks, and the use of speech technology for programming and games.
Jinjuan Feng, Shaojian Zhu, Ruimin Hu, Andrew Sears
ASSETS3
2007 A Novel Intra/Inter Mode Decision Algorithm for H.264/AVC Based on Spatio-temporal Correlation
Qiong Liu 0001, Shengfeng Ye, Ruimin Hu, Zhen Han 0002
MMM (1)3
2006 Introduction to AVS Audio
Haojun Ai, Shuixian Chen, Ruimin Hu
J. Comput. Sci. Technol.3
2005 AVS Generic Audio Coding
abstract
AVS Audio Coding Standard is the first standard for Hi-Fi audio in China. The framework of AVS Audio was introduced. Many key technologies are described in details, including long/short window switch decision based on energy and unpredictability, integer MDCT for lossless time-frequency transform, square polar stereo coding, and context-dependent bitplane coding for scalable entropy coding. The informal subject test result is given between AVS audio codec and several dominating audio codecs. It is shown that AVS audio codec is enough for Hi-Fi audio applications.
Ruimin Hu, Shuixian Chen, Haojun Ai, Naixue Xiong
PDCAT1