Min Xu 0001

dblp:09/0-1 · DBLP profile ↗
← Back
148ranked-venue papers
14as first author
61since 2021 · last 2026
0000-0001-9581-8849ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 84 · 14 first-author · 21 since 2021Artificial intelligence and machine learning · 52 · 31 since 2021Computer networks · 11 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Security and privacy · 2
YearPublicationVenuePosition
2026 PromptHG: Prompt-Enhanced Heterogeneous Graph for Personalized News Recommendation
Hai-Dang Kieu, Delvin Ce Zhang, Qiang Wu 0001, Min Xu 0001, Dung D. Le
ECIR (1)5
2026 MEGG: replay via maximally extreme GGscore in incremental learning for neural recommendation models
abstract
Abstract Neural collaborative filtering (NCF)-based recommendation models have been widely adopted in practical recommender systems due to their effectiveness. However, these models are typically developed under the static deep learning paradigm, where training is conducted on fixed datasets with the implicit assumption of a static data distribution. This approach is ill-suited for dynamic environments, such as those encountered in real-world platforms, where user preferences and collaborative filtering patterns evolve continuously. To address this limitation, incremental learning-a paradigm designed to integrate new knowledge while preserving previously learned information-emerges as a promising alternative. Despite its potential, the direct application of conventional incremental learning methods, which are prevalent in domains like computer vision and natural language processing, is hindered by unique challenges in recommender systems. These include the distinct task paradigm, data complexity, and sparsity issues. Moreover, existing incremental learning approaches tailored for neural recommendation models remain scarce and often suffer from limited generalizability. To bridge this gap, we propose an innovative experience replay-based incremental learning framework specifically designed for neural recommendation models, termed Replay Samples with Maximally Extreme GGscore (MEGG). At the core of MEGG is a novel metric, the GGscore, which quantifies the influence of individual samples on model training. By selectively replaying samples with the most extreme GGscores, our method effectively mitigates catastrophic forgetting, thereby maintaining high predictive performance over time. A key advantage of MEGG lies in its data-centric nature, which renders it agnostic to the underlying model architecture. This ensures broad applicability across various neural recommendation models and seamless integration with existing incremental learning frameworks to further enhance performance. Extensive experiments conducted on three neural recommendation models across four benchmark datasets demonstrate the superior effectiveness of MEGG compared to state-of-the-art methods. Furthermore, additional evaluations highlight its scalability, efficiency, and robustness. The implementation of MEGG will be made publicly available upon acceptance.
Yunxiao Shi, Shuo Yang 0006, Haimin Zhang 0001, Li Wang 0064, Yongze Wang, Qiang Wu 0001, Min Xu 0001
Data Min. Knowl. Discov.7
2026 Probing, priors, and teaching: A framework to segment any bone with partial supervision
Tianyou Liang, Min Xu 0001
Expert Syst. Appl.4
2026 FedPCL-CDR: A federated prototype-based contrastive learning framework for privacy-preserving cross-domain recommendation
Li Wang 0064, Qiang Wu 0001, Min Xu 0001
Neural Networks3
2026 Noisy Correspondence Rectification in Multimodal Clustering Space for Cross-Modal Matching
abstract
As one of the most fundamental techniques in multimodal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose BiCro++ (Improved Bidirectional Cross-modal Similarity Consistency). This module can be integrated into existing cross-modal matching models, enhancing their robustness against noisy data through self-adaptive soft labels that dynamically reflect the true correspondence of data pairs. The basic idea of BiCro++ is motivated by that - taking image-text matching as an example - similar images should have similar textual descriptions and vice versa. This bidirectional similarity consistency can be directly translated into soft labels as a self-supervision signal to train the matching model. To further refine soft label quality, BiCro++ first introduces a Diagonal-Dominance Purification process to identify reliable anchor points from noisy dataset as the reference for soft label estimation. Then it employs a Hybrid-level Codebook Alignment mechanism that establishes enhanced consistency in bidirectional cross-modal similarity. The experiments on three popular cross-modal matching datasets show that our method significantly improves the noise-robustness of various matching models, and surpasses the state-of-the-art method by an average of 5.3%, 3.1% and 6.4% in terms of recall, respectively.
Shuo Yang 0006, Yancheng Long, Zeke Xie, Hongxun Yao, Min Xu 0001, Liqiang Nie
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Differential Encoding for Improved Representation Learning Over Graphs
abstract
Combining the message-passing paradigm with the global attention mechanism has emerged as an effective framework for learning over graphs. The message-passing paradigm and the global attention mechanism basically generate embeddings of nodes by taking the sum of information from a node's local neighbourhood and from the entire graph, respectively. However, this simple summation aggregation approach fails to distinguish between the information from a node itself or from the node's neighbours. Therefore, there exists information lost at each layer of embedding generation, and this information lost could be accumulated and become more serious in deeper model layers. In this paper, we present a differential encoding method to address the issue of information lost. Instead of simply taking the sum to aggregate local or global information, we explicitly encode the difference between the information from a node itself and that from the node's local neighbours (or from the rest of the entire graph nodes). The obtained differential encoding is then combined with the original aggregated representation to generate the updated node embedding. By combining differential encodings, the representational ability of generated node embeddings is improved, and therefore the model performance is improved. The differential encoding method is empirically evaluated on different graph tasks on seven benchmark datasets. The results show that it is a general method that improves the message-passing update and the global attention update, advancing the state-of-the-art performance for graph representation learning on these benchmark datasets.
Haimin Zhang 0001, Jiahao Xia 0001, Min Xu 0001
IEEE Trans. Big Data3
2026 It Takes Two: Multi-Frequency Perception With Complementary Fusion Network for Complex Scene Segmentation
abstract
Complex scene segmentation aims to segment objects with intricate details or those concealed within the background. Despite significant advancements, a persistent challenge remains: accurately identifying object edges in backgrounds with high inherent similarity and complex structures. To address this, we identify the prevalent spectral bias in image segmentation, where networks preferentially learn low-frequency information, as a key impediment to recognizing and learning object edges, which are rich in high-frequency details. To mitigate this bias, we propose MCNet, a segmentation framework designed to promote balanced frequency learning. MCNet comprises two primary components: multi-frequency perception (MP), which independently captures high-frequency details and low-frequency structural components of objects, and complementary fusion (CF), which intelligently fuses these distinct frequency features through learnable, adaptive mechanisms. Crucially, MCNet employs a novel frequency-aware consistency adversarial loss to explicitly guide the learning across different frequency bands. MCNet effectively integrates MP and CF, enhancing the detection of high-frequency details and low-frequency structures, thereby alleviating challenges posed by spectral bias. We evaluate the proposed method on complex scene segmentation tasks, including camouflaged object detection and dichotomous image segmentation. Through extensive comparisons with 31 existing methods across 8 benchmark datasets, we demonstrate the superiority of the proposed method.
Jin Zhang 0021, Ruiheng Zhang 0001, Zhe Cao 0001, Lixin Xu 0001, Xi Chen 0090, Min Xu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 StructGS: Adaptive Spherical Harmonics and Rendering Enhancements for Superior 3D Gaussian Splatting
abstract
Recent advancements in 3D reconstruction coupled with neural rendering techniques have greatly improved the creation of photo-realistic 3D scenes, influencing both academic research and industry applications. The technique of 3D Gaussian Splatting and its variants incorporate the strengths of both primitive-based and volumetric representations, achieving superior rendering quality. While 3D Geometric Scattering (3DGS) and its variants have advanced the field of 3D representation, they fall short in capturing the stochastic properties of non-local structural information during the training process. Additionally, the initialisation of spherical functions in 3DGS-based methods often fails to engage higher-order terms in early training rounds, leading to unnecessary computational overhead as training progresses. Furthermore, current 3DGS-based approaches require training on higher resolution images to render higher resolution outputs, significantly increasing memory demands and prolonging training durations. We introduce StructGS, a framework that enhances 3D Gaussian Splatting (3DGS) for improved novel-view synthesis in 3D reconstruction. StructGS innovatively incorporates a patch-based SSIM loss, dynamic spherical harmonics initialisation and a Multi-scale Residual Network (MSRN) to address the above-mentioned limitations, respectively. Our framework significantly reduces computational redundancy, enhances detail capture and supports high-resolution rendering from low-resolution inputs. Experimentally, StructGS demonstrates superior performance over state-of-the-art (SOTA) models, achieving higher quality and more detailed renderings with fewer artifacts. (The link to the code will be made available after publication.).
Zexu Huang, Min Xu 0001, Stuart W. Perry
IEEE Trans. Multim.2
2026 Classifier Enhancement Using Extended Context and Domain Experts for Semantic Segmentation
abstract
Prevalent semantic segmentation methods generally adopt a vanilla classifier to categorize each pixel into specific classes. Although such a classifier learns global information from the training data, this information is represented by a set of fixed parameters (weights and biases). However, each image has a different class distribution, which prevents the classifier from addressing the unique characteristics of individual images. At the dataset level, class imbalance leads to segmentation results being biased towards majority classes, limiting the model's effectiveness in identifying and segmenting minority class regions. In this paper, we propose an Extended Context-Aware Classifier (ECAC) that dynamically adjusts the classifier using global (dataset-level) and local (image-level) contextual information. Specifically, we leverage a memory bank to learn dataset-level contextual information of each class, incorporating the class-specific contextual information from the current image to improve the classifier for precise pixel labeling. Additionally, a teacher-student network paradigm is adopted, where the domain expert (teacher network) dynamically adjusts contextual information with ground truth and transfers knowledge to the student network. Comprehensive experiments illustrate that the proposed ECAC can achieve state-of-the-art performance across several datasets, including ADE20K, COCO-Stuff10K, and Pascal-Context.
Huadong Tang, Youpeng Zhao 0002, Min Xu 0001, Jun Wang 0001, Qiang Wu 0001
IEEE Trans. Multim.3
2026 A Multi-Modal Prompt-Tuning Framework for Non-Overlapping Multi-Domain Recommendation
abstract
Cross-domain recommendation (CDR) aims to enhance recommendation accuracy in data-sparse domains by transferring knowledge from data-rich domains. Most existing CDR methods conduct knowledge transfer based on overlapping users or items to address the user cold-start problems, including few-shot (i.e., users with sparse interactions) and zero-shot (i.e., users with no interactions) scenarios. However, in real-world scenarios, such overlap is often sparse or non-existent, limiting the effectiveness of these approaches. To overcome this challenge, we propose a novelMulti-modalPrompt-tuningFramework forNon-overlappingMulti-DomainRecommendation (MPF-NMDR). MPF-NMDR transfers knowledge across non-overlapping domains, enhancing recommendation performance in both few-shot and zero-shot scenarios. Specifically, we first pre-train the MPF-NMDR framework on data from all domains to capture users' generalized cross-domain preferences, which are learned through the generalized multi-modal interest mining module. We then conduct prompt-tuning with domain, user, and item prompts in the target domain to capture distinctions among various domains, users, and items. In this process, only the prompt parameters are fine-tuned, while all other parameters remain frozen, enabling the model to capture the distinctions among domains, users, and items while preserving the cross-domain knowledge. Extensive experiments on Amazon and Douban review datasets validate the superior performance of MPF-NMDR compared to SOTA baselines. We release our code athttps://github.com/Lili1013/MPF_NMDR.
Li Wang 0064, Shoujin Wang, Qiang Wu 0001, Min Xu 0001
IEEE Trans. Multim.4
2026 Beyond KAN: Introducing KarSein for Adaptive High-Order Feature Interaction Modeling in CTR Prediction
abstract
Modeling high-order feature interactions is crucial for Click-Through Rate (CTR) prediction, yet traditional approaches typically predefine a maximum interaction order and exhaustively enumerate feature combinations up to that order. This paradigm depends heavily on prior domain knowledge to delimit the interaction space and incurs substantial computational overhead. As a result, conventional CTR models face a persistent tension between enriching representations with complex high-order interactions and keeping computation tractable. To address this dual challenge, this study introduces the Kolmogorov–Arnold Represented Sparse Efficient Interaction Network (KarSein). Drawing inspiration from the learnable activation mechanism in the Kolmogorov–Arnold Network (KAN), KarSein leverages this mechanism to adaptively transform low-order basic features into high-order feature interactions, offering a novel approach to feature interaction modeling. KarSein extends the capabilities of KAN by introducing a more efficient architecture that significantly reduces computational costs while accommodating 2D embedding vectors as feature inputs. Furthermore, it overcomes the limitation of KAN’s its inability to spontaneously capture multiplicative relationships among features. Extensive experiments highlight the superiority of KarSein, demonstrating its ability to surpass not only the vanilla implementation of KAN in CTR prediction tasks but also other baseline methods. Remarkably, KarSein achieves exceptional predictive accuracy while maintaining a highly compact parameter size and minimal computational overhead. Moreover, KarSein retains the key advantages of KAN, such as strong interpretability and structural sparsity. As the first systematic adaptation of KAN to CTR prediction, KarSein offers a practical, parameter-efficient, and interpretable alternative for modeling complex feature interactions in large-scale recommendation systems.
Yunxiao Shi, Wujiang Xu, Haimin Zhang 0001, Qiang Wu 0001, Min Xu 0001
ACM Trans. Inf. Syst.5
2025 Segment Any Bone in CT with Partial Supervision
abstract
Automatic bone segmentation is a fundamental task supporting various clinical practices. Conventional methods in this field rely heavily on dense annotations, which incurs substantial labeling labor and expense. Recent efforts have been made to reduce the labeling workload through semi-supervised and weakly supervised learning. However, methods under these two paradigms usually assume that all objects of interest e.g., bones, are covered by labels. In this work, we explore a less studied problem setting that assumes only partially labeled bone CT data. To tackle the supervision bias brought by incomplete annotations, we design a three-stage learning method that automatically detects unlabeled bones while being robust to their various shape. Extensive experiments are conducted on the curated dataset to test the proposed method and promising performance is observed. To the best of our knowledge, this is the first work on the partially supervised bone segmentation problem.
Tianyou Liang, Min Xu 0001
ICASSP4
2025 STPM: Spatial-Temporal Point Mamba for Activity Recognition Using mmWave Radar Point Clouds
abstract
Human activity recognition using millimeter-wave radar point clouds has emerged as a promising visual privacy-preserving sensing paradigm, transitioning from multi-domain Doppler analysis to point cloud-based methods for richer spatial information. However, existing approaches face critical challenges in modeling temporal dependencies between consecutive frames and simultaneously capturing both local geometric structures and global spatial relationships. To address these challenges, we propose STPM (Spatial-Temporal Point Mamba), a novel framework that extends traditional State Space Models through three key innovations: (1) a bidirectional selective mechanism that captures comprehensive temporal dependencies while maintaining linear memory complexity, (2) a queue-based temporal processing strategy with theoretical guarantees for preventing error accumulation, and (3) a hierarchical grouping strategy that effectively models both local geometric details and global spatial contexts. Through extensive evaluations on RadHAR and MM-Fi datasets, STPM achieves state-of-the-art performance with 98.14% and 95.69% accuracy respectively, while reducing memory consumption by 35% compared to transformer-based alternatives. Extensive experiments demonstrate its effectiveness in distinguishing semantically similar but functionally distinct motions.
Yingru Chen, Haimin Zhang 0001, Min Xu 0001
ICME4
2025 RLD-GS: Reinforcement Learning-Driven Gaussian Splatting for High-Fidelity Neural Rendering
Zexu Huang, Min Xu 0001, Stuart W. Perry
ICONIP (2)2
2025 Mitigating Knowledge Discrepancies among Multiple Datasets for Task-agnostic Unified Face Alignment
abstract
Abstract Despite the similar structures of human faces, existing face alignment methods cannot learn unified knowledge from multiple datasets with different landmark annotations. The limited training samples in a single dataset commonly result in fragile robustness in this field. To mitigate knowledge discrepancies among different datasets and train a task-agnostic unified face alignment (TUFA) framework, this paper presents a strategy to unify knowledge from multiple datasets. Specifically, we calculate a mean face shape for each dataset. To explicitly align these mean shapes on an interpretable plane based on their semantics, each shape is then incorporated with a group of semantic alignment embeddings. The 2D coordinates of these aligned shapes can be viewed as the anchors of the plane. By encoding them into structure prompts and further regressing the corresponding facial landmarks using image features, a mapping from the plane to the target faces is finally established, which unifies the learning target of different datasets. Consequently, multiple datasets can be utilized to boost the generalization ability of the model. The successful mitigation of discrepancies also enhances the efficiency of knowledge transferring to a novel dataset, significantly boosts the performance of few-shot face alignment. Additionally, the interpretable plane endows TUFA with a task-agnostic characteristic, enabling it to locate landmarks unseen during training in a zero-shot manner. Extensive experiments are carried on seven benchmarks and the results demonstrate an impressive improvement in face alignment brought by knowledge discrepancies mitigation. The code is available at https://github.com/Jiahao-UTS/TUFA
Jiahao Xia 0001, Min Xu 0001, Wenjian Huang 0001, Jianguo Zhang 0001, Haimin Zhang 0001, Chunxia Xiao
Int. J. Comput. Vis.2
2025 Causal disentanglement for regulating social influence bias in social recommendation
Li Wang 0064, Min Xu 0001, Quangui Zhang, Yunxiao Shi, Qiang Wu 0001
Neurocomputing2
2025 User Reidentification Through mmWave Radio Imaging
abstract
User re-identification (Re-ID) plays a crucial role in the research fields of wireless sensing and Internet of things. However, existing vision-based Re-ID approaches often face privacy concerns and are susceptible to variations in lighting conditions. In response to these challenges, in this paper, we propose a novel radio imaging based user Re-ID scheme with a commercial off-the-shelf (COTS) millimeter-wave (mmWave) radar, called RImID. RImID capitalizes on radar signals to initially track individuals in motion. To achieve high-resolution images of tracked subjects, we propose an advanced mmWave Inverse Synthetic Aperture Radar (ISAR) imaging algorithm, which incorporates focusing and nonlinear motion compensation techniques. Importantly, as individuals exhibit unique body movement patterns, RImID captures these distinctions and encodes them within the radio images. Subsequently, RImID utilizes a well-designed convolutional neural network (CNN) that takes these radio images as inputs to identify each registered person. We have implemented RImID using a mmWave radar, specifically the TI IWR1843BOOST, and have collected a comprehensive dataset of over 25.6k mmWave frames for CNN training. Our experimental results confirm the effectiveness of RImID, achieving a remarkable recognition accuracy exceeding 98 percent for 10 individuals and 93 percent for 30 individuals.
Min Xu 0001, Yongze Wang, Jian (Andrew) Zhang
IEEE Internet Things J.2
2025 Passive Human Tracking With WiFi Point Clouds
abstract
Integrated sensing and communication (ISAC) technology empowers WiFi to function as both sensors for wireless sensing and communication devices for data exchange. Currently, achieving accurate object tracking with commercial WiFi devices is still challenging due to the limited bandwidth, a small number of antennas, and the clock asynchronization in a bi-static setup. Many existing methods achieve tracking only via extracting a dominant Doppler frequency shift (DFS) from a moving person. However, since the human body is nonrigid, various body parts generate different DFSs, and different subcarriers can exhibit varying Doppler characteristics in a multipath environment. This work presents WiDFS2.0, an enhanced real-time tracking scheme that leverages the micro-Doppler effect to extract multiple signal features from various body parts of a moving person, represented as WiFi point clouds. Each point cloud consists of Doppler, Angle of Arrival, Range, and signal-to-noise ratio. We design a novel signal processing chain to extract the WiFi point clouds. Then, we refine these point clouds and implement an extended Kalman filter-based algorithm to track the person’s trajectory. Our experiments demonstrate that WiDFS2.0 can achieve real-time tracking with a median position error of 0.55 m, while determining the presence of a moving person with over 98% accuracy during tracking.
Jian (Andrew) Zhang, Haimin Zhang 0001, Min Xu 0001, Y. Jay Guo
IEEE Internet Things J.4
2025 Visible-Infrared Person Re-Identification With Real-World Label Noise
abstract
In recent years, growing needs for advanced security and traffic management have significantly heightened the prominence of the visible-infrared person re-identification community (VI-ReID), garnering considerable attention. A critical challenge in VI-ReID is the performance degradation attributable to label noise, an issue that becomes even more pronounced in cross-modal scenarios due to an increased likelihood of data confusion. While previous methods have achieved notable successes, they often overlook the complexities of instance-dependent and real-world noise, creating a disconnect from the practical applications of person re-identification. To bridge this gap, our research analyzes the primary sources of label noise in real-world settings, which include a) instantiated identities, b) blurry infrared images, and c) annotators’ errors. In response to these challenges, we develop a Robust Hybrid Loss function (RHL) that enables targeted recognition and retrieval optimization through a more fine-grained division of the noisy dataset. The proposed method categorises data into three sets: clean, obviously noisy, and indistinguishably noisy, with bespoke loss calculations for each category. The identification loss is structured to address the varied nature of these sets specifically. For the retrieval sub-task, we utilize an enhanced triplet loss, adept at handling noisy correspondences. Furthermore, to empirically validate our method, we have re-annotated a real-world dataset, SYSU-Real. Our experiments on SYSU-MM01 and RegDB, conducted under various noise ratios of random and instance-dependent label noise, demonstrate the generalized robustness and effectiveness of our proposed approach.
Ruiheng Zhang 0001, Zhe Cao 0001, Yan Huang 0023, Shuo Yang 0006, Lixin Xu 0001, Min Xu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Federated User Preference Modeling for Privacy-Preserving Cross-Domain Recommendation
abstract
Cross-domain recommendation (CDR) aims to address the data-sparsity problem by transferring knowledge across domains. Existing CDR methods generally assume that the user-item interaction data is shareable between domains, which leads to privacy leakage. Recently, some privacy-preserving CDR (PPCDR) models have been proposed to solve this problem. However, they primarily transfer simple representations learned only from user-item interaction histories, overlooking other useful side information, leading to inaccurate user preferences. Additionally, they transfer differentially private user-item interaction matrices or embeddings across domains to protect privacy. However, these methods offer limited privacy protection, as attackers may exploit external information to infer the original data. To address these challenges, we propose a novel Federated User Preference Modeling (FUPM) framework. In FUPM, first, a novel comprehensive preference exploration module is proposed to learn users' comprehensive preferences from both interaction data and additional data including review texts and potentially positive items. Next, a private preference transfer module is designed to first learn differentially private local and global prototypes, and then privately transfer the global prototypes using a federated learning strategy. These prototypes are generalized representations of user groups, making it difficult for attackers to infer individual information. Extensive experiments on four CDR tasks conducted on the Amazon and Douban datasets validate the superiority of FUPM over SOTA baselines.
Li Wang 0064, Shoujin Wang, Quangui Zhang, Qiang Wu 0001, Min Xu 0001
IEEE Trans. Multim.5
2025 FasterSal: Robust and Real-Time Single-Stream Architecture for RGB-D Salient Object Detection
abstract
RGB-D Salient Object Detection (SOD) aims to segment the most prominent areas and objects in a given pair of RGB and depth images. Most current models adopt a dual-stream structure to extract information from both RGB and depth images. However, this leads to an exponential increase in the number of parameters and computations in the model. Moreover, the discrepancy between RGB pretrained and the 3D geometric relationships in depth maps present a challenge for the encoder in capturing spatial structural details. These issues impact the model's accuracy in locating salient objects and distinguishing edge details. To address these, we propose a novel early feature fusion network, named FasterSal, which enables more efficient RGB-D SOD. FasterSal uses a single stream structure to receive RGB images and depth maps, extracting features based on the 3D geometric relationships in the depth map while fully leveraging the pretrained RGB encoder. This approach effectively avoids the inconsistencies between depth modality and the RGB pretrained encoder. It also significantly reduces the number of network parameters while maintaining efficient feature encoding capabilities. To achieve finer edge learning, the detail-aware loss and texture enhancement module are introduced. These modules are designed to extract latent details in high-frequency component features and to enhance the edge learning capability of the model using distance information. Experimental results on several benchmark datasets confirm the effectiveness and superiority of our method over the state-of-the-art approaches, achieving a good balance between performance and speed with only 3.4 million parameters and a CPU operating speed of 63 FPS.
Jing Zhang 0037, Ruiheng Zhang 0001, Lixin Xu 0001, Xiankai Lu, Yushu Yu, Min Xu 0001, He Zhao 0002
IEEE Trans. Multim.6
2024 Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems
abstract
Retrieval-augmented generation (RAG) techniques leverage the in-context learning capabilities of large language models (LLMs) to produce more accurate and relevant responses. Originating from the simple ‘retrieve-then-read’ approach, the RAG framework has evolved into a highly flexible and modular paradigm. A critical component, the Query Rewriter module, enhances knowledge retrieval by generating a search-friendly query. This method aligns input questions more closely with the knowledge base. Our research identifies opportunities to enhance the Query Rewriter module to Query Rewriter+ by generating multiple queries to overcome the Information Plateaus associated with a single query and by rewriting questions to eliminate Ambiguity, thereby clarifying the underlying intent. We also find that current RAG systems exhibit issues with Irrelevant Knowledge; to overcome this, we propose the Knowledge Filter. These two modules are both based on the instruction-tuned Gemma-2B model, which together enhance response quality. The final identified issue is Redundant Retrieval; we introduce the Memory Knowledge Reservoir and the Retriever Trigger to solve this. The former supports the dynamic expansion of the RAG system’s knowledge base in a parameter-free manner, while the latter optimizes the cost for accessing external knowledge, thereby improving resource utilization and response efficiency. These four RAG modules synergistically improve the response quality and efficiency of the RAG system. The effectiveness of these modules has been validated through experiments and ablation studies across six common QA datasets. The source code can be accessed at https://github.com/Ancientshi/ERM4.
Yunxiao Shi, Xing Zi, Zijing Shi, Haimin Zhang 0001, Qiang Wu 0001, Min Xu 0001
ECAI6
2024 Center-bridged Interaction Fusion for hyperspectral and LiDAR classification
Lu Huo, Jiahao Xia 0001, Leijie Zhang, Haimin Zhang 0001, Min Xu 0001
Neurocomputing5
2024 CAA: Class-Aware Affinity calculation add-on for semantic segmentation
abstract
Leveraging contextual dependencies is a commonly used technique to enhance the performance of image segmentation. However, existing solutions do not effectively catch the class-level association between the pixels along the boundary across the objects of the different classes but focus more on the local pixel-to-pixel relation. This work proposes a Class-Aware Affinity module (CAA) that considers both pixel-to-pixel relation and pixel-to-class association. We try to argue that the pixel-to-pixel relations still catch the relation (e.g. similarity, attention, or affiliation) on the local texture level. At the same time, it should also consider the association between the pixel and the class context produced by the given image. Pixel-to-class association can best reveal the co-occurrent dependency on the semantic level between the given pixels and their nearby context. Such pixel-to-class association combined with the pixel-to-pixel relations aggregating the local texture information will best mitigate the confusion caused in the boundary regions across the objects of the different classes. Moreover, the proposed framework can serve as a generic add-on to be integrated with the existing image segmentation solution to boost the current performance. Equipped with CAA, we achieve promising performance against the existing work with 54.59% mIoU on ADE20K, 49.96% mIoU on COCO-Stuff10k, and 64.38% mIoU on Pascal-Context.
Huadong Tang, Youpeng Zhao 0002, Chaofan Du, Min Xu 0001, Qiang Wu 0001
Knowl. Based Syst.4
2024 A privacy-preserving framework with multi-modal data for cross-domain recommendation
abstract
Cross-domain recommendation (CDR) aims to enhance the recommendation accuracy in a target domain with sparse data by leveraging rich information in a source domain, thereby addressing the data-sparsity problem. Some existing CDR methods highlight the advantages of extracting domain-common and domain-specific features to learn comprehensive user and item representations. However, these methods cannot effectively disentangle these components, as they often rely on simple user-item historical interaction information (such as ratings, clicks, and browsing), neglecting the rich multi-modal features. In addition, they do not protect user-sensitive data from potential leakage during knowledge transfer between domains. To address these challenges, we propose a P rivacy- P reserving Framework with M ulti- M odal Data for C ross- D omain R ecommendation, called P2M2-CDR. Specifically, we first design a multi-modal disentangled encoder that utilizes multi-modal information to disentangle more informative domain-common and domain-specific embeddings. Furthermore, we introduce a privacy-preserving decoder to mitigate user privacy leakage during knowledge transfer. Local differential privacy (LDP) is used to obfuscate disentangled embeddings before the inter-domain exchange, thereby enhancing privacy protection. To ensure both consistency and differentiation among these obfuscated disentangled embeddings, we incorporate contrastive learning-based domain-inter and domain-intra losses. Extensive experiments conducted on six CDR tasks from two real-world datasets demonstrate that P2M2-CDR outperforms other state-of-the-art single- and cross-domain baselines. The code is available at https://github.com/Lili1013/P2M2-CDR .
Li Wang 0064, Lei Sang 0001, Quangui Zhang, Qiang Wu 0001, Min Xu 0001
Knowl. Based Syst.5
2024 Unsupervised Part Discovery via Dual Representation Alignment
abstract
Object parts serve as crucial intermediate representations in various downstream tasks, but part-level representation learning still has not received as much attention as other vision tasks. Previous research has established that Vision Transformer can learn instance-level attention without labels, extracting high-quality instance-level representations for boosting downstream tasks. In this paper, we achieve unsupervised part-specific attention learning using a novel paradigm and further employ the part representations to improve part discovery performance. Specifically, paired images are generated from the same image with different geometric transformations, and multiple part representations are extracted from these paired images using a novel module, named PartFormer. These part representations from the paired images are then exchanged to improve geometric transformation invariance. Subsequently, the part representations are aligned with the feature map extracted by a feature map encoder, achieving high similarity with the pixel representations of the corresponding part regions and low similarity in irrelevant regions. Finally, the geometric and semantic constraints are applied to the part representations through the intermediate results in alignment for part-specific attention learning, encouraging the PartFormer to focus locally and the part representations to explicitly include the information of the corresponding parts. Moreover, the aligned part representations can further serve as a series of reliable detectors in the testing phase, predicting pixel masks for part discovery. Extensive experiments are carried out on four widely used datasets, and our results demonstrate that the proposed method achieves competitive performance and robustness due to its part-specific attention.
Jiahao Xia 0001, Wenjian Huang 0001, Min Xu 0001, Jianguo Zhang 0001, Haimin Zhang 0001, Ziyu Sheng, Dong Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Vital Sign Monitoring in Dynamic Environment via mmWave Radar and Camera Fusion
abstract
Contact-free vital sign monitoring, which uses wireless signals for recognizing human vital signs (i.e, breath and heartbeat), is an attractive solution to health and security. However, the subject’s body movement and the change in actual environments can result in inaccurate frequency estimation of heartbeat and respiratory. In this paper, we propose a robust mmWave radar and camera fusion system for monitoring vital signs, which can perform consistently well in dynamic scenarios, e.g., when some people move around the subject to be tracked, or a subject waves his/her arms and marches on the spot. Three major processing modules are developed in the system, to enable robust sensing. First, we utilize a camera to assist a mmWave radar to accurately localize the subjects of interest. Second, we exploit the calculated subject position to form transmitting and receiving beamformers, which can improve the reflected power from the targets and weaken the impact of dynamic interference. Third, we propose a weighted multi-channel Variational Mode Decomposition (WMC-VMD) algorithm to separate the weak vital sign signals from the dynamic ones due to subject’s body movement. Experimental results show that, the 90th percentile errors in respiration rate (RR) and heartbeat rate (HR) are less than 0.5 RPM (respirations per minute) and 6 BPM (beats per minute), respectively.
Jian (Andrew) Zhang, Haimin Zhang 0001, Min Xu 0001
IEEE Trans. Mob. Comput.5
2024 Towards High-Quality Photorealistic Image Style Transfer
abstract
Preserving important textures of the content image and achieving prominent style transfer results remains a challenge in the field of image style transfer. This challenge arises from the entanglement between color and texture during the style transfer process. To address this challenge, we propose an end-to-end network that incorporates adaptive weighted least squares (AWLS) filter, iterative least squares (ILS) filter, and channel separation. Given a content image ($\mathcal {C}$) and a reference style image ($\mathcal {S}$), we begin by separating the RGB channels and utilizing ILS filter to decompose them into structure and texture layers. We then perform style transfer on the structural layers using WCT$^{2}$(incorporating wavelet pooling and unpooling techniques for whitening and coloring transforms) in the R, G, and B channels, respectively. We address the texture distortion caused by WCT$^{2}$with a texture enhancing (TE) module in the structural layer. Furthermore, we propose an estimating and compensating for the structure loss (ECSL) module. In the ECSL module, with the AWLS filter and the ILS filter, we estimate the texture loss caused by TE, convert the loss of the structural layer to the loss of the texture layer, and compensate for the loss in the texture layer. The final structural layer and the texture layer are merged into the channel style transfer results in the separated R, G, and B channels into the final style transfer result. Thereby, this enables a more complete texture preservation and a significant style transfer process. To evaluate our method, we utilize quantitative experiments using various metrics, including NIQE, AG, SSIM, PSNR, and a user study. The experimental results demonstrate the superiority of our approach over the previous state-of-the-art methods.
Haimin Zhang 0001, Gang Fu 0003, Caoqing Jiang, Fei Luo 0004, Chunxia Xiao, Min Xu 0001
IEEE Trans. Multim.7
2024 SSFG: Stochastically Scaling Features and Gradients for Regularizing Graph Convolutional Networks
abstract
Graph convolutional networks (GCNs) have been successfully applied in various graph-based tasks. In a typical graph convolutional layer, node features are updated by aggregating neighborhood information. Repeatedly applying graph convolutions can cause the oversmoothing issue, i.e., node features at deep layers converge to similar values. Previous studies have suggested that oversmoothing is one of the major issues that restrict the performance of GCNs. In this article, we propose a stochastic regularization method to tackle the oversmoothing problem. In the proposed method, we stochastically scale features and gradients (SSFG) by a factor sampled from a probability distribution in the training procedure. By explicitly applying a scaling factor to break feature convergence, the oversmoothing issue is alleviated. We show that applying stochastic scaling at the gradient level is complementary to that applied at the feature level to improve the overall performance. Our method does not increase the number of trainable parameters. When used together with ReLU, our SSFG can be seen as a stochastic ReLU activation function. We experimentally validate our SSFG regularization method on three commonly used types of graph networks. Extensive experimental results on seven benchmark datasets for four graph-based tasks demonstrate that our SSFG regularization is effective in improving the overall performance of the baseline graph networks. The code is available at https://github.com/vailatuts/SSFG-regularization.
Haimin Zhang 0001, Min Xu 0001, Guoqiang Zhang 0003, Kenta Niwa
IEEE Trans. Neural Networks Learn. Syst.2
2024 Learning Graph Representations Through Learning and Propagating Edge Features
abstract
Graph convolutional networks have achieved considerable success in various graph domain tasks. Recently, numerous types of graph convolutional networks have been developed. A typical rule for learning a node's feature in these graph convolutional networks is to aggregate node features from the node's local neighborhood. However, in these models, the interrelation information between adjacent nodes is not well-considered. This information could be helpful to learn improved node embeddings. In this article, we present a graph representation learning framework that generates node embeddings through learning and propagating edge features. Instead of aggregating node features from a local neighborhood, we learn a feature for each edge and update a node's representation by aggregating local edge features. The edge feature is learned from the concatenation of the edge's starting node feature, the input edge feature, and the edge's end node feature. Unlike node feature propagation-based graph networks, our model propagates different features from a node to its neighbors. In addition, we learn an attention vector for each edge in aggregation, enabling the model to focus on important information in each feature dimension. By learning and aggregating edge features, the interrelation between a node and its neighboring nodes is integrated in the aggregated feature, which helps learn improved node embeddings in graph representation learning. Our model is evaluated on graph classification, node classification, graph regression, and multitask binary graph classification on eight popular datasets. The experimental results demonstrate that our model achieves improved performance compared with a wide variety of baseline models.
Haimin Zhang 0001, Jiahao Xia 0001, Guoqiang Zhang 0003, Min Xu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity Consistency
abstract
As one of the most fundamental techniques in multi-modal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose a general framework called BiCro (Bidirectional Cross-modal similarity consistency), which can be easily integrated into existing cross-modal matching models and improve their robustness against noisy data. Specifically, BiCro aims to estimate soft labels for noisy data pairs to reflect their true correspondence degree. The basic idea of BiCro is motivated by that – taking image-text matching as an example – similar images should have similar textual descriptions and vice versa. Then the consistency of these two similarities can be recast as the estimated soft labels to train the matching model. The experiments on three popular cross-modal matching datasets demonstrate that our method significantly improves the noise-robustness of various matching models, and surpass the state-of-the-art by a clear margin. The code is available at https://github.com/xu5zhao/BiCro.
Shuo Yang 0006, Zhaopan Xu, Kai Wang 0036, Yang You 0001, Hongxun Yao, Tongliang Liu, Min Xu 0001
CVPR7
2023 Dataset Pruning: Reducing Training Data by Examining Generalization Influence
Shuo Yang 0006, Zeke Xie, Hanyu Peng, Min Xu 0001, Mingming Sun 0001, Ping Li 0001
ICLR4
2023 Learning enhanced features and inferring twice for fine-grained image classification
abstract
Abstract Fine-Grained Visual Categorization (FGVC) aims to distinguish between extremely similar subordinate-level categories within the same basic-level category. Existing research has proven the great importance of the discriminative features in FGVC but ignored the contributions for correct classification from other features, and the extracted features always contain more information about the obvious regions but less about subtle regions. In this paper, firstly, a novel module named forcing module is proposed to force the network to extract more diverse features for FGVC, which generates a suppression mask based on the class activation maps to suppress the most distinguishable regions, so as to force the network to extract other secondary distinguishable features as the final features. The forcing module consists of the original branch and the forcing branch. The original branch focuses on the primary discriminative regions while the forcing branch focuses on secondary discriminative regions. Secondly, in order to solve the problem that information of small-scale distinguishable features is lost seriously after multi-layer down-sampling, according to the class activation maps of the first prediction, the object is cropped and scaled as the second input. To reduce the prediction error, the first and second prediction probabilities are fused as the final prediction result. Experimental results indicate that the proposed method not only outperforms the baseline model by a large margin (3.7%, 5.9%, 3.1% respectively) on CUB-200-2011, Stanford-Cars, and FGVC-Aircraft, but also achieves state-of-the-art performance on FGVC-Aircraft.
Xuan Nie, Bosong Chai, Qiyu Liao, Min Xu 0001
Multim. Tools Appl.5
2023 Robust Face Alignment via Inherent Relation Learning and Uncertainty Estimation
abstract
Human tends to locate the facial landmarks with heavy occlusion by their relative position to the easily identified landmarks. The clue is defined as the landmark inherent relation while it is ignored by most existing methods. In this paper, we present Dynamic Sparse Local Patch Transformer (DSLPT), a novel face alignment framework for the inherent relation learning and uncertainty estimation. Unlike most existing methods that regress facial landmarks directly from global features, the DSLPT first generates a rough representation of each landmark from a local patch cropped from the feature map and then adaptively aggregates them by a case dependent inherent relation. Finally, the DSLPT predicts the coordinate and uncertainty of each landmark by regressing their probability distribution from the output features. Moreover, we introduce a coarse-to-fine framework to incorporate with DSLPT for an improved result. In the framework, the position and size of each patch are determined by the probability distribution of the corresponding landmark predicted in the previous stage. The dynamic patches will ensure a fine-grained landmark representation for inherent relation learning so that a rough prediction result can gradually converge to the target facial landmarks. We integrate the coarse-to-fine model into an end-to-end training pipeline and carry out experiments on the mainstream benchmarks. The results demonstrate that the DSLPT achieves state-of-the-art performance with much less computational complexity. The codes and models are available at https://github.com/Jiahao-UTS/DSLPT.
Jiahao Xia 0001, Min Xu 0001, Haimin Zhang 0001, Jianguo Zhang 0001, Wenjian Huang 0001, Hu Cao, Shiping Wen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 A Parametrical Model for Instance-Dependent Label Noise
abstract
In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited on class-dependent label-noise (wherein all samples in a clean class share the same label transition matrix). However, the CLTM cannot handle the more common instance-dependent label-noise well (wherein the clean-to-noisy label transition matrix needs to be estimated at the instance level by considering the input quality). Motivated by the fact that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-Label Transition Matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have no uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, this work proposes a parametrical model for estimating the instance-dependent label-noise transition matrix by employing a deep neural network, leading to better generalization and superior classification performance.
Shuo Yang 0006, Songhua Wu, Erkun Yang, Bo Han 0003, Yang Liu 0018, Min Xu 0001, Gang Niu 0001, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Adversarial Heterogeneous Graph Neural Network for Robust Recommendation
abstract
Recommendation systems play a vital role in identifying the hidden interactions between users and items in online social networks. Recently, graph neural networks (GNNs) have exhibited significant performance gains by modeling the information propagation process in graph-structured data for a recommendation. However, existing GNN-based methods do not have broad applicability to heterogeneous graphs that integrate auxiliary data with diverse types. Moreover, graph structures are susceptible to noise and even unnoticed malicious perturbations, as perturbations from connected nodes can create cumulative effects on a target node in the graph. To enhance the robustness and generalization of GNN-based recommendations, we propose a new optimization model named Adversarial Heterogeneous Graph Neural Network for RECommendation (AHGNNRec). First, AHGNNRec learns user and item embeddings by exploring the distinct contributions of various types of interactions between users and items using a hierarchical heterogeneous graph neural network (HGNN). Second, to produce more robust embeddings for recommendations, we employ the adversarial training (AT) method to optimize the HGNN layers. AT is a min-max optimization training process where the generated adversarial fake nodes from normal nodes with intentional perturbations try to maximally deteriorate the recommendation performance. Following this, we learn about these adversarial user or item nodes by minimizing the impact of an additional regularization term for the recommendation. The experimental outcomes on two real-world benchmark datasets demonstrate the effectiveness of AHGNNRec.
Lei Sang 0001, Min Xu 0001, Shengsheng Qian, Xindong Wu 0001
IEEE Trans. Comput. Soc. Syst.2
2023 Single-Target Real-Time Passive WiFi Tracking
abstract
Device-free human tracking is an essential ingredient for ubiquitous wireless sensing. Recent passive WiFi tracking systems face the challenges of inaccurate separation of dynamic human components and time-consuming estimation of multi-dimensional signal parameters. In this work, we present a scheme namedWiFiDopplerFrequencyShift (WiDFS), which can achieve single-target real-time passive tracking using channel state information (CSI) collected from commercial-off-the-shelf (COTS) WiFi devices. We consider the typical system setup including a transmitter with a single antenna and a receiver with three antennas; while our scheme can be readily extended to another setup. To remove the impact of transceiver asynchronization, we first apply CSI cross-correlation between each RX antenna pair. We then combine them to estimate a Doppler frequency shift (DFS) in a short-time window. After that, we leverage the DFS estimate to separate dynamic human components from CSI self-correlation terms of each antenna, thereby separately calculating angle-of-arrival (AoA) and human reflection distance for tracking. In addition, a hardware calibration algorithm is presented to refine the spacing between RX antennas and eliminate the hardware-related phase differences between them. A prototype demonstrates that WiDFS can achieve real-time tracking with a median position error of 72.32 cm in multipath-rich environments.
Jian (Andrew) Zhang, Min Xu 0001, Y. Jay Guo
IEEE Trans. Mob. Comput.3
2023 Multiscale Emotion Representation Learning for Affective Image Recognition
abstract
Recognition of emotions conveyed in images has attracted increasing research attention. Recent studies show that leveraging local affective regions helps to improve the recognition performance. However, these studies do not consider features from the broad context of the local affective regions, which could provide useful information for learning improved emotion representations. In this paper, we present a region-based multiscale network that learns features for the local affective region as well as the broad context for affective image recognition. The proposed network consists of an affective region detection module and a multiscale feature learning module. The class activation mapping method is used to generate pseudo affective regions from a pretrained deep neural network to train the detection module. For the affective region outputted by the detection module, three-scale features are extracted and then encoded by a kernel-based graph attention network for final emotion classification. We show that integrating features from the broad context is effective in improving the recognition performance. We experimentally evaluate the proposed network for both multi-class emotion recognition and binary sentiment classification on different benchmark datasets. The experimental results demonstrate that the proposed network achieves improved or comparable performance as compared to previous state-of-the-art models.
Haimin Zhang 0001, Min Xu 0001
IEEE Trans. Multim.2
2023 Recognition of Emotions in User-Generated Videos through Frame-Level Adaptation and Emotion Intensity Learning
abstract
Recognition of emotions in user-generated videos has attracted considerable research attention. Most existing approaches focus on learning frame-level features and fail to consider frame-level emotion intensities which are critical for video representation. In this research, we aim to extract frame-level features and emotion intensities through transferring emotional information from an image emotion dataset. To achieve this goal, we propose an end-to-end network for joint emotion recognition and intensity learning with unsupervised adversarial adaptation. The proposed network consists of a classification stream, an intensity learning stream and an adversarial adaptation module. The classification stream is used to generate pseudo intensity maps with the class activation mapping method to train the intensity learning subnetwork. The intensity learning stream is built upon an improved feature pyramid network in which features from different scales are cross-connected. The adversarial adaptation module is employed to reduce the domain difference between the source dataset and target video frames. By aligning cross domain features, we enable our network to learn on the source data while generalizing to video frames. Finally, we apply a weighted sum pooling method to frame-level features and emotion intensities to generate video-level features. We evaluate the proposed method on two benchmark datasets,i.e.,VideoEmotion-8 and Ekman-6. The experimental results show that the proposed method achieves improved performance compared to previous state-of-the-art methods.
Haimin Zhang 0001, Min Xu 0001
IEEE Trans. Multim.2
2022 Sparse Local Patch Transformer for Robust Face Alignment and Landmarks Inherent Relation Learning
abstract
Heatmap regression methods have dominated face alignment area in recent years while they ignore the inherent relation between different landmarks. In this paper, we propose a Sparse Local Patch Transformer (SLPT) for learning the inherent relation. The SLPT generates the representation of each single landmark from a local patch and aggregates them by an adaptive inherent relation based on the attention mechanism. The subpixel coordinate of each landmark is predicted independently based on the aggregated feature. Moreover, a coarse-to-fine framework is further introduced to incorporate with the SLPT, which enables the initial landmarks to gradually converge to the target facial landmarks using fine-grained features from dynamically resized local patches. Extensive experiments carried out on three popular benchmarks, including WFLW, 300W and COFW, demonstrate that the proposed method works at the state-of-the-art level with much less computational complexity by learning the inherent relation between facial landmarks. The code is available at the project website11https://github.com/Jiahao-UTS/SLPT-master.
Jiahao Xia 0001, Weiwei Qu, Wenjian Huang 0001, Jianguo Zhang 0001, Min Xu 0001
CVPR6
2022 Objects in Semantic Topology
Shuo Yang 0006, Peize Sun, Yi Jiang 0009, Xiaobo Xia, Ruiheng Zhang 0001, Zehuan Yuan, Changhu Wang, Ping Luo 0002, Min Xu 0001
ICLR9
2022 Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network
abstract
In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited to learn a clean label classifier by employing the noisy data. Motivated by that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-label transition matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have less uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, we estimate the BLTM parametrically by employing a deep neural network, leading to better generalization and superior classification performance.
Shuo Yang 0006, Erkun Yang, Bo Han 0003, Yang Liu 0018, Min Xu 0001, Gang Niu 0001, Tongliang Liu
ICML5
2022 An efficient multitask neural network for face alignment, head pose estimation and face tracking
abstract
While Convolutional Neural Networks (CNNs) have significantly boosted the performance of face related algorithms, maintaining accuracy and efficiency simultaneously in practical use remains challenging. The state-of-the-art methods employ deeper networks for better performance, which makes it less practical for mobile applications because of more parameters and higher computational complexity. Therefore, we propose an efficient multitask neural network, Alignment & Tracking & Pose Network (ATPN) for face alignment, face tracking and head pose estimation . Specifically, to achieve better performance with fewer layers for face alignment, we introduce a shortcut connection between shallow-layer and deep-layer features. We find the shallow-layer features are highly correspond to facial boundaries that can provide the structural information of face and it is crucial for face alignment. Moreover, we generate a cheap heatmap based on the face alignment result and fuse it with features to improve the performance of the other two tasks. Based on the heatmap, the network can utilize both geometric information of landmarks and appearance information for head pose estimation. The heatmap also provides attention clues for face tracking. The face tracking task also saves us the face detection procedure for each frame, which also significantly boost the real-time capability for video-based tasks. We experimentally validate ATPN on four benchmark datasets, WFLW, 300VW, WIDER Face and 300W-LP. The experimental results demonstrate that it achieves better performance with much less parameters and lower computational complexity compared to other light models.
Jiahao Xia 0001, Haimin Zhang 0001, Shiping Wen 0001, Shuo Yang 0006, Min Xu 0001
Expert Syst. Appl.5
2022 Distribution-Aware Margin Calibration for Semantic Segmentation in Images
Litao Yu, Zhibin Li 0002, Min Xu 0001, Yongsheng Gao 0001, Jiebo Luo 0001, Jian Zhang 0002
Int. J. Comput. Vis.3
2022 Accurate AoA Estimation for RFID Tag Array With Mutual Coupling
abstract
Angle-of-Arrival (AoA) estimation is an important problem in passive radio-frequency identification (RFID) systems. Affixing an RFID tag array to an object enables to acquire its orientation information. However, the electromagnetic interaction between the tags can induce mutual coupling interference, distorting the RFID fingerprint measurements used for AoA estimation. Moreover, RFID reader modes with radio-frequency (RF) noise-tolerant Miller encoding can induce$\pi $-radians phase jump. In this article, we propose a scheme called RF-Mirror that can resolve the mutual coupling and phase jump problems and achieve accurate AoA estimation for an array with two or more tags. First, we characterize the impact of mutual coupling on a tag’s signal fingerprint and develop novel RSSI/phase-distance models. We then develop new experimental methods and signal processing techniques to verify the effectiveness of the proposed models. Based on the validated models, we develop new AoA estimation algorithms for tag arrays that deal with the mutual coupling effect explicitly. We provide extensive experimental results, which demonstrate that RF-Mirror can achieve significantly improved performance compared to baseline schemes, with median AoA estimation errors of 11.65° and 6.29° for two- and four-tag arrays, respectively.
Jian (Andrew) Zhang, Fu Xiao 0001, Min Xu 0001
IEEE Internet Things J.4
2022 Bridging the Gap Between Few-Shot and Many-Shot Learning via Distribution Calibration
abstract
A major gap between few-shot and many-shot learning is the data distribution empirically oserved by the model during training. In few-shot learning, the learned model can easily become over-fitted based on the biased distribution formed by only a few training examples, while the ground-truth data distribution is more accurately uncovered in many-shot learning to learn a well-generalized model. In this paper, we propose to calibrate the distribution of these few-sample classes to be more unbiased to alleviate such an over-fitting problem. The distribution calibration is achieved by transferring statistics from the classes with sufficient examples to those few-sample classes. After calibration, an adequate number of examples can be sampled from the calibrated distribution to expand the inputs to the classifier. Specifically, we assume every dimension in the feature representation from the same class follows a Gaussian distribution so that the mean and the variance of the distribution can borrow from that of similar classes whose statistics are better estimated with an adequate number of samples. Extensive experiments on three datasets,miniImageNet,tieredImageNet, and CUB, show that a simple linear classifier trained using the features sampled from our calibrated distribution can outperform the state-of-the-art accuracy by a large margin. Besides the favorable performance, the proposed method also exhibits high flexibility by showing consistent accuracy improvement when it is built on top of any off-the-shelf pretrained feature extractors and classification models without extra learnable parameters. The visualization of these generated features demonstrates that our calibrated distribution is an accurate estimation thus the generalization ability gain is convincing. We also establish a generalization error bound for the proposed distribution-calibration-based few-shot learning, which consists of thedistribution assumption error, thedistribution approximation error, and theestimation error. This generalization error bound theoretically justifies the effectiveness of the proposed method.
Shuo Yang 0006, Songhua Wu, Tongliang Liu, Min Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Category attention transfer for efficient fine-grained visual categorization
abstract
Fine-Grained Visual Categorization (FGVC) aims at distinguishing subordinate-level categories with subtle interclass differences. Although previous research shows the impressive effectiveness of the recurrent multi-attention models and the second-order feature encoding, they often require an enormous amount of both computation and memory space, making them inadequate for mobile applications. This paper proposed a Category Attention Transfer CNN (CAT-CNN) to address the efficiency issue in solving FGVC problems. We transfer part attention knowledge from a very large-scale FGVC network to a small but efficient network to significantly improve its presentation ability. Using the proposed CAT-CNN, the accuracy of the efficient networks, such as ShuffleNet, MobilieNet, and EfficientNet, can be improved by up to 5.7% on the CUB-2011-200 dataset without increasing computation complexity or memory cost. Our experiments show that the proposed CAT-CNN can be applied to multiple structures to enhance their performance. With a single efficient network structure and single inference, the proposed CAT-MobileNet-large-1.0 and the CAT-EfficientNet-b0 can achieve accuracies of 86.5% and 86.7%, respectively, on the CUB-2011-200 dataset, which is close to or better than the results from state-of-the-art methods using large scale networks and multiple inferences, and make FGVC feasible on mobile devices.
Qiyu Liao, Dadong Wang, Min Xu 0001
Pattern Recognit. Lett.3
2022 Deep-IRTarget: An Automatic Target Detector in Infrared Imagery Using Dual-Domain Feature Extraction and Allocation
abstract
Recently, convolutional neural networks (CNNs) have brought impressive improvements for object detection. However, detecting targets in infrared images still remains challenging, because the poor texture information, low resolution and high noise levels of the thermal imagery restrict the feature extraction ability of CNNs. In order to deal with these difficulties in the feature extraction, we propose a novel backbone network named Deep-IRTarget, composing of a frequency feature extractor, a spatial feature extractor and a dual-domain feature resource allocation model. Hypercomplex Infrared Fourier Transform is developed to calculate the infrared intensity saliency by designing hypercomplex representations in the frequency domain, while a convolutional neural network is invoked to extract feature maps in the spatial domain. Features from the frequency domain and spatial domain are stacked to construct Dual-domain features. To efficiently integrate and recalibrate them, we propose a Resource Allocation model for Features (RAF). The well-designed channel attention block and position attention block are used in RAF to respectively extract interdependent relationships among channel and position dimensions, and capture channel-wise and position-wise contextual information. Extensive experiments are conducted on three challenging infrared imagery databases. We achieve 10.14%, 9.1% and 8.05% improvement on mAP scores, compared to the current state of the art method on MWIR, BITIR and WCIR respectively.
Ruiheng Zhang 0001, Lixin Xu 0001, Zhengyu Yu, Chengpo Mu, Min Xu 0001
IEEE Trans. Multim.6
2021 Single-View 3D Object Reconstruction From Shape Priors in Memory
abstract
Existing methods for single-view 3D object reconstruction directly learn to transform image features into 3D representations. However, these methods are vulnerable to images containing noisy backgrounds and heavy occlusions because the extracted image features do not contain enough information to reconstruct high-quality 3D shapes. Humans routinely use incomplete or noisy visual cues from an image to retrieve similar 3D shapes from their memory and reconstruct the 3D shape of an object. Inspired by this, we propose a novel method, named Mem3D, that explicitly constructs shape priors to supplement the missing information in the image. Specifically, the shape priors are in the forms of "image-voxel" pairs in the memory network, which is stored by a well-designed writing strategy during training. We also propose a voxel triplet loss function that helps to retrieve the precise 3D shapes that are highly related to the input image from shape priors. The LSTM-based shape encoder is introduced to extract information from the retrieved 3D shapes, which are useful in recovering the 3D shape of an object that is heavily occluded or in complex environments. Experimental results demonstrate that Mem3D significantly improves reconstruction quality and performs favorably against state-of-the-art methods on the ShapeNet and Pix3D datasets.
Shuo Yang 0006, Min Xu 0001, Haozhe Xie, Stuart W. Perry, Jiahao Xia 0001
CVPR2
2021 Recognizing 3D Orientation of a Two-RFID-Tag Labeled Object in Multipath Environments Using Deep Transfer Learning
abstract
State-of-the-art battery-free RFID systems attach multiple RFID tags to an object and exploit their RF phase to estimate its three-dimensional (3D) orientation. However, the measured RF phase may be inaccurate because each tag's signal fingerprint (i.e., RSSI and RF Phase) is distorted by multipath interference and electromagnetic interaction between neighboring tags. In this paper, we propose RF-Orien3D that minimizes these interferences for accurate 3D orientation recognition only using two RFID tags. The electromagnet interference modifies the radiation pattern and modulation factor of each tag in the two-element tag array, which can be estimated to compensate for the distortion in RFID fingerprints. To deal with the multipath impact, we simulate multipath noise to generate huge amounts of RFID fingerprints and use them to pre-train a convolutional neural network (CNN). Then we only collect dozens of actual samples to fine-tune the CNN for multipath-tolerant orientation recognition. The experiments show RF-Orien3D recognizes a two-tag labeled object's 2D orientation with the angular error of about 16° and its 3D orientation (azimuth and elevation) with the errors of about 29° and 11° in low/rich multipath scenarios.
Min Xu 0001, Fu Xiao 0001
ICDCS2
2021 Free Lunch for Few-shot Learning: Distribution Calibration
Shuo Yang 0006, Min Xu 0001
ICLR3
2021 Knowledge graph enhanced neural collaborative recommendation
Lei Sang 0001, Min Xu 0001, Shengsheng Qian, Xindong Wu 0001
Expert Syst. Appl.2
2021 GEME: Dual-stream multi-task GEnder-based micro-expression recognition
Xuan Nie, Madhumita A. Takalkar, Mengyang Duan, Haimin Zhang 0001, Min Xu 0001
Neurocomputing5
2021 Knowledge Graph enhanced Neural Collaborative Filtering with Residual Recurrent Network
Lei Sang 0001, Min Xu 0001, Shengsheng Qian, Xindong Wu 0001
Neurocomputing2
2021 LGAttNet: Automatic micro-expression detection using dual-stream local and global attentions
Madhumita A. Takalkar, Selvarajah Thuseethan, Sutharshan Rajasegarar, Zenon Chaczko, Min Xu 0001, John Yearwood
Knowl. Based Syst.5
2021 Graph neural networks with multiple kernel ensemble attention
Haimin Zhang 0001, Min Xu 0001
Knowl. Based Syst.2
2021 Noise Augmented Double-Stream Graph Convolutional Networks for Image Captioning
abstract
Image captioning, aiming at generating natural sentences to describe image contents, has received significant attention with remarkable improvements in recent advances. The problem nevertheless is not trivial for cross-modal training due to the two challenges: 1) image detectors often consider only salient areas in an image and seldom explore the rich background context; 2) the language model is highly vulnerable to small but intentional perturbation attacks. To alleviate these issues, we propose the Noise Augmented Double-stream Graph Convolutional Networks (NADGCN) that novelly exploits the additional background context and enhances the generalization of the language model. Technically, NADGCN capitalizes on grid-stream GCN as a supplementary to the region stream, following the recipe that a rescaled grid graph can encode the relationship across grid areas over the full image rather than salient areas only. Moreover, we devise a noise module and integrate into the double-stream GCN to augment the capability of the basic generator. Such noise module introduces adaptive noise into the Recurrent Neural Networks (RNN) and is learnt through regarding the module as an agent with a stochastic Gaussian policy in Reinforcement Learning (RL). Extensive experiments on MSCOCO validate the design of the grid-stream GCN and the noise agent, and our generator outperforms the comparative baselines clearly.
Lingxiang Wu, Min Xu 0001, Lei Sang 0001, Ting Yao 0003, Tao Mei 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 4-D Flight Trajectory Prediction With Constrained LSTM Network
abstract
The increasing aviation activities pose a challenge to ensure a safe and orderly flight. Trajectory prediction is one of the most important forecasting tasks in Air Traffic Management. Accurate prediction is reasonable for safe and orderly flight tasks in civil aviation monitoring. Points of interests play an important role in most land traffic prediction algorithms due to their abilities in positioning and marking. Compared with land traffic, the sparse way-points and shared airways make it difficult for flight trajectory prediction. A constrained Long Short-Term Memory network for flight trajectory prediction is proposed in this paper. According to the dynamic characteristics of the aircraft, we propose three kinds of constraints to climbing, cruising, and descending/approaching phases, in particular, they are Top of climb, Way-points, and Runway direction, correspondingly. Our model is able to keep long-term dependencies with dynamic physical constraints. Density-Based Spatial Clustering of Applications with Noise and Linear Least Squares are used in data segmentation and preprocessing. Sliding windows help maintain the continuity of trajectory. Four-dimensional spatial-temporal trajectory set consisting of spatial position and timestamps is used to prove the efficiency of our approach. Multiple ADS-B ground stations contribute to our experimental dataset. The widely used Long Short-Term Memory network, Markov Model, weighted Markov Model, Support Vector Machine, and Kalman Filter are used for comparison. Quantitative analysis demonstrates that our model outperforms the above-mentioned state-of-the-art models, and lays a good foundation for decision-making in different scenarios.
Min Xu 0001, Quan Pan 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Computer Vision-Assisted 3D Object Localization via COTS RFID Devices and a Monocular Camera
abstract
In most RFID localization systems, acquiring a reader antenna's position at each sampling time is challenging, especially for those antenna-carrying robot or drone systems with unpredictable trajectories. In this article, we present RF-MVO that fuses RFID and computer vision for stationary RFID localization in 3D space by attaching a light-weight 2D monocular camera to two reader antennas in parallel. First, the existing monocular visual odometry only recovers a camera/antenna trajectory in the camera view from 2D images. By combining it with RF phase, we design a model to estimate a scale factor for real-world trajectory transformation, along with spatial directions of an RFID tag relative to a virtual antenna array due to the mobility of each antenna. Then we propose a novel RFID localization algorithm that does not require exhaustively searching all possible positions within the pre-specified region. Second, to speed up the searching process and improve localization accuracy, we propose a coarse-to-fine optimization algorithm. Third, we introduce the concept of horizontal dilution of precision (HDOP) to measure the confidence level of localization results. Our experiments demonstrate the effectiveness of proposed algorithms and show RF-MVO can achieve 6.23 cm localization error.
Min Xu 0001, Ning Ye 0004, Fu Xiao 0001, Ruchuan Wang 0001, Haiping Huang
IEEE Trans. Mob. Comput.2
2021 Context-Dependent Propagating-Based Video Recommendation in Multimodal Heterogeneous Information Networks
abstract
With the emergence of online social networks (OSNs), video recommendation has come to play a crucial role in mitigating the semantic gap between users and videos. Conventional approaches to video recommendation primarily focus on exploiting content features or simple user-video interactions to model the users’ preferences. Although these methods have achieved promising results, they fail to model the complex video context interdependency, which is obscure/hidden in heterogeneous auxiliary data from OSNs. In this paper, we study the problem of video recommendation in Heterogeneous Information Networks (HINs) due to its excellence in characterizing heterogeneous and complex context information. We propose a Context-Dependent Propagating Recommendation network (CDPRec) to obtain accurate video embedding and capture global context cues among videos in HINs. The CDPRec can iteratively propagate the contexts of a video along links in a graph-structured HIN and explore multiple types of dependencies among the surrounding video nodes. Then, each video is represented as the composition of the multimodal content feature and global dependency structure information using an attention network. The learned video embedding with sequential based recommendation are jointly optimized for the final rating prediction. Experimental results on real-world YouTube video recommendation scenarios demonstrate the effectiveness of the proposed methods compared with strong baselines.
Lei Sang 0001, Min Xu 0001, Shengsheng Qian, Matt Martin, Peter Li, Xindong Wu 0001
IEEE Trans. Multim.2
2021 Weakly Supervised Emotion Intensity Prediction for Recomi/tmi40.htmlgnition of Emotions in Images
abstract
Recognition of emotions in images is attracting increasing research attention. Recent studies show that using local region information helps to improve the recognition performance. Intuitively, emotion intensity maps provide more detailed information than image regions. Inspired by this intuition, we propose an end-to-end deep neural network for image emotion recognition leveraging emotion intensity learning. The proposed network is composed of a first classification stream, an intensity prediction stream and a second classification stream. The intensity prediction stream is built on top of the feature pyramid network to extract multilevel features. The class activation mapping technique is used to generate pseudo intensity maps from the first classification stream to guide the proposed network for emotion intensity learning. The predicted intensity map is integrated into the second classification stream for final emotion recognition. The three streams are trained cooperatively to improve the performance. We evaluate the proposed network for both emotion recognition and sentiment classification on different benchmark datasets. The experimental results demonstrate that the proposed network achieves improved performance compared to previous state-of-the-art approaches.
Haimin Zhang 0001, Min Xu 0001
IEEE Trans. Multim.2
2020 Infrared Target Detection Using Intensity Saliency And Self-Attention
abstract
Infrared target detection is essential for many computer vision tasks. Generally, the IR images present common infrared characteristics, such as poor texture information, low resolution, and high noise. However, these characteristics are ignored in the existing detection methods, making them fail in real-world scenarios. In this paper, we take infrared intensity into account and propose a novel backbone network named Deep-IRTarget. We first extract infrared intensity saliency by a convolution with a Gaussian kernel filtering the images in the frequency domain. We then propose the triple self-attention network to further extract spatial domain image saliency by selectively emphasize interdependent semantic features in each channel. Jointly exploiting infrared characteristics in the frequency domain and the overall semantic interdependencies in the spatial domain, the proposed Deep-IRTarget outperforms existing methods in real-world Infrared target detection tasks. Experimental results on two infrared imagery datasets demonstrate the superiorly of our model.
Ruiheng Zhang 0001, Min Xu 0001, Yaxin Shi, Jian Fan, Chengpo Mu, Lixin Xu 0001
ICIP2
2020 Multi-camera Sports Players 3D Localization with Identification Reasoning
abstract
Multi-camera sports players 3D localization is always a challenging task due to heavy occlusions in crowded sports scenes. Traditional methods can only provide players' locations without identifying information. Existing methods of localization may cause ambiguous detection and unsatisfactory precision and recall, especially when heavy occlusions occur. To solve this problem, we propose a generic localization method by providing distinguishable results that have the probabilities of locations being occupied by players with unique ID labels. We design the algorithms with a multi-dimensional Bayesian model to create a Probabilistic and Identified Occupancy Map (PIOM). By using this model, we jointly apply deep learning-based object segmentation and identification to obtain sports players' probable positions and their likely identification labels. This approach not only provides players 3D locations but also gives their ID information that is distinguishable from others. Experimental results demonstrate that our method outperforms the previous localization approaches with reliable and distinguishable outcomes.
Yukun Yang 0003, Ruiheng Zhang 0001, Wanneng Wu, Min Xu 0001
ICPR5
2020 Data-Driven Characterization and Detection of COVID-19 Themed Malicious Websites
abstract
COVID-19 has hit hard on the global community, and organizations are working diligently to cope with the new norm of "work from home". However, the volume of remote work is unprecedented and creates opportunities for cyber attackers to penetrate home computers. Attackers have been leveraging websites with COVID-19 related names, dubbed COVID-19 themed malicious websites. These websites mostly contain false information, fake forms, fraudulent payments, scams, or malicious payloads to steal sensitive information or infect victims' computers. In this paper, we present a data-driven study on characterizing and detecting COVID-19 themed malicious websites. Our characterization study shows that attackers are agile and are deceptively crafty in designing geolocation targeted websites, often leveraging popular domain registrars and top-level domains. Our detection study shows that the Random Forest classifier can detect COVID-19 themed malicious websites based on the lexical and WHOIS features defined in this paper, achieving a 98% accuracy and 2.7% false-positive rate.
Mir Mehedi Ahsan Pritom, Kristin M. Schweitzer, Raymond M. Bateman, Min Xu 0001, Shouhuai Xu
ISI4
2020 Characterizing the Landscape of COVID-19 Themed Cyberattacks and Defenses
abstract
COVID-19 (Coronavirus) hit the global society and economy with a big surprise. In particular, work-from-home has become a new norm for employees. Despite the fact that COVID-19 can equally attack innocent people and cyber criminals, it is ironic to see surges in cyberattacks leveraging COVID-19 as a theme, dubbed COVID-19 themed cyberattacks or COVID-19 attacks for short, which represent a new phenomenon that has yet to be systematically understood. In this paper, we make a first step towards fully characterizing the landscape of these attacks, including their sophistication via the Cyber Kill Chain model. We also explore the solution space of defenses against these attacks.
Mir Mehedi Ahsan Pritom, Kristin M. Schweitzer, Raymond M. Bateman, Min Xu 0001, Shouhuai Xu
ISI4
2020 RF-Mirror: Mitigating Mutual Coupling Interference in Two-Tag Array Labeled RFID Systems
abstract
Recent RFID systems start attaching a tag array consisting of two or more tags on an object to deal with polarization mismatch and RF phase periodicity for battery-free sensing and localization. The multi-tag solution can also provide target orientation estimation. However, when these tags are closely spaced apart, mutual coupling will be induced, producing the unexpected changes in reported RSSI and RF phase. In this paper, we present RF-Mirror that enables compensating the distortion in a two-tag array labeled RFID system. The system would output the accurate difference in tag-to-antenna distances between two tags, which is a fundamental parameter in previous works for use. Firstly, we model the backscatter signal of a responding tag in a two-tag scenario, and then formulate novel RSSI- and RF phase-distance models with coupling terms. Secondly, we design an algorithm to characterize the coupling effect on tag gain by fusing RSSI and RF phase. Thirdly, we design a decoupling algorithm based on an observation that tag mutual coupling is independent of the position of a tag array relative to a reader antenna. Our experiments show the effectiveness of our models and RF-Mirror achieves the decoupling error of 0.197 cm in calculating the tag-to-antenna distance difference.
Min Xu 0001, Ning Ye 0004, Haiping Huang, Ruchuan Wang 0001, Fu Xiao 0001
SECON2
2020 Multi-camera 3D ball tracking framework for sports video
abstract
Accurate ball tracking in sports is vital for automatic sports analysis yet it is challenging mainly due to the small size and occlusions. This study proposes a novel multi‐camera 3D ball tracking (MBT) framework for sports video. The proposed framework consists of four parts: 2D ball detection, 2D ball tracking, 3D position fusion, and 3D ball tracking. In 2D aspect, the multi‐scale features are introduced to enhance the 2D ball detection, and the 2D ball tracking is also improved by exploring cross‐view information to handle the occlusion and timely updating tracking model with detection results to alleviate the problem of tracking drift. For 3D ball, a novel 3D position fusion method is proposed to optimise the ball position and the 3D ball tracking approach with improved Kalman filter is finally applied to ensure a smooth 3D ball trajectory. Moreover, compared to the existing products in commercial, the proposed framework does not require any special equipment and is thus low cost. Extensive experiments for 2D and 3D ball on a public dataset demonstrate that the proposed framework is robust to ball tracking in sports video, even in the presence of environmental interferences, substantial occlusions, and even calibration errors.
Wanneng Wu, Min Xu 0001, Qiaokang Liang, Li Mei
IET Image Process.2
2020 LSTM-Cubic A*-based auxiliary decision support system in air traffic management
Quan Pan 0001, Min Xu 0001
Neurocomputing3
2020 Improving the generalization performance of deep networks by dual pattern learning with adversarial adaptation
abstract
In this paper, we present a dual pattern learning network architecture with adversarial adaptation (DPLAANet). Unlike conventional networks, the proposed network has two input branches and two loss functions. This architecture forces the network to learn robust features by analysing dual inputs. The dual input structure allows the network to have a considerably large number of image pairs, which can help address the overfitting issue due to limited training data . In addition, we propose to associate the two input branches with two random interest values during training. As a stochastic regularization technique, this method can improve the generalization performance . Moreover, we introduce to use the adversarial training approach to reduce the domain difference between fused image features and single image features . Extensive experiments on CIFAR-10, CIFAR-100, FI-8, the Google commands dataset, and MNIST demonstrate that our DPLAANets exhibit better performance than the baseline networks. The experimental results on subsets of CIFAR-10, CIFAR-100, and MNIST demonstrate that DPLAANets have a good generalization performance on small datasets. The proposed architecture can be easily extended to have more than two input branches. The experimental results on subsets of MNIST show that the architecture with three branches outperforms two branches when the training set is extremely small.
Haimin Zhang 0001, Min Xu 0001
Knowl. Based Syst.2
2020 Manifold feature integration for micro-expression recognition
Madhumita A. Takalkar, Min Xu 0001, Zenon Chaczko
Multim. Syst.2
2020 Learning Multi-level Deep Representations for Image Emotion Classification
Tianrong Rao, Min Xu 0001
Neural Process. Lett.3
2020 Multi-camera multi-player tracking with deep player identification in sports video
Ruiheng Zhang 0001, Lingxiang Wu, Yukun Yang 0003, Wanneng Wu, Yueqiang Chen, Min Xu 0001
Pattern Recognit.6
2020 A Distributed and Anonymous Data Collection Framework Based on Multilevel Edge Computing Architecture
abstract
Industrial Internet of Things applications demand trustworthiness in terms of quality of service (QoS), security, and privacy, to support the smooth transmission of data. To address these challenges, in this article, we propose a distributed and anonymous data collection (DaaC) framework based on a multilevel edge computing architecture. This framework distributes captured data among multiple level-one edge devices (LOEDs) to improve the QoS and minimize packet drop and end-to-end delay. Mobile sinks are used to collect data from LOEDs and upload to cloud servers. Before data collection, the mobile sinks are registered with a level-two edge-device to protect the underlying network. The privacy of mobile sinks is preserved through group-based signed data collection requests. Experimental results show that our proposed framework improves QoS through distributed data transmission. It also helps in protecting the underlying network through a registration scheme and preserves the privacy of mobile sinks through group-based data collection requests.
Muhammad Usman 0015, Mian Ahmad Jan, Alireza Jolfaei, Min Xu 0001, Xiangjian He, Jinjun Chen
IEEE Trans. Ind. Informatics4
2020 Recall What You See Continually Using GridLSTM in Image Captioning
abstract
The goal of image captioning is to automatically describe an image with a sentence, and the task has attracted research attention from both the computer vision and natural-language processing research communities. The existing encoder-decoder model and its variants, which are the most popular models for image captioning, use the image features in three ways: first, they inject the encoded image features into the decoder only once at the initial step, which does not enable the rich image content to be explored sufficiently while gradually generating a text caption; second, they concatenate the encoded image features with text as extra inputs at every step, which introduces unnecessary noise; and, third, they using an attention mechanism, which increases the computational complexity due to the introduction of extra neural nets to identify the attention regions. Different from the existing methods, in this paper, we propose a novel network, Recall Network, for generating captions that are consistent with the images. The recall network selectively involves the visual features by using a GridLSTM and, thus, is able to recall image contents while generating each word. By importing the visual information as the latent memory along the depth dimension LSTM, the decoder is able to admit the visual features dynamically through the inherent LSTM structure without adding any extra neural nets or parameters. The Recall Network efficiently prevents the decoder from deviating from the original image content. To verify the efficiency of our model, we conducted exhaustive experiments on full and dense image captioning. The experimental results clearly demonstrate that our recall network outperforms the conventional encoder-decoder model by a large margin and that it performs comparably to the state-of-the-art methods.
Lingxiang Wu, Min Xu 0001, Jinqiao Wang, Stuart W. Perry
IEEE Trans. Multim.2
2020 Image to Modern Chinese Poetry Creation via a Constrained Topic-aware Model
abstract
Artificial creativity has attracted increasing research attention in the field of multimedia and artificial intelligence. Despite the promising work on poetry/painting/music generation, creating modern Chinese poetry from images, which can significantly enrich the functionality of photo-sharing platforms, has rarely been explored. Moreover, existing generation models cannot tackle three challenges in this task: (1) Maintaining semantic consistency between images and poems; (2) preventing topic drift in the generation; (3) avoidance of certain words appearing frequently. These three points are even common challenges in other sequence generation tasks. In this article, we propose a Constrained Topic-aware Model (CTAM) to create modern Chinese poetries from images regarding the challenges above. Without image-poetry paired dataset, we construct a visual semantic vector to embed visual contents via image captions. For the topic-drift problem, we propose a topic-aware poetry generation model. Additionally, we design an Anti-frequency Decoding (AFD) scheme to constrain high-frequency characters in the generation. Experimental results show that our model achieves promising performance and is effective in poetry’s readability and semantic consistency.
Lingxiang Wu, Min Xu 0001, Shengsheng Qian, Jianwei Cui 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Airborne Object Detection Using Hyperspectral Imaging: Deep Learning Review
Thuy T. Pham, Madhumita A. Takalkar, Min Xu 0001, Dinh Thai Hoang, H. A. Truong, Eryk Dutkiewicz, Stuart W. Perry
ICCSA (1)3
2019 Structured Modeling of Joint Deep Feature and Prediction Refinement for Salient Object Detection
abstract
Recent saliency models extensively explore to incorporate multi-scale contextual information from Convolutional Neural Networks (CNNs). Besides direct fusion strategies, many approaches introduce message-passing to enhance CNN features or predictions. However, the messages are mainly transmitted in two ways, by feature-to-feature passing, and by prediction-to-prediction passing. In this paper, we add message-passing between features and predictions and propose a deep unified CRF saliency model . We design a novel cascade CRFs architecture with CNN to jointly refine deep features and predictions at each scale and progressively compute a final refined saliency map. We formulate the CRF graphical model that involves message-passing of feature-feature, feature-prediction, and prediction-prediction, from the coarse scale to the finer scale, to update the features and the corresponding predictions. Also, we formulate the mean-field updates for joint end-to-end model training with CNN through back propagation. The proposed deep unified CRF saliency model is evaluated over six datasets and shows highly competitive performance among the state of the arts.
Yingyue Xu, Dan Xu 0002, Xiaopeng Hong, Wanli Ouyang, Rongrong Ji, Min Xu 0001, Guoying Zhao 0001
ICCV6
2019 Improving Micro-expression Recognition Accuracy Using Twofold Feature Extraction
Madhumita A. Takalkar, Haimin Zhang 0001, Min Xu 0001
MMM (1)3
2019 AAANE: Attention-Based Adversarial Autoencoder for Multi-scale Network Embedding
Lei Sang 0001, Min Xu 0001, Shengsheng Qian, Xindong Wu 0001
PAKDD (3)2
2019 Multi-level region-based Convolutional Neural Network for image emotion classification
Tianrong Rao, Haimin Zhang 0001, Min Xu 0001
Neurocomputing4
2019 Multi-modal multi-view Bayesian semantic embedding for community question answering
Lei Sang 0001, Min Xu 0001, Shengsheng Qian, Xindong Wu 0001
Neurocomputing2
2019 Error Concealment for Cloud-Based and Scalable Video Coding of HD Videos
abstract
The encoding of HD videos faces two challenges: requirements for a strong processing power and a large storage space. One time-efficient solution addressing these challenges is to use a cloud platform and to use a scalable video coding technique to generate multiple video streams with varying bit-rates. Packet-loss is very common during the transmission of these video streams over the Internet and becomes another challenge. One solution to address this challenge is to retransmit lost video packets, but this will create end-to-end delay. Therefore, it would be good if the problem of packet-loss can be dealt with at the user's side. In this paper, we present a novel system that encodes and stores the videos using the Amazon cloud computing platform, and recover lost video frames on user side using a new Error Concealment (EC) technique. To efficiently utilize the computation power of a user's mobile device, the EC is performed based on a multiple-thread and parallel process. The simulation results clearly show that, on average, our proposed EC technique outperforms the traditional Block Matching Algorithm (BMA) and the Frame Copy (FC) techniques.
Muhammad Usman 0015, Xiangjian He, Kin-Man Lam 0001, Min Xu 0001, Syed Mohsin Matloob Bokhari, Jinjun Chen, Mian Ahmad Jan
IEEE Trans. Cloud Comput.4
2018 RF-MVO: Simultaneous 3D Object Localization and Camera Trajectory Recovery Using RFID Devices and a 2D Monocular Camera
abstract
Most of the existing RFID-based localization systems cannot well locate RFID-tagged objects in a 3D space. Limited robot-based RFID solutions require reader antennas to be carried by a robot moving along an already-known trajectory at a constant speed. As the first attempt, this paper presents RF-MVO, which fuses battery-free RFID and monocular visual odometry to locate stationary RFID tags in a 3D space and recover an unknown trajectory of reader antennas binding with a 2D monocular camera. The proposed hybrid system exhibits three unique features. Firstly, since the trajectory of a 2D monocular camera can only be recovered up to an unknown scale factor, RF-MVO combines the relative-scale camera trajectory with depth-enabled RF phase to estimate an absolute scale factor and spatially incident angles of an RFID tag. Secondly, we propose a joint optimization algorithm consisting of coarse-to-fine angular refinement, 3D tag localization and parameter nonlinear optimization, to improve real-time performance. Thirdly, RF-MVO can determine the effect of relative tag-antenna geometry on the estimation precision, providing optimal tag positions and absolute scale factors. Our experiments show that RF-MVO can achieve 6.23cm tag localization accuracy in a 3D space and 0.0158 absolute scale factor estimation accuracy for camera trajectory recovery.
Min Xu 0001, Ning Ye 0004, Ruchuan Wang 0001, Haiping Huang
ICDCS2
2018 LSTM-based Flight Trajectory Prediction
abstract
Safety ranks the first in Air Traffic Management (ATM). Accurate trajectory prediction can help ATM to forecast potential dangers and effectively provide instructions for safely traveling. Most trajectory prediction algorithms work for land traffic, which rely on points of interest (POIs) and are only suitable for stationary road condition. Compared with land traffic prediction, flight trajectory prediction is very difficult because way-points are sparse and the flight envelopes are heavily affected by external factors. In this paper, we propose a flight trajectory prediction model based on a Long Short-Term Memory (LSTM) network. The four interacting layers of a repeating module in an LSTM enables it to connect the long-term dependencies to present predicting task. Applying sliding windows in LSTM maintains the continuity and avoids compromising the dynamic dependencies of adjacent states in the long-term sequences, which helps to improve accuracy of trajectory prediction. Taking time dimension into consideration, both 3-D (time stamp, latitude and longitude) and 4-D (time stamp, latitude, longitude and altitude) trajectories are predicted to prove the efficiency of our approach. The dataset we use was collected by ADS-B ground stations. We evaluate our model by widely used measurements, such as the mean absolute error (MAE), the mean relative error (MRE), the root mean square error (RMSE) and the dynamic warping time (DWT) methods. As Markov Model is the most popular in time series processing, comparisons among Markov Model (MM), weighted Markov Model (wMM) and our model are presented. Our model outperforms the existing models (MM and wMM) and provides a strong basis for abnormal detection and decision-making.
Min Xu 0001, Quan Pan 0001, Bing Yan 0001, Haimin Zhang 0001
IJCNN2
2018 ASMMC-MMAC 2018: The Joint Workshop of 4th the Workshop on Affective Social Multimedia Computing and first Multi-Modal Affective Computing of Large-Scale Multimedia Data Workshop
abstract
Affective social multimedia computing is an emergent research topic for both affective computing and multimedia research communities. Social multimedia is fundamentally changing how we communicate, interact, and collaborate with other people in our daily lives. Social multimedia contains much affective information. Effective extraction of affective information from social multimedia can greatly help social multimedia computing (e.g., processing, index, retrieval, and understanding). Besides, with the rapid development of digital photography and social networks, people get used to sharing their lives and expressing their opinions online. As a result, user-generated social media data, including text, images, audios, and videos, grow rapidly, which urgently demands advanced techniques on the management, retrieval, and understanding of these data.
Dong-Yan Huang, Sicheng Zhao, Björn W. Schuller, Hongxun Yao, Jianhua Tao 0001, Min Xu 0001, Lei Xie 0001, Qingming Huang
ACM Multimedia6
2018 Appearance features in Encoding Color Space for visual surveillance
Lingxiang Wu, Min Xu 0001, Guibo Zhu, Jinqiao Wang, Tianrong Rao
Neurocomputing2
2018 Generating affective maps for images
Tianrong Rao, Min Xu 0001
Multim. Tools Appl.2
2018 A survey: facial micro-expression recognition
Madhumita A. Takalkar, Min Xu 0001, Qiang Wu 0001, Zenon Chaczko
Multim. Tools Appl.2
2018 A Joint Framework for QoS and QoE for Video Transmission over Wireless Multimedia Sensor Networks
abstract
With the emergence of Wireless Multimedia Sensor Networks (WMSNs), the distribution of multimedia contents have now become a reality. Without proper management, the transmission of multimedia data over WMSNs affects the performance of networks due to excessive packet-drop. The existing studies on Quality of Service (QoS) mostly deal with simple Wireless Sensor Networks (WSNs) and as such do not account for an increasing number of sensor nodes and an increasing volume of data. In this paper, we propose a novel framework to support QoS in WMSNs along with a light-weight Error Concealment (EC) scheme. The EC schemes play a vital role to enhance Quality of Experience (QoE) by maintaining an acceptable quality at the receiving ends. The main objectives of the proposed framework are to maximize the network throughput and to cover-up the effects produced by dropped video packets. To control the data-rate, Scalable High efficiency Video Coding (SHVC) is applied at multimedia sensor nodes with variable Quantization Parameters (QPs). Multi-path routing is exploited to support real-time video transmission. Experimental results show that the proposed framework can efficiently adjust large volumes of video data under certain network distortions and can effectively conceal lost video frames by producing better objective measurements.
Muhammad Usman 0015, Ning Yang 0003, Mian Ahmad Jan, Xiangjian He, Min Xu 0001, Kin-Man Lam 0001
IEEE Trans. Mob. Comput.5
2018 Recognition of Emotions in User-Generated Videos With Kernelized Features
abstract
Recognition of emotions in user-generated videos has attracted increasing research attention. Most existing approaches are based on spatial features extracted from video frames. However, due to the broad affective gap between spatial features of images and high-level emotions, the performance of existing approaches is restricted. To bridge the affective gap, we propose recognizing emotions in user-generated videos with kernelized features. We reformulate the equation of the discrete Fourier transform as a linear kernel function and construct a polynomial kernel function based on the linear kernel. The polynomial kernel is applied to spatial features of video frames to generate kernelized features. Compared with spatial features, kernelized features show superior discriminative capability. Moreover, we are the first to apply the sparse representation method to reduce the impact of noise contained in videos; this method helps contribute to performance improvement. Extensive experiments are conducted on two challenging benchmark datasets, that is, VideoEmotion-8 and Ekman-6. The experimental results demonstrate that the proposed method achieves state-of-the-art performance.
Haimin Zhang 0001, Min Xu 0001
IEEE Trans. Multim.2
2017 Dependency Exploitation: A Unified CNN-RNN Approach for Visual Emotion Recognition
abstract
Visual emotion recognition aims to associate images with appropriate emotions. There are different visual stimuli that can affect human emotion from low-level to high-level, such as color, texture, part, object, etc. However, most existing methods treat different levels of features as independent entity without having effective method for feature fusion. In this paper, we propose a unified CNN-RNN model to predict the emotion based on the fused features from different levels by exploiting the dependency among them. Our proposed architecture leverages convolutional neural network (CNN) with multiple layers to extract different levels of features with in a multi-task learning framework, in which two related loss functions are introduced to learn the feature representation. Considering the dependencies within the low-level and high-level features, a new bidirectional recurrent neural network (RNN) is proposed to integrate the learned features from different layers in the CNN model. Extensive experiments on both Internet images and art photo datasets demonstrate that our method outperforms the state-of-the-art methods with at least 7% performance improvement.
Xinge Zhu, Liang Li 0003, Weigang Zhang, Tianrong Rao, Min Xu 0001, Qingming Huang, Dong Xu 0001
IJCAI5
2017 User relationship strength modeling for friend recommendation on Instagram
Dongyan Guo, Jingsong Xu, Jian Zhang 0002, Min Xu 0001, Xiangjian He
Neurocomputing4
2017 Hierarchically Supervised Deconvolutional Network for Semantic Video Segmentation
Jing Liu 0001, Yong Li 0034, Jun Fu 0005, Min Xu 0001, Hanqing Lu
Pattern Recognit.5
2017 Who Are Your "Real" Friends: Analyzing and Distinguishing Between Offline and Online Friendships From Social Multimedia Data
abstract
The Internet has extended the physical boundary of people's social circles to manage an inordinate number of online friends. It is recognized that only a fraction of these online friends are also known with each other in offline circumstances, i.e., the offline friends. An important type of offline friend, onsite offline friend, is defined and addressed in this paper. We explores the possibility of utilizing users' online photo sharing-related behaviors and network topologies to analyze and distinguish between online and onsite offline friendships. Different from traditional social science studies which rely on survey-based data, we employ users' tagged people on the shared Instagram photos as the ground-truth for onsite offline friends. This enables a large-scale and objective analysis and experimental evaluation, which compares between different factors and identifies the features that are key to onsite offline friend identification.
Dongyuan Lu, Jitao Sang 0001, Zhineng Chen, Min Xu 0001, Tao Mei 0001
IEEE Trans. Multim.4
2016 Multi-scale blocks based image emotion classification using multiple instance learning
abstract
Emotional factors usually affect users' preferences for and evaluations of images. Although affective image analysis attracts increasing attention, there are still three major challenges remaining: 1) it is difficult to classify an image into a single emotion type since different regions within an image can represent different emotions; 2) there is a gap between low-level features and high-level emotions and 3) it is difficult to collect a training set of reliable emotional image content. To address these three issues, we propose an emotion classification method based on multi-scale blocks using Multiple Instance Learning (MIL). We firstly extract blocks of an image at multiple scales using different image segmentation methods pyramid segmentation and simple linear iterative clustering (SLIC) and represent each block using the bag-of-visual-words (BoVW) method. Then, to bridge the “affective gap”, probabilistic latent semantic analysis (pLSA) is employed to estimate the latent topic distribution as a mid-level representation of each block. Finally, MIL, which reduces the need for exact labelling, is employed to classify the dominant emotion type of the image. Experiments carried out on three widely used datasets demonstrate that our proposed method with S-LIC effectively improves the state-of-the-art results of image emotion classification 5.1% on average.
Tianrong Rao, Min Xu 0001, Jinqiao Wang, Ian S. Burnett
ICIP2
2016 Modeling temporal information using discrete fourier transform for recognizing emotions in user-generated videos
abstract
With the widespread of user-generated Internet videos, emotion recognition in those videos attracts increasing research efforts. However, most existing works are based on framelevel visual features and/or audio features, which might fail to model the temporal information, e.g. characteristics accumulated along time. In order to capture video temporal information, in this paper, we propose to analyse features in frequency domain transformed by discrete Fourier transform (DFT features). Frame-level features are firstly extract by a pre-trained deep convolutional neural network (CNN). Then, time domain features are transferred and interpolated into DFT features. CNN and DFT features are further encoded and fused for emotion classification. By this way, static image features extracted from a pre-trained deep CNN and temporal information represented by DFT features are jointly considered for video emotion recognition. Experimental results demonstrate that combining DFT features can effectively capture temporal information and therefore improve emotion recognition performance. Our approach has achieved a state-of-the-art performance on the largest video emotion dataset (VideoEmotion-8 dataset), improving accuracy from 51.1% to 55.6%.
Haimin Zhang 0001, Min Xu 0001
ICIP2
2016 Person re-identification via rich color-gradient feature
abstract
Person re-identification refers to match the same pedestrian across disjoint views in non-overlapping camera networks. Lots of local and global features in the literature are put forward to solve the matching problem, where color feature is robust to viewpoint variance and gradient feature provides a rich representation robust to illumination change. However, how to effectively combine the color and gradient features is an open problem. In this paper, to effectively leverage the color-gradient property in multiple color spaces, we propose a novel Second Order Histogram feature (SOH) for person reidentification in large surveillance dataset. Firstly, we utilize discrete encoding to transform commonly used color space into Encoding Color Space (ECS), and calculate the statistical gradient features on each color channel. Then, a second order statistical distribution is calculated on each cell map with a spatial partition. In this way, the proposed SOH feature effectively leverages the statistical property of gradient and color as well as reduces the redundant information. Finally, a metric learned by KISSME [1] with Mahalanobis distance is used for person matching. Experimental results on three public datasets, VIPeR, CAVIAR and CUHK01, show the promise of the proposed approach.
Lingxiang Wu, Jinqiao Wang, Guibo Zhu, Min Xu 0001, Hanqing Lu
ICME4
2016 ActiveAd: A novel framework of linking ad videos to online products
Jinqiao Wang, Min Xu 0001, Hanqing Lu, Ian S. Burnett
Neurocomputing2
2016 A unified model sharing framework for moving object detection
Yingying Chen 0003, Jinqiao Wang, Min Xu 0001, Xiangjian He, Hanqing Lu
Signal Process.3
2016 Adaptive Content Condensation Based on Grid Optimization for Thumbnail Image Generation
abstract
An ideal thumbnail generator should effectively condense unimportant regions and keep the important content undeformed, completed, and at a proper scale, i.e., accuracy, completeness, and sufficiency. Each retargeting method has its own advantage for resizing arbitrary images. However, they often ignore the completeness and sufficiency for information presentation in thumbnails. In this paper, we formulate thumbnail generation as an image content condensation problem and propose a unified grid optimization framework to fuse multiple operators. From the view of accuracy, completeness, and sufficiency for information presentation, we exploit complementary relationships among three condensation operators and fuse them into a unified grid-based convex programming problem, which could be solved simultaneously and efficiently through numerical optimization. Besides warping energy to preserve the geometric structure of important objects, we put forward two grid-based energy terms to keep the completeness of important objects and retain them at a proper size. Finally, an adaptive procedure is proposed to dynamically adjust the contribution of loss functions for achieving optimal content condensation. Both qualitative and quantitative comparison results demonstrate that the proposed method achieves an excellent tradeoff among accuracy, completeness, and sufficiency of information preservation. The experimental results show that our approach is obviously superior to the state-of-the-art techniques.
Jinqiao Wang, Yingying Chen 0003, Tao Mei 0001, Min Xu 0001, La Zhang, Hanqing Lu
IEEE Trans. Circuits Syst. Video Technol.5
2016 Frame Interpolation for Cloud-Based Mobile Video Streaming
abstract
Cloud-based High Definition (HD) video streaming is becoming popular day by day. On one hand, it is important for both end users and large storage servers to store their huge amount of data at different locations and servers. On the other hand, it is becoming a big challenge for network service providers to provide reliable connectivity to the network users. There have been many studies over cloud-based video streaming for Quality of Experience (QoE) for services like YouTube. Packet losses and bit errors are very common in transmission networks, which affect the user feedback over cloud-based media services. To cover up packet losses and bit errors, Error Concealment (EC) techniques are usually applied at the decoder/receiver side to estimate the lost information. This paper proposes a time-efficient and quality-oriented EC method. The proposed method considers H.265/HEVC based intra-encoded videos for the estimation of whole intra-frame loss. The main emphasis in the proposed approach is the recovery of Motion Vectors (MVs) of a lost frame in real-time. To boost-up the search process for the lost MVs, a bigger block size and searching in parallel are both considered. The simulation results clearly show that our proposed method outperforms the traditional Block Matching Algorithm (BMA) by approximately 2.5 dB and Frame Copy (FC) by up to 12 dB at a packet loss rate of 1%, 3%, and 5% with different Quantization Parameters (QPs). The computational time of the proposed approach outperforms the BMA by approximately 1788 seconds.
Muhammad Usman 0015, Xiangjian He, Kin-Man Lam 0001, Min Xu 0001, Syed Mohsin Matloob Bokhari, Jinjun Chen
IEEE Trans. Multim.4
2016 Improving Visual Saliency Computing With Emotion Intensity
abstract
Saliency maps that integrate individual feature maps into a global measure of visual attention are widely used to estimate human gaze density. Most of the existing methods consider low-level visual features and locations of objects, and/or emphasize the spatial position with center prior. Recent psychology research suggests that emotions strongly influence human visual attention. In this paper, we explore the influence of emotional content on visual attention. On top of the traditional bottom-up saliency map generation, our saliency map is generated in cooperation with three emotion factors, i.e., general emotional content, facial expression intensity, and emotional object locations. Experiments, carried out on National University of Singapore Eye Fixation (a public eye tracking data set), demonstrate that incorporating emotion does improve the quality of visual saliency maps computed by bottom-up approaches for the gaze density estimation. Our method increases about 0.1 on an average of area under the curve of receiver operation characteristic curve, compared with the four baseline bottom-up approaches (Itti's, attention based on information maximization, saliency using natural, and graph-based vision saliency).
Min Xu 0001, Jinqiao Wang, Tianrong Rao, Ian S. Burnett
IEEE Trans. Neural Networks Learn. Syst.2
2015 A Survey of Applying Machine Learning Techniques for Credit Rating: Existing Models and Open Issues
Min Xu 0001, Özgür Tolga Pusatli
ICONIP (2)2
2015 Learning Multi-view Deep Features for Small Object Retrieval in Surveillance Scenarios
abstract
With the explosive growth of surveillance videos, object retrieval has become a significant task for security monitoring. However, visual objects in surveillance videos are usually of small size with complex light conditions, view changes and partial occlusions, which increases the difficulty level of efficiently retrieving objects of interest in a large-scale dataset. Although deep features have achieved promising results on object classification and retrieval and have been verified to contain rich semantic structure property, they lack of adequate color information, which is as crucial as structure information for effective object representation. In this paper, we propose to leverage discriminative Convolutional Neural Network (CNN) to learn deep structure and color feature to form an efficient multi-view object representation. Specifically, we utilize CNN trained on ImageNet to abstract rich semantic structure information. Meanwhile, we propose a CNN model supervised by 11 color names to extract deep color features. Compared with traditional color descriptors, deep color features can capture the common color property across different illumination conditions. Then, the complementary multi-view deep features are encoded into short binary codes by Locality-Sensitive Hash (LSH) and fused to retrieve objects. Retrieval experiments are performed on a dataset of 100k objects extracted from multi-camera surveillance videos. Comparison results with several popular visual descriptors show the effectiveness of the proposed approach.
Haiyun Guo, Jinqiao Wang, Min Xu 0001, Zhengjun Zha, Hanqing Lu
ACM Multimedia3
2015 Community Detection Based on Links and Node Features in Social Networks
Fengli Zhang, Min Xu 0001, Xiangjian He
MMM (1)4
2015 Survey of Error Concealment techniques: Research directions and open issues
abstract
Error Concealment (EC) techniques use either spatial, temporal or a combination of both types of information to recover the data lost in transmitted video. In this paper, existing EC techniques are reviewed, which are divided into three categories, namely Intra-frame EC, Inter-frame EC, and Hybrid EC techniques. We first focus on the EC techniques developed for the H.264/AVC standard. The advantages and disadvantages of these EC techniques are summarized with respect to the features in H.264. Then, the EC algorithms are also analyzed. These EC algorithms have been recently adopted in the newly introduced H.265/HEVC standard. A performance comparison between the classic EC techniques developed for H.264 and H.265 is performed in terms of the average PSNR. Lastly, open issues in the EC domain are addressed for future research consideration.
Muhammad Usman 0015, Xiangjian He, Min Xu 0001, Kin-Man Lam 0001
PCS3
2015 NIF-based seam carving for image resizing
Dongyan Guo, Jundi Ding, Jinhui Tang 0001, Min Xu 0001, Chunxia Zhao
Multim. Syst.4
2015 A camera motion histogram descriptor for video shot classification
Muhammad Abul Hasan, Min Xu 0001, Xiangjian He, Yi Wang 0037
Multim. Tools Appl.2
2014 Estimate Gaze Density by Incorporating Emotion
abstract
Gaze density estimation has attracted many research efforts in the past years. The factors considered in the existing methods include low level feature saliency, spatial position, and objects. Emotion, as an important factor driving attention, has not been taken into account. In this paper, we are the first to estimate gaze density through incorporating emotion. To estimate the emotion intensity of each position in an image, we consider three aspects, generic emotional content, facial expression intensity, and emotional objects. Generic emotional content is estimated by using Multiple instance learning, which is employed to train an emotion detector from weakly labeled images. Facial expression intensity is estimated by using a ranking method. Emotional objects are detected, by taking blood/injury and worm/snake as examples. Finally, emotion intensity, low level feature saliency, and spatial position, are fused, through a linear support vector machine, to estimate gaze density. The performance is tested on public eye tracking dataset. Experimental results indicate that incorporating emotion does improve the performance of gaze density estimation.
Min Xu 0001, Xiangjian He, Jinqiao Wang
ACM Multimedia2
2014 Mask Assisted Object Coding with Deep Learning for Object Retrieval in Surveillance Videos
abstract
Retrieving visual object from a large-scale video dataset is one of multimedia research focuses but a challenging task due to imprecise object extraction and partial occlusion. This paper presents a novel approach to efficiently encode and retrieve visual objects, which addresses some practical complications in surveillance videos. Specifically, we take advantage of the mask information to assist object representation, and develop an encoding method by utilizing highly nonlinear mapping with a deep neural network. Furthermore, we add some occluded noise into the learning process to enhance the robustness of dealing with background noise and partial occlusions. A real-life surveillance video data containing over 10 million objects are built to evaluate the proposed approach. Experimental results show our approach significantly outperforms state-of-the-art solutions for object retrieval in large-scale video dataset.
Kezhen Teng, Jinqiao Wang, Min Xu 0001, Hanqing Lu
ACM Multimedia3
2014 A three-level framework for affective content analysis and its case studies
Min Xu 0001, Jinqiao Wang, Xiangjian He, Jesse S. Jin, Suhuai Luo, Hanqing Lu
Multim. Tools Appl.1
2014 A hybrid domain enhanced framework for video retargeting with spatial-temporal importance and 3D grid optimization
Jinqiao Wang, Min Xu 0001, Xiangjian He, Hanqing Lu, Doan B. Hoang
Signal Process.2
2014 CAMHID: Camera Motion Histogram Descriptor and Its Application to Cinematographic Shot Classification
abstract
In this paper, we propose a nonparametric camera motion descriptor for video shot classification. In the proposed method, a motion vector field (MVF) is constructed for each consecutive video frame by computing the motion vector (MV) of each macroblock. Then, the MVFs are divided into a number of local region of equal size. Next, the inconsistent/noisy MVs of each local region are eliminated by a motion consistency analysis. The remaining MVs of each local region from a number of consecutive frames are further collected for a compact representation. Initially, a matrix is formed using the MVs. Then, the matrix is decomposed using a singular value decomposition technique to represent the dominant motion. Finally, the angle of the most variance retaining principal component is computed and quantized to represent the motion of a local region by using a histogram. In order to represent the global camera motion, the local histograms are combined. The effectiveness of the proposed motion descriptor for video shot classification is tested by using a support vector machine. First, the proposed camera motion descriptors for video shots classification are computed on a video data set consisting of regular camera motion patterns (e.g., pan, zoom, tilt, static). Then, we apply the camera motion descriptors with an extended set of features to the classification of cinematographic shots. The experimental results show that the proposed shot level camera motion descriptor has a strong discriminative capability to classify different camera motion patterns of different videos effectively. We also show that our approach outperforms state-of-the-art methods.
Muhammad Abul Hasan, Min Xu 0001, Xiangjian He, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2014 Mobile Landmark Search with 3D Models
abstract
Landmark search is crucial to improve the quality of travel experience. Smart phones make it possible to search landmarks anytime and anywhere. Most of the existing work computes image features on smart phones locally after taking a landmark image. Compared with sending original image to the remote server, sending computed features saves network bandwidth and consequently makes sending process fast. However, this scheme would be restricted by the limitations of phone battery power and computational ability. In this paper, we propose to send compressed (low resolution) images to remote server instead of computing image features locally for landmark recognition and search. To this end, a robust 3D model based method is proposed to recognize query images with corresponding landmarks. Using the proposed method, images with low resolution can be recognized accurately, even though images only contain a small part of the landmark or are taken under various conditions of lighting, zoom, occlusions and different viewpoints. In order to provide an attractive landmark search result, a 3D texture model is generated to respond to a landmark query. The proposed search approach, which opens up a new direction, starts from a 2D compressed image query input and ends with a 3D model search result.
Weiqing Min, Changsheng Xu, Min Xu 0001, Xian Xiao, Bing-Kun Bao
IEEE Trans. Multim.3
2013 Semantically-Based Human Scanpath Estimation with HMMs
abstract
We present a method for estimating human scan paths, which are sequences of gaze shifts that follow visual attention over an image. In this work, scan paths are modeled based on three principal factors that influence human attention, namely low-level feature saliency, spatial position, and semantic content. Low-level feature saliency is formulated as transition probabilities between different image regions based on feature differences. The effect of spatial position on gaze shifts is modeled as a Levy flight with the shifts following a 2D Cauchy distribution. To account for semantic content, we propose to use a Hidden Markov Model (HMM) with a Bag-of-Visual-Words descriptor of image regions. An HMM is well-suited for this purpose in that 1) the hidden states, obtained by unsupervised learning, can represent latent semantic concepts, 2) the prior distribution of the hidden states describes visual attraction to the semantic concepts, and 3) the transition probabilities represent human gaze shift patterns. The proposed method is applied to task-driven viewing processes. Experiments and analysis performed on human eye gaze data verify the effectiveness of this method.
Dong Xu 0001, Qingming Huang, Wen Li 0001, Min Xu 0001, Stephen Lin 0001
ICCV5
2013 Hierarchical affective content analysis in arousal and valence dimensions
Min Xu 0001, Changsheng Xu, Xiangjian He, Jesse S. Jin, Suhuai Luo, Yong Rui
Signal Process.1
2013 Context-Aware Video Retargeting via Graph Model
abstract
Video retargeting is a crowded but challenging research area. In order to maximally comfort the viewers' watching experience, the most challenging issue is how to retain the spatial shape of important objects while ensure temporal smoothness and coherence. Existing retargeting techniques deal with these spatial-temporal requirements individually, which preserve the spatial geometry and temporal coherence for each region. However, the spatial-temporal property of the video content should be context-relevant, i.e., the regions belonging to the same object are supposed to undergo uniform spatial-temporal transformation. Regardless of the contextual information, the divide-and-rule strategy of existing techniques usually incurs various spatial-temporal artifacts. In order to achieve satisfactory spatial-temporal coherent video retargeting, in this paper, a novel context-aware solution is proposed via graph model. First, we employ a grid-based warping framework to preserve the spatial structure and temporal motion trend at the unit of grid cell. Second, we propose a graph-based motion layer partition algorithm to estimate motions of different regions, which simultaneously provides the evaluation of contextual relationship between grid cells while estimating the motions of regions. Third, complementing the salience-based spatial-temporal information preservation, two novel context constraints are encoded for encouraging the grid cells of the same object to undergo uniform spatial and temporal transformation, respectively. Finally, we formulate the objective function as a quadratic programming problem. Our method achieves a satisfactory spatial-temporal coherence while maximally avoiding the influence of artifacts. In addition, the grid-cell-wise motion estimation could be calculated every few frames, which obviously improves the speed. Experimental results and comparisons with state-of-the-art methods demonstrate the effectiveness and efficiency of our approach.
Jinqiao Wang, Min Xu 0001, Hanqing Lu
IEEE Trans. Multim.3
2012 Efficient Clothing Retrieval with Semantic-Preserving Visual Phrases
Jianlong Fu, Jinqiao Wang, Zechao Li, Min Xu 0001, Hanqing Lu
ACCV (2)4
2012 Fusing Warping, Cropping, and Scaling for Optimal Image Thumbnail Generation
Jinqiao Wang, Min Xu 0001, Hanqing Lu
ACCV (4)3
2012 Shot Classification Using Domain Specific Features for Movie Management
Muhammad Abul Hasan, Min Xu 0001, Xiangjian He, Ling Chen 0006
DASFAA (2)2
2012 On splitting dataset: Boosting Locally Adaptive Regression Kernels for car localization
abstract
In this paper, we study the impact of learning an Adaboost classifier with small sample set (i.e., with fewer training examples). In particular, we make use of car localization as an underlying application, because car localization can be widely used to various real world applications. In order to evaluate the performance of Adaboost learning with a few examples, we simply apply Adaboost learning to a recently proposed feature descriptor - Locally Adaptive Regression Kernel (LARK). As a type of state-of-the-art feature descriptor, LARK is robust against illumination changes and noises. More importantly, we use LARK because its spatial property is also favorable for our purpose (i.e., each patch in the LARK descriptor corresponds to one unique pixel in the original image). In addition to learning a detector from the entire training dataset, we also split the original training dataset into several sub-groups and then we train one detector for each sub-group. We compare those features associated using the detector of each sub-group with that of the detector learnt with the entire training dataset and propose improvements based on the comparison results. Our experimental results indicate that the Adaboost learning is only successful on a small dataset when those learnt features simultaneously satisfy two conditions that: 1. features are learnt from the Region of Interest (ROI), and 2. features are sufficiently far away from each other.
Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Min Xu 0001
ICARCV4
2012 Enhanced 3-D Modeling for Landmark Image Classification
abstract
Landmark image classification is a challenging task due to the various circumstances, e.g., illumination, viewpoint, zoom in/out and occlusion under which landmark images are taken. Most existing approaches utilize features extracted from the whole image including both landmark and non-landmark areas. However, non-landmark areas introduce redundant and noisy information. In this paper, we propose a novel approach to improve landmark image classification consisting of three steps. First, an attention-based 3-D reconstruction method is proposed to reconstruct sparse 3-D landmark models. Second, the sparse 3-D models are projected onto iconic images in order to identify images of the hot regions. For a landmark, hot regions are parts of a landmark which attract photographers' attention and are popularly captured in photos. These hot region images are later used to enhance reconstructed sparse 3-D models. Third, the landmark regions are obtained through mapping the enhanced 3-D models to landmark images. A k-dimensional tree (kd-tree) is then constructed for each landmark based on scale invariant feature transform (SIFT) features extracted from the landmark area to classify unlabeled images into pre-defined landmark categories. The proposed method is evaluated using 291 661 images of 51 landmarks. Experiments of comparison indicate that our method outperforms bag-of-words (BoW) based approach 18.5% and method of spatial-pyramid-matching using sparse-coding (ScSPM) 8.4%.
Xian Xiao, Changsheng Xu, Jinqiao Wang, Min Xu 0001
IEEE Trans. Multim.4
2011 Cascade-Based License Plate Localization with Line Segment Features and Haar-Like Features
abstract
AdaBoost classifiers with Haar-like features are widely used for license plate (LP) localization. However, it normally requires high-dimensional Haar-like features which cause extremely high computational cost. In this paper, a rejection cascade was built for LP localization with reduced Haar-like features. We first introduced line segment features as pre-input of Haar-like features for AdaBoost to eliminate more than 70% of the background in an image. Line segment features, including density, directionality and regularity, were extracted from line segments, which were detected by applying Hough Transform on an edge image. Later, AdaBoost classifiers with Haar-like features were further applied to identify the exact location of license plates. Our method dramatically reduced the demanded dimensions of Haar-like features, therefore saved much time in AdaBoost training stage. By comparing our method with methods of only using Haar-like features and only using line segment features, experimental results demonstrated that our proposed method achieved the best detection rate with significantly reduced dimensions of Haar-like features.
Min Xu 0001, Jesse S. Jin, Suhuai Luo
ICIG2
2011 Using context saliency for movie shot classification
abstract
Movie shot classification is vital but challenging task due to various movie genres, different movie shooting techniques and much more shot types than other video domain. Variety of shot types are used in movies in order to attract audiences attention and enhance their watching experience. In this pa per, we introduce context saliency to measure visual attention distributed in keyframes for movie shot classification. Different from traditional saliency maps, context saliency map is generated by removing redundancy from contrast saliency and incorporating geometry constrains. Context saliency is later combined with color and texture features to generate feature vectors. Support Vector Machine (SVM) is used to classify keyframes into pre-defined shot classes. Different from the existing works of either performing in a certain movie genre or classifying movie shot into limited directing semantic classes, the proposed method has three unique features: 1) context saliency significantly improves movie shot classification; 2) our method works for all movie genres; 3) our method deals with the most common types of video shots in movies. The experimental results indicate that the proposed method is effective and efficient for movie shot classification.
Min Xu 0001, Jinqiao Wang, Muhammad Abul Hasan, Xiangjian He, Changsheng Xu, Hanqing Lu, Jesse S. Jin
ICIP1
2010 Visual attention based small object segmentation in natual images
abstract
Small object segmentation is a challenging task in image processing and computer vision. In this paper we propose a visual attention based segmentation approach to segment interesting objects with small size in natural images. Different from traditional methods which use the single feature vectors, visual attention analysis is used on local and global features to extract the region of interesting objects. Within the region selected by visual attention analysis, Gaussian Mixture Model (GMM) is applied to further locate the object region. By incorporation of visual attention analysis into object segmentation, the proposed approach is able to narrow the searching region for object segmentation so as to increase the segmentation accuracy and reduce the computational complex. Experimental results demonstrate that the proposed approach is efficient for object segmentation in natural images, especially for small objects. The proposed method outperforms traditional GMM based segmentation significantly.
Wen Guo 0003, Changsheng Xu, Songde Ma, Min Xu 0001
ICIP4
2010 A close-up detection method for movies
abstract
Close-up (CU) is a photographic technique which tightly frames a person or an object. In movies, it is applied to guide audience attention and to evoke audience emotion. In this paper, we detect face CU, object CU, and lean of movies, which are widely used to romance emotions. A lean consists of shots in a sequence, with a close-up shot as focus. A set of features are extracted by considering movie making techniques and human attention for CU detection. The features are average saliency, color entropy, color variance, face height, skin area, and texture scales. These features are tested through statistical hypothesis test to be significantly discriminating for CUs. Then, Support Vector Machine (SVM) is applied on these features to detect face CU and object CU. Based on the face CU and object CU detection result, lean is further detected by investigating the changing of the face/object size. Lean detection is of challenge due to the technique of montage. We solve this problem through color similarity estimation and SIFT point matching. Experimental results on four full length movies verify the effectiveness of the proposed method.
Min Xu 0001, Qingming Huang, Jesse S. Jin, Shuqiang Jiang, Changsheng Xu
ICIP2
2008 Hierarchical movie affective content analysis based on arousal and valence features
abstract
Emotional factors directly reflect audiences' attention, evaluation and memory. Affective contents analysis not only create an index for users to access their interested movie segments, but also provide feasible entry for video highlights. Most of the work focus on emotion type detection. Besides emotion type, emotion intensity is also a significant clue for users to find their interested content. For some film genres (Horror, Action, etc), the segments with high emotion intensity have the most possibilities to be video highlights. In this paper, we propose a hierarchical structure for emotion categories and analyze emotion intensity and emotion type by using arousal and valence related features hierarchically. Firstly, High, Medium and Low are detected as emotion intensity levels by using fuzzy c-mean clustering on arousal features. Fuzzy clustering provides a mathematical model to represent vagueness, which is close to human perception. After that, valence related features are used to detect emotion types (Anger, Sad, Fear, Happy and Neutral). Considering video is continuous time series data and the occurrence of a certain emotion is affected by recent emotional history, Hidden Markov Models (HMMs) are used to capture the context information. Experimental results shows the movie segments with high emotion intensity cover over 80% of the movie highlights in Horror and Action movies and the hierarchical method outperforms the one-step method on emotion type detection. Meanwhile, it is flexible for user to pick up their favorite affective content by choosing both emotion intensity levels and emotion types.
Min Xu 0001, Jesse S. Jin, Suhuai Luo, Ling-Yu Duan
ACM Multimedia1
2008 Comparison analysis on supervised learning based solutions for sports video categorization
abstract
Due to the wide viewer-ship and high commercial potentials, recently, sports video analysis attracts extensive research efforts. One of the main tasks in sports video analysis is to identify sports genres i.e. sports video categorization. Most of the existing work focus on mapping content-based features to sports genres by using supervised learning methods. Moreover, video data sets seeks efficient data reduction methods due to the large size and noisy data. It lacks comparison analysis on the implementation and performance of these methods. In this paper, the research is carried out by using four dominant machine learning algorithms, namely Decision Tree, Support Vector Machine, K Nearest Neighbor and Naive Bayesian, and comparing their performance on a high dimensional feature set which selected by some feature selection tools such as Correlation-based Feature Selection (CFS), Principal Components Analysis (PCA) and Relief. Experimental results shows that Support Vector Machine (SVM) and k-NN are not sensitive to reduction of training sets. Moreover, three different feature reduction methods perform very differently with respect to four different tools.
Min Xu 0001, Mira Park 0001, Suhuai Luo, Jesse S. Jin
MMSP1
2008 Audio keywords generation for sports video analysis
abstract
Sports video has attracted a global viewership. Research effort in this area has been focused on semantic event detection in sports video to facilitate accessing and browsing. Most of the event detection methods in sports video are based on visual features. However, being a significant component of sports video, audio may also play an important role in semantic event detection. In this paper, we have borrowed the concept of the “keyword” from the text mining domain to define a set of specific audio sounds. These specific audio sounds refer to a set of game-specific sounds with strong relationships to the actions of players, referees, commentators, and audience, which are the reference points for interesting sports events. Unlike low-level features, audio keywords can be considered as a mid-level representation, able to facilitate high-level analysis from the semantic concept point of view. Audio keywords are created from low-level audio features with learning by support vector machines. With the help of video shots, the created audio keywords can be used to detect semantic events in sports video by Hidden Markov Model (HMM) learning. Experiments on creating audio keywords and, subsequently, event detection based on audio keywords have been very encouraging. Based on the experimental results, we believe that the audio keyword is an effective representation that is able to achieve satisfying results for event detection in sports video. Application in three sports types demonstrates the practicality of the proposed method.
Min Xu 0001, Changsheng Xu, Ling-Yu Duan, Jesse S. Jin, Suhuai Luo
ACM Trans. Multim. Comput. Commun. Appl.1
2007 Efficient sampling of training set in large and noisy multimedia data
abstract
As the amount of multimedia data is increasing day-by-day thanks to less expensive storage devices and increasing numbers of information sources, machine learning algorithms are faced with large-sized and noisy datasets. Fortunately, the use of a good sampling set for training influences the final results significantly. But using a simple random sample (SRS) may not obtain satisfactory results because such a sample may not adequately represent the large and noisy dataset due to its blind approach in selecting samples. The difficulty is particularly apparent for huge datasets where, due to memory constraints, only very small sample sizes are used. This is typically the case for multimedia applications, where data size is usually very large. In this article we propose a new and efficient method to sample of large and noisy multimedia data. The proposed method is based on a simple distance measure that compares the histograms of the sample set and the whole set in order to estimate the representativeness of the sample. The proposed method deals with noise in an elegant manner which SRS and other methods are not able to deal with. We experiment on image and audio datasets. Comparison with SRS and other methods shows that the proposed method is vastly superior in terms of sample representativeness, particularly for small sample sizes although time-wise it is comparable to SRS, the least expensive method in terms of time.
Surong Wang, Manoranjan Dash, Liang-Tien Chia, Min Xu 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2006 An Event-Driven Sports Video Adaptation for the MPEG-21 DIA Framework
abstract
We present an event-driven video adaptation system in this paper. Events are detected by audio/video analysis and annotated by the description schemes (DSs) provided by MPEG-7 multimedia description schemes (MDSs). And then, adaptation take account of users' preference of events and network characteristic to adapt video by event selection and frame dropping as following three steps: 1) the event information is parsed from MPEG-7 annotation XML file together with bitstream to generate generic bitstream syntax description (gBSD), 2) users' preference, network characteristic and adaptation QoS (AQoS) are considered for making adaptation decision, 3) adaptation engine automatically parses adaptation decisions and gBSD to achieve adaptation. Different from most existing adaptation work, the system adapts video by interesting events according to users' preference. To achieve a generic adaptation solution, the system is developed following MPEG-7 and MPEG-21 standards. gBSD based adaptation avoids complex video computation. 30 students from various departments test the system with satisfaction. Although, the system is tested on basketball video adaptation so far, it is easy to extend to other video domains
Min Xu 0001, Jiaming Li 0003, Yiqun Hu, Liang-Tien Chia, Bu-Sung Lee, Deepu Rajan, Jianfei Cai 0001
ICME1
2006 Event on demand with MPEG-21 video adaptation system
abstract
In this paper, we present an event-on-demand (EoD)video adaptation system. The proposed system supports users in deciding their events of interest and considers network conditions to adapt video source by event selection and frame dropping.Firstly, events are detected by audio/video analysis and annotated by the description schemes (DSs)provided by MPEG-7 Multimedia Description Schemes (MDSs). And then, to achieve a generic adaptation solution, the adaptation is developed following MPEG-21 Digital Item Adaptation (DIA)framework. We look at early release of the MPEG-21 Reference Software on XML generation and develop our own system for EoD video adaptation in three steps:1) the event information is parsed from MPEG-7 annotation XML file together with bitstream to generate generic Bitstream Syntax Description (gBSD). 2) Users' preference, Network Characteristic and Adaptation QoS (AQoS) are considered for making adaptation decision. 3) adaptation engine automatically parses adaptation decisions and gBSD to achieve adaptation.Unlike most existing adaptation work, the system adapts video of events with interest according to users' preference. Implementation following MPEG-7 and MPEG-21 standards provides a generic video adaptation solution. gBSD based adaptation avoids complex video computation. 30 students from various departments were invited to test the system and their responses has been positive.
Min Xu 0001, Jiaming Li 0003, Liang-Tien Chia, Yiqun Hu, Bu-Sung Lee, Deepu Rajan, Jesse S. Jin
ACM Multimedia1
2006 Affective content detection in sitcom using subtitle and audio
abstract
From a personalized media point of view, many users favor a flexible tool to quickly browse the affective content in a video. Such affective content may cause audiences' strong reactions or special emotional experiences, such as anger, sadness, fear, joy and love. This paper attempts to extract affective content for digital videos by analyzing the subtitle files of DVD/DivX videos and utilize audio event to assist affective content detection. Firstly, videos are segmented by dialogue script partition. Compared to traditional video shot, video segmented by scripts is not affected by camera changes and shooting angles and easy to include video segments with compact content. Secondly, emotion-related vocabularies in video script are detected to locate affective video content. Using script to directly access video content avoids complex video analysis. Thirdly, audio event detection is utilized to assist affective content detection. Compared with traditional video semantic analysis, affective content analysis puts much more emphasis on the audience's reactions and emotions. Initial experiments are carried on sitcom videos because its simple video structure provides useful domain knowledge. The experimental results demonstrate that subtitle file analysis and audio event detection provides effective and efficient clues to determine the emotional content of the videos.
Min Xu 0001, Liang-Tien Chia, Haoran Yi, Deepu Rajan
MMM1
2006 Efficient data reduction in multimedia data
Surong Wang, Manoranjan Dash, Liang-Tien Chia, Min Xu 0001
Appl. Intell.4
2005 EASIER Sampling for Audio Event Identification
abstract
An audio event refers to some specific audio sound which plays important role for video content analysis. In our previous work [3], we have established audio event identification as an audio classification task. Due to the large size of audio database, representative samples are necessary for training the classifier. However, the commonly used random selection of training samples is often not adequate in selecting representative samples. In this paper we present EASIER sampling algorithm to select those data which more efficiently represent audio data characters for audio event identifier training. EASIER deterministically produces a subsample whose “distance” from the complete database is minimal. Experiments in the context of audio event identification show that EASIER outperforms simple random sampling significantly.
Surong Wang, Min Xu 0001, Liang-Tien Chia, Manoranjan Dash
ICME2
2005 Affective content analysis in comedy and horror videos by audio emotional event detection
abstract
We study the problem of affective content analysis. In this paper, we think of affective contents as those video/audio segments, which may cause an audience's strong reactions or special emotional experiences, such as laughing or fear. Those emotional factors are related to the users' attention, evaluation, and memories of the content. The modeling of affective effects depends on the video genres. In this work, we focus on comedy and horror films to extract the affective content by detecting a set of so-called audio emotional events (AEE) such as laughing, horror sounds, etc. Those AEE can be modeled by various audio processing techniques, and they can directly reflect an audience's emotion. We use the AEE as a clue to locate corresponding video segments. Domain knowledge is more or less employed at this stage. Our experimental dataset consists of 40-minutes comedy video and 40-minutes horror film. An average recall and precision of above 90% is achieved. It is shown that, in addition to rich visual information, an appropriate usage of special audios is an effective way to assist affective content analysis.
Min Xu 0001, Liang-Tien Chia, Jesse S. Jin
ICME1
2005 A unified framework for semantic shot classification in sports video
abstract
The extensive amount of multimedia information available necessitates content-based video indexing and retrieval methods. Since humans tend to use high-level semantic concepts when querying and browsing multimedia databases, there is an increasing need for semantic video indexing and analysis. For this purpose, we present a unified framework for semantic shot classification in sports video, which has been widely studied due to tremendous commercial potentials. Unlike most existing approaches, which focus on clustering by aggregating shots or key-frames with similar low-level features, the proposed scheme employs supervised learning to perform a top-down video shot classification. Moreover, the supervised learning procedure is constructed on the basis of effective mid-level representations instead of exhaustive low-level features. This framework consists of three main steps: 1) identify video shot classes for each sport; 2) develop a common set of motion, color, shot length-related mid-level representations; and 3) supervised learning of the given sports video shots. It is observed that for each sport we can predefine a small number of semantic shot classes, about 5-10, which covers 90%-95% of broadcast sports video. We employ nonparametric feature space analysis to map low-level features to mid-level semantic video shot attributes such as dominant object (a player) motion, camera motion patterns, and court shape, etc. Based on the fusion of those mid-level shot attributes, we classify video shots into the predefined shot classes, each of which has clear semantic meanings. With this framework we have achieved good classification accuracy of 85%-95% on the game videos of five typical ball type sports (i.e., tennis, basketball, volleyball, soccer, and table tennis) with over 5500 shots of about 8 h. With correctly classified sports video shots, further structural and temporal analysis, such as event detection, highlight extraction, video skimming, and table of content, will be greatly facilitated.
Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu, Jesse S. Jin
IEEE Trans. Multim.2
2004 Mean shift based video segment representation and applications to replay detection
abstract
Effective and efficient representation of the low-level features of groups of frames or shots is an important yet challenging task for video analysis and retrieval. Key frame-based representation is limited by the difficulties in shot boundary detection of gradual transitions, and the variety of methods of key frame extraction. In this paper, we employ the mean shift-based mode seeking function to develop a new approach for compact representation of the video segment. The proposed video representation is motivated by recognizing that, on the global level, humans perceive images only as a combination of the few most prominent colors. We exploit the spatiotemporal mode seeking in feature space to simulate the "subjectivity" of human decisions in video segment retrieval and identification. The effectiveness of the video representation and matching scheme is shown by initial experiments on replay detection in broadcast sports videos.
Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu
ICASSP (5)2
2004 Mean shift based nonparametric motion characterization
Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu
ICIP2
2004 Nonparametric motion model with applications to camera motion pattern classification
abstract
Motion information is a powerful cue for visual perception. In the context of video indexing and retrieval, motion content serves as a useful source for compact video representation. There has been a lot of literature about parametric motion models. However, it is hard to secure a proper parametric assumption in a wide range of video scenarios. Diverse camera shots and frequent occurrences of bad optical flow estimation motivate us to develop nonparametric motion models. In this paper, we employ the mean shift procedure to propose a novel nonparametric motion representation. With this compact representation, various motion characterization tasks can be achieved by machine learning. Such a learning mechanism can not only capture the domain-independent parametric constraints, but also acquire the domain-dependent knowledge to tolerate the influence of bad dense optical flow vectors or block-based MPEG motion vector fields (MVF). The proposed nonparametric motion model has been applied to camera motion pattern classification on 23191 MVF extracted from MPEG-7 dataset.
Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu
ACM Multimedia2
2004 Nonparametric motion model
abstract
Motion information is a powerful cue for visual perception. In the context of video indexing and retrieval, motion content serves as a useful source for compact video representation. There has been a lot of literature about parametric motion models. However, it is hard to secure a proper parametric assumption in a wide range of video scenarios. Diverse camera shots and frequent occurrences of improper optical flow estimation or block matching motivate us to develop nonparametric motion models. In this demonstration, we present a novel nonparametric motion model. The unique features mainly include: 1) Instead of computationally expensive and vulnerable parametric regression our proposed model bases the motion characterization on the classification of motion patterns; 2) we employ machine learning to capture the knowledge of recognizing camera motion patterns from bad motion vector fields (MVF); and 3) with the mean shift filtering our proposed motion representation elegantly incorporates the spatial-range information for noise removal and discontinuity preserving smoothing of MVF. Promising results have been achieved on two tasks: 1) camera motion pattern recognition on 23191 MVFs and 2) recognition of the intensity of motion activity on 622 video segments culled from the MPEG-7 dataset.
Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu
ACM Multimedia2
2004 Audio keyword generation for sports video analysis
abstract
Semantic sports video analysis has attracted many research interests and audio cues have been shown to play an important role in semantics inference. To facilitate event detection using audio information, we have introduced the concept of audio keyword (e.g. excited/plain commentator speech, excited/plain audience sound, etc.) to describe the game-specific sound associated with an event. In our previous work, we have designed a hierarchical Support Vector Machine (SVM) classifier for audio keyword identification. However, there are two inherent weaknesses: 1) a frame-based SVM classifier does not incorporate any contextual information; 2) a robust recognizer relies on large amounts of training data in the case of different sports games videos. In this demo, we present a flexible Hidden Markov Model (HMM)-based audio keyword generation system. This is motivated by the successful story of applying HMM in speech recognition. Unlike the frame-based SVM classification followed by a major voting, our HMM-based system treats an audio keyword as a continuous time series data and employs hidden states transition to capture contexts. Moreover, our system introduces an adaptation mechanism to tune the initial HMM models (obtained from available training data) to improve performance by a small number of data from a new sports game video. Promising results has been demonstrated on the tennis, soccer and basketball videos with the total length of 2 hours.
Min Xu 0001, Ling-Yu Duan, Liang-Tien Chia, Changsheng Xu
ACM Multimedia1
2003 A fusion scheme of visual and auditory modalities for event detection in sports video
abstract
We propose an effective fusion scheme of visual and auditory modalities to detect events in sports video. The proposed scheme is built upon semantic shot classification, where we classify video shots into several major or interesting classes, each of which has clear semantic meanings. Among major shot classes we perform classification of the different auditory signal segments (i.e. silence, hitting ball, applause, commentator speech) with the goal of detecting events with strong semantic meaning. For instance, for tennis video, we have identified five interesting events: serve, reserve, ace, return, and score. Since we have developed a unified framework for semantic shot classification in sports videos and a set of audio mid-level representation with supervised learning methods, the proposed fusion scheme can be easily adapted to a new sports game. We are extending this fusion scheme to three additional typical sports videos: basketball, volleyball and soccer. Correctly detected sports video events will greatly facilitate further structural and temporal analysis, such as sports video skimming, table of content, etc.
Min Xu 0001, Ling-Yu Duan, Changsheng Xu, Qi Tian 0002
ICASSP (3)1
2003 A fusion scheme of visual and auditory modalities for event detection in sports video
abstract
In this paper, we propose an effective fusion scheme of visual and auditory modalities to detect events in sports video. The proposed scheme is built upon semantic shot classification, where we classify video shots into several major or interesting classes, each of which has clear semantic meanings. Among major shot classes we perform classification of the different auditory signal segments (i.e. silence, hitting ball, applause, commentator speech) with the goal of detecting events with strong semantic meaning. For instance, for tennis video, we have identified five interesting events: serve, reserve, ace, return, and score. Since we have developed a unified framework for semantic shot classification in sports videos and a set of audio mid-level representation with supervised learning methods, the proposed fusion scheme can be easily adapted to a new sports game. We are extending this fusion scheme to three additional typical sports videos: basketball, volleyball and soccer. Correctly detected sports video events will greatly facilitate further structural and temporal analysis, such as sports video skimming, table of content, etc.
Min Xu 0001, Ling-Yu Duan, Changsheng Xu, Qi Tian 0002
ICME1
2003 Creating audio keywords for event detection in soccer video
abstract
This paper presents a novel framework called audio keywords to assist event detection in soccer video. Audio keyword is a middle-level representation that can bridge the gap between low-level features and high-level semantics. Audio keywords are created from low-level audio features by using support vector machine learning. The created audio keywords can be used to detect semantic events in soccer video by applying a heuristic mapping. Experiments of audio keywords creation and event detection based on audio keywords have illustrated promising results. According to the experimental results, we believe that audio keyword is an effective representation that is able to achieve more intuitionistic result for event detection in sports video compared with the method of event detection directly based on low-level features.
Min Xu 0001, Namunu Chinthaka Maddage, Changsheng Xu, Mohan Kankanhalli, Qi Tian 0002
ICME1
2003 A mid-level representation framework for semantic sports video analysis
abstract
Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-level representation framework for semantic sports video analysis. The mid-level representation layer is introduced between the low-level audiovisual processing and high-level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video. Categories and Subject Descriptors H.3.1 [Content Analysis and Indexing]: abstracting methods, indexing methods.
Ling-Yu Duan, Min Xu 0001, Tat-Seng Chua, Qi Tian 0002, Changsheng Xu
ACM Multimedia2
2003 Nonparametric color characterization using mean shift
abstract
Color is very useful in locating and recognizing objects that occur in artificial environments. The color histogram has shown its efficiency and advantages as a general tool for various applications, such as content-based image retrieval and video browsing, object indexing and location, and video segmentation. However, due to the lack of any spatial and context information, the histogram is not robust and effective for color characterization (e.g. dominant color) in large video databases. In this paper, we propose a nonparametric color characterization model using mean shift procedure, with an emphasis on spatio-temporal consistency. Experimental results suggest that the color characterization model is much more effective for video indexing and browsing, particularly in the domain of structured video (e.g. sports video).
Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu
ACM Multimedia2
2002 A unified framework for semantic shot classification in sports videos
abstract
In this demonstration, we present a unified framework for semantic shot classification in sports videos. Unlike previous approaches, which focus on clustering by aggregating shots with similar low-level features, the proposed scheme makes use of domain knowledge of specific sport to perform a top-down video shot classification, including identification of video shots classes for each sport, and supervised learning and classification of given sports video with low-level and middle-level features extracted from the sports video. It's observed that for each sport we can predefine a small number of semantic shot classes, 5--10, which cover 90 to 95 % of sports broadcasting video. With supervised learning method, we can map the low-level features to middle-level semantic video shot attributes such as dominant object motion (a player), camera motion patterns, and court shape, etc. On the basis of the appropriate fusion of those middle-level shot attributes, we classify video shots into the predefined video shot classes, each of which has a clear semantic meaning. The proposed method has been tested over 3 types of sports videos: tennis, basketball, and soccer. Good classification results ranging from 80~95% have been achieved. The proposed framework provides a generic solution for sports video semantic shot classification, which can be adapted to a new sport type easily. With correctly classified sports video shots further structural and temporal analysis will be greatly facilitated.
Ling-Yu Duan, Min Xu 0001, Xiao-Dong Yu, Qi Tian 0002
ACM Multimedia2