EDBT 2026 Demo / reviewers in the wild / expert
Lin Qi 0001
dblp:82/3116-1
· DBLP profile ↗
75ranked-venue papers
0as first author
36since 2021 · last 2025
0000-0002-4558-6741ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 15 since 2021Artificial intelligence and machine learning · 19 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Systems, architecture and hardware · 6 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EntroFormer: An entropy-based sparse vision transformer for real-time semantic segmentation
Song Wang 0008, Lin Wu 0001, Deyin Liu, Lei Gao 0001, Lin Qi 0001, Guanghui Wang 0001 |
Comput. Vis. Image Underst. | 6 |
| 2025 | MrgaNet: Multi-scale recursive gated aggregation network for tracheoscopy images
Tie Yun, Dalong Zhang, Fenghui Liu, Lin Qi 0001 |
Image Vis. Comput. | 5 |
| 2025 | Dynamic gesture recognition using 3D central difference separable residual LSTM coordinate attention networks
Lichuan Geng, Tie Yun, Lin Qi 0001, Chengwu Liang |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Hybrid Learning Module-Based Transformer for Multitrack Music Generation With Music TheoryabstractIn recent years, multitrack music generation has garnered significant attention in both academic and industrial spheres for its versatile utilization of various instruments in collaborative settings. The primary challenge lies in achieving a harmonious balance within individual tracks and fostering effective collaboration across multiple tracks. To address this issue, this article introduces a pioneering hybrid learning encoder architecture. Each music track's encoder is implemented as an independent transformer architecture, preserving self-attention mechanisms within a single track and interattention mechanisms between different tracks. The resulting features are then seamlessly integrated into the decoder through concatenation. Of particular significance, previous multitrack music generation efforts have predominantly operated under unconditional settings, yielding music that lacks practical value due to noncompliance with established music theory principles. Recognizing this limitation, the article proposes a novel approach to multitrack music generation guided by music theory rules. Employing reinforcement learning techniques, the decoder-generated music serves as the initial state. Positive feedback is provided when the generated music adheres to music theory rules; conversely, negative feedback is applied to compel the multitrack music to align with widely accepted music theory principles. Finally, comprehensive simulation validation is conducted on both the publicly available LMD dataset and the self-constructed MUT dataset. The plethora of experimental results overwhelmingly corroborates the efficacy of the proposed methodology. Tie Yun, Xin Guo 0005, Jiessie Tie, Lin Qi 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | A Deep Semantic Segmentation Network With Semantic and Contextual RefinementsabstractSemantic segmentation is a fundamental task in multimedia processing, which can be used for analyzing, understanding, editing contents of images and videos, among others. To accelerate the analysis of multimedia data, existing segmentation researches tend to extract semantic information by progressively reducing the spatial resolutions of feature maps. However, this approach introduces a misalignment problem when restoring the resolution of high-level feature maps. In this paper, we design a Semantic Refinement Module (SRM) to address this issue within the segmentation network. Specifically, SRM is designed to learn a transformation offset for each pixel in the upsampled feature maps, guided by high-resolution feature maps and neighboring offsets. By applying these offsets to the upsampled feature maps, SRM enhances the semantic representation of the segmentation network, particularly for pixels around object boundaries. Furthermore, a Contextual Refinement Module (CRM) is presented to capture global context information across both spatial and channel dimensions. To balance dimensions between channel and space, we aggregate the semantic maps from all four stages of the backbone to enrich channel context information. The efficacy of these proposed modules is validated on three widely used datasets—Cityscapes, Bdd100 K, and ADE20K—demonstrating superior performance compared to state-of-the-art methods. Additionally, this paper extends these modules to a lightweight segmentation network, achieving an mIoU of 82.5% on the Cityscapes validation set with only 137.9 GFLOPs. Deyin Liu, Lin Wu 0001, Song Wang 0008, Xin Guo 0005, Lin Qi 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | One-Stage Lightweight Network of Object Detection for Rectangular Panoramic Images
Yingying Lu, Tie Yun, Lin Qi 0001 |
ICIC (7) | 3 |
| 2024 | Multitrack Emotion-Based Music Generation Network Using Continuous Symbolic FeaturesabstractMusicians need an efficient composition process to yield a multitude of musical pieces. To enhance the artistry and emotional relevance of AI-generated music, a novel Multi-Track Emotion Music Generation model (MEMG) is proposed. MEMG minimizes reliance on emotion-labeled datasets, instead employing musical features such as conditional pitch histograms to signify music emotion attributes. By analyzing music features and emotional responses, MEMG facilitates controlled music generation with diverse emotions, where specific emotions are mapped based on Russell’s psychological model. A novel hybrid feedback reward network is constructed, integrating music theory rules, discriminator, and musical emotion interpretation. Incorporating music theory and embedding music emotion as input parameters during generation. Objective music theory evaluation and subjective auditory testing validate MEMG’s competitiveness in composing varied emotional music, coupled with exemplary adherence to musical rules. Tie Yun, Lin Qi 0001 |
ICME | 6 |
| 2024 | Multiple Query-based Multi-omics Fusion Algorithm for Cancer Classification and Metastasis PredictionabstractThe research on cancer classification and metastasis prediction based on multi-omics data has emerged with the development of high-throughput sequencing techniques and artificial intelligence. However, effective fusion among omics data remains a challenging problem, besides the ’curse of dimensionality’ resulting from high-dimensional features and small sample sizes. The transform-based fusion methods commonly used in computer vision and natural language processing require abundant samples for learning, which can lead to overfitting problems when applied in the biomedical field. Thus, this paper proposes a multiple query-based fusion algorithm with fewer parameters to learn, while retaining effective integration abilities for multiple omics data. It consists of interaction and fusion modules, which can prevent the model from training too aggressively and effectively mitigate overfitting, thus enabling the fusion of small datasets. Furthermore, this paper introduces a multi-task learning model based on it for cancer classification and metastasis prediction tasks. The proposed fusion algorithm can generalize to multiple omics data and maintain stable parameter counts. Experimental results demonstrate that the proposed method outperforms other state-of-the-art supervised multi-omics integrative methods in various cancer classification applications using DNA methylation, RNA expression, and miRNA expression data. Tie Yun, Fenghui Liu, Lin Qi 0001 |
IJCNN | 4 |
| 2024 | Dynamic Gesture Recognition Using a Spatio-Temporal Adaptive Multiscale Convolutional Transformer NetworkabstractThe Convolutional Neural Network has established itself as a cornerstone in gesture recognition systems. Nonetheless, its capacity to capture global contextual features is limited, presenting hurdles for dynamic gesture recognition. This paper introduces the Spatio-temporal Adaptive Multiscale Convolutional Transformer Network (SAMCT), a novel solution poised to propel the domain of dynamic gesture recognition forward. By integrating adaptive multiscale analysis with the transformer’s adeptness in processing sequential data, SAMCT refines the feature extraction process, providing a nuanced comprehension of gesture dynamics. The network incorporates a Spatiotemporal Adaptive Module (SAM), leveraging it to alleviate temporal redundancy and enabling gesture region selection for effective extraction of essential channels and spatial information from local features. To address CNNs’ constraints in extracting global contextual features, we propose a Multiscale Convolutional Transformer Network (MCT) that synergistically merges convolution and transformer networks to extract multiscale features of dynamic gestures. Our experimental evaluation using the Isogd dataset demonstrates that SAMCT surpasses other state-of-the-art recognition networks, showcasing superior recognition performance. Tie Yun, Lin Qi 0001, Chengwu Liang |
IJCNN | 3 |
| 2024 | ECA-SLAM: A Visual SLAM Utilizing Depth Features and Attention MechanismabstractSimultaneous Localization and Mapping (SLAM) has a broad application prospect in robot navigation, autonomous driving, and augmented reality. While visual SLAM systems have made significant progress, they still face challenges in complex environments due to the complexity and time variability of the real environment. Traditional methods have limitations in feature extraction, matching, and mapping, and the effect of localization and mapping can be greatly reduced due to scene changes. In response to these challenges, this paper proposes an improved visual SLAM method based on deep learning. We use a lightweight Mobilenetv2 network for feature extraction and introduce the Hardswish activation function to improve the performance further. Meanwhile, the ECA-Net attention mechanism is introduced into the feature extraction network to enhance its ability to capture key image features accurately. In addition, we use multi-task distillation to train our network to reduce the complexity of the model. Extensive experimental evaluation on multiple datasets reveals that the proposed method has significantly improved accuracy and robustness, particularly in challenging environments, demonstrating superior performance. Index Terms—SLAM, ECA-Net, deep learning, feature extraction Wenjie Deng, Tie Yun, Lin Qi 0001 |
IJCNN | 3 |
| 2024 | A Complementary Action Recognition Network based on Conv-TransformerabstractAs one of the important research directions in the field of computer vision, action recognition has extensive application value in today's internet. Since traditional convolutional neural networks performed well in processing local features, but the ability to process global information and performance on large-scale datasets is weaker. The transformer-based model can efficiently model global features through the attention mechanism, but it cannot capture the information of dynamic changes in spatial and temporal, and its ability to process local information is relatively weak. To do this, this paper proposed a Conv-Transformer Residual Network(CTRN) for action recognition, used a unique Conv-Transformer Residual Module(CTRM) to extract important features that combine global contextual and local information efficiently. Furthermore this paper designed a loss function consisting of pixel loss, cross-entropy loss and CIOU_Loss, allows action recognition networks to simultaneously optimise performance on pixel-level prediction, action classification and action localisation for more accurate action recognition. This network makes full use of the local modeling capabilities of the cnn and the advantages of the global modeling of the Transformer, which not only effectively reduces the complexity of the model, but also breaks through the limitations of the Transformer's lack of spatial sensing bias. Training the proposed model in an end-to-end manner and conducted extensive experiments on two typical datasets, Kinetics-400 and Something and Something V2, and outperformed some representative state-of-the-art models in both qualitative and quantitative evaluations. Linfu Liu, Lin Qi 0001, Tie Yun, Chengwu Liang |
IJCNN | 2 |
| 2024 | Edge-labeling based modified gated graph network for few-shot learning
Peixiao Zheng, Xin Guo 0005, Enqing Chen, Lin Qi 0001, Ling Guan |
Pattern Recognit. | 4 |
| 2023 | DCNet: Densely Connected Neural Network with Global Context for Abnormalities Prediction in Bronchoscopy ImagesabstractBronchoscopy plays an important role in the examination of lower respiratory tract diseases. However, the large number of medical images produced by bronchoscopy makes it a time-consuming and labor-intensive task for doctors to analyze. Therefore, developing computer-aided diagnostic algorithms to assist doctors in lung image analysis is of great significance.In this paper, we propose a novel densely connected convolutional network called DCNet with local and global receptive fields for the automatic prediction of bronchoscopy images. Our experimental results on the clinical bronchoscopy image analysis dataset demonstrate that our proposed method outperforms state-of-the-art algorithms in distinguishing malignant bronchoscopy images, with an accuracy rate of 92.81%. And DCNet achieved 99.31% accuracy on the public data set called Kvasir-Capsule. This model may be a valuable tool to assist medical staff in the analysis of bronchoscopy images. Tie Yun, Fenghui Liu, Lin Qi 0001, Ziqing Hu |
BIBM | 4 |
| 2023 | CTCM: Clustering based on three correlation matrices for multi-omics data integration and cancer subtype identificationabstractLarge-scale sequencing data is used by biologists to understand the biological systems and molecular mechanisms of disease. One challenge is to effectively use valuable information from different cancer omics to produce more accurate and reliable subtypes. We propose a clustering strategy based on three correlation matrices (CTCM) to identify cancer subtypes in multi-omics data. Connectivity matrix, similarity matrix and resampling matrix collect useful data information from different perspectives. The connection matrix divides the raw data into stable subtypes by adding noise to simulate the systematic error of the sequencing platform. The similar matrix use Gaussian kernel functions to construct connections between samples as the "skeleton" of the whole. The resampling matrix adapts to the explosive growth of data by sampling subsets. For each omics data, we combine the connectivity matrix and resampling matrix with the similarity matrix to generate an iterative version. Iterating the three relationship matrices for each omics produces a fusion matrix that is used for spectral clustering to identify cancer subtypes. Compared with six other state-of-the-art multi-omics clustering methods, CTCM achieves excellent performance on the benchmark data sets of simulation and TCGA databases. The method is general enough to replace existing unsupervised clustering techniques outside the scope of biomedical research to integrate multiple types of data. Tie Yun, Dalong Zhang, Fenghui Liu, Lin Qi 0001 |
BIBM | 5 |
| 2023 | MOVNG: Applied a Novel Sparse Fusion Representation into GTCN for Pan-Cancer Classification and Biomarker Identification
Tie Yun, Fenghui Liu, Dalong Zhang, Lin Qi 0001 |
ICIC (1) | 5 |
| 2023 | Intelligence Evaluation of Music Composition Based on Music Knowledge
Tie Yun, Lin Qi 0001 |
ICIC (5) | 5 |
| 2023 | A Feature Refinement Module for Light-Weight Semantic Segmentation NetworkabstractLow computational complexity and high segmentation accuracy are both essential to the real-world semantic segmentation tasks. However, to speed up the model inference, most existing approaches tend to design light-weight networks with a very limited number of parameters, leading to a considerable degradation in accuracy due to the decrease of the representation ability of the networks. To solve the problem, this paper proposes a novel semantic segmentation method to improve the capacity of obtaining semantic information for the light-weight network. Specifically, a feature refinement module (FRM) is proposed to extract semantics from multi-stage feature maps generated by the backbone and capture non-local contextual information by utilizing a transformer block. On Cityscapes and Bdd100K datasets, the experimental results demonstrate that the proposed method achieves a promising trade-off between accuracy and computational cost, especially for Cityscapes test set where 80.4% mIoU is achieved and only 214.82 GFLOPs are required. Xin Guo 0005, Song Wang 0008, Peixiao Zheng, Lin Qi 0001 |
ICIP | 5 |
| 2023 | Music Question Answering Based on Aesthetic ExperienceabstractMusic understanding has always been considered to be the work of experts. Ordinary people have insufficient aesthetic experience when facing music. We put forward the topic of music question and answer based on aesthetic experience to help people understand music more comprehensively. We summarized the relevant external characteristics of music elements that people are most concerned about in the field of music analysis, and generated a structured music aesthetic experience database. Based on this, we constructed the AMQA-dataset. We propose a new method of music graph representation, which can improve the reasonability and interpretability of music question answering system. We have verified our idea on the AMQA dataset, and achieved experimental results comparable to those of humans. Wenhao Gao 0004, Tie Yun, Lin Qi 0001 |
IJCNN | 4 |
| 2023 | Dynamic Gesture Recognition Based on 3D Central Difference Separable Residual LSTM Coordinate Attention Networks
Tie Yun, Lin Qi 0001, Chengwu Liang |
PRCV (8) | 3 |
| 2022 | PNF: a novel method based on connectivity and similarity for data integration and cancer subtypingabstractThe development of high-throughput technology enables measurements of many types of omics data, but the meaningful integration of different data types still is a significant challenge. Another difficult and vital challenge is the discovery of cancer molecular subtypes with relevant clinical differences. Here we propose a novel method, called perturbation network fusion clustering (PNF) for multi-omics data integration and cancer subtyping, which can address these two challenges. We creatively combine the connectivity and similarity of patient pairs, first adding statistical knowledge to the similarity network. Adopting perturbation clustering to get the probability (i.e., connectivity) that any two patients are grouped into a cluster on each type of data, and then calculating the similarity between any two patients using the Gaussian kernel function for each data type. Next using connectivity and similarity matrices we generate multiple stable and strong similarity kernels. Finally, we use the similarity network fusion strategy to fuse similarity kernel from each omics data and spectral clustering to discover cancer subtypes with survival differences. The method is validated on simulated data and six cancer datasets from The Cancer Genome Atlas (TCGA) nearly a thousand patient samples, including gene expression, microRNA, and DNA methylation data. PNF accurately identifies known cancer subtypes and novel subgroups of patients with significantly different survival profiles. The method without any prior knowledge is general enough to replace existing multi-omics data integration and unsupervised clustering methods outside the scope of biomedicine research. Fenghui Liu, Lin Qi 0001, Tie Yun |
BIBM | 4 |
| 2022 | WINMLP: Quantum & Involution Inspire False Positive Reduction in Lung Nodule Detection
Fenghui Liu, Lin Qi 0001, Tie Yun |
ICONIP (3) | 3 |
| 2022 | Fused Multiscale Cost Volume for Efficient Stereo MatchingabstractStereo matching is an important part of reconstructing depth information and 3D scenes. However, a important problem is improving accuracy and efficiency at the same time. We propose an end-to-end convolutional neural network model to estimate disparity from stereo images. To improve the accuracy of feature extraction in complex regions, we use atrous spatial pyramid pooling module to extract and fuse features, and atrous convolutions of different rates explicitly adjust the receptive field of the filter. To balance the accuracy of the algorithm and GPU memory consumption, we design a fused multiscale cost volume and the fusion strategy to improve the efficiency in estimating the disparity, which reduces the amount of memory consumption while ensuring the accuracy of the algorithm. Our model is trained and evaluated on SceneFlow and KITTI datasets. Experimental results demonstrate that our method achieve a better accuracy even in weak textures areas, while GPU memory consumption and runtime are reduced. Aihao Cui, Lin Qi 0001, Tie Yun |
MMSP | 2 |
| 2022 | RDDP: Reliable Detection and Description of Interest PointsabstractRecent works have paid much attention on learning repeatable maps for detection and description of interest points. However, there are always some components which is repeatable but not discriminative in reality. Training on the whole images with these components will lead to poor matching performance. Thus, we propose a self-supervised network with a new branch which can eliminate the negative influence of unreliable image components. We also design a loss function including a part calculated with Average Precision (AP) to make full use of the reliability branch. Evaluations on HPatches dataset show that the proposed method achieves competitive results compared with state-of-the-art methods. Yuning Gao, Lin Qi 0001, Tie Yun |
SMC | 2 |
| 2021 | Face Recognition Based on Panoramic VideoabstractRecently, face recognition has been widely used in lives, but it can still only maintain good results in specific scenarios. In different tasks, changes in the external environment will cause diversified changes in face images, such as panoramic video. The barrel distortion of the panoramic images generated by the fisheye lens will cause the face images to be deformed to varying degrees around the fisheye image. A new model DCMNet is proposed in this paper and a panoramic face dataset PIFI is created based on the panoramic video system. After experiments are conducted on different datasets, the results show that our model has better results on popular datasets and panoramic dataset. Tie Yun, Lin Qi 0001, Rui Zhang 0010, Juanjuan Cai |
FG | 3 |
| 2021 | 2D-FRFT Based Frequency Shift-Invariant Digital Image EncryptionabstractIn this paper, we study the property of frequency shift in two-dimensional Fractional Fourier Transform (2D-FRFT) domain. Based on the mathematical verification and computer simulations, it is demonstrated that the magnitude of reconstruction from phase information satisfies frequency shift-invariant in 2D-FRFT domain, improving the robustness of image encryption. Experiments are implemented to verify the effectiveness of this property against the frequency shift attack in 2D-FRFT domain, improving the robustness of image encryption. Lei Gao 0001, Lin Qi 0001, Ling Guan |
ICASSP | 2 |
| 2021 | Local Feature Descriptors with Deep Hypersphere LearningabstractRecent works have demonstrated the power of L2normalization in local feature descriptor learning. While the descriptors are typically learned in the Euclidean space, the similarity between descriptors is often evaluated on a unit hypersphere due to the post-processing of L2normalization for descriptors, which creates a gap between the training stage and the usage stage of feature descriptors. To bridge the gap, we propose a hyperspherical descriptor learning model, where the whole network is projected onto the hyperspherical space. In addition, a squared angular triplet loss is designed to enable the proposed hyperspherical model to learn angularly discriminative descriptors. Experiments on UBC dataset show that the proposed hyperspherical descriptor outperforms its Euclidean counterparts and the state-of-the-art methods on the feature matching task. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 4 |
| 2021 | Discriminative Patch Descriptor Learning With Focal Triplet Loss FunctionabstractThis paper proposes a focal triplet loss function for discriminative patch descriptor learning. The standard triplet loss function usually restrains the distance difference between the matching samples and the non-matching ones. However, along with the training procedure, the majority of triplets in each batch tend to satisfy the constraint of the loss function and produce low loss values, leading to a masquerade that the model is well-trained. To address this problem, the focal triplet loss function is proposed in this paper to weaken the impact of the easy triplets and focus training on the hard ones. By emphasizing the importance of hard triplets on the model training, the proposed loss forces the descriptor vectors with fixed dimension to carry more discriminative information from the patches. With the benefits of the focal mechanism, the proposed method achieves better performance compared to the state-of-the-art on UBC dataset for image matching task. Furthermore, to demonstrate the effectiveness of the proposed method, we extend the focal triplet loss on the cross-model retrieval task. The experimental results indicate that the proposed method can also be used to improve visual-semantic embedding learning. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 4 |
| 2021 | Edge-Labeling Based Directed Gated Graph Network for Few-Shot LearningabstractExisting graph-network-based few-shot learning methods obtain similarity between nodes through a convolution neural network (CNN). However, the CNN is designed for image data with spatial information rather than vector form node feature. In this paper, we proposed an edge-labeling-based directed gated graph network (DGGN) for few-shot learning, which utilizes gated recurrent units to implicitly update the similarity between nodes. DGGN is composed of a gated node aggregation module and an improved gated recurrent unit (GRU) based edge update module. Specifically, the node update module adopts a gate mechanism using activation of edge feature, making a learnable node aggregation process. Besides, improved GRU cells are employed in the edge update procedure to compute the similarity between nodes. Further, this mechanism is beneficial to gradient backpropagation through the GRU sequence across layers. Experiment results conducted on two benchmark datasets show that our DGGN achieves a comparable performance to the-state-of-art methods. Peixiao Zheng, Xin Guo 0005, Lin Qi 0001 |
ICIP | 3 |
| 2021 | Context-Aware Based Visual-Audio Feature Fusion for Emotion RecognitionabstractVideo emotion recognition is a significant branch in the field of emotion computing. However, traditional recognition works mainly focus on human features, ignoring the contextual clues of video scenes and objects. In our work, we propose a context-aware framework for bi-modal video emotion recognition. Unlike existing methods that directly extract features of the entire video frame, we extract key frames and key regions of videos to obtain emotional cues contained in video scenes and objects. Specifically, for visual stream, the hierarchical Bidirectional Long-Short Term Memory (Bi-LSTM) is applied to summarize video scenes and find key frames that mostly contribute to video emotion; Meantime, we introduce the Region Proposal Network (RPN) to extract corresponding features of object regions in video frames and construct the emotional similarity graph. After using the Feedforward Neural Network (FNN) to assign different weight coefficients to different regions, the Graph Convolutional Network (GCN) is used to reason about the connections between key regions. Moreover, the context information of the frame-level Log-Mel spectrum fragments supplement the visual information. Finally, we fuse the visual and acoustics features by adaptive gated multimodal fusion module for video emotion classification. We conduct experiments on Video Emotion-8 and Ekman-6 datasets. The experimental results demonstrate that our model achieves better classification accuracy than several baseline models. Huijie Cheng, Tie Yun, Lin Qi 0001 |
IJCNN | 3 |
| 2021 | Research on Panoramic Stereo Live Streaming Based on the Virtual RealityabstractVirtual Reality (VR) technology has been successfully applied in many fields, such as education, tourism, games, etc. Some researchers hope to combine VR to watch panoramic stereo video and bring a brand new living experience in the field of live broadcasting. However, in practical applications, there are many problems such as panoramic video real-time stitching and high-resolution video compression coding, so it is still difficult to construct a panoramic stereo living video system. In this paper, we design and implement a panoramic stereo video living system. The system based on the eight-eye panoramic annular lens, the 360-Degree binocular video is stitched. Compared with the traditional coding algorithm, H.265 coding algorithm is used to compress the video, which improves the compression rate by more than fifty percent. Then the video stream is pushed to the cloud for forwarding. The receiving end combines VR helmet to watch panoramic stereo video live, bringing viewers a brand new viewing experience. Mingyao Zheng, Tie Yun, Lin Qi 0001, Yuning Gao |
ISCAS | 4 |
| 2021 | A two-stream heterogeneous network for action recognition based on skeleton and RGB modalitiesabstractRecent years, skeleton based action recognition with graph convolutional network (GCN) has achieved great success. However, since skeleton data only includes human body joints coordinates, other key information on actions is missing such as the subtle motion of hands, the objects the human is interacting, leading to an unsatisfactory performance. In this respect, the RGB data can offer help to recognize actions that skeleton-based methods have limitations on. In this work, we propose a novel two-stream heterogeneous network consisting of GCN and CNN networks for action recognition. Specifically, the GCN network takes the skeletal sequence as input to exploit skeleton information. For the RGB video, the CNN model, ResNet (2+1)D, is adapted to exploit RGB information. Afterwards, the discriminant canonical correlation analysis (DCCA) method is utilized to integrate the output feature maps from the skeleton and RGB streams, resulting in improved performance. Experimental results on the large-scale dataset NTU RGB+D show that the proposed model outperforms state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISM | 4 |
| 2021 | Integrating vertex and edge features with Graph Convolutional Networks for skeleton-based action recognition
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
Neurocomputing | 4 |
| 2021 | Auto-encoder based structured dictionary learning for visual classification
Deyin Liu, Chengwu Liang, Shaokang Chen, Tie Yun, Lin Qi 0001 |
Neurocomputing | 5 |
| 2021 | A discriminant kernel entropy-based framework for feature representation learning
Lei Gao 0001, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | The Property of Frequency Shift in 2D-FRFT Domain With Application to Image EncryptionabstractThe Fractional Fourier Transform (FRFT) has been playing a unique and increasingly important role in signal and image processing. In this letter, we investigate the property of frequency shift in two-dimensional FRFT (2D-FRFT) domain. It is shown that the magnitude of image reconstruction from phase information is frequency shift-invariant in 2D-FRFT domain, enhancing the robustness of image encryption, an important multimedia security task. Experiments are conducted to demonstrate the effectiveness of this property against the frequency shift attack, improving the robustness of image encryption. Lei Gao 0001, Lin Qi 0001, Ling Guan |
IEEE Signal Process. Lett. | 2 |
| 2021 | A Multi-Stream Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action RecognitionabstractRecently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN with HCRF to retain the human skeleton structure information even during the classification stage. Our model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework which takes the relative coordinate of the joints and bone direction as two static feature streams, and the temporal displacements between two consecutive frames as the dynamic feature stream. Experimental results on three challenging benchmarks (NTU RGB+D, N-UCLA, SYSU) show the superior performance of the proposed model over state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
IEEE Trans. Multim. | 4 |
| 2020 | Negative Label Guided Discriminative Canonical Correlation Analysis for Semi-Supervised and Semi-Paired LearningabstractSemi-supervised learning is a popular trend for learning based methods in recent years, as it fully exploits both the labeled and unlabeled samples in a dataset. This paper sets itself apart from most existing semi-supervised learning algorithms, which only use the exact labels of data already known. We take the negative label as side information to guide the process of semi-supervised learning. Two types of supervision information are regarded as negative label; the first type indicates that a sample definitely does not belong to a specific category, and the second indicates that two samples come from different views, and cannot have a one to one correspondence. By reasonably assuming that nearby points should have similar class indicators, the data labels are propagated under the negative label and the geometric structure revealed by both labeled and unlabeled points. Specifically, we predict one to one pair information by utilizing the neighbor information of samples, under the guidance of the negative pair label. Extensive experiments on several datasets demonstrate the effectiveness of our proposed method. Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 4 |
| 2020 | A Vertex-Edge Graph Convolutional Network for Skeleton-Based Action RecognitionabstractThe Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability of exploiting the joint information from the graph structure of the skeleton data. Recently, as a strong and complementary modality for action recognition, the bone information from skeleton data has attracted more attention. However, most existing GCN-based methods extract the bone and joint features with two separate GCN networks, ignoring the dependencies between joints and bones. In this work, a Vertex-Edge Graph Convolutional Network(VE-GCN) is proposed to reveal the information across joints, bones and their relationships simultaneously. In addition, we learn the additional connections among joints and bones for various action samples besides the natural connections of the skeleton. Then we conduct the convolution operation on joints and their neighbors based on these additional connections. Moreover, the conditional random field (CRF) is utilized as the loss function to achieve improved performance. Experimental results on two large-scale datasets NTU RGB+D and NTU RGB+D 120 show that the proposed model outperforms state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISCAS | 4 |
| 2020 | Auto-Encoder based Structured Dictinoary LearningabstractDictionary learning and deep learning are two popular representation learning paradigms, which can be combined to boost the classification task. However, existing combination methods often learn multiple dictionaries embedded in a cascade of layers, and a specialized classifier accordingly. This may inattentively lead to overfitting and high computational cost. In this paper, we present a novel deep auto-encoding architecture to learn only a dictionary for classification. To empower the dictionary with discrimination, we construct the dictionary with class-specific sub-dictionaries, and introduce supervision by imposing category constraints. The proposed framework is inspired by a sparse optimization method, namely Iterative Shrinkage Thresholding Algorithm, which characterizes the learning process by the forward-propagation based optimization w.r.t the dictionary only, reducing the number of parameters to learn and the computational cost dramatically. Extensive experiments demonstrate the effectiveness of our method in image classification. Deyin Liu, Yuanbo Lin Wu, Qichang Hu, Lin Qi 0001 |
MMSP | 5 |
| 2020 | Medi-Care AI: Predicting medications from billing codes via robust recurrent neural networks
Deyin Liu, Lin Wu 0001, Xue Li 0001, Lin Qi 0001 |
Neural Networks | 4 |
| 2020 | Multi-task image set classification via joint representation with class-level sparsity and intra-task low-rankness
Deyin Liu, Tie Yun, Lin Qi 0001 |
Pattern Recognit. Lett. | 4 |
| 2020 | Weighted hybrid fusion with rank consistency
Song Wang 0008, Xin Guo 0005, Tie Yun, Ivan Lee 0001, Lin Qi 0001, Ling Guan |
Pattern Recognit. Lett. | 5 |
| 2020 | Deep Local Feature Descriptor Learning With Dual Hard Batch ConstructionabstractLocal feature descriptor learning aims to represent distinctive images or patches with the same local features, where their representation is invariant under different types of deformation. Recent studies have demonstrated that descriptor learning based on Convolutional Neural Network (CNN) is able to improve the matching performance significantly. However, they tend to ignore the importance of sample selection during the training process, leading to unstable quality of descriptors and learning efficiency. In this paper, a dual hard batch construction method is proposed to sample the hard matching and non-matching examples for training, improving the performance of the descriptor learning on different tasks. To construct the dual hard training batches, the matching examples with the minimum similarity are selected as the hard positive pairs. For each positive pair, the most similar non-matching example is then sampled from the generated hard positive pairs in the same batch as the corresponding negative. By sampling the hard positive pairs and the corresponding hard negatives, the hard batches are produced to force the CNN model to learn the descriptors with more efforts. In addition, based on the above dual hard batch construction, an ℓ22 triplet loss function is built for optimizing the training model. Specifically, we analyze the superiority of the ℓ22 loss function when dealing with hard examples, and also demonstrate it in the experiments. With the benefits of the proposed sampling strategy and the ℓ22 triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmarks for different matching tasks. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
IEEE Trans. Image Process. | 4 |
| 2019 | A Novel Weighted Hybrid Multi-View Fusion Algorithm for Semi-Supervised ClassificationabstractSemi-supervised learning aims to improve the learning performance with very limited label information. To dig more available information from the collected data, we propose a weighted hybrid multi-view feature fusion approach for semi-supervised classification problem. Specifically, under the rank consistency constraint for labels predicted by view-specific learners, the proposed method estimates the optimal fusion weight for each learner to balance the incomparable square losses on different views. In this case, the learners with more powerful prediction capability are pushed to have higher weights during the fusion process. Experimental results on 6 real-world datasets demonstrate the effectiveness of the proposed technique. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 4 |
| 2019 | Local Feature Descriptor Learning with a Dual Hard Sampling StrategyabstractLocal feature descriptor learning based on Convolutional Neural Network (CNN) has demonstrated its capability to generate descriptors with high quality. While extensive studies focused on mining hard non-matching examples to improve descriptor learning performance, a random sampling strategy is adopted for matching examples. In this paper, a dual hard sampling strategy based on the triplet loss function is proposed to generate the hard matching and non-matching examples for training. To start with, a pair of matching examples with the maximum distance for each class are selected as the positive pair. For each positive pair, their closest non-matching example is then sampled from the generated positive pairs with other classes as the corresponding negative. Based on the above dual hard sampling strategy, a novel triplet loss function is presented for optimization. With the benefits of the proposed sampling strategy and the novel triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmark for local feature matching. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 4 |
| 2019 | Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action RecognitionabstractRecently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN and HCRF to retain the human skeleton structure information during the classification stage. The proposed model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework that takes the relative coordinates of the joints and bone direction as two static feature streams and the temporal displacements as the dynamic feature stream. Experimental results on two challenging benchmarks (NTU RGB+D, N-UCLA) show the superior performance of the proposed model over state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISM | 4 |
| 2019 | Region-of-Interest Detection Based on Statistical Distinctiveness for Panchromatic Remote Sensing ImagesabstractRegion-of-interest (ROI) detection plays a significant role in the analysis and interpretation of remote sensing images (RSI), due to the huge size of satellite images and their explosive growth in quantity. However, when applied to panchromatic RSI directly, traditional saliency models cannot achieve satisfying performance for two reasons: one is the computational efficiency decrease caused by the huge image size; the other is the absence of color information for panchromatic RSI. Thus, in this letter, an ROI detection model based on statistical distinctiveness (SD) is proposed for saliency analysis and ROIs detection in panchromatic RSI. The proposed SD model incorporates both the lower order SD (LSD) and the higher order SD (HSD), in order to identify regions of interest that are highly distinctive from the rest of the scene. Finally, the saliency map is determined by fusing cue maps obtained by calculating LSD locally and HSD globally. Experimental results show that our approach achieves promising results when compared with existing state-of-the-art saliency detection models. Guichi Liu, Lin Qi 0001, Tie Yun, Long Ma 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Hyperspectral Image Classification Using Kernel Fused Representation via a Spatial-Spectral Composite Kernel With Ideal RegularizationabstractTo adequately exploit spectral, spatial, and label information of the given hyperspectral data, a kernel fused representation-based classifier via a spatial-spectral composite kernel with ideal regularization (CKIR) method is proposed in this letter. Specifically, the learned CKIR is embedded into the kernel version of representation-based classifiers, i.e., kernel sparse representation-based classifier (KSRC) and kernel collaborative representation-based classifier (KCRC), to obtain more discriminative representation coefficients. Furthermore, to benefit from both sparsity and data correlation in representation, KSRC and KCRC are combined in the CKIR-based residual domain to further enhance the discriminative ability of the proposed classifier. The experimental results on two real hyperspectral images demonstrate that the proposed method outperforms the other state-of-the-art classifiers. Guichi Liu, Lin Qi 0001, Tie Yun, Long Ma 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | The Labeled Multiple Canonical Correlation Analysis for Information FusionabstractThe objective of multimodal information fusion is to mathematically analyze information carried in different sources and create a new representation that will be more effectively utilized in pattern recognition and other multimedia information processing tasks. In this paper, we introduce a new method for multimodal information fusion and representation based on the Labeled Multiple Canonical Correlation Analysis (LMCCA). By incorporating class label information of the training samples, the proposed LMCCA ensures that the fused features carry discriminative characteristics of the multimodal information representations and are capable of providing superior recognition performance. We implement a prototype of LMCCA to demonstrate its effectiveness on handwritten digit recognition, face recognition, and object recognition utilizing multiple features, bimodal human emotion recognition involving information from both audio and visual domains. The generic nature of LMCCA allows it to take as input features extracted by any means, including those by deep learning (DL) methods. Experimental results show that the proposed method enhanced the performance of both statistical machine learning methods, and methods based on DL. Lei Gao 0001, Rui Zhang 0010, Lin Qi 0001, Enqing Chen, Ling Guan |
IEEE Trans. Multim. | 3 |
| 2019 | Mobile-Traffic-Aware Offloading for Energy- and Spectral-Efficient Large-Scale D2D-Enabled Cellular NetworksabstractThis paper investigates how to enhance the energy and spectral efficiency (ESE) performance of large-scale cellular networks by offloading mobile traffic with the aid of device-to-device (D2D) communication. By appropriately exploiting the D2D-based mobile-traffic offloading mechanism, the users' behaviors and the specific network operating conditions, we develop an ESE evaluation framework for large-scale D2D-enabled cellular networks. This framework enables us to characterize the explicit relationship between the network's ESE and the offloading parameters as well as to quantify the influence of the users' behavior. Explicitly, we quantify the effects of the mobile-traffic intensity, the users' quality of service requirements as well as the base station density and other cellular system parameters on the achievable ESE. Tractable closed-form ESE-expressions are derived for a pair of spectrum sharing schemes, namely, D2D overlay and underlay in-band modes. Furthermore, we apply the analytical results to derive an optimal D2D-enabled mobile-traffic offloading scheme for the D2D overlay cellular networks to maximize the network's ESE under a specific maximal cellular user outage and D2D transmitter power constraint. The numerical and simulation results are provided to verify our modeling accuracy and to demonstrate the impact of the system parameters on the achievable ESE. Guogang Zhao, Sheng Chen 0001, Lin Qi 0001, Lajos Hanzo |
IEEE Trans. Wirel. Commun. | 3 |
| 2018 | Real-World Field Snail Detection and TrackingabstractWith the development of computer vision and machine learning, smart farming is becoming more popular and more important in agricultural industries. In this paper, we design and develop a snail detection and tracking system for real-world application. In this approach, deep learning is adopted to detect snails in the real-world. This can make full use of the computer's computing power to analyze big data and reduce researchers' workload. We have set up a snail dataset according to the video collected in the field environment, and we use a Faster R-CNN based algorithm to detect snails. Experiments show that this method can achieve good detection results. On this basis, by analyzing snail data sets, we optimized Faster R-CNN based algorithm according to the characteristics of snail's smaller size. These two methods are used by setting different anchor scale sizes and combining shallower features for detection. As a result, we improve the performance of snail detection in field conditions. We also adopt a linear Kalman filter as tracker to link objects into each trajectories. Ivan Lee 0001, Tie Yun, Jinhai Cai, Lin Qi 0001 |
ICARCV | 5 |
| 2018 | Incremental generalized multiple maximum scatter difference with applications to feature extraction
Ning Zheng 0003, Xin Guo 0005, Tie Yun, Nan Dong, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Joint intermodal and intramodal correlation preservation for semi-paired learning
Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
Pattern Recognit. | 4 |
| 2018 | 3D Human Action Recognition Using a Single Depth Feature and Locality-Constrained Affine Subspace CodingabstractThis paper addresses the problem of recognizing human actions from depth videos. We propose a depth-based local descriptor and affine subspace coding representation with locality-constrained affine subspace coding (LASC) for 3D action recognition. First, each depth video sequence is divided into a set of subsequences (i.e., multi-scale sub-actions) based on the normalized motion energy vector. Next, depth motion map-based gradient local auto-correlation features are employed to capture the shape information and motion cues of each sub-action. In order to obtain discriminative and compact representation, we extract the local high-order information of the depth video using LASC. Through experiments, we show that the use of LASC exhibits better performance compared with existing methods such as locality-constrained linear coding. We compared LASC with the state-of-the-art methods based on similar principle, using features extracted from a single modality, on four datasets, and with those using multiple features or nonlinear recognition machines. The results on four datasets clearly show the effectiveness of the proposed method. Chengwu Liang, Lin Qi 0001, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Discriminative Multiple Canonical Correlation Analysis for Information FusionabstractIn this paper, we propose the discriminative multiple canonical correlation analysis (DMCCA) for multimodal information analysis and fusion. DMCCA is capable of extracting more discriminative characteristics from multimodal information representations. Specifically, it finds the projected directions, which simultaneously maximize the within-class correlation and minimize the between-class correlation, leading to better utilization of the multimodal information. In the process, we analytically demonstrate that the optimally projected dimension by DMCCA can be quite accurately predicted, leading to both superior performance and substantial reduction in computational cost. We further verify that canonical correlation analysis (CCA), multiple canonical correlation analysis (MCCA) and discriminative canonical correlation analysis (DCCA) are special cases of DMCCA, thus establishing a unified framework for canonical correlation analysis. We implement a prototype of DMCCA to demonstrate its performance in handwritten digit recognition and human emotion recognition. Extensive experiments show that DMCCA outperforms the traditional methods of serial fusion, CCA, MCCA, and DCCA. Lei Gao 0001, Lin Qi 0001, Enqing Chen, Ling Guan |
IEEE Trans. Image Process. | 2 |
| 2017 | Heterogeneous Features Fusion with Collaborative Representation Learning for 3D Action RecognitionabstractHuman action recognition of depth sensors has drawn wide attentions in computer vision and multimedia processing areas. In contrast to simple periodic actions, irrelevant actions or sharing sub-actions between different classes of two-person non-periodic interactions make this task challenging. This paper presents heterogeneous features fusion with Collaborative Representation (CR) to address this challenge. Two effective high dimensional low-level features are developed from depth image sequence and skeleton pose sequence respectively. In the Canonical Correlations Analysis (CCA) feature space of these two features, Collaborative Representation (CR) is learned and adopted as the final high-level discriminative representation. Experiments on two depth action datasets (SBU Kinect-Interaction and MSR Action 3D) show that the proposed method is superior to the state-of-the-art methods compared, including some recent deep learning based methods. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2017 | Motion energy guided multi-scale heterogeneous features for 3D action recognitionabstractThis paper is to address the problem of human action recognition in depth sequences. The actions with various speeds and shared sub-actions make the recognition challenging. A new feature set, consisting of two heterogeneous features are proposed to address this challenge. Specifically, we propose an adaptive normalized action motion energy based on the depth video. Guided by this multi-scale energy vector, depth sequence and skeleton pose sequence are divided respectively into two sets of subsequences with multiple scales (i.e., multi-scale sub-actions). Then in depth modality, based on the depth sub-sequence, Depth Motion Maps (DMMs) based Histogram Oriented Gradient (HOG) features are employed to capture the shape information and motion cues. In skeleton modality, based on the pose sub-sequence, pose dynamics using skeleton information are extracted. In order to obtain discriminative and compact representation, the Collaborative Representation (CR) learning scheme based classifier is adopted. Experiments on two datasets show the effectiveness of the proposed method. Chengwu Liang, Lin Qi 0001, Ling Guan |
VCIP | 2 |
| 2016 | Information fusion based on kernel entropy component analysis in discriminative canonical correlation space with application to audio emotion recognitionabstractAs an information fusion tool, Kernel Entropy Component Analysis (KECA) is realized by using descriptor of information entropy and optimized by entropy estimation. However, as an unsuper-vised method, it merely puts the information or features from different channels together without considering their intrinsic structures and relations. In this paper, we introduce an enhanced version of KECA for information fusion, KECA in Discriminative Canonical Correlation Space (DCCS). Not only the intrinsic structures and discriminative representations are considered, but also the natural representations of input data are revealed by entropy estimation, leading to improved recognition accuracy. The effectiveness of the proposed solution is evaluated through experiments on two audio emotion databases. Experimental results show that the proposed solution outperforms the existing methods based on similar principles. Lei Gao 0001, Lin Qi 0001, Ling Guan |
ICASSP | 2 |
| 2016 | A Novel Discriminative Framework Integrating Kernel Entropy Component Analysis and Discriminative Multiple Canonical Correlation for Information FusionabstractThe effective interpretation and integration of multiple information content are important for the efficacious utilisation of multimedia in a wide variety of application context. The major challenge in information fusion lies in the difficulty of identifying the complementary and discriminatory representations from individual channels or data sources. In this paper, we propose a novel framework integrating kernel entropy-estimation and discriminative multiple canonical correlation (DMCC) to address this challenge. Not only the distribution and complementary representations of input data are revealed by entropy estimation, but also the discriminative representations are considered by DMCC, achieving improved recognition accuracy. The effectiveness of the proposed method is demonstrated on two audio emotion databases. Experimental results show that it outperforms the existing methods based on similar principles. Lei Gao 0001, Ling Guan, Lin Qi 0001, Enqing Chen |
ISM | 3 |
| 2016 | Semi-Supervised and Semi-Paired Graph Regularized Multiset Canonical Correlation AnalysisabstractMultiset canonical correlation analysis (MCCA) plays a key role in analyzing linear correlations among multimodal data. However, when facing semi-supervised and semi-paired multimodal data which widely exist in real world. MCCA normally performs poorly because it requires paired information among different models. At the same time, it fails to exploit the discriminative information due to the fact that it is an unsupervised dimension reduction method. In this paper, we propose a novel algorithm, named semi-supervised semi-paired graph regularized multiset canonical correlation analysis(SSGMCCA). SSGMCCA employs a small amount of paired data to perform MCCA and simultaneously utilize both the global structural information captured from the unlabeled data and the local structural information captured from the labeled data to compensate the limited paired. As a result, SSGMCCA can find the directions which not only maximal correlation among the multimodal data but also maximal separability of the labeled data. Experimental results illustrate the effectiveness of the proposed algorithm. Xin Guo 0005, Lin Qi 0001, Ling Guan |
ISM | 2 |
| 2016 | 3D Action Recognition Using Depth-Based Feature and Locality-Constrained Affine Subspace CodingabstractWe propose a 3D action recognition algorithm which uses depth-based Gradient Local Auto-Correlations (GLAC) feature and Locality-constrained Affine Subspace Coding (LASC) to improve the discriminative ability of human actions in spatio-temporal subsequences of 3D depth videos. First, each entire depth video sequence is divided automatically into a set of subsequences (i.e., multi-scale sub-actions) by the normalized motion energy vector. Next Depth Motion Maps (DMMs) based GLAC features are employed to capture the shape information and motion cues of each sub-action. In order to obtain a more compact and discriminative representation, LASC is then proposed to encode the features extracted from the depth video. We show that the use of LASC exhibits better performance compared to existing methods such as Locality-constrained Linear Coding (LLC). On all three datasets we obtain competitive results compared to fifteen methods, while using fewer features and less complex models. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2016 | Improving Action Recognition Using Collaborative Representation of Local Depth Map FeatureabstractBased on depth information, this letter introduces a new local depth map feature describing local spatiotemporal details of human motion and a collaborative representation for classification with regularized least squares. By extracting a multilayered depth motion feature and then applying a multiscale Histograms of Oriented Gradient (HOG) descriptor to it, the proposed feature characterizes the local temporal change of human motion and the local spatial structure (appearance) of an action. Instead of class-specific dictionary, the test action sample is represented collaboratively by the common shared dictionary. Moreover, we present an analytical solution of collaborative representation, which is independent of the query and can be precalculated as a projection matrix, leading to low computational cost in recognition. The evaluations on MSRAction3D and MSRGesture3D datasets demonstrate its effectiveness. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
IEEE Signal Process. Lett. | 3 |
| 2015 | Distribution of FRFT Coefficients of Natural Images
Guichi Liu, Lin Qi 0001 |
ICIG (2) | 3 |
| 2015 | Sparsity preserving multiple canonical correlation analysis with visual emotion recognition to multi-feature fusionabstractSparsity preserving projections (SPP) aim to preserve the sparse reconstructive relationship among the data and have been successfully applied to face recognition. The projections are invariant to rotations, rescalings, and translations of the data, and more importantly, they contain natural discriminating information even without class labels. Based on the concept of SSP, it presents a new method for multi-feature information fusion based on the Sparsity Preserving Multiple Canonical Correlation Analysis (SPMCCA), which can preserve the sparse reconstructive relationship of the data for recognition from multi-feature information representation. We implement a prototype of SPM-CCA with the application to visual-based human emotion recognition. Experimental results show that the proposed method outperforms the traditional methods of serial fusion, Canonical Correlation Analysis (CCA), Multiple Canonical Correlation Analysis (MCCA) and recently proposed Sparsity Preserving Canonical Correlation Analysis (SPCCA). Lei Gao 0001, Lin Qi 0001, Ling Guan |
ICIP | 2 |
| 2015 | Two-dimensional discriminant multi-manifolds locality preserving projection for facial expression recognitionabstractIn this paper, we assume that samples of different expressions reside on different manifolds and propose a novel human emotion recognition framework named two-dimensional discriminant multi-manifolds locality preserving projection (2D-DMLPP). 2D-DMLPP focuses on salient regions which reflect the significant variation from facial expression images so that it can learn an expression-specific model from salient patches rather than that of subject-specific. Furthermore, conventional manifold learning methods ignore the variation among nearby samples from the same class, leading to serious overfitting. We construct three adjacency graphs to model the margin and information, including diversity and similarity of salient patches from the same expression, and then incorporate the information and margin into dimensionality reduction function. Several experiments show that the proposed method significantly improves the recognition performance of facial expression recognition. Ning Zheng 0003, Xin Guo 0005, Lin Qi 0001, Ling Guan |
ISCAS | 3 |
| 2015 | Information Fusion of Audio Emotion Recognition Based on Kernel Entropy Component Analysis in Canonical Correlation SpaceabstractKernel Entropy Component Analysis(KECA), an effective information fusion tool, is realized using descriptor of information entropy and optimized by entropy estimation. However, it merely put the information or data from different channels together to achieve the information fusion without considering their intrinsic structures and relations. In this paper, we enhance the performance of KECA by introducing KECA in Canonical Correlation Space (CCS) or KECA+CCS. Not only the intrinsic structures and relations are considered in CCS, but also the nature of input data are revealed by entropy estimation. It improves the recognition accuracy effectively. The effectiveness of the proposed method is evaluated through experimentation on two audio-based emotion databases. The results show that the proposed method outperforms the existing methods based on similar principles. Lei Gao 0001, Lin Qi 0001, Ling Guan |
ISM | 2 |
| 2015 | A Novel Semi-Supervised Dimensionality Reduction Framework for Multi-manifold LearningabstractIn pattern recognition, traditional single manifold assumption can hardly guarantee the best classification performance, since the data from multiple classes does not lie on a single manifold. When the dataset contains multiple classes and the structure of the classes are different, it is more reasonable to assume each class lies on a particular manifold. In this paper, we propose a novel framework of semi-supervised dimensionality reduction for multi-manifold learning. Within this framework, methods are derived to learn multiple manifold corresponding to multiple classes in a data set, including both the labeled and unlabeled examples. In order to connect each unlabeled point to the other points from the same manifold, a similarity graph construction, based on sparse manifold clustering, is introduced when constructing the neighbourhood graph. Experimental results verify the advantages and effectiveness of this new framework. Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2015 | Action recognition using multi-layer Depth Motion maps and Sparse Dictionary LearningabstractIn this paper, we propose a new spatio-temporal feature based method for human action recognition using depth image sequence. Fist, Layered Depth Motion maps (LDM) are utilized to capture the temporal motion feature. Next, multi-scale HOG descriptors are computed on LDM to characterize the structural information of actions. Then sparse coding is applied for feature representation. Extending Sparse fisher Discriminative Dictionary Learning (SDDL) model and its corresponding classification scheme are also introduced. In SDDL model, the sub-dictionary is updated class by class, leading to class-specific compact discriminative dictionaries. The proposed method is evaluated on public MSR Action3D datasets and demonstrates great performance, especially in cross subject test. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
MMSP | 3 |
| 2015 | Spatio-Temporal Pyramid Model based on depth maps for action recognitionabstractThis paper presents a novel human action recognition method by using depth maps. Each depth frame in a depth video sequence is projected onto three orthogonal Cartesian planes. Under each projection view, we divide the entire depth maps into several sub-actions. The absolute difference between two consecutive projected maps is accumulated through a depth video (several sub-actions) sequence to form a Depth Motion Map (DMM) to describe the dynamic feature of an action. Also the difference within the threshold between two consecutive projected maps is calculated through the entire depth video to form another kind of Depth Static Map (DSM) to describe the static feature. Collectively, we call them Temporal Pyramid of Depth Model (TPDM). Then Spatial Pyramid Histograms of Oriented Gradient (SPHOG) is computed from the TPDM for the representation of an action. For classification, we apply support vector machine (SVM) to classify the proposed descriptorsbased on MSR Action3D dataset. Experimental results demonstrates the effectiveness of our proposed method. Haining Xu, Enqing Chen, Chengwu Liang, Lin Qi 0001, Ling Guan |
MMSP | 4 |
| 2015 | Advanced weight graph transformation matching algorithmabstractAn efficient and accurate point matching algorithm named advanced weight graph transformation matching (AWGTM) is proposed in this study. Instead of relying only on the elimination of dubious matches, the method iteratively reserve correspondences which have a small angular distance between two nearest‐neighbour graphs. The proposed algorithm is compared against weight graph transformation matching (WGTM) and graph transformation matching (GTM). Experimental results demonstrate the superior performance in eliminating outliers and reserving inliers of AWGTM algorithm under various conditions for images, such as duplication of patterns and non‐rigid deformation of objects. An execution time comparison is also presented, where AWGTM shows the best results for high outlier rates. Song Wang 0008, Xin Guo 0005, Xiaomin Mu, Yahong Huo, Lin Qi 0001 |
IET Comput. Vis. | 5 |
| 2014 | Incremental GMMSD2 with applications to feature extractionabstractThe generalized MMSD (GMMSD) is considered an efficient implementation of MMSD to extract discriminative information. However, a significant issue with the implementation of GMMSD is the complete recomputation of the training process when new training samples are presented. In this paper, we propose an alternative solution for feature extraction using the principles of GMMSD, which we call GMMSD2. GMMSD2 only requires the computation of centroid matrix, and it can overcome computational cost by applying efficient QR-updating techniques when new training samples are presented. Our experiments on FERET database demonstrate that incremental version of GMMSD2 eliminates the complete recomputation of the training process when new training samples are available, leading to significantly reduced computational cost. Ning Zheng 0003, Lin Qi 0001, Ling Guan |
ISCAS | 2 |
| 2014 | Generalized multiple maximum scatter difference feature extraction using QR decomposition
Ning Zheng 0003, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | Discriminative Multiple Canonical Correlation Analysis for Multi-feature Information FusionabstractThis paper presents a novel approach for multi-feature information fusion. The proposed method is based on the Discriminative Multiple Canonical Correlation Analysis (DMCCA), which can extract more discriminative characteristics for recognition from multi-feature information representation. It represents the different patterns among multiple subsets of features identified by minimizing the Frobenius norm. We will demonstrate that the Canonical Correlation Analysis (CCA), the Multiple Canonical Correlation Analysis (MCCA), and the Discriminative Canonical Correlation Analysis (DCCA) are special cases of the DMCCA. The effectiveness of the DMCCA is demonstrated through experimentation in speaker recognition and speech-based emotion recognition. Experimental results show that the proposed approach outperforms the traditional methods of serial fusion, CCA, MCCA and DCCA. Lei Gao 0001, Lin Qi 0001, Enqing Chen, Ling Guan |
ISM | 2 |
| 2012 | 2D-FRFT Based Rotation Invariant Digital Image WatermarkingabstractThe extraction of rotation invariant representation is important for many signal processing problems such as image analysis, computer vision, and pattern recognition. In this paper, we present a systematic analysis of the Two-Dimensional Fractional Fourier Transform (2D-FRFT), and show that under certain conditions, the 2D-FRFT technique possesses the attractive property of rotation invariance. Based on our analysis, we proposed a novel digital image watermarking method which combines 2D chirp signal with the addition and rotation invariant properties of 2D-FRFT to achieve improved robustness and security. The effectiveness of the proposed solution is demonstrated through experiments. Lei Gao 0001, Lin Qi 0001, Shouyi Yang, Yongjin Wang, Tie Yun, Ling Guan |
ISM | 2 |
| 2012 | Generalized MMSD feature extraction using QR decompositionabstractMultiple Maximum scatter difference (MMSD) discriminant criterion is an effective feature extraction method that computes the discriminant vectors from both the range of the between-class scatter matrix and the null space of the within-class scatter matrix. However, singular value decomposition (SVD) of two times is involved in MMSD, making this method impractical for high dimensional data. In this paper, we propose a novel method for feature extraction and classification based on MMSD criterion, called generalized MMSD (GMMSD), which employs QR decomposition rather than SVD. Unlike MMSD, GMMSD does not require the computation of the whole scatter matrix. Instead, it computes the discriminant vectors from both the range of whitenizated input data matrix and the null space of the within-class scatter matrix. We evaluate the effectiveness of the GMMSD method in terms of classification accuracy in the reduced dimensional space. Our experiments on two facial expression databases demonstrate that the GMMSD method provides favorable performance in terms of both recognition accuracy and computational efficiency. Ning Zheng 0003, Lin Qi 0001, Lei Gao 0001, Ling Guan |
VCIP | 2 |