EDBT 2026 Demo / reviewers in the wild / expert
Tie Yun
dblp:66/1748 · also Yun Tie
· DBLP profile ↗
61ranked-venue papers
11as first author
30since 2021 · last 2025
0000-0002-8258-6206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 11 since 2021Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PTRS: Parallax-Tolerant Robust Omnidirectional Deep Image StitchingabstractThe rapid advancement of technologies such as virtual and augmented reality has garnered substantial attention from both academia and industry. As the foundation of immersive multimedia content, omnidirectional image generation necessitates the development of robust and efficient image stitching algorithms. Unlike planar image stitching, omnidirectional image stitching involves stitching images captured by binocular or quadrinocular camera systems, introducing lager parallaxes (e.g., 90° or even 180°), which lead to pronounced distortions and hard-to-remove artifacts. Conventional planar image stitching methods based on optical flow typically employ unidirectional models. However, the intrinsic spatial relationships between sub-images in omnidirectional image necessitate the use of bidirectional optical flow. Existing approaches often estimate optical flow for each branch independently through simple feature mapping, neglecting the latent correlations between bidirectional flow. To address these challenges, we propose using a seam-driven approach to replace the traditional weighted blending strategy to effectively minimize artifacts. Additionally, we incorporate an advanced attention mechanism to establish a "bridge" that enables joint optimization of the two optical flow branches. Finally, we lightweight the model at a small cost of precision. Experimental results demonstrate that our method comprehensively outperforms the baseline model and generates natural omnidirectional images. Tie Yun, Dalong Zhang, Lei Shi 0001, Chenliang Ma |
IJCNN | 2 |
| 2025 | MCMG: Multi-level Controllable Music Generation Model Based on Fine-grained ControlabstractThe task of controlled music generation has been well developed, but the lack of modeling of music control attributes and neglect of music structure affect the quality of music creation. In order to solve the above problems, we first classify music into three levels according to its essentially different nature and extract the corresponding interpretable control attributes. Then by adding the control attributes to the music representation, a multi-level connection between music generation process and human composition is established. Finally, we propose a multi-level controllable music generation model with fine-grained control (MCMG). This model considers the structural relationships between music levels, enabling the music generation process highly controllable. The experiments show that our generative model not only improves on the basic music metric compared to the baseline, but also performs well on the controllability metric. Jingge Zhao, Qingwen Zhou, Tie Yun |
SMC | 6 |
| 2025 | Deep image stitching for panoramic stereoscopic live broadcast system
Tie Yun, Dalong Zhang, Yuning Gao |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | MrgaNet: Multi-scale recursive gated aggregation network for tracheoscopy images
Tie Yun, Dalong Zhang, Fenghui Liu, Lin Qi 0001 |
Image Vis. Comput. | 2 |
| 2025 | Dynamic gesture recognition using 3D central difference separable residual LSTM coordinate attention networks
Lichuan Geng, Tie Yun, Lin Qi 0001, Chengwu Liang |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | A Novel Vision Neural Network for Pan-Cancer Classification by Constructing Somatic Mutation Map With Feature Selection OptimizationabstractDue to the high fatality rate of cancer, timely detection and treatment in the early stage of cancer is very important. In this study, a method for constructing gene mutation maps based on the principle of RGB three-channel image was proposed to realize the dimensional transformation of somatic mutation data, making it suitable for the current image classification model. In order to better capture the global features of the mutation map, this paper proposes a network model M-MNet based on the inverted residual module and the multi-head self-attention module, which can effectively extract both local and global features, and the classification effect is better than that of the existing network models. Tie Yun, Dalong Zhang, Chenxu Quan |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Hybrid Learning Module-Based Transformer for Multitrack Music Generation With Music TheoryabstractIn recent years, multitrack music generation has garnered significant attention in both academic and industrial spheres for its versatile utilization of various instruments in collaborative settings. The primary challenge lies in achieving a harmonious balance within individual tracks and fostering effective collaboration across multiple tracks. To address this issue, this article introduces a pioneering hybrid learning encoder architecture. Each music track's encoder is implemented as an independent transformer architecture, preserving self-attention mechanisms within a single track and interattention mechanisms between different tracks. The resulting features are then seamlessly integrated into the decoder through concatenation. Of particular significance, previous multitrack music generation efforts have predominantly operated under unconditional settings, yielding music that lacks practical value due to noncompliance with established music theory principles. Recognizing this limitation, the article proposes a novel approach to multitrack music generation guided by music theory rules. Employing reinforcement learning techniques, the decoder-generated music serves as the initial state. Positive feedback is provided when the generated music adheres to music theory rules; conversely, negative feedback is applied to compel the multitrack music to align with widely accepted music theory principles. Finally, comprehensive simulation validation is conducted on both the publicly available LMD dataset and the self-constructed MUT dataset. The plethora of experimental results overwhelmingly corroborates the efficacy of the proposed methodology. Tie Yun, Xin Guo 0005, Jiessie Tie, Lin Qi 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | One-Stage Lightweight Network of Object Detection for Rectangular Panoramic Images
Yingying Lu, Tie Yun, Lin Qi 0001 |
ICIC (7) | 2 |
| 2024 | Multitrack Emotion-Based Music Generation Network Using Continuous Symbolic FeaturesabstractMusicians need an efficient composition process to yield a multitude of musical pieces. To enhance the artistry and emotional relevance of AI-generated music, a novel Multi-Track Emotion Music Generation model (MEMG) is proposed. MEMG minimizes reliance on emotion-labeled datasets, instead employing musical features such as conditional pitch histograms to signify music emotion attributes. By analyzing music features and emotional responses, MEMG facilitates controlled music generation with diverse emotions, where specific emotions are mapped based on Russell’s psychological model. A novel hybrid feedback reward network is constructed, integrating music theory rules, discriminator, and musical emotion interpretation. Incorporating music theory and embedding music emotion as input parameters during generation. Objective music theory evaluation and subjective auditory testing validate MEMG’s competitiveness in composing varied emotional music, coupled with exemplary adherence to musical rules. Tie Yun, Lin Qi 0001 |
ICME | 4 |
| 2024 | Multiple Query-based Multi-omics Fusion Algorithm for Cancer Classification and Metastasis PredictionabstractThe research on cancer classification and metastasis prediction based on multi-omics data has emerged with the development of high-throughput sequencing techniques and artificial intelligence. However, effective fusion among omics data remains a challenging problem, besides the ’curse of dimensionality’ resulting from high-dimensional features and small sample sizes. The transform-based fusion methods commonly used in computer vision and natural language processing require abundant samples for learning, which can lead to overfitting problems when applied in the biomedical field. Thus, this paper proposes a multiple query-based fusion algorithm with fewer parameters to learn, while retaining effective integration abilities for multiple omics data. It consists of interaction and fusion modules, which can prevent the model from training too aggressively and effectively mitigate overfitting, thus enabling the fusion of small datasets. Furthermore, this paper introduces a multi-task learning model based on it for cancer classification and metastasis prediction tasks. The proposed fusion algorithm can generalize to multiple omics data and maintain stable parameter counts. Experimental results demonstrate that the proposed method outperforms other state-of-the-art supervised multi-omics integrative methods in various cancer classification applications using DNA methylation, RNA expression, and miRNA expression data. Tie Yun, Fenghui Liu, Lin Qi 0001 |
IJCNN | 2 |
| 2024 | Dynamic Gesture Recognition Using a Spatio-Temporal Adaptive Multiscale Convolutional Transformer NetworkabstractThe Convolutional Neural Network has established itself as a cornerstone in gesture recognition systems. Nonetheless, its capacity to capture global contextual features is limited, presenting hurdles for dynamic gesture recognition. This paper introduces the Spatio-temporal Adaptive Multiscale Convolutional Transformer Network (SAMCT), a novel solution poised to propel the domain of dynamic gesture recognition forward. By integrating adaptive multiscale analysis with the transformer’s adeptness in processing sequential data, SAMCT refines the feature extraction process, providing a nuanced comprehension of gesture dynamics. The network incorporates a Spatiotemporal Adaptive Module (SAM), leveraging it to alleviate temporal redundancy and enabling gesture region selection for effective extraction of essential channels and spatial information from local features. To address CNNs’ constraints in extracting global contextual features, we propose a Multiscale Convolutional Transformer Network (MCT) that synergistically merges convolution and transformer networks to extract multiscale features of dynamic gestures. Our experimental evaluation using the Isogd dataset demonstrates that SAMCT surpasses other state-of-the-art recognition networks, showcasing superior recognition performance. Tie Yun, Lin Qi 0001, Chengwu Liang |
IJCNN | 2 |
| 2024 | ECA-SLAM: A Visual SLAM Utilizing Depth Features and Attention MechanismabstractSimultaneous Localization and Mapping (SLAM) has a broad application prospect in robot navigation, autonomous driving, and augmented reality. While visual SLAM systems have made significant progress, they still face challenges in complex environments due to the complexity and time variability of the real environment. Traditional methods have limitations in feature extraction, matching, and mapping, and the effect of localization and mapping can be greatly reduced due to scene changes. In response to these challenges, this paper proposes an improved visual SLAM method based on deep learning. We use a lightweight Mobilenetv2 network for feature extraction and introduce the Hardswish activation function to improve the performance further. Meanwhile, the ECA-Net attention mechanism is introduced into the feature extraction network to enhance its ability to capture key image features accurately. In addition, we use multi-task distillation to train our network to reduce the complexity of the model. Extensive experimental evaluation on multiple datasets reveals that the proposed method has significantly improved accuracy and robustness, particularly in challenging environments, demonstrating superior performance. Index Terms—SLAM, ECA-Net, deep learning, feature extraction Wenjie Deng, Tie Yun, Lin Qi 0001 |
IJCNN | 2 |
| 2024 | A Complementary Action Recognition Network based on Conv-TransformerabstractAs one of the important research directions in the field of computer vision, action recognition has extensive application value in today's internet. Since traditional convolutional neural networks performed well in processing local features, but the ability to process global information and performance on large-scale datasets is weaker. The transformer-based model can efficiently model global features through the attention mechanism, but it cannot capture the information of dynamic changes in spatial and temporal, and its ability to process local information is relatively weak. To do this, this paper proposed a Conv-Transformer Residual Network(CTRN) for action recognition, used a unique Conv-Transformer Residual Module(CTRM) to extract important features that combine global contextual and local information efficiently. Furthermore this paper designed a loss function consisting of pixel loss, cross-entropy loss and CIOU_Loss, allows action recognition networks to simultaneously optimise performance on pixel-level prediction, action classification and action localisation for more accurate action recognition. This network makes full use of the local modeling capabilities of the cnn and the advantages of the global modeling of the Transformer, which not only effectively reduces the complexity of the model, but also breaks through the limitations of the Transformer's lack of spatial sensing bias. Training the proposed model in an end-to-end manner and conducted extensive experiments on two typical datasets, Kinetics-400 and Something and Something V2, and outperformed some representative state-of-the-art models in both qualitative and quantitative evaluations. Linfu Liu, Lin Qi 0001, Tie Yun, Chengwu Liang |
IJCNN | 3 |
| 2023 | DCNet: Densely Connected Neural Network with Global Context for Abnormalities Prediction in Bronchoscopy ImagesabstractBronchoscopy plays an important role in the examination of lower respiratory tract diseases. However, the large number of medical images produced by bronchoscopy makes it a time-consuming and labor-intensive task for doctors to analyze. Therefore, developing computer-aided diagnostic algorithms to assist doctors in lung image analysis is of great significance.In this paper, we propose a novel densely connected convolutional network called DCNet with local and global receptive fields for the automatic prediction of bronchoscopy images. Our experimental results on the clinical bronchoscopy image analysis dataset demonstrate that our proposed method outperforms state-of-the-art algorithms in distinguishing malignant bronchoscopy images, with an accuracy rate of 92.81%. And DCNet achieved 99.31% accuracy on the public data set called Kvasir-Capsule. This model may be a valuable tool to assist medical staff in the analysis of bronchoscopy images. Tie Yun, Fenghui Liu, Lin Qi 0001, Ziqing Hu |
BIBM | 2 |
| 2023 | CTCM: Clustering based on three correlation matrices for multi-omics data integration and cancer subtype identificationabstractLarge-scale sequencing data is used by biologists to understand the biological systems and molecular mechanisms of disease. One challenge is to effectively use valuable information from different cancer omics to produce more accurate and reliable subtypes. We propose a clustering strategy based on three correlation matrices (CTCM) to identify cancer subtypes in multi-omics data. Connectivity matrix, similarity matrix and resampling matrix collect useful data information from different perspectives. The connection matrix divides the raw data into stable subtypes by adding noise to simulate the systematic error of the sequencing platform. The similar matrix use Gaussian kernel functions to construct connections between samples as the "skeleton" of the whole. The resampling matrix adapts to the explosive growth of data by sampling subsets. For each omics data, we combine the connectivity matrix and resampling matrix with the similarity matrix to generate an iterative version. Iterating the three relationship matrices for each omics produces a fusion matrix that is used for spectral clustering to identify cancer subtypes. Compared with six other state-of-the-art multi-omics clustering methods, CTCM achieves excellent performance on the benchmark data sets of simulation and TCGA databases. The method is general enough to replace existing unsupervised clustering techniques outside the scope of biomedical research to integrate multiple types of data. Tie Yun, Dalong Zhang, Fenghui Liu, Lin Qi 0001 |
BIBM | 2 |
| 2023 | MOVNG: Applied a Novel Sparse Fusion Representation into GTCN for Pan-Cancer Classification and Biomarker Identification
Tie Yun, Fenghui Liu, Dalong Zhang, Lin Qi 0001 |
ICIC (1) | 2 |
| 2023 | Intelligence Evaluation of Music Composition Based on Music Knowledge
Tie Yun, Lin Qi 0001 |
ICIC (5) | 2 |
| 2023 | Music Question Answering Based on Aesthetic ExperienceabstractMusic understanding has always been considered to be the work of experts. Ordinary people have insufficient aesthetic experience when facing music. We put forward the topic of music question and answer based on aesthetic experience to help people understand music more comprehensively. We summarized the relevant external characteristics of music elements that people are most concerned about in the field of music analysis, and generated a structured music aesthetic experience database. Based on this, we constructed the AMQA-dataset. We propose a new method of music graph representation, which can improve the reasonability and interpretability of music question answering system. We have verified our idea on the AMQA dataset, and achieved experimental results comparable to those of humans. Wenhao Gao 0004, Tie Yun, Lin Qi 0001 |
IJCNN | 3 |
| 2023 | Dynamic Gesture Recognition Based on 3D Central Difference Separable Residual LSTM Coordinate Attention Networks
Tie Yun, Lin Qi 0001, Chengwu Liang |
PRCV (8) | 2 |
| 2022 | PNF: a novel method based on connectivity and similarity for data integration and cancer subtypingabstractThe development of high-throughput technology enables measurements of many types of omics data, but the meaningful integration of different data types still is a significant challenge. Another difficult and vital challenge is the discovery of cancer molecular subtypes with relevant clinical differences. Here we propose a novel method, called perturbation network fusion clustering (PNF) for multi-omics data integration and cancer subtyping, which can address these two challenges. We creatively combine the connectivity and similarity of patient pairs, first adding statistical knowledge to the similarity network. Adopting perturbation clustering to get the probability (i.e., connectivity) that any two patients are grouped into a cluster on each type of data, and then calculating the similarity between any two patients using the Gaussian kernel function for each data type. Next using connectivity and similarity matrices we generate multiple stable and strong similarity kernels. Finally, we use the similarity network fusion strategy to fuse similarity kernel from each omics data and spectral clustering to discover cancer subtypes with survival differences. The method is validated on simulated data and six cancer datasets from The Cancer Genome Atlas (TCGA) nearly a thousand patient samples, including gene expression, microRNA, and DNA methylation data. PNF accurately identifies known cancer subtypes and novel subgroups of patients with significantly different survival profiles. The method without any prior knowledge is general enough to replace existing multi-omics data integration and unsupervised clustering methods outside the scope of biomedicine research. Fenghui Liu, Lin Qi 0001, Tie Yun |
BIBM | 5 |
| 2022 | WINMLP: Quantum & Involution Inspire False Positive Reduction in Lung Nodule Detection
Fenghui Liu, Lin Qi 0001, Tie Yun |
ICONIP (3) | 4 |
| 2022 | Fused Multiscale Cost Volume for Efficient Stereo MatchingabstractStereo matching is an important part of reconstructing depth information and 3D scenes. However, a important problem is improving accuracy and efficiency at the same time. We propose an end-to-end convolutional neural network model to estimate disparity from stereo images. To improve the accuracy of feature extraction in complex regions, we use atrous spatial pyramid pooling module to extract and fuse features, and atrous convolutions of different rates explicitly adjust the receptive field of the filter. To balance the accuracy of the algorithm and GPU memory consumption, we design a fused multiscale cost volume and the fusion strategy to improve the efficiency in estimating the disparity, which reduces the amount of memory consumption while ensuring the accuracy of the algorithm. Our model is trained and evaluated on SceneFlow and KITTI datasets. Experimental results demonstrate that our method achieve a better accuracy even in weak textures areas, while GPU memory consumption and runtime are reduced. Aihao Cui, Lin Qi 0001, Tie Yun |
MMSP | 3 |
| 2022 | RDDP: Reliable Detection and Description of Interest PointsabstractRecent works have paid much attention on learning repeatable maps for detection and description of interest points. However, there are always some components which is repeatable but not discriminative in reality. Training on the whole images with these components will lead to poor matching performance. Thus, we propose a self-supervised network with a new branch which can eliminate the negative influence of unreliable image components. We also design a loss function including a part calculated with Average Precision (AP) to make full use of the reliability branch. Evaluations on HPatches dataset show that the proposed method achieves competitive results compared with state-of-the-art methods. Yuning Gao, Lin Qi 0001, Tie Yun |
SMC | 3 |
| 2022 | Deep learning based audio and video cross-modal recommendationabstractWith the rapid development of the Internet multimedia industry, audio and video cross-modal retrieval and recommendation have become a compelling topic in deep learning. However, most methods still have certain problems. For example, it takes a lot of effort to find suitable background music for the video, and the recommendations mode of the audio-visual database platform are too singular. This article will take the matching problem of video and music as the starting point in optimizing the music retrieval mode. To improve the multimedia website recommendation mechanism, a video and audio cross-domain recommendation model KATLN is proposed. We have incorporated the ideas of domain adaptation and attention mechanism, obtained rich item feature representations, captured users’ finer-grained interest preferences, and improved the accuracy of modeling. Through experiments, it can be seen that, compared with the existing mainstream algorithms, the KATLN proposed in this paper has outstanding performance and is more prominent in accurately grasping user preferences. Tie Yun, Chu Jie Jiessie Tie |
SMC | 1 |
| 2021 | Face Recognition Based on Panoramic VideoabstractRecently, face recognition has been widely used in lives, but it can still only maintain good results in specific scenarios. In different tasks, changes in the external environment will cause diversified changes in face images, such as panoramic video. The barrel distortion of the panoramic images generated by the fisheye lens will cause the face images to be deformed to varying degrees around the fisheye image. A new model DCMNet is proposed in this paper and a panoramic face dataset PIFI is created based on the panoramic video system. After experiments are conducted on different datasets, the results show that our model has better results on popular datasets and panoramic dataset. Tie Yun, Lin Qi 0001, Rui Zhang 0010, Juanjuan Cai |
FG | 2 |
| 2021 | Local Feature Descriptors with Deep Hypersphere LearningabstractRecent works have demonstrated the power of L2normalization in local feature descriptor learning. While the descriptors are typically learned in the Euclidean space, the similarity between descriptors is often evaluated on a unit hypersphere due to the post-processing of L2normalization for descriptors, which creates a gap between the training stage and the usage stage of feature descriptors. To bridge the gap, we propose a hyperspherical descriptor learning model, where the whole network is projected onto the hyperspherical space. In addition, a squared angular triplet loss is designed to enable the proposed hyperspherical model to learn angularly discriminative descriptors. Experiments on UBC dataset show that the proposed hyperspherical descriptor outperforms its Euclidean counterparts and the state-of-the-art methods on the feature matching task. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 3 |
| 2021 | Discriminative Patch Descriptor Learning With Focal Triplet Loss FunctionabstractThis paper proposes a focal triplet loss function for discriminative patch descriptor learning. The standard triplet loss function usually restrains the distance difference between the matching samples and the non-matching ones. However, along with the training procedure, the majority of triplets in each batch tend to satisfy the constraint of the loss function and produce low loss values, leading to a masquerade that the model is well-trained. To address this problem, the focal triplet loss function is proposed in this paper to weaken the impact of the easy triplets and focus training on the hard ones. By emphasizing the importance of hard triplets on the model training, the proposed loss forces the descriptor vectors with fixed dimension to carry more discriminative information from the patches. With the benefits of the focal mechanism, the proposed method achieves better performance compared to the state-of-the-art on UBC dataset for image matching task. Furthermore, to demonstrate the effectiveness of the proposed method, we extend the focal triplet loss on the cross-model retrieval task. The experimental results indicate that the proposed method can also be used to improve visual-semantic embedding learning. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 3 |
| 2021 | Context-Aware Based Visual-Audio Feature Fusion for Emotion RecognitionabstractVideo emotion recognition is a significant branch in the field of emotion computing. However, traditional recognition works mainly focus on human features, ignoring the contextual clues of video scenes and objects. In our work, we propose a context-aware framework for bi-modal video emotion recognition. Unlike existing methods that directly extract features of the entire video frame, we extract key frames and key regions of videos to obtain emotional cues contained in video scenes and objects. Specifically, for visual stream, the hierarchical Bidirectional Long-Short Term Memory (Bi-LSTM) is applied to summarize video scenes and find key frames that mostly contribute to video emotion; Meantime, we introduce the Region Proposal Network (RPN) to extract corresponding features of object regions in video frames and construct the emotional similarity graph. After using the Feedforward Neural Network (FNN) to assign different weight coefficients to different regions, the Graph Convolutional Network (GCN) is used to reason about the connections between key regions. Moreover, the context information of the frame-level Log-Mel spectrum fragments supplement the visual information. Finally, we fuse the visual and acoustics features by adaptive gated multimodal fusion module for video emotion classification. We conduct experiments on Video Emotion-8 and Ekman-6 datasets. The experimental results demonstrate that our model achieves better classification accuracy than several baseline models. Huijie Cheng, Tie Yun, Lin Qi 0001 |
IJCNN | 2 |
| 2021 | Research on Panoramic Stereo Live Streaming Based on the Virtual RealityabstractVirtual Reality (VR) technology has been successfully applied in many fields, such as education, tourism, games, etc. Some researchers hope to combine VR to watch panoramic stereo video and bring a brand new living experience in the field of live broadcasting. However, in practical applications, there are many problems such as panoramic video real-time stitching and high-resolution video compression coding, so it is still difficult to construct a panoramic stereo living video system. In this paper, we design and implement a panoramic stereo video living system. The system based on the eight-eye panoramic annular lens, the 360-Degree binocular video is stitched. Compared with the traditional coding algorithm, H.265 coding algorithm is used to compress the video, which improves the compression rate by more than fifty percent. Then the video stream is pushed to the cloud for forwarding. The receiving end combines VR helmet to watch panoramic stereo video live, bringing viewers a brand new viewing experience. Mingyao Zheng, Tie Yun, Lin Qi 0001, Yuning Gao |
ISCAS | 2 |
| 2021 | Auto-encoder based structured dictionary learning for visual classification
Deyin Liu, Chengwu Liang, Shaokang Chen, Tie Yun, Lin Qi 0001 |
Neurocomputing | 4 |
| 2020 | Negative Label Guided Discriminative Canonical Correlation Analysis for Semi-Supervised and Semi-Paired LearningabstractSemi-supervised learning is a popular trend for learning based methods in recent years, as it fully exploits both the labeled and unlabeled samples in a dataset. This paper sets itself apart from most existing semi-supervised learning algorithms, which only use the exact labels of data already known. We take the negative label as side information to guide the process of semi-supervised learning. Two types of supervision information are regarded as negative label; the first type indicates that a sample definitely does not belong to a specific category, and the second indicates that two samples come from different views, and cannot have a one to one correspondence. By reasonably assuming that nearby points should have similar class indicators, the data labels are propagated under the negative label and the geometric structure revealed by both labeled and unlabeled points. Specifically, we predict one to one pair information by utilizing the neighbor information of samples, under the guidance of the negative pair label. Extensive experiments on several datasets demonstrate the effectiveness of our proposed method. Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 3 |
| 2020 | A Style-Specific Music Composition Neural Network
Tie Yun, Shouxun Liu |
Neural Process. Lett. | 2 |
| 2020 | Multi-task image set classification via joint representation with class-level sparsity and intra-task low-rankness
Deyin Liu, Tie Yun, Lin Qi 0001 |
Pattern Recognit. Lett. | 3 |
| 2020 | Weighted hybrid fusion with rank consistency
Song Wang 0008, Xin Guo 0005, Tie Yun, Ivan Lee 0001, Lin Qi 0001, Ling Guan |
Pattern Recognit. Lett. | 3 |
| 2020 | Deep Local Feature Descriptor Learning With Dual Hard Batch ConstructionabstractLocal feature descriptor learning aims to represent distinctive images or patches with the same local features, where their representation is invariant under different types of deformation. Recent studies have demonstrated that descriptor learning based on Convolutional Neural Network (CNN) is able to improve the matching performance significantly. However, they tend to ignore the importance of sample selection during the training process, leading to unstable quality of descriptors and learning efficiency. In this paper, a dual hard batch construction method is proposed to sample the hard matching and non-matching examples for training, improving the performance of the descriptor learning on different tasks. To construct the dual hard training batches, the matching examples with the minimum similarity are selected as the hard positive pairs. For each positive pair, the most similar non-matching example is then sampled from the generated hard positive pairs in the same batch as the corresponding negative. By sampling the hard positive pairs and the corresponding hard negatives, the hard batches are produced to force the CNN model to learn the descriptors with more efforts. In addition, based on the above dual hard batch construction, an ℓ22 triplet loss function is built for optimizing the training model. Specifically, we analyze the superiority of the ℓ22 loss function when dealing with hard examples, and also demonstrate it in the experiments. With the benefits of the proposed sampling strategy and the ℓ22 triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmarks for different matching tasks. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
IEEE Trans. Image Process. | 3 |
| 2019 | A Novel Weighted Hybrid Multi-View Fusion Algorithm for Semi-Supervised ClassificationabstractSemi-supervised learning aims to improve the learning performance with very limited label information. To dig more available information from the collected data, we propose a weighted hybrid multi-view feature fusion approach for semi-supervised classification problem. Specifically, under the rank consistency constraint for labels predicted by view-specific learners, the proposed method estimates the optimal fusion weight for each learner to balance the incomparable square losses on different views. In this case, the learners with more powerful prediction capability are pushed to have higher weights during the fusion process. Experimental results on 6 real-world datasets demonstrate the effectiveness of the proposed technique. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 3 |
| 2019 | Local Feature Descriptor Learning with a Dual Hard Sampling StrategyabstractLocal feature descriptor learning based on Convolutional Neural Network (CNN) has demonstrated its capability to generate descriptors with high quality. While extensive studies focused on mining hard non-matching examples to improve descriptor learning performance, a random sampling strategy is adopted for matching examples. In this paper, a dual hard sampling strategy based on the triplet loss function is proposed to generate the hard matching and non-matching examples for training. To start with, a pair of matching examples with the maximum distance for each class are selected as the positive pair. For each positive pair, their closest non-matching example is then sampled from the generated positive pairs with other classes as the corresponding negative. Based on the above dual hard sampling strategy, a novel triplet loss function is presented for optimization. With the benefits of the proposed sampling strategy and the novel triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmark for local feature matching. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2019 | Semantic based autoencoder-attention 3D reconstruction network
Long Ye, Wei Zhong 0001, Tie Yun, Qin Zhang 0009 |
Graph. Model. | 5 |
| 2019 | Region-of-Interest Detection Based on Statistical Distinctiveness for Panchromatic Remote Sensing ImagesabstractRegion-of-interest (ROI) detection plays a significant role in the analysis and interpretation of remote sensing images (RSI), due to the huge size of satellite images and their explosive growth in quantity. However, when applied to panchromatic RSI directly, traditional saliency models cannot achieve satisfying performance for two reasons: one is the computational efficiency decrease caused by the huge image size; the other is the absence of color information for panchromatic RSI. Thus, in this letter, an ROI detection model based on statistical distinctiveness (SD) is proposed for saliency analysis and ROIs detection in panchromatic RSI. The proposed SD model incorporates both the lower order SD (LSD) and the higher order SD (HSD), in order to identify regions of interest that are highly distinctive from the rest of the scene. Finally, the saliency map is determined by fusing cue maps obtained by calculating LSD locally and HSD globally. Experimental results show that our approach achieves promising results when compared with existing state-of-the-art saliency detection models. Guichi Liu, Lin Qi 0001, Tie Yun, Long Ma 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Hyperspectral Image Classification Using Kernel Fused Representation via a Spatial-Spectral Composite Kernel With Ideal RegularizationabstractTo adequately exploit spectral, spatial, and label information of the given hyperspectral data, a kernel fused representation-based classifier via a spatial-spectral composite kernel with ideal regularization (CKIR) method is proposed in this letter. Specifically, the learned CKIR is embedded into the kernel version of representation-based classifiers, i.e., kernel sparse representation-based classifier (KSRC) and kernel collaborative representation-based classifier (KCRC), to obtain more discriminative representation coefficients. Furthermore, to benefit from both sparsity and data correlation in representation, KSRC and KCRC are combined in the CKIR-based residual domain to further enhance the discriminative ability of the proposed classifier. The experimental results on two real hyperspectral images demonstrate that the proposed method outperforms the other state-of-the-art classifiers. Guichi Liu, Lin Qi 0001, Tie Yun, Long Ma 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Joint optimization of routing and VM resource allocation for multimedia cloud
Wenqiang Gong, Jinyao Yan, Xiaoming Nan, Tie Yun |
Multim. Syst. | 4 |
| 2018 | Real-World Field Snail Detection and TrackingabstractWith the development of computer vision and machine learning, smart farming is becoming more popular and more important in agricultural industries. In this paper, we design and develop a snail detection and tracking system for real-world application. In this approach, deep learning is adopted to detect snails in the real-world. This can make full use of the computer's computing power to analyze big data and reduce researchers' workload. We have set up a snail dataset according to the video collected in the field environment, and we use a Faster R-CNN based algorithm to detect snails. Experiments show that this method can achieve good detection results. On this basis, by analyzing snail data sets, we optimized Faster R-CNN based algorithm according to the characteristics of snail's smaller size. These two methods are used by setting different anchor scale sizes and combining shallower features for detection. As a result, we improve the performance of snail detection in field conditions. We also adopt a linear Kalman filter as tracker to link objects into each trajectories. Ivan Lee 0001, Tie Yun, Jinhai Cai, Lin Qi 0001 |
ICARCV | 3 |
| 2018 | Design of linear-phase nonsubsampled nonuniform directional filter bank with arbitrary directional partitioning
Long Ye, Tie Yun, Wei Zhong 0001, Qin Zhang 0009 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Identifying facial expression using adaptive sub-layer compensation based feature extraction
Xin Guo 0005, Tie Yun, Long Ye, Jinyao Yan |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Multimedia analysis with collective intelligence
Meng Wang 0001, Qi Tian 0001, Abdulmotaleb El Saddik, Mathias Lux, Tie Yun |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Incremental generalized multiple maximum scatter difference with applications to feature extraction
Ning Zheng 0003, Xin Guo 0005, Tie Yun, Nan Dong, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Joint intermodal and intramodal correlation preservation for semi-paired learning
Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
Pattern Recognit. | 3 |
| 2015 | A new audiovisual emotion recognition system using entropy-estimation-based multimodal information fusionabstractWe present a novel audiovisual emotion recognition solution using multimodal information fusion based on entropy estimation. Considering the limitations of existing methods, we propose a new dual-level fusion framework which consists of feature level fusion module based on kernel entropy component analysis and score level fusion module based on maximum correntropy criterion. In our system, audio and visual channels are utilized to detect and classify emotional states for intelligent human computer interfaces. Our extensive experimental study on eNTERFACE database and RML database demonstrates the feasibility of the proposed multimodal emotion recognition framework based on integrated analysis of speech and facial expression. The experimental results show that the proposed methods are capable of providing improved performance. The comparison with other methods shows that the proposed two-stage fusion platform outperforms the traditional algorithms in terms of both accuracy and reliability. Zhibing Xie, Tie Yun, Ling Guan |
ISCAS | 2 |
| 2015 | A Novel Semi-Supervised Dimensionality Reduction Framework for Multi-manifold LearningabstractIn pattern recognition, traditional single manifold assumption can hardly guarantee the best classification performance, since the data from multiple classes does not lie on a single manifold. When the dataset contains multiple classes and the structure of the classes are different, it is more reasonable to assume each class lies on a particular manifold. In this paper, we propose a novel framework of semi-supervised dimensionality reduction for multi-manifold learning. Within this framework, methods are derived to learn multiple manifold corresponding to multiple classes in a data set, including both the labeled and unlabeled examples. In order to connect each unlabeled point to the other points from the same manifold, a similarity graph construction, based on sparse manifold clustering, is introduced when constructing the neighbourhood graph. Experimental results verify the advantages and effectiveness of this new framework. Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 2 |
| 2013 | Human emotion recognition using the adaptive sub-layer-compensation based facial edge detectionabstractIn this paper, we propose an adaptive sub-layer compensation (ASLC) based facial edge detection for human emotion recognition. We modify the Marr-Hildreth edge detector with Wiener filtering, sub-layer compensation and hysteresis analysis to compensate the negative effects of the Laplacian of Gaussian (LoG) operator such as image degradation, high response to unwanted details, and disconnected edges. We investigate a Discriminative Isomap (D-Isomap) based approach that combines the ASLC feature and deformable elastic body spline (EBS) feature for the final decision. RML emotion database and Cohn-Kanade database are used for the experiment and the results demonstrate the effectiveness of the proposed method. Tie Yun, Anastasios N. Venetsanopoulos, Ling Guan |
ISCAS | 2 |
| 2013 | Human emotional state recognition using real 3D visual features from Gabor library
Tie Yun, Ling Guan |
Pattern Recognit. | 1 |
| 2013 | A Deformable 3-D Facial Expression Model for Dynamic Human Emotional State RecognitionabstractAutomatic emotion recognition from facial expression is one of the most intensively researched topics in affective computing and human-computer interaction. However, it is well known that due to the lack of 3-D feature and dynamic analysis the functional aspect of affective computing is insufficient for natural interaction. In this paper, we present an automatic emotion recognition approach from video sequences based on a fiducial point controlled 3-D facial model. The facial region is first detected with local normalization in the input frames. The 26 fiducial points are then located on the facial region and tracked through the video sequences by multiple particle filters. Depending on the displacement of the fiducial points, they may be used as landmarked control points to synthesize the input emotional expressions on a generic mesh model. As a physics-based transformation, elastic body spline technology is introduced to the facial mesh to generate a smooth warp that reflects the control point correspondences. This also extracts the deformation feature from the realistic emotional expressions. Discriminative Isomap-based classification is used to embed the deformation feature into a low dimensional manifold that spans in an expression space with one neutral and six emotion class centers. The final decision is made by computing the nearest class center of the feature space. Tie Yun, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Human emotion recognition using a deformable 3D facial expression modelabstractAutomatic emotion recognition from facial expression is one of the most intensively researched topics in affective computing and human-computer interaction. However, due to the lack of 3D feature and dynamic analysis the functional aspect of affective computing is insufficient for natural interaction. This paper presents an automatic emotion recognition approach from video sequences based on a fiducial point controlled 3D facial model. As a physics-based transformation, elastic body spline technology is applied on a facial mesh to generate a smooth warp that reflects the control point corresponding to the displacement of fiducial points. It also extracts the deformation feature from the realistic emotional expressions. Discriminative Isomap based classification is used to embed the deformation feature into a low dimensional manifold that spans in an expression space with one neutral and six emotion class centers. The final decision is made by computing the Nearest Class Center of the feature space. Tie Yun, Ling Guan |
ISCAS | 1 |
| 2012 | 2D-FRFT Based Rotation Invariant Digital Image WatermarkingabstractThe extraction of rotation invariant representation is important for many signal processing problems such as image analysis, computer vision, and pattern recognition. In this paper, we present a systematic analysis of the Two-Dimensional Fractional Fourier Transform (2D-FRFT), and show that under certain conditions, the 2D-FRFT technique possesses the attractive property of rotation invariance. Based on our analysis, we proposed a novel digital image watermarking method which combines 2D chirp signal with the addition and rotation invariant properties of 2D-FRFT to achieve improved robustness and security. The effectiveness of the proposed solution is demonstrated through experiments. Lei Gao 0001, Lin Qi 0001, Shouyi Yang, Yongjin Wang, Tie Yun, Ling Guan |
ISM | 5 |
| 2010 | Fiducial point tracking for facial expression using multiple particle filters with kernel correlation analysisabstractDetecting and tracking fiducial points successfully can generate necessary dynamic and deformable information for facial image interpretation tasks with numerous potential applications. In this paper we propose an automatic fiducial points tracking method using multiple Differential Evolution Markov Chain (DE-MC) particle filters with kernel correlation techniques. Fiducial points are initialized through the scale invariant feature based detectors. By taking the advantage of the ability to approximate complicated proposal distributions, multiple DE-MC particle filters are applied for fiducial points tracking by building a path connecting sampling with measurements, based on the fact that the posteriori depends on both the previous state and the current observation. A Kernel correlation analysis approach is proposed to find the detection likelihood with maximization of the similarity criterion between the target points and the candidate points. Sampling efficiency is improved and computational time is substantially reduced by making use of the intermediate results obtained in particle allocation. Tie Yun, Ling Guan |
ICIP | 1 |
| 2010 | Human emotion recognition using real 3D visual features from Gabor libraryabstractEmotional state recognition is an important component for efficient human-computer interaction. Most existing works address this problem using 2D features, but they are sensitive to head pose, clutter, and variations in lighting conditions. The general 3D based methods only consider geometric information for feature extraction. In this paper, we present a real 3D visual features based method for human emotion recognition. 3D geometric information plus colour/density information of the facial expressions are extracted by 3D Gabor library to construct visual feature vectors. The filter's scale, orientation, and shape of the library are specified according to the appearance patterns of the 3D facial expressions. An improved kernel canonical correlation analysis (IKCCA) algorithm is proposed for final decision. From training samples, the semantic ratings that describe the different facial expressions are computed by IKCCA to generate a seven dimensional semantic expression vector. It is applied for learning the correlation with different testing samples. According to this correlation, we estimate the associated expression vector and perform expression classification. From experiment results, our proposed method demonstrates impressive performance. Tie Yun, Ling Guan |
MMSP | 1 |
| 2009 | Multimedia multimodal methodologiesabstractThis paper outlines several multimedia systems that utilize a multimodal approach. These systems include audiovisual based emotion recognition, image and video retrieval, and face and head tracking. Data collected from diverse sources/sensors are employed to improve the accuracy of correctly detecting, classifying, identifying, and tracking of a desired object or target. It is shown that the integration of multimodality data will be more efficient and potentially more accurate than if the data was acquired from a single source. A number of cutting-edge applications for multimodal systems will be discussed. An advanced assistance robot using the multimodal systems will be presented. Ling Guan, Paisarn Muneesawang, Yongjin Wang, Rui Zhang 0010, Tie Yun, Adrian Bulzacki, Muhammad Talal Ibrahim |
ICME | 5 |
| 2009 | Toward natural and efficient human computer interactionabstractEffective detection, recognition, interpretation, and analysis of human physiological and behavioral characteristics are of fundamental importance in the design and development of intelligent human computer interaction (HCI) systems. This paper illustrates the issues and challenges in the design of such systems through two real examples, emotion recognition and face detection. In particular, we focus on audiovisual based bimodal emotion recognition, face detection in crowded scene, and facial fiducial points detection. The integration of these systems is expected to produce more robust and stronger performance, and provide more natural and friendly man-machine interaction. Ling Guan, Yongjin Wang, Tie Yun |
ICME | 3 |
| 2009 | Automatic fiducial points detection for facial expressions using scale invariant featureabstractDetecting fiducial points successfully in facial images or video sequences can play an important role in numerous facial image interpretation tasks such as face detection and identification, facial expression recognition, emotion recognition, and face image database management. In this paper we propose an automatic and robust method of facial fiducial point's detection for facial expressions analysis in video sequences using scale invariant feature based Adaboost classifiers. Face region is first located using the face detector with local normalization and optimal adaptive correlation technique. Candidate points are then selected over the face region using local scale-space extrema detection. The scale invariant feature for each candidate point is extracted for further examination. We choose 26 fiducial points on the face region from training samples to build the fiducial point detectors with Adaboost classifiers. All the candidate points in the test samples are examined through these detectors. Finally, all the 26 facial fiducial points are located on each frame of the test samples. Cohn-Kanade database and Mind Reading DVD are used for experiment. The results show that our method achieves a good performance of 90.69% average recognition rate. Tie Yun, Ling Guan |
MMSP | 1 |
| 2009 | Automatic face detection in video sequences using local normalization and optimal adaptive correlation techniques
Tie Yun, Ling Guan |
Pattern Recognit. | 1 |
| 2008 | Local Normalization with Optimal Adaptive Correlation for Automatic and Robust Face Detection on Video SequencesabstractThis paper proposes an automatic and robust method to detect human faces from video sequences that combines feature extraction and face detection based on local normalization, Gabor wavelets transform and Adaboost algorithm. The key step and the main contribution of this work is the incorporation of a normalization technique based on local histograms with optimal adaptive correlation (OAC) technique to alleviate a common problem in conventional face detection methods: inconsistent performance due to the sensitivity to illumination variations such as local shadowing, noise and occlusion. This approach uses a cascade of classifiers to adopt a coarse-to-fine strategy to achieve higher detection rate with lower false positives. The experimental results demonstrate a significant performance improvement by local normalization over method without normalizations in real video sequences with a wide range of facial variations in color, position, scale, and varying lighting conditions. Tie Yun, Ling Guan |
ISM | 1 |