EDBT 2026 Demo / reviewers in the wild / expert
Xin Guo 0005
dblp:17/1430-5
· DBLP profile ↗
24ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-4153-4642ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-authorComputer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Low-Complexity Noncoherent Uplink for IoT Devices in Massive MIMO SystemsabstractThis paper presents a noncoherent massive multiple-input multiple-output (MIMO) uplink scheme designed for low-complexity Internet of Things (IoT) devices. We consider a scenario where a single device equipped with two antennas transmits short packets to a base station with a large antenna array. To minimize latency and avoid pilot overhead, we introduce a structured space-time constellation design. Our key innovation is a parametric constellation framework that jointly encodes information in amplitude, mixing angle, and phase within an Alamouti structure. By optimizing the Kullback-Leibler (KL) divergence, we formulate the design as a tractable problem of selecting parameters for geometric and arithmetic sequences, rather than a complex search over all possible constellations. This results in a scheme specified by a few parameters, enabling both minimal storage and a low-complexity sequential maximum likelihood detector. Theoretical and numerical results demonstrate that our design achieves a larger minimum KL divergence and a superior symbol error rate compared to benchmarks, including Riemannian distance-based codes, KL-optimized single-input multiple-output schemes, and noncoherent pulse amplitude modulation, particularly in the high-SNR regime, where it eliminates error floors. Shuangzhi Li 0001, Xin Guo 0005 |
IEEE Internet Things J. | 3 |
| 2025 | Hybrid Learning Module-Based Transformer for Multitrack Music Generation With Music TheoryabstractIn recent years, multitrack music generation has garnered significant attention in both academic and industrial spheres for its versatile utilization of various instruments in collaborative settings. The primary challenge lies in achieving a harmonious balance within individual tracks and fostering effective collaboration across multiple tracks. To address this issue, this article introduces a pioneering hybrid learning encoder architecture. Each music track's encoder is implemented as an independent transformer architecture, preserving self-attention mechanisms within a single track and interattention mechanisms between different tracks. The resulting features are then seamlessly integrated into the decoder through concatenation. Of particular significance, previous multitrack music generation efforts have predominantly operated under unconditional settings, yielding music that lacks practical value due to noncompliance with established music theory principles. Recognizing this limitation, the article proposes a novel approach to multitrack music generation guided by music theory rules. Employing reinforcement learning techniques, the decoder-generated music serves as the initial state. Positive feedback is provided when the generated music adheres to music theory rules; conversely, negative feedback is applied to compel the multitrack music to align with widely accepted music theory principles. Finally, comprehensive simulation validation is conducted on both the publicly available LMD dataset and the self-constructed MUT dataset. The plethora of experimental results overwhelmingly corroborates the efficacy of the proposed methodology. Tie Yun, Xin Guo 0005, Jiessie Tie, Lin Qi 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | A Deep Semantic Segmentation Network With Semantic and Contextual RefinementsabstractSemantic segmentation is a fundamental task in multimedia processing, which can be used for analyzing, understanding, editing contents of images and videos, among others. To accelerate the analysis of multimedia data, existing segmentation researches tend to extract semantic information by progressively reducing the spatial resolutions of feature maps. However, this approach introduces a misalignment problem when restoring the resolution of high-level feature maps. In this paper, we design a Semantic Refinement Module (SRM) to address this issue within the segmentation network. Specifically, SRM is designed to learn a transformation offset for each pixel in the upsampled feature maps, guided by high-resolution feature maps and neighboring offsets. By applying these offsets to the upsampled feature maps, SRM enhances the semantic representation of the segmentation network, particularly for pixels around object boundaries. Furthermore, a Contextual Refinement Module (CRM) is presented to capture global context information across both spatial and channel dimensions. To balance dimensions between channel and space, we aggregate the semantic maps from all four stages of the backbone to enrich channel context information. The efficacy of these proposed modules is validated on three widely used datasets—Cityscapes, Bdd100 K, and ADE20K—demonstrating superior performance compared to state-of-the-art methods. Additionally, this paper extends these modules to a lightweight segmentation network, achieving an mIoU of 82.5% on the Cityscapes validation set with only 137.9 GFLOPs. Deyin Liu, Lin Wu 0001, Song Wang 0008, Xin Guo 0005, Lin Qi 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Few-shot defect classification via feature aggregation based on graph neural networkabstractThe effectiveness of deep learning models is greatly dependent on the availability of a vast amount of labeled data. However, in the realm of surface defect classification, acquiring and annotating defect samples proves to be quite challenging. Consequently, accurately predicting defect types with only a limited number of labeled samples has emerged as a prominent research focus in recent years. Few-shot learning, which leverages a restricted sample set in the support set, can effectively predict the categories of unlabeled samples in the query set. This approach is particularly well-suited for defect classification scenarios. In this article, we propose a transductive few-shot surface defect classification method, which using both the instance-level relations and distribution-level relations in each few-shot learning task. Furthermore, we calculate class center features in transductive manner and incorporate them into the feature aggregation operation to rectify the positioning of edge samples in the mapping space. This adjustment aims to minimize the distance between samples of the same category, thereby mitigating the influence of unlabeled samples at category boundary on classification accuracy . Experimental results on the public dataset show the outstanding performance of our proposed approach compared to the state-of-the-art methods in the few-shot learning settings. Our code is available at https://github.com/Harry10459/CIDnet . Peixiao Zheng, Xin Guo 0005, Enqing Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Edge-labeling based modified gated graph network for few-shot learning
Peixiao Zheng, Xin Guo 0005, Enqing Chen, Lin Qi 0001, Ling Guan |
Pattern Recognit. | 2 |
| 2024 | Parametric channel estimation for RIS-assisted mmWave MIMO-OFDM systems with low pilot overhead
Shuangzhi Li 0001, Xin Guo 0005, Gangtao Han, Jian-Kang Zhang 0001 |
Signal Process. | 3 |
| 2023 | A Feature Refinement Module for Light-Weight Semantic Segmentation NetworkabstractLow computational complexity and high segmentation accuracy are both essential to the real-world semantic segmentation tasks. However, to speed up the model inference, most existing approaches tend to design light-weight networks with a very limited number of parameters, leading to a considerable degradation in accuracy due to the decrease of the representation ability of the networks. To solve the problem, this paper proposes a novel semantic segmentation method to improve the capacity of obtaining semantic information for the light-weight network. Specifically, a feature refinement module (FRM) is proposed to extract semantics from multi-stage feature maps generated by the backbone and capture non-local contextual information by utilizing a transformer block. On Cityscapes and Bdd100K datasets, the experimental results demonstrate that the proposed method achieves a promising trade-off between accuracy and computational cost, especially for Cityscapes test set where 80.4% mIoU is achieved and only 214.82 GFLOPs are required. Xin Guo 0005, Song Wang 0008, Peixiao Zheng, Lin Qi 0001 |
ICIP | 2 |
| 2023 | Multi-channel and multi-scale separable dilated convolutional neural network with attention mechanism for flue-cured tobacco classification
Zhong Zhang 0012, Xin Guo 0005 |
Neural Comput. Appl. | 4 |
| 2021 | Local Feature Descriptors with Deep Hypersphere LearningabstractRecent works have demonstrated the power of L2normalization in local feature descriptor learning. While the descriptors are typically learned in the Euclidean space, the similarity between descriptors is often evaluated on a unit hypersphere due to the post-processing of L2normalization for descriptors, which creates a gap between the training stage and the usage stage of feature descriptors. To bridge the gap, we propose a hyperspherical descriptor learning model, where the whole network is projected onto the hyperspherical space. In addition, a squared angular triplet loss is designed to enable the proposed hyperspherical model to learn angularly discriminative descriptors. Experiments on UBC dataset show that the proposed hyperspherical descriptor outperforms its Euclidean counterparts and the state-of-the-art methods on the feature matching task. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 2 |
| 2021 | Discriminative Patch Descriptor Learning With Focal Triplet Loss FunctionabstractThis paper proposes a focal triplet loss function for discriminative patch descriptor learning. The standard triplet loss function usually restrains the distance difference between the matching samples and the non-matching ones. However, along with the training procedure, the majority of triplets in each batch tend to satisfy the constraint of the loss function and produce low loss values, leading to a masquerade that the model is well-trained. To address this problem, the focal triplet loss function is proposed in this paper to weaken the impact of the easy triplets and focus training on the hard ones. By emphasizing the importance of hard triplets on the model training, the proposed loss forces the descriptor vectors with fixed dimension to carry more discriminative information from the patches. With the benefits of the focal mechanism, the proposed method achieves better performance compared to the state-of-the-art on UBC dataset for image matching task. Furthermore, to demonstrate the effectiveness of the proposed method, we extend the focal triplet loss on the cross-model retrieval task. The experimental results indicate that the proposed method can also be used to improve visual-semantic embedding learning. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 2 |
| 2021 | Edge-Labeling Based Directed Gated Graph Network for Few-Shot LearningabstractExisting graph-network-based few-shot learning methods obtain similarity between nodes through a convolution neural network (CNN). However, the CNN is designed for image data with spatial information rather than vector form node feature. In this paper, we proposed an edge-labeling-based directed gated graph network (DGGN) for few-shot learning, which utilizes gated recurrent units to implicitly update the similarity between nodes. DGGN is composed of a gated node aggregation module and an improved gated recurrent unit (GRU) based edge update module. Specifically, the node update module adopts a gate mechanism using activation of edge feature, making a learnable node aggregation process. Besides, improved GRU cells are employed in the edge update procedure to compute the similarity between nodes. Further, this mechanism is beneficial to gradient backpropagation through the GRU sequence across layers. Experiment results conducted on two benchmark datasets show that our DGGN achieves a comparable performance to the-state-of-art methods. Peixiao Zheng, Xin Guo 0005, Lin Qi 0001 |
ICIP | 2 |
| 2021 | Constellation Design for Noncoherent Massive SIMO Systems in URLLC ApplicationsabstractIn this paper, we concern the uplink of a massive single-input multiple-output enabled ultra-reliable low-latency communication system, in which a single-antenna transmitter aims to timely and reliably send data to a receiver equipped with a large number of antennas over Rayleigh fading channels. For such a scenario, to eliminate the considerable overhead caused by channel estimation, we adopt a noncoherent maximum-likelihood (ML) receiver, which is known to be optimal in terms of average symbol-error rate for equiprobable discrete input signals. We propose a two-dimensional noncoherent constellation design framework to enhance the reliability of the considered system. Specifically, our design principle is to maximize the minimum Kullback-Leibler divergence between the conditional distributions induced by different transmitted signals under average power constraint for any given transmission rate. The resulting optimization problem is shown to be a challenging mixed discrete-continuous problem. We manage to solve the problem by deliberately designing optimal bit allocation and optimal constellation structure as a function of signal-to-noise ratios. We then unveil that the proposed constellation can facilitate efficient ML detection with low computational complexity. Finally, simulation results illustrate that the proposed scheme has a superior error performance than conventional training-based schemes and existing energy detection designs. Shuangzhi Li 0001, Dong Zheng 0003, He Henry Chen, Xin Guo 0005 |
IEEE Trans. Commun. | 4 |
| 2020 | Negative Label Guided Discriminative Canonical Correlation Analysis for Semi-Supervised and Semi-Paired LearningabstractSemi-supervised learning is a popular trend for learning based methods in recent years, as it fully exploits both the labeled and unlabeled samples in a dataset. This paper sets itself apart from most existing semi-supervised learning algorithms, which only use the exact labels of data already known. We take the negative label as side information to guide the process of semi-supervised learning. Two types of supervision information are regarded as negative label; the first type indicates that a sample definitely does not belong to a specific category, and the second indicates that two samples come from different views, and cannot have a one to one correspondence. By reasonably assuming that nearby points should have similar class indicators, the data labels are propagated under the negative label and the geometric structure revealed by both labeled and unlabeled points. Specifically, we predict one to one pair information by utilizing the neighbor information of samples, under the guidance of the negative pair label. Extensive experiments on several datasets demonstrate the effectiveness of our proposed method. Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 1 |
| 2020 | Weighted hybrid fusion with rank consistency
Song Wang 0008, Xin Guo 0005, Tie Yun, Ivan Lee 0001, Lin Qi 0001, Ling Guan |
Pattern Recognit. Lett. | 2 |
| 2020 | Deep Local Feature Descriptor Learning With Dual Hard Batch ConstructionabstractLocal feature descriptor learning aims to represent distinctive images or patches with the same local features, where their representation is invariant under different types of deformation. Recent studies have demonstrated that descriptor learning based on Convolutional Neural Network (CNN) is able to improve the matching performance significantly. However, they tend to ignore the importance of sample selection during the training process, leading to unstable quality of descriptors and learning efficiency. In this paper, a dual hard batch construction method is proposed to sample the hard matching and non-matching examples for training, improving the performance of the descriptor learning on different tasks. To construct the dual hard training batches, the matching examples with the minimum similarity are selected as the hard positive pairs. For each positive pair, the most similar non-matching example is then sampled from the generated hard positive pairs in the same batch as the corresponding negative. By sampling the hard positive pairs and the corresponding hard negatives, the hard batches are produced to force the CNN model to learn the descriptors with more efforts. In addition, based on the above dual hard batch construction, an ℓ22 triplet loss function is built for optimizing the training model. Specifically, we analyze the superiority of the ℓ22 loss function when dealing with hard examples, and also demonstrate it in the experiments. With the benefits of the proposed sampling strategy and the ℓ22 triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmarks for different matching tasks. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
IEEE Trans. Image Process. | 2 |
| 2019 | A Novel Weighted Hybrid Multi-View Fusion Algorithm for Semi-Supervised ClassificationabstractSemi-supervised learning aims to improve the learning performance with very limited label information. To dig more available information from the collected data, we propose a weighted hybrid multi-view feature fusion approach for semi-supervised classification problem. Specifically, under the rank consistency constraint for labels predicted by view-specific learners, the proposed method estimates the optimal fusion weight for each learner to balance the incomparable square losses on different views. In this case, the learners with more powerful prediction capability are pushed to have higher weights during the fusion process. Experimental results on 6 real-world datasets demonstrate the effectiveness of the proposed technique. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 2 |
| 2019 | Local Feature Descriptor Learning with a Dual Hard Sampling StrategyabstractLocal feature descriptor learning based on Convolutional Neural Network (CNN) has demonstrated its capability to generate descriptors with high quality. While extensive studies focused on mining hard non-matching examples to improve descriptor learning performance, a random sampling strategy is adopted for matching examples. In this paper, a dual hard sampling strategy based on the triplet loss function is proposed to generate the hard matching and non-matching examples for training. To start with, a pair of matching examples with the maximum distance for each class are selected as the positive pair. For each positive pair, their closest non-matching example is then sampled from the generated positive pairs with other classes as the corresponding negative. Based on the above dual hard sampling strategy, a novel triplet loss function is presented for optimization. With the benefits of the proposed sampling strategy and the novel triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmark for local feature matching. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 2 |
| 2018 | Identifying facial expression using adaptive sub-layer compensation based feature extraction
Xin Guo 0005, Tie Yun, Long Ye, Jinyao Yan |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Incremental generalized multiple maximum scatter difference with applications to feature extraction
Ning Zheng 0003, Xin Guo 0005, Tie Yun, Nan Dong, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Joint intermodal and intramodal correlation preservation for semi-paired learning
Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
Pattern Recognit. | 1 |
| 2016 | Semi-Supervised and Semi-Paired Graph Regularized Multiset Canonical Correlation AnalysisabstractMultiset canonical correlation analysis (MCCA) plays a key role in analyzing linear correlations among multimodal data. However, when facing semi-supervised and semi-paired multimodal data which widely exist in real world. MCCA normally performs poorly because it requires paired information among different models. At the same time, it fails to exploit the discriminative information due to the fact that it is an unsupervised dimension reduction method. In this paper, we propose a novel algorithm, named semi-supervised semi-paired graph regularized multiset canonical correlation analysis(SSGMCCA). SSGMCCA employs a small amount of paired data to perform MCCA and simultaneously utilize both the global structural information captured from the unlabeled data and the local structural information captured from the labeled data to compensate the limited paired. As a result, SSGMCCA can find the directions which not only maximal correlation among the multimodal data but also maximal separability of the labeled data. Experimental results illustrate the effectiveness of the proposed algorithm. Xin Guo 0005, Lin Qi 0001, Ling Guan |
ISM | 1 |
| 2015 | Two-dimensional discriminant multi-manifolds locality preserving projection for facial expression recognitionabstractIn this paper, we assume that samples of different expressions reside on different manifolds and propose a novel human emotion recognition framework named two-dimensional discriminant multi-manifolds locality preserving projection (2D-DMLPP). 2D-DMLPP focuses on salient regions which reflect the significant variation from facial expression images so that it can learn an expression-specific model from salient patches rather than that of subject-specific. Furthermore, conventional manifold learning methods ignore the variation among nearby samples from the same class, leading to serious overfitting. We construct three adjacency graphs to model the margin and information, including diversity and similarity of salient patches from the same expression, and then incorporate the information and margin into dimensionality reduction function. Several experiments show that the proposed method significantly improves the recognition performance of facial expression recognition. Ning Zheng 0003, Xin Guo 0005, Lin Qi 0001, Ling Guan |
ISCAS | 2 |
| 2015 | A Novel Semi-Supervised Dimensionality Reduction Framework for Multi-manifold LearningabstractIn pattern recognition, traditional single manifold assumption can hardly guarantee the best classification performance, since the data from multiple classes does not lie on a single manifold. When the dataset contains multiple classes and the structure of the classes are different, it is more reasonable to assume each class lies on a particular manifold. In this paper, we propose a novel framework of semi-supervised dimensionality reduction for multi-manifold learning. Within this framework, methods are derived to learn multiple manifold corresponding to multiple classes in a data set, including both the labeled and unlabeled examples. In order to connect each unlabeled point to the other points from the same manifold, a similarity graph construction, based on sparse manifold clustering, is introduced when constructing the neighbourhood graph. Experimental results verify the advantages and effectiveness of this new framework. Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 1 |
| 2015 | Advanced weight graph transformation matching algorithmabstractAn efficient and accurate point matching algorithm named advanced weight graph transformation matching (AWGTM) is proposed in this study. Instead of relying only on the elimination of dubious matches, the method iteratively reserve correspondences which have a small angular distance between two nearest‐neighbour graphs. The proposed algorithm is compared against weight graph transformation matching (WGTM) and graph transformation matching (GTM). Experimental results demonstrate the superior performance in eliminating outliers and reserving inliers of AWGTM algorithm under various conditions for images, such as duplication of patterns and non‐rigid deformation of objects. An execution time comparison is also presented, where AWGTM shows the best results for high outlier rates. Song Wang 0008, Xin Guo 0005, Xiaomin Mu, Yahong Huo, Lin Qi 0001 |
IET Comput. Vis. | 2 |