EDBT 2026 Demo / reviewers in the wild / expert
Xuemei Xie
dblp:06/5645
· DBLP profile ↗
86ranked-venue papers
4as first author
50since 2021 · last 2027
0000-0001-7857-0845ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 3 first-author · 26 since 2021Artificial intelligence and machine learning · 33 · 1 first-author · 22 since 2021Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | OWseg-WE: An effective method for open-world semantic segmentation of ginkgo biloba pests and diseases
Yuebo Zhou, Junfeng Ye, Xuemei Xie, Yurong Sun |
Expert Syst. Appl. | 3 |
| 2026 | Adaptive Semantic Compression and Transmission for Cognitive Knowledge Coordination in a Hierarchical LLM-Agents SystemabstractThe rapid development of Large Language Model (LLM) agents has facilitated the advancement of multi-agent systems, where cognitive knowledge sharing is crucial for the execution of complex tasks. However, achieving the synchronization of the cognitive Knowledge base (KB) among agents under restricted wireless resources remains a challenge, especially in dynamic real-time environments. Therefore, we propose a hierarchical LLM agent system that consists of a high-level Cluster Brain (CB) and multiple Lower-level LLM Agents (LLAs). The cognitive KB of each LLA is represented in the form of a Knowledge Graph (KG). To improve the efficiency of transmitting cognitive KB updates from LLAs to CB, a KG compression framework named MED-EmPress is proposed, which adaptively compresses the semantic features of cognitive KB by applying dimensionality reduction and binary quantization, and then a joint optimization problem of Semantic Compression and Resource Allocation (SCRA) is formulated to maximize semantic fidelity of the cognitive KB being transmitted. To solve this problem, a hierarchical SCRA algorithm is designed to decouple the SCRA problem into two subproblems, which involve dynamically allocating wireless resources and rationally choosing the semantic compression ratio of cognitive KB. The goal is to maximize the system’s semantic fidelity. The evaluation results demonstrate that the MED-EmPress framework reduces the size of the cognitive updates that need to be transmitted by 96%, with only a loss of 3.6% in the entity alignment task. Furthermore, the proposed adaptive compression and transmission scheme improves semantic fidelity by 94% compared to existing methods when wireless resources are severely limited. Xinju He, Jiayi Liu 0001, Xuemei Xie, Guangming Shi |
IEEE Internet Things J. | 4 |
| 2026 | Multi-layer graph constraint dictionary pair learning for image classification
Guangming Shi, Weisheng Dong, Xuemei Xie |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | Parse graph-based visual-language interaction for human pose estimation
Shibang Liu, Xuemei Xie, Guangming Shi |
Pattern Recognit. | 2 |
| 2026 | SceneReasoner: Sufficient Embodied Scene Understanding From Limited Perception by Explicit Functional AssociationabstractSufficient embodied scene understanding serves as the foundation for embodied agents to perceive, interpret, and solve scene-related questions in a scene. Such understanding is often constrained by limited perception, which can be summarized into two aspects: (1) The lack of perceptual abilities in the agent. (2) The scene itself is incomplete. Existing large models-based methods attempt to leverage implicit knowledge to overcome these limitations, but this lacks interpretability and controllability. Inspired by human explainable associative thinking, we propose SceneReasoner framework, which imposes explicit functional associative rules on LLMs to guide the process of the scene understanding. This framework mines deeper functional relationships between objects, enabling the agent to gain sufficient scene understanding from limited perception in a controllable manner. Specifically, SceneReasoner employs an associative knowledge base to provide such rules from two aspects. (1) Functional complementarity of objects in a scene. For instance, in a computer workspace, if the agent perceives a monitor and other unclear objects, it can first analyze the function of the monitor in this area (content display), and then infer the presence of other related objects (e.g., a mouse and keyboard for content input). (2) Commonality of objects in a scene. For example, in a picture-hanging task, when the hammer in the scene is missing, the agent needs to first identify the hammer’s attributes (hard and applying force), and then associate a suitable substitute in the scene (e.g., a hard wrench). Due to such explicit functional association, the agent can rapidly form a sufficient scene understanding and effectively solve scene-related questions. Experimental results demonstrate that such explicit association augmented with functional reasoning can significantly enhance agents’ scene understanding under limited perception. It improves perceptual quality by 9.75% and scene reasoning ability by 21.42% compared with other methods. Xiukun Liu, Ning Lan, Xuemei Xie, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Hierarchical Scene Graph Generation With Coarse-to-Fine ReasoningabstractScene graph generation denotes the process of parsing a visual input into a graph representation for downstream reasoning tasks. Most existing research models input scenes as flat scene graphs, capturing solely on horizontal relationships between two objects, while disregarding vertical dependencies. Such flat scene graphs merely describe direct relationships between objects, and cannot represent higher-level semantics. The lack of hierarchical relationships also limits the performance of scene graphs in downstream reasoning tasks. In this work, we propose hierarchical scene graph generation, a new problem task that requires the model to generate scene graphs with clear hierarchy. To achieve this goal, a hierarchical scene graph generation with coarse-to-fine reasoning framework is proposed. It contains three stages: first, summarize the main caption of the scene; second, describe the main visual elements related to the theme; and finally, add secondary elements with background. To facilitate training, a high-quality hierarchical scene graph dataset that comprises 40 K image-scene description pairs is constructed. Utilize rich knowledge and powerful multi-turn conversation capabilities of multi-modal large language models, the proposed framework can not only achieve better semantic understanding and relational modeling, but also seamlessly integrating its relational modeling capabilities to enhance visual-language tasks. Extensive experimental results demonstrate that our method constructs an effective scene representation, outperforming the state-of-the-art by 3.35% in recall while generating significantly fewer average triplets 8.7vs. 9.2. The results on downstream tasks also indicate that hierarchical scene graph significantly contributes to visual-language interactive reasoning. Zefang Han, Xiukun Liu, Xuemei Xie, Jianxiu Yang, Guangming Shi |
IEEE Trans. Multim. | 3 |
| 2026 | CMANet: Context-Aware Mutual Attention Network for Referring Image Segmentation
Xiong Pan, Xuemei Xie, Jianxiu Yang, Xiaodan Song, Guangming Shi |
IEEE Trans. Multim. | 2 |
| 2025 | Affine Transformation-Based Generative Face Video CompressionabstractIn this paper, we propose a generative face video compression framework based on affine transformations to better represent large movements without parameter transmission. It mainly consists of an encoder and decoder, and our encoder is similar to the one in [1]. Intra frame are compressed by the existing encoder, while subsequent inter frames are compressed into compact inter frame features. In the decoder, feature alignment is first established to map the decoded intra frame and inter frame features into the same domain. The aligned features are then combined with the appearance features extracted by the appearance encoder from the intra frame and fed into the coarse-fine affine transform module to establish motion estimation and compensation. The coarse affine transform focuses on global motion, while the fine affine transform deals with local motion, such as lip motion. Finally, the transformed features are fed into the image generation module to obtain the final reconstruction results. Xihua Lin, Xiaodan Song, Xuguang Zuo, Dahua Gao, Xuemei Xie, Guangming Shi |
DCC | 6 |
| 2025 | A Transmitter-Model Unaware Generative Image Compression Framework for Semantic CommunicationabstractUnlike traditional bit-level data transmission methods, semantic communication focuses on conveying the meaning behind the data. Though promising results have been achieved, existing end-to-end learning-based semantic communication frameworks often require a synchronization of deep models between the transmitter and the receiver. Such design leads to tens of thousands models to be stored at receiver since different manufactures may optimize their own models. To address this problem, we propose a novel model-unaware generative image compression framework for semantic communication. It features at employing human-understandable multi-modality representations as an intermediate layer to enhance information transmission efficiency and semantic consistency. Our framework introduces a mask-based rate-distortion optimization module, which effectively removes low-relevance information for image generation and reduces the bit rate while maintaining semantic consistency. Experimental results demonstrate that the framework can still reconstruct high-quality images at very low bit rates, showcasing its potential for applications in modern communication systems. Rongcan Zheng, Xiaodan Song, Xuguang Zuo, Minxi Yang, Dahua Gao, Xuemei Xie |
ICASSP | 6 |
| 2025 | A Fine-Grained Pose-Aware Pattern Discovering ConvNet for Human Pose Estimation
Xuemei Xie |
ICIC (11) | 2 |
| 2025 | Visual Relationships Are Different: Appropriate Way To Predict Each RelationshipabstractVisual relationships are different, and their types play an important role in visual scene understanding. However, most of the existing works ignore the different types of visual relationships and adopt the unified approach to learn all visual relationships. It not only limits the flexibility of the model, but also causes fuzzy relationship representation. To address this problem, we deeply study four visual relationship types. Among them, the geometric reflects the spatial interaction, the semantic reflects an action, and the possessive and misc indicates an intrinsic correlation. And then, we propose a novel method- Types Determine Methods (TDM) - which designs different learning strategies according to the relationship types to infer visual relationships. Experiments demonstrate that our approach achieves superior or competitive performance over previous methods, validating its effectiveness. Zhenhua Lei, Xuemei Xie |
ICME | 2 |
| 2025 | Multi-Scale Adaptive Skeleton Transformer for action recognition
Xiaotian Wang 0001, Zhifu Zhao, Guangming Shi, Xuemei Xie, Xiang Jiang 0011 |
Comput. Vis. Image Underst. | 5 |
| 2025 | Compressing Vision Transformer from the View of Model Property in Frequency Domain
Zhenyu Wang 0008, Xuemei Xie, Hao Luo 0004, Weisheng Dong, Yongxu Liu 0001, Fan Wang 0019, Guangming Shi |
Int. J. Comput. Vis. | 2 |
| 2025 | Hierarchical language description knowledge base for LLM-based human pose estimation
Xuemei Xie |
Neurocomputing | 2 |
| 2025 | Language-guided temporal primitive modeling for skeleton-based action recognition
Qingzhe Pan, Xuemei Xie |
Neurocomputing | 2 |
| 2025 | Mixed-scale cross-modal fusion network for referring image segmentationabstractReferring image segmentation aims to segment the target by a given language expression. Recently, the bottom-up fusion network utilizes language features to highlight the most relevant regions during the visual encoder stage. However, it is not comprehensive that establish only the relationship between pixels and words. To alleviate this problem, we propose a mixed-scale cross-modal fusion method that widens the interaction between vision and language. Specially, at each stage, pyramid pooling is used to augment visual perception and improve the interaction between visual and linguistic features, thereby highlighting relevant regions in the visual data. Additionally, we employ a simple multi-scale feature fusion module to effectively combine multi-scale aligned features. Experiments conducted on Standard RIS benchmarks demonstrate that the proposed method achieves favorable performance against state-of-the- art approaches. Moreover, we conducted experiments on different visual backbones respectively, and the proposed method yielded better and significantly improved performance results. Xiong Pan, Xuemei Xie, Jianxiu Yang |
Neurocomputing | 2 |
| 2025 | Recognizing human-object interactions in videos with the supervision of natural language
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
Neural Networks | 2 |
| 2025 | Visual Environment-Interactive Planning for Embodied Complex-Question AnsweringabstractThis study focuses on Embodied Complex-Question Answering task, which means the embodied robot need to understand human questions with intricate structures and abstract semantics. The core of this task lies in making appropriate plans based on the perception of the visual environment. Existing methods often generate plans in a once-for-all manner,$i.e$., one-step planning. Such approach rely on large models, without sufficient understanding of the environment. Considering multi-step planning, the framework for formulating plans in a sequential manner is proposed in this paper. To ensure the ability of our framework to tackle complex questions, we create a structured semantic space, where hierarchical visual perception and chain expression of the question essence can achieve iterative interaction. This space makes sequential task planning possible. Within the framework, we first parse human natural language based on a visual hierarchical scene graph, which can clarify the intention of the question. Then, we incorporate external rules to make a plan for current step, weakening the reliance on large models. Every plan is generated based on feedback from visual perception, with multiple rounds of interaction until an answer is obtained. This approach enables continuous feedback and adjustment, allowing the robot to optimize its action strategy. To test our framework, we contribute a new dataset with more complex questions. Experimental results demonstrate that our approach performs excellently and stably on complex tasks. And also, the feasibility of our approach in real-world scenarios has been established, indicating its practical applicability. Ning Lan, Baoshan Ou, Xuemei Xie, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Local Uncertainty Energy Transfer for Active Domain AdaptationabstractActive Domain Adaptation (ADA) improves knowledge transfer efficiency from the labeled source domain to the unlabeled target domain by selecting a few target sample labels. However, most existing active sampling methods ignore the local uncertainty of neighbors in the target domain,making it easier to pick out anomalous samples that are detrimental to the model. To address this problem, we present a new approach to active domain adaptation called Local Uncertainty Energy Transfer (LUET), which integrates active learning of local uncertainty confusion and energy transfer alignment constraints into a unified framework. First, in the active learning module, the uncertainty difficult and representative samples from the target domain are selected through local uncertainty energy selection and entropy-weighted class confusion selection. And the active learning strategy based on local uncertainty energy will avoid selecting anomalous samples in the target domain. Second, for the discrimination issue caused by domain shift, we use a global and local energy-transfer alignment constraint module to eliminate the domain gap and improve accuracy. Finally, we used negative log-likelihood loss for supervised learning of source domains and query samples. With the introduction of sample-based energy metrics, the active learning strategy is more closely with the domain alignment. Experiments on multiple domain-adaptive datasets have demonstrated that our LUET can achieve outstanding results and outperform existing state-of-the-art approaches. Guangming Shi, Weisheng Dong, Xin Li 0005, Xuemei Xie |
IEEE Trans. Image Process. | 6 |
| 2025 | Learning Pyramid-Structured Long-Range Dependencies for 3D Human Pose EstimationabstractAction coordination in human structure is indispensable for the spatial constraints of 2D joints to recover 3D pose. Usually, action coordination is represented as a long-range dependence among body parts. However, there are two main challenges in modeling long-range dependencies. First, joints should not only be constrained by other individual joints but also be modulated by the body parts. Second, existing methods make networks deeper to learn dependencies between non-linked parts. They introduce uncorrelated noise and increase the model size. In this paper, we utilize a pyramid structure to better learn potential long-range dependencies. It can capture the correlation across joints and groups, which complements the context of the human sub-structure. In an effective cross-scale way, it captures the pyramid-structured long-range dependence. Specifically, we propose a novel Pyramid Graph Attention (PGA) module to capture long-range cross-scale dependencies. It concatenates information from various scales into a compact sequence, and then computes the correlation between scales in parallel. Combining PGA with graph convolution modules, we develop a Pyramid Graph Transformer (PGFormer) for 3D human pose estimation, which is a lightweight multi-scale transformer architecture. It encapsulates human sub-structures into self-attention by pooling. Extensive experiments show that our approach achieves lower error and smaller model size than state-of-the-art methods on Human3.6 M and MPI-INF-3DHP datasets. Xuemei Xie, Yutong Zhong, Guangming Shi |
IEEE Trans. Multim. | 2 |
| 2025 | Reconfiguring Satellite CDNs With Dynamic Uncertain User Requests Based on Multi-Agent DRLabstractThe Satellite-Terrestrial Integrated Network (STIN) is a key paradigm to achieve global coverage and ubiquitous connection in the 6G era. Integrating the network slice technology based on Software Defined Networking (SDN) and Network Function Virtualization (NFV) into STIN is recognized as an effective solution to achieve a rapid flexible service provisioning. Specifically, the Content Delivery Network (CDN) service, which is storage resource intensive, is suitable to be deployed on STIN to provide a global range content service suppply. Most existing research on the resource deployment of STIN focuses on static requests, ignoring the dynamic changes in requests caused by the high-speed movement of satellites. Especially, the reconfiguration of CDN slices and the variation of STIN are asynchronous in different time scales: this makes the reconfiguration of CDN slices to cope with faster changing user requests a challenging task. In this paper, we adopt a periodic reconfiguration strategy to configure CDN slices on the edge LEO satellite network of STIN in discrete time intervals. Within each time interval, we formulate the reconfiguration optimization problem to cope with the dynamically changing user requests, wherein the Stochastic Network Calculus (SNC) is used to measure the deployment performance within this time interval. Then, we describe the optimization problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP), and propose a reconfiguration algorithm based on Multi-agent Deep Reinforcement Learning (MADRL) to determine the optimal adjustment strategy. Finally, intensive simulations are implemented to verify the performance of the algorithm. Compared with the baselines, the QoS of the proposed algorithm is increased by about 5.8%, while the operation cost and reconfiguration cost are decreased by about 8.5% and 35.3% respectively. Jiayi Liu 0001, Xuemei Xie, Guangming Shi |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2025 | Prompt-Based Invertible Mapping Alignment for Unsupervised Domain AdaptationabstractLarge pre-trained vision-language models (VLMs) like CLIP have shown great potential for solving the unsupervised domain adaptation (UDA) problem. Existing prompt learning for UDA based on the unsupervised-trained VLMs requires distribution alignment between source and target domains in the common space for both vision and language branches. However, it is difficult for rough cross-domain alignment to maintain the discriminative semantic structure of both domains. Besides, the coarse features with non-informative noises due to ignoring the pseudo-label noises may cause failures to concentrate on precise semantics alignment. In this work, we propose a Prompt-Based Invertible Mapping Alignment (PIMA) method to incorporate discriminative domain knowledge into prompt learning, which is featured with refined cross-domain alignment in two separate space with a well-kept structure. Specifically, we design an invertible neural network-based homeomorphism mapping, and then achieve distribution alignment through such invertible mapping for connecting source and target visual feature space, which can preserve the data semantic structure. For better semantic alignment in vision-language space, we develop cross-modal implicit contrastive learning module to regularize non-informative features, which aims to find the low-rankness of implicit representation space. We conducted extensive experiments on three benchmark datasets to prove the advantages of our proposed PIMA over state-of-the-art methods. Xiaodan Song, Xuemei Xie |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | BVT-IMA: Binary Vision Transformer with Information-Modified AttentionabstractAs a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized models suffer from serious performance drops. In this paper, an attention shifting is observed in the binary multi-head self-attention module, which can influence the information fusion between tokens and thus hurts the model performance. From the perspective of information theory, we find a correlation between attention scores and the information quantity, further indicating that a reason for such a phenomenon may be the loss of the information quantity induced by constant moduli of binarized tokens. Finally, we reveal the information quantity hidden in the attention maps of binary vision transformers and propose a simple approach to modify the attention values with look-up information tables so that improve the model performance. Extensive experiments on CIFAR-100/TinyImageNet/ImageNet-1k demonstrate the effectiveness of the proposed information-modified attention on binary vision transformers. Zhenyu Wang 0008, Hao Luo 0004, Xuemei Xie, Fan Wang 0019, Guangming Shi |
AAAI | 3 |
| 2024 | GRA-Net: Group response attention for deep learning
Zhenyuan Wang, Xuemei Xie, Xiaodan Song, Jianxiu Yang |
Neurocomputing | 2 |
| 2024 | STDM-transformer: Space-time dual multi-scale transformer network for skeleton-based action recognition
Zhifu Zhao, Jianan Li 0003, Xuemei Xie, Xiaotian Wang 0001, Guangming Shi |
Neurocomputing | 4 |
| 2024 | Human Pose Estimation via Parse Graph of Body StructureabstractWhen observing a person’s body, humans can extract the structured representation of the body called a parse graph, which includes the hierarchical decompositions from the entire body to parts and primitives and the context relations by horizontal links between the body parts. This ability helps humans better locate body structures at different levels. In order for the model to have this ability for single-person pose estimation, we design a hierarchical network to model the context relations and hierarchical structure in the parse graph of body structure by convolutional neural networks. It overcomes the problem that most methods ignore one of the context relations and hierarchical structure in the parse graph. Our network contains bottom-up and top-down stages. In the bottom-up stage, the structural features of the hierarchy are captured from primitives to parts and the entire body. Then in the top-down stage, with the context information of each body part, the structural features of the body parts are refined separately rather than together from the entire body to parts and primitives. Experiments show that our model enhances the reasonableness of predictions and achieves superior results on the CrowdPose, COCO keypoint detection and MPII human pose datasets. Shibang Liu, Xuemei Xie, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Glimpse and Zoom: Spatio-Temporal Focused Dynamic Network for Skeleton-Based Action RecognitionabstractGCN-based methods have achieved remarkable performance in skeleton-based action recognition. However, existing methods have not explicitly attempted to remove temporal and spatial redundancy that might introduce additional computational costs. Inspired by the fact that humans always tend to glimpse at overall motion and then zoom into the most important spatio-temporal regions, we propose a Spatio Temporal Focused Dynamic Network (STFD-Net) trained with reinforcement learning for skeleton-based action recognition. Specifically, we first propose a global extractor with Skeleton Pooling Module (SPM) to enable the network to focus on overall motion information with a refined skeleton structure. Then, a local extractor, containing pair-wise part partition, tubelet proposal network, and Partition-Grouped Module (PGM), is proposed to extract local motion details as a complement to the overall motion information. Finally, the dynamic classifier utilizes a recurrent neural network to dynamically terminate the process once the network is adequately confident. Extensive experiments have demonstrated that the proposed network achieves SOTA level performance with lower computational cost on the NTU 60 and NTU 120 dataset. Zhifu Zhao, Jianan Li 0003, Xiaotian Wang 0001, Xuemei Xie, Wanxin Zhang, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | TransVQA: Transferable Vector Quantization Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the unlabeled target domain. Most existing domain adaptation methods are based on convolutional neural networks (CNNs) to learn cross-domain invariant features. Inspired by the success of transformer architectures and their superiority to CNNs, we propose to combine the transformer with UDA to improve their generalization properties. In this paper, we present a novel model named Trans ferable V ector Q uantization A lignment for Unsupervised Domain Adaptation (TransVQA), which integrates the Transferable transformer-based feature extractor (Trans), vector quantization domain alignment (VQA), and mutual information weighted maximization confusion matrix (MIMC) of intra-class discrimination into a unified domain adaptation framework. First, TransVQA uses the transformer to extract more accurate features in different domains for classification. Second, TransVQA, based on the vector quantization alignment module, uses a two-step alignment method to align the extracted cross-domain features and solve the domain shift problem. The two-step alignment includes global alignment via vector quantization and intra-class local alignment via pseudo-labels. Third, for intra-class feature discrimination problem caused by the fuzzy alignment of different domains, we use the MIMC module to constrain the target domain output and increase the accuracy of pseudo-labels. The experiments on several datasets of domain adaptation show that TransVQA can achieve excellent performance and outperform existing state-of-the-art methods. Weisheng Dong, Xin Li 0005, Guangming Shi, Xuemei Xie |
IEEE Trans. Image Process. | 6 |
| 2024 | Mobility-Aware MEC Planning With a GNN-Based Graph Partitioning FrameworkabstractMobile service continuity is essential important to ensure that user sessions and services will survive user mobility. The 5G enhances its mobility management by providing the flexibility and offering three types of Session and Service Continuity (SSC) modes to address various service continuity requirements. Multi-access edge computing (MEC) is a type of widely adopted network architecture that delivers network services from the boundary of the mobile network by provisioning a set of edge servers. Determining an optimum planning of MEC edge servers, which involves determining edge servers appropriate geographical positions and their serving areas, is a precondition for more efficient service provisioning and better usage of network resources. In this work, we investigate the MEC servers planning problem by considering the management cost for maintaining MEC service continuity. The problem is formulated as a graph partitioning problem to partition the RAN graph with minimum SSC management costs and balanced MEC servers workloads. Then, we adapt a generalizable approximate Graph Partitioning framework which leverages on Graph Neural Network (GNN) to embed the RAN network spacial feature and on Multilayer Perceptron (MLP) for graph partitioning. Based on the framework, we propose a MEC server planning algorithm named MECP-GAP. Finally, we evaluate MECP-GAP with extensive simulations and real network data. Comparing to several baselines, MECP-GAP achieves better performance with lower running time. Jiayi Liu 0001, Zhongyi Xu, Xuefang Liu, Xuemei Xie, Guangming Shi |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2024 | Graph Transformer and LSTM Attention for VNF Multi-Step Workload Prediction in SFCabstractThe knowledge on a Service Function Chain’s (SFC’s) resource requirements is an indispensable prerequisite for proactive resource provisioning and run-time management of the SFC. However, due to the intrinsic dynamics in network environment, accurate resource requirements and workloads prediction for the Virtual Network Functions (VNFs) of a SFC, especially in a large time scale, is a non-trivial challenge. In the literature, existing works largely neglect the application-level relationship of VNFs in improving prediction accuracy, and few work investigates the multi-step prediction. In this work, we propose a deep-learning-based multi-step prediction model for accurate workload prediction for SFC VNFs in a dynamic network environment. We first demonstrate that predictability can be improved by taking into account application-level dependency by calculating the spatial conditional entropy of adjacent VNFs workloads. Then, the prediction model, named Graph Transformer Networks and sequence-to-sequence LSTM with Attention (GTN-LA) is introduced, which utilizes the Graph Transformer as the encoder to capture the application-level dependencies among VNFs, and the LSTM with attention as the decoder to extract the temporal dependencies within the time varying load information. Finally, GTN-LA is validated through intensive evaluation with a real SFC workload dataset by comparing towards several baselines. Jiayi Liu 0001, Xuemei Xie, Guangming Shi |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2023 | Gradient Corner Pooling for Keypoint-Based Object DetectionabstractDetecting objects as multiple keypoints is an important approach in the anchor-free object detection methods while corner pooling is an effective feature encoding method for corner positioning. The corners of the bounding box are located by summing the feature maps which are max-pooled in the x and y directions respectively by corner pooling. In the unidirectional max pooling operation, the features of the densely arranged objects of the same class are prone to occlusion. To this end, we propose a method named Gradient Corner Pooling. The spatial distance information of objects on the feature map is encoded during the unidirectional pooling process, which effectively alleviates the occlusion of the homogeneous object features. Further, the computational complexity of gradient corner pooling is the same as traditional corner pooling and hence it can be implemented efficiently. Gradient corner pooling obtains consistent improvements for various keypoint-based methods by directly replacing corner pooling. We verify the gradient corner pooling algorithm on the dataset and in real scenarios, respectively. The networks with gradient corner pooling located the corner points earlier in the training process and achieve an average accuracy improvement of 0.2%-1.6% on the MS-COCO dataset. The detectors with gradient corner pooling show better angle adaptability for arrayed objects in the actual scene test. Xuemei Xie, Mingxuan Yu, Jiakai Luo, Chengwei Rao, Guangming Shi |
AAAI | 2 |
| 2023 | Improving Robotic Tactile Localization Super-resolution via Spatiotemporal Continuity Learning and Overlapping Air ChambersabstractHuman hand has amazing super-resolution ability in sensing the force and position of contact and this ability can be strengthened by practice. Inspired by this, we propose a method for robotic tactile super-resolution enhancement by learning spatiotemporal continuity of contact position and a tactile sensor composed of overlapping air chambers. Each overlapping air chamber is constructed of soft material and seals the barometer inside to mimic adapting receptors of human skin. Each barometer obtains the global receptive field of the contact surface with the pressure propagation in the hyperelastic seal overlapping air chambers. Neural networks with causal convolution are employed to resolve the pressure data sampled by barometers and to predict the contact position. The temporal consistency of spatial position contributes to the accuracy and stability of positioning. We obtain an average super-resolution (SR) factor of over 2500 with only four physical sensing nodes on the rubber surface (0.1 mm in the best case on 38 × 26 mm²), which outperforms the state-of-the-art. The effect of time series length on the location prediction accuracy of causal convolution is quantitatively analyzed in this article. We show that robots can accomplish challenging tasks such as haptic trajectory following, adaptive grasping, and human-robot interaction with the tactile sensor. This research provides new insight into tactile super-resolution sensing and could be beneficial to various applications in the robotics field. Xuemei Xie, Guangming Shi |
AAAI | 3 |
| 2023 | Learning Primitive-Aware Discriminative Representations for Few-Shot Learning
Jianpeng Yang, Yuhang Niu, Xuemei Xie, Guangming Shi |
ICONIP (2) | 3 |
| 2023 | OCSKB: An Object Component Sketch Knowledge Base for Fast 6D Pose Estimationabstract6D pose estimation from a single RGB image is a fundamental task in computer vision. In most methods of instance-level or category-level 6D pose estimation, accurate CAD models or point cloud models are indispensable part. It is not easy to quickly obtain the models of these everyday objects. To address this issue, we present a part-level object component sketch knowledge base which consists of 270 real-world object sketch models of 30 categories. Objects are disassembled into geometry components with spatial relationship according to their functions and structures, and convert them into three basic spatial structures: frustum, circular truncated cone, and sphere. We present a fast pipeline for sketch modeling with our tool. The average time for this method to build a simple model for everyday objects is about 2 minutes. Additionally, we leverage the geometric information and spatial relationships inherent in the multiple viewpoint projection maps of these sketch bases to develop a rapid inference framework for 6D pose estimation. The interpretable steps in our framework gradually retrieve and activate valid solutions in the discrete 6D pose space. Extensive experiments in real-world environments have demonstrated that our method can reliably and robustly estimate the 6D pose of objects, even without access to accurate CAD or point cloud models. Furthermore, our method achieves state-of-the-art performance, operating at a speed of 90 frames per second using parallel computing on GPU. Guangming Shi, Xuemei Xie, Mingxuan Yu, Chengwei Rao, Jiakai Luo |
ACM Multimedia | 3 |
| 2023 | Structure guided network for human pose estimation
Xuemei Xie, Bo'ao Li, Fu Li 0002 |
Appl. Intell. | 2 |
| 2023 | MADPL-net: Multi-layer attention dictionary pair learning network for image classification
Guangming Shi, Weisheng Dong, Xuemei Xie |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Few-shot human-object interaction video recognition with transformers
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
Neural Networks | 2 |
| 2023 | View-Normalized and Subject-Independent Skeleton Generation for Action RecognitionabstractSkeleton-based action recognition has attracted great interest in computer vision. For this task, a challenging problem concerns the large intraclass variances of skeleton data, which are mainly caused by diverse viewpoints and subjects, and greatly increase the difficulty of modeling actions through a network. To address the above problem, we propose a variance reduction (VaRe) framework for skeleton-based action recognition, which consists of a view-normalization generative adversarial network (VN-GAN), a subject-independent network (SINet) and a classification network. First, the VN-GAN is responsible for reducing view-induced intraclass variances. Specifically, this network, comprising a generator and a discriminator, is aimed at learning a mapping from a diverse-view skeleton distribution to a unified-view skeleton distribution in an unsupervised manner, thereby generating a view-normalized skeleton. Second, taking the view-normalized skeleton as input, the SINet focuses on reducing the influences of the personal habits of subjects on action recognition. To generate SI skeleton data, the SINet automatically adjusts the human pose according to the human kinematic structure under a classification loss constraint. Finally, without the interference of view- and subject-induced variances, the classification network can concentrate more on learning discriminative action features to predict classes. Furthermore, by combining the joint and bone modalities, the proposed framework achieves competitive performance on three benchmarks: NTU RGB+D, NTU-120 RGB+D and Northwestern-UCLA Multiview Action 3D. Qingzhe Pan, Zhifu Zhao, Xuemei Xie, Jianan Li 0003, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Filter Clustering for Compressing CNN Model With Better Feature DiversityabstractAs a practical approach for compressing convolutional neural networks (CNNs), network pruning has been rapidly developed in recent years. The conventional methods prune inactive filters permanently from models to reduce the width of each layer and then train the pruned model until convergence. However, such methods have limitations in that: (1) The activation-based pruning criteria ignore the correlation between filters, leading to attenuation in types of features; (2) The permanent filter removal restricts the architecture of models in the subsequent training so that reducing the chances of learning more features; (3) The single-width compression may generate narrow layers that block the information flow, resulting in limited feature capacity in the next layers and hard optimization. These limitations reduce the feature diversity in the pruned model and thus lead to sub-optimal model quality. In this paper, a compression method named filter clustering is proposed to rectify the problem of poor feature diversity in traditional pruning and achieve better model quality from three perspectives. Firstly, to maintain the variety of features after pruning, we treat the model compression as a clustering task and merge filters with similar outputs, rather than removing inactive filters. Specifically, a handy estimation approach is designed to convert the similarity of the output into filter similarity, which liberates the measurement from sampling numerous images. Secondly, to increase the probability of learning more features during training, we propose a periodic training and clustering pipeline, which creates a larger optimization space by dynamically exploring different sub-model architectures. Finally, to prevent the feature capacity from being influenced by the narrow layers, we introduce and leverage a fusible anti-blocking branch to smoothly remove such layers. Extensive experiments demonstrate that the proposed method can achieve compact models with better feature diversity and reduce 1%~15% more calculations than the previous methods while maintaining performance. Zhenyu Wang 0008, Xuemei Xie, Qinghang Zhao, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Correlation filters based on spatial-temporal Gaussion scale mixture modelling for visual tracking
Guangming Shi, Weisheng Dong, Tianzhu Zhang 0001, Jinjian Wu, Xuemei Xie, Xin Li 0005 |
Neurocomputing | 6 |
| 2022 | Detecting human-object interactions in videos by modeling the trajectory of objects and human skeleton
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
Neurocomputing | 2 |
| 2022 | Soft focal loss: Evaluating sample quality for dense object detection
Zhenyuan Wang, Xuemei Xie, Jianxiu Yang, Guangming Shi |
Neurocomputing | 2 |
| 2022 | Language-guided graph parsing attention network for human-object interaction recognition
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Bayesian Correlation Filter Learning With Gaussian Scale Mixture Model for Visual TrackingabstractCorrelation filters (CF), a popular tool for visual tracking, suffer from unwanted boundary effects due to the periodic assumption needed for FFT implementation. To address this issue, spatially regularized discriminative correlation filters (SRDCF) have been proposed by introducing a weighting matrix to the regularization term. However, the existing design of spatial weighting matrix is often heuristic and non-adaptive. Inspired by recent advances in joint discrimination and reliability learning for correlation tracking, we propose a principled Bayesian correlation filter learning method using Gaussian scale mixture (GSM) model. The key idea is to decompose each CF coefficient into the product of a positive scalar multiplier and a Gaussian random variable. Treating positive multipliers as weighting coefficients, GSM-based modeling of CFs leads to a spatially adaptive regularization strategy with improved capability of handling various appearance-related uncertainty factors (e.g., scale variation, out-of-plane rotation, and motion blur). Moreover, by imposing a sparse prior over the multipliers, we can jointly learn multipliers and CFs under a unified Bayesian estimation framework. Structured GSM model allows us to better exploit the spatial correlations among CFs and further improve the tracking performance. Experimental results on OTB-2013, OTB-2015, Temple Color-128, VOT-2016, and VOT-2017 show that our tracking method performs favorably when compared with current state-of-the-art methods. Guangming Shi, Tianzhu Zhang 0001, Weisheng Dong, Jinjian Wu, Xuemei Xie, Xin Li 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | View-normalized Skeleton Generation for Action RecognitionabstractSkeleton-based action recognition has attracted great interest due to low cost of skeleton data acquisition and high robustness to external conditions. A challenging problem of skeleton-based action recognition is the large intra-class gap caused by various viewpoints of skeleton data, which makes the action modeling difficult for network. To alleviate this problem, a feasible solution is to utilize label supervised methods to learn a view-normalization model. However, since the skeleton data in real scenes is acquired from diverse viewpoints, it is difficult to obtain the corresponding view-normalized skeleton as label. Therefore, how to learn a view-normalization model without the supervised label is the key to solving view-variance problem. To this end, we propose a view normalization-based action recognition framework, which is composed of view-normalization generative adversarial network (VN-GAN) and classification network. For VN-GAN, the model is designed to learn the mapping from diverse-view distribution to normalized-view distribution. In detail, it is implemented by graph convolution, where the generator predicts the transformation angles for view normalization and discriminator classifies the real input samples from the generated ones. For classification network, view-normalized data is processed to predict the action class. Without the interference of view variances, classification network can extract more discriminative feature of action. Furthermore, by combining the joint and bone modalities, the proposed method reaches the state-of-the-art performance on NTU RGB+D and NTU-120 RGB+D datasets. Especially in NTU-120 RGB+D, the accuracy is improved by 3.2% and 2.3% under cross-subject and cross-set criteria, respectively. Qingzhe Pan, Zhifu Zhao, Xuemei Xie, Jianan Li 0003, Guangming Shi |
ACM Multimedia | 3 |
| 2021 | Scene Semantic Guidance for Object Detection
Xuemei Xie |
PRCV (1) | 2 |
| 2021 | Knowledge embedded GCN for skeleton-based two-person interaction recognition
Jianan Li 0003, Xuemei Xie, Qingzhe Pan, Zhifu Zhao, Guangming Shi |
Neurocomputing | 2 |
| 2021 | Knowledge-guided semantic computing network
Guangming Shi, Dahua Gao, Jie Lin 0008, Xuemei Xie, Danhua Liu |
Neurocomputing | 5 |
| 2021 | Domain-aware Stacked AutoEncoders for zero-shot learning
Jianqiang Song, Guangming Shi, Xuemei Xie, Qingtao Wu, Mingchuan Zhang |
Neurocomputing | 3 |
| 2021 | Blind Image Quality Assessment With Active InferenceabstractBlind image quality assessment (BIQA) is a useful but challenging task. It is a promising idea to design BIQA methods by mimicking the working mechanism of human visual system (HVS). The internal generative mechanism (IGM) indicates that the HVS actively infers the primary content (i.e., meaningful information) of an image for better understanding. Inspired by that, this paper presents a novel BIQA metric by mimicking the active inference process of IGM. Firstly, an active inference module based on the generative adversarial network (GAN) is established to predict the primary content, in which the semantic similarity and the structural dissimilarity (i.e., semantic consistency and structural completeness) are both considered during the optimization. Then, the image quality is measured on the basis of its primary content. Generally, the image quality is highly related to three aspects, i.e., the scene information (content-dependency), the distortion type (distortion-dependency), and the content degradation (degradation-dependency). According to the correlation between the distorted image and its primary content, the three aspects are analyzed and calculated respectively with a multi-stream convolutional neural network (CNN) based quality evaluator. As a result, with the help of the primary content obtained from the active inference and the comprehensive quality degradation measurement from the multi-stream CNN, our method achieves competitive performance on five popular IQA databases. Especially in cross-database evaluations, our method achieves significant improvements. Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie, Guangming Shi, Weisi Lin |
IEEE Trans. Image Process. | 5 |
| 2020 | Active Inference of GAN for No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) is a challenging task. It is a promising idea to design NR-IQA algorithms by mimicking how human visual system (HVS) works. The internal generative mechanism (IGM) indicates that HVS actively infers the primary content of an image for better understanding. Inspired by that, a novel NR-IQA method with active inference is proposed in this paper. First, a generative adversarial network (GAN) is proposed to predict the primary content of a distorted image, in which two IGM-inspired constraints are considered during the optimization. Next, based on the correlation between the distorted image and its primary content, different degradations (i.e., the content/distortion-/structure-dependency degradation) are measured simultaneously with a multi-stream convolutional neural network (CNN) for NR-IQA. Benefit from the primary content obtained from GAN and the multiple degradations measurement of CNN, our method achieves the state-of-the-art on five public IQA databases. Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie |
ICME | 5 |
| 2020 | Multi-level Prediction with Graphical Model for Human Pose Estimation
Xuemei Xie, Lihua Ma, Jiang Du 0011, Guangming Shi |
PRCV (2) | 2 |
| 2020 | A Multi-level Equilibrium Clustering Approach for Unsupervised Person Re-identification
Fangyu Wang, Zhenyu Wang 0008, Xuemei Xie, Guangming Shi |
PRCV (3) | 3 |
| 2020 | Network pruning using sparse learning and genetic algorithm
Zhenyu Wang 0008, Fu Li 0002, Guangming Shi, Xuemei Xie, Fangyu Wang |
Neurocomputing | 4 |
| 2020 | SGM-Net: Skeleton-guided multimodal network for action recognition
Jianan Li 0003, Xuemei Xie, Qingzhe Pan, Zhifu Zhao, Guangming Shi |
Pattern Recognit. | 2 |
| 2019 | Zero-Shot Learning Using Stacked Autoencoder with Manifold RegularizationsabstractZero-shot learning (ZSL), which focuses on transferring the knowledge from the seen classes to unseen ones, has attracted more and more attention in the computer vision community. Exploring the relationships among the spaces of visual representation, semantic description and label information is a key to the success of ZSL. In this paper, we propose a novel approach by using a two-layer Stacked AutoEncoder (StAE) with manifold regularizations to construct the tight relations of different spaces, where the first-layer encoder aims to project a visual feature vector into the semantic space, and the second-layer encoder connects the semantic description of a sample with its label directly. Meanwhile, the decoders seek to reconstruct the visual representation from label information and semantic description successively. Besides, two manifold regularizers are integrated in the stacked autoencoder, which captures the manifold structures residing in the different spaces effectively. Compared with the previous related works, the proposed approach is a more general framework and has stronger transfer ability from seen classes to unseen classes. Extensive experiments on the benchmark datasets clearly demonstrate that our StAE performs significantly better than the state-of-the-arts. Jianqiang Song, Guangming Shi, Xuemei Xie, Dahua Gao |
ICIP | 3 |
| 2019 | Channel Feature Enhanced Detector for Small Ball Detection
Shambel Ferede, Xuemei Xie, Jiang Du 0011, Guangming Shi |
PRCV (1) | 2 |
| 2019 | A Real-Time Rock-Paper-Scissor Hand Gesture Recognition System Based on FlowNet and Event Camera
Xuemei Xie, Jinjian Wu, Guangming Shi |
PRCV (1) | 1 |
| 2019 | Fully convolutional measurement network for compressive sensing image reconstruction
Jiang Du 0011, Xuemei Xie, Chenye Wang, Guangming Shi |
Neurocomputing | 2 |
| 2019 | SISRSet: Single image super-resolution subjective evaluation test and objective quality assessment
Guangming Shi, Wenfei Wan, Jinjian Wu, Xuemei Xie, Weisheng Dong, Hong Ren Wu |
Neurocomputing | 4 |
| 2019 | Visualizing and understanding of learned compressive sensing with residual network
Zhifu Zhao, Xuemei Xie, Chenye Wang, Wan Liu 0001, Guangming Shi, Jiang Du 0011 |
Neurocomputing | 2 |
| 2019 | Blind image quality assessment with semantic information
Weiping Ji, Jinjian Wu, Guangming Shi, Wenfei Wan, Xuemei Xie |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Multi-layer discriminative dictionary learning with locality constraint for image classification
Jianqiang Song, Xuemei Xie, Guangming Shi, Weisheng Dong |
Pattern Recognit. | 2 |
| 2019 | ROI-CSNet: Compressive sensing network for ROI-aware image recovery
Zhifu Zhao, Xuemei Xie, Chenye Wang, Siying Mao, Wan Liu 0001, Guangming Shi |
Signal Process. Image Commun. | 2 |
| 2018 | Full Image Recover for Block-Based Compressive SensingabstractCompressive sensing (CS) theory is able to acquire measurements of a scene at sub-Nyquist rate and recover the scene image from these under-sampled measurements. Recent years, CS has been improved greatly for the application of deep learning technology. In conventional methods, block-based mechanism is used to recover images from measurements, which usually causes block effect in reconstructed images. In this paper, we propose a novel CNN-based network for CS to solve this problem. In the measurement part, the input is measured block by block to acquire the measurements. While in the recovery part, all the measurements from one image are used simultaneously to reconstruct the full image. Different from previous methods recovering images block by block, the proposed framework rebuilds the structure information destroyed in the measurement part. Block effect is removed accordingly. Experiments show that there is no block effect at all in the reconstructed images. On a standard dataset our method has significant improvements in reconstruction results compared with existing state-of-the-art methods. Xuemei Xie, Chenye Wang, Jiang Du 0011, Guangming Shi |
ICME | 1 |
| 2018 | Color Image Reconstruction with Perceptual Compressive SensingabstractWe propose a novel compressive sensing framework for color images. Recently, compressive sensing (CS) has gain its popularity with the development of deep learning. To our best knowledge, existing methods all deal with RGB images channel by channel. This brings redundancy of measurements. In this paper, we do a breakthrough work. Instead of recovering RGB images channel by channel uniformly, we adopt non-uniform sampling in different channels in YCbCr color space. The luminance component takes up more measurements while the other channels take up less in the proposed framework. It greatly enhances the performance on CS for color images. Moreover, perceptual loss gives a powerful ability to better capture the structure information. We give the measurement rate at 2% as an example in the experiments, and the results show the proposed method outperforms all the existing methods with better structure of images. Jiang Du 0011, Xuemei Xie, Chenye Wang, Guangming Shi |
ICPR | 2 |
| 2018 | Perceptual Compressive Sensing
Jiang Du 0011, Xuemei Xie, Chenye Wang, Guangming Shi |
PRCV (3) | 2 |
| 2018 | Gated Feature Pyramid Network for Object Detection
Xuemei Xie, Quan Liao, Lihua Ma |
PRCV (4) | 1 |
| 2018 | Exploiting class-wise coding coefficients: Learning a discriminative dictionary for pattern classification
Jianqiang Song, Xuemei Xie, Guangming Shi, Weisheng Dong |
Neurocomputing | 2 |
| 2018 | Image Super-Resolution With Parametric Sparse Model LearningabstractRecovering a high-resolution (HR) image from its low-resolution (LR) version is an ill-posed inverse problem. Learning accurate prior of HR images is of great importance to solve this inverse problem. Existing super-resolution (SR) methods either learn a non-parametric image prior from training data (a large set of LR/HR patch pairs) or estimate a parametric prior from the LR image analytically. Both methods have their limitations: the former lacks flexibility when dealing with different SR settings; while the latter often fails to adapt to spatially varying image structures. In this paper, we propose to take a hybrid approach toward image SR by combining those two lines of ideas - that is, a parametric sparse prior of HR images is learned from the training set as well as the input LR image. By exploiting the strengths of both worlds, we can more accurately recover the sparse codes and therefore HR image patches than conventional sparse coding approaches. Experimental results show that the proposed hybrid SR method significantly outperforms existing model-based SR methods and is highly competitive to current state-of-the-art learning-based SR methods in terms of both subjective and objective image qualities. Weisheng Dong, Xuemei Xie, Guangming Shi, Jinjian Wu, Xin Li 0005 |
IEEE Trans. Image Process. | 3 |
| 2018 | Robust Foreground Estimation via Structured Gaussian Scale Mixture ModelingabstractRecovering the background and foreground parts from video frames has important applications in video surveillance. Under the assumption that the background parts are stationary and the foreground are sparse, most of existing methods are based on the framework of robust principal component analysis (RPCA), i.e., modeling the background and foreground parts as a low-rank and sparse matrices, respectively. However, in realistic complex scenarios, the conventional norm sparse regularizer often fails to well characterize the varying sparsity of the foreground components. How to select the sparsity regularizer parameters adaptively according to the local statistics is critical to the success of the RPCA framework for background subtraction task. In this paper, we propose to model the sparse component with a Gaussian scale mixture (GSM) model. Compared with the conventional norm, the GSM-based sparse model has the advantages of jointly estimating the variances of the sparse coefficients (and hence the regularization parameters) and the unknown sparse coefficients, leading to significant estimation accuracy improvements. Moreover, considering that the foreground parts are highly structured, a structured extension of the GSM model is further developed. Specifically, the input frame is divided into many homogeneous regions using superpixel segmentation. By characterizing the set of sparse coefficients in each homogeneous region with the same GSM prior, the local dependencies among the sparse coefficients can be effectively exploited, leading to further improvements for background subtraction. Experimental results on several challenging scenarios show that the proposed method performs much better than most of existing background subtraction methods in terms of both performance and speed. Guangming Shi, Weisheng Dong, Jinjian Wu, Xuemei Xie |
IEEE Trans. Image Process. | 5 |
| 2017 | No-reference image quality assessment with orientation selectivity mechanismabstractNo-reference (NR) image quality assessment (IQA) technology is greatly required in quality-orientated visual signal processing systems. However, without the guidance of the reference information, it is still a great challenge for NR IQA to perform consistent with the subjective perception. Researches on cognitive neuroscience state that the human visual system (HVS) presents substantially orientation selectivity mechanism, within which the visual structures are extracted in the local receptive fields for scene understanding. Inspired by this mechanism, a set of orientation selectivity based visual patterns are designed. By analyzing the quality degradation on those patterns, a novel visual pattern degradation based NR IQA method is proposed. Experimental results on large databases demonstrate that the proposed method outperforms the existing NR IQA methods. Jinjian Wu, Man Zhang 0007, Guangming Shi, Xuemei Xie, Weisi Lin |
ICIP | 4 |
| 2017 | Bag-of-words feature representation for blind image quality assessment with local quantized pattern
Xuemei Xie, Yazhong Zhang, Jinjian Wu, Guangming Shi, Weisheng Dong |
Neurocomputing | 1 |
| 2017 | A hierarchical multiplier-free architecture for HEVC transform
Chunxiao Fan 0002, Fu Li 0002, Guangming Shi, Fei Qi 0001, Xuemei Xie, Dandan Jiao |
Multim. Tools Appl. | 6 |
| 2017 | An AR based fast mode decision for H.265/HEVC intra coding
Fu Li 0002, Dandan Jiao, Guangming Shi, Chunxiao Fan 0002, Xuemei Xie |
Multim. Tools Appl. | 6 |
| 2017 | Estimation of directions of arrival of multiple distributed sources for nested array
Guangming Shi, Xuemei Xie |
Signal Process. | 3 |
| 2017 | Off-grid DOA estimation under nonuniform noise via variational sparse Bayesian learning
Xuemei Xie, Guangming Shi |
Signal Process. | 2 |
| 2017 | Mixed Noise Removal via Laplacian Scale Mixture Modeling and Nonlocal Low-Rank ApproximationabstractRecovering the image corrupted by additive white Gaussian noise (AWGN) and impulse noise is a challenging problem due to its difficulties in an accurate modeling of the distributions of the mixture noise. Many efforts have been made to first detect the locations of the impulse noise and then recover the clean image with image in painting techniques from an incomplete image corrupted by AWGN. However, it is quite challenging to accurately detect the locations of the impulse noise when the mixture noise is strong. In this paper, we propose an effective mixture noise removal method based on Laplacian scale mixture (LSM) modeling and nonlocal low-rank regularization. The impulse noise is modeled with LSM distributions, and both the hidden scale parameters and the impulse noise are jointly estimated to adaptively characterize the real noise. To exploit the nonlocal self-similarity and low-rank nature of natural image, a nonlocal low-rank regularization is adopted to regularize the denoising process. Experimental results on synthetic noisy images show that the proposed method outperforms existing mixture noise removal methods. Weisheng Dong, Xuemei Xie, Guangming Shi, Xiang Bai |
IEEE Trans. Image Process. | 3 |
| 2016 | Learning Parametric Sparse Models for Image Super-ResolutionabstractLearning accurate prior knowledge of natural images is of great importance for single image super-resolution (SR). Existing SR methods either learn the prior from the low/high-resolution patch pairs or estimate the prior models from the input low-resolution (LR) image. Specifically, high-frequency details are learned in the former methods. Though effective, they are heuristic and have limitations in dealing with blurred LR images; while the latter suffers from the limitations of frequency aliasing. In this paper, we propose to combine those two lines of ideas for image super-resolution. More specifically, the parametric sparse prior of the desirable high-resolution (HR) image patches are learned from both the input low-resolution (LR) image and a training image dataset. With the learned sparse priors, the sparse codes and thus the HR image patches can be accurately recovered by solving a sparse coding problem. Experimental results show that the proposed SR method outperforms existing state-of-the-art methods in terms of both subjective and objective image qualities. Weisheng Dong, Xuemei Xie, Guangming Shi, Xin Li 0005, Donglai Xu |
NIPS | 3 |
| 2015 | Learning Parametric Distributions for Image Super-Resolution: Where Patch Matching Meets Sparse CodingabstractExisting approaches toward Image super-resolution (SR) is often either data-driven (e.g., based on internet-scale matching and web image retrieval) or model-based (e.g., formulated as an Maximizing a Posterior estimation problem). The former is conceptually simple yet heuristic, while the latter is constrained by the fundamental limit of frequency aliasing. In this paper, we propose to develop a hybrid approach toward SR by combining those two lines of ideas. More specifically, the parameters underlying sparse distributions of desirable HR image patches are learned from a pair of LR image and retrieved HR images. Our hybrid approach can be interpreted as the first attempt of reconciling the difference between parametric and nonparametric models for low-level vision tasks. Experimental results show that the proposed hybrid SR method performs much better than existing state-of-the-art methods in terms of both subjective and objective image qualities. Weisheng Dong, Guangming Shi, Xuemei Xie |
ICCV | 4 |
| 2015 | Reduced-reference image quality assessment based on entropy differences in DCT domainabstractReduced-reference image quality assessment (RR-IQA) algorithm aims to automatically evaluate the image quality using only partial information about the reference image. In this paper, we propose a new RR-IQA metric by employing the entropy features of each frequency band in the DCT domain. It is well known that human eyes have different sensitivity to different bands, and distortions on each band result in individual quality degradations. Therefore, we suggest to separately compute the visual information degradations on different band for quality assessment. The degradations on each DCT band are firstly analyzed according to the entropy difference. And then, the quality score is obtained using the weighted sum of the entropy difference of each band from low frequency to high frequency. Experimental results on several public image databases show that the proposed method uses limited reference data (8 values) and performs highly consistent with human perception. Yazhong Zhang, Jinjian Wu, Guangming Shi, Xuemei Xie |
ISCAS | 4 |
| 2013 | Compressive modulation in digital communicationabstractBandwidth efficiency is one of the most important indicators to measure different modulation schemes in digital communication systems. The waveforms of existing modulation schemes are all separated in time domain, making it difficult for them to improve in bandwidth efficiency. Compressive Sensing (CS) theory shows that it is possible to reconstruct original signals in aliasing measurements. In this paper, we propose a Compressive Modulation scheme by combining CS theory and traditional BPSK pattern. Theoretic analysis and experimental results show that the bandwidth efficiency can be highly improved by using the proposed scheme. Yingyu Li, Guangming Shi, Xuemei Xie, Chongyu Chen |
ISCAS | 3 |
| 2011 | Nonuniform Directional Filter Banks With Arbitrary Frequency PartitioningabstractDirectional filter banks (DFBs) are highly desired in directional representation of images. In this correspondence, we propose a 2-D nonsubsampled nonuniform directional filter bank (NUDFB) and its design method. The proposed NUDFB has nonuniform wedge-shaped subbands and allows arbitrary frequency partitioning schemes. It can extract directional information according to the directional distribution of images. This attractive advantage cannot be achieved by the existing directional transforms. The design method of the proposed NUDFB is based upon the pseudopolar Fourier transform. By utilizing the geometry property of the pseudopolar grid, we employ a 1-D nonsubsampled nonuniform filter bank to obtain a set of nonuniform wedge-shaped subbands. During the design process, only 1-D operations are involved and, thus, the difficulty encountered in the design of 2-D fan filters is avoided. To demonstrate the potential of the proposed NUDFB, an example on image directional decomposition is given. Lili Liang, Guangming Shi, Xuemei Xie |
IEEE Trans. Image Process. | 3 |
| 2011 | High-Resolution Imaging Via Moving Random Exposure and Its SimulationabstractIn this correspondence, we introduce a new imaging method to obtain high-resolution (HR) images. The image acquisition is performed in two stages, compressive measurement and optimization reconstruction. In order to reconstruct HR images by a small number of sensors, compressive measurements are made. Specifically, compressive measurements are made by a low-resolution (LR) camera with randomly fluttering shutter, which can be viewed as a moving random exposure pattern. In the optimization reconstruction stage, the HR image is computed by different models according to the prior knowledge of scenes. The proposed imaging method offers a new way of acquiring HR images of essentially static scenes when the camera resolution is limited by severe constraints such as cost, battery capacity, memory space, transmission bandwidth, etc. and when the prior knowledge of scenes is available. The simulation results demonstrate the effectiveness of the proposed imaging method. Guangming Shi, Dahua Gao, Xiaoxia Song, Xuemei Xie, Danhua Liu |
IEEE Trans. Image Process. | 4 |
| 2006 | On the theory and design of a class of recombination nonuniform filter banks with low-delay FIR and IIR filtersabstractThis paper studies the theory and design of a class of recombination nonuniform FBs (RNFB) with low-delay (LD) FIR and IIR filters. The conditions for suppressing the spurious response and achieving a good frequency characteristic for these LD FIR/IIR RNFBs are developed. The proposed LD FIR RNFBs have a lower system delay than their linear-phase counterparts, at the expense of slight increase in phase distortion of the analysis filters and arithmetic complexity. By model reducing the LD FIR uniform FBs by the modified model reduction method, an IIR RNFB with a similar characteristic can be readily obtained. A design example is given illustrate the effectiveness of the proposed method S. S. Yin, S. C. Chan 0001, Xuemei Xie |
ISCAS | 3 |
| 2000 | Theory and design of a class of cosine-modulated non-uniform filter banksabstractIn this paper, the theory and design of a class of PR cosine-modulated nonuniform filter bank is proposed. It is based on a structure previously proposed by Cox (1986), where the outputs of a uniform filter bank are combined or merged by means of the synthesis section of another filter bank with smaller channel number. Simplifications are imposed on this structure so that the design procedure can be considerably simplified. Due to the use of cosine modulated filter banks as the original and recombination filter banks, excellent filter quality and low design and implementation complexities can be achieved. Problems with these merging techniques such as spectrum inversion, equivalent filter representations and protrusion cancellation are also addressed. As the merging is performed after the decimation, the arithmetic complexity is lower than other conventional approaches. Design examples show that PR nonuniform filter banks with high stopband attenuation and low design and implementation complexities can be obtained by the proposed method. S. C. Chan 0001, Xuemei Xie, Tony Tung Ip Yuk |
ICASSP | 2 |