Shan An

dblp:176/3512 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0001-7796-6952ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 9 since 2021Systems, architecture and hardware · 7 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Edge-guided semi-supervised 6D pose estimation with cross-domain alignment for robotic grasping
Jiajie Wen, Shan An
Expert Syst. Appl.6
2026 Dexterous Manipulation Through Imitation Learning: A Survey
abstract
Dexterous manipulation, which refers to the ability of a robotic hand or multi-fingered end-effector to skillfully control, reorient, and manipulate objects through precise, coordinated finger movements and adaptive force modulation, enables complex interactions similar to human hand dexterity. With recent advances in robotics and machine learning, there is a growing demand for these systems to operate in complex and unstructured environments. Traditional model-based approaches struggle to generalize across tasks and object variations due to the high dimensionality and complex contact dynamics of dexterous manipulation. Although model-free methods such as reinforcement learning (RL) show promise, they require extensive training, large-scale interaction data, and carefully designed rewards for stability and effectiveness. Imitation learning (IL) offers an alternative by allowing robots to acquire dexterous manipulation skills directly from expert demonstrations, capturing fine-grained coordination and contact dynamics while bypassing the need for explicit modeling and large-scale trial-and-error. This survey provides an overview of dexterous manipulation methods based on imitation learning, details recent advances, and addresses key challenges in the field. Additionally, it explores potential research directions to enhance IL-driven dexterous manipulation. Our goal is to offer researchers and practitioners a comprehensive introduction to this rapidly evolving domain.
Shan An, Chao Tang 0001, Yuning Zhou, Tengyu Liu, Fangqiang Ding, Shufang Zhang, Yao Mu 0001, Ran Song 0001, Wei Zhang 0021, Zeng-Guang Hou, Hong Zhang 0013
IEEE Trans Autom. Sci. Eng.1
2026 MBE-UNet: Multi-Branch Boundary Enhanced U-Net for Ultrasound Segmentation
abstract
Accurately capturing object areas in medical images is crucial for the clinical diagnosis and treatment of diseases. Due to the inherent low contrast and blurry edges in ultrasound images, most existing CNN-based methods often yield unsatisfactory segmentation results, making ultrasound image segmentation a challenging task. This paper introduces a novel multi-branch boundary enhanced network (MBE-UNet) for automatic ultrasound image segmentation. This method can accurately segment targets and delineate boundaries simultaneously using a multi-branch network. First, a global pyramid attention module (GPAM) is designed to capture multi-scale contextual information. Second, we embed a boundary cascade module (BCM) in the main branch to ensure the network focuses on edge information flow and generates relatively desirable boundaries. Finally, a boundary feature fusion module (BFM) is used to integrate boundary and region information, obtaining a boundary enhanced region map. The visual results and quantitative analysis demonstrate that the proposed MBE-UNet outperforms classical segmentation networks on three publicly available ultrasound datasets.
Qing Qin, Ziwei Lin, Guangyuan Gao, Chunxiao Han, Yingmei Qin, Shan An, Yanqiu Che
IEEE J. Biomed. Health Informatics8
2025 TinyMIG: Transferring Generalization from Vision Foundation Models to Single-Domain Medical Imaging
abstract
Medical imaging faces significant challenges in single-domain generalization (SDG) due to the diversity of imaging devices and the variability among data collection centers. To address these challenges, we propose \textbf{TinyMIG}, a framework designed to transfer generalization capabilities from vision foundation models to medical imaging SDG. TinyMIG aims to enable lightweight specialized models to mimic the strong generalization capabilities of foundation models in terms of both global feature distribution and local fine-grained details during training. Specifically, for global feature distribution, we propose a Global Distribution Consistency Learning strategy that mimics the prior distributions of the foundation model layer by layer. For local fine-grained details, we further design a Localized Representation Alignment method, which promotes semantic alignment and generalization distillation between the specialized model and the foundation model. These mechanisms collectively enable the specialized model to achieve robust performance in diverse medical imaging scenarios. Extensive experiments on large-scale benchmarks demonstrate that TinyMIG, with extremely low computational cost, significantly outperforms state-of-the-art models, showcasing its superior SDG capabilities. All the code and model weights will be publicly available.
Hongyan Xu 0002, Yichao Cao, Xiu Su, Tianfa Li, Shan An, Haogang Zhu
ICML7
2025 HyperGraph ROS: An Open-Source Robot Operating System for Hybrid Parallel Computing based on Computational HyperGraph
abstract
This paper presents HyperGraph ROS, an open-source robot operating system that unifies intra-process, inter-process, and cross-device computation into a computational hypergraph for efficient message passing and parallel execution. In order to optimize communication, HyperGraph ROS dynamically selects the optimal communication mechanism while maintaining a consistent API. For intra-process messages, Intel-TBB Flow Graph is used with C++ pointer passing, which ensures zero memory copying and instant delivery. Meanwhile, inter-process and cross-device communication seamlessly switch to ZeroMQ. When a node receives a message from any source, it is immediately activated and scheduled for parallel execution by Intel-TBB. The computational hypergraph consists of nodes represented by TBB flow graph nodes and edges formed by TBB pointer-based connections for intra-process communication, as well as ZeroMQ links for inter-process and cross-device communication. This structure enables seamless distributed parallelism. Additionally, HyperGraph ROS provides ROS-like utilities such as a parameter server, a coordinate transformation tree, and visualization tools. Evaluation in diverse robotic scenarios demonstrates significantly higher transmission and throughput efficiency compared to ROS 2. Our work is available at https://github.com/wujiazheng2020a/hyper_graph_ros.
Shufang Zhang, Jiazheng Wu, Shan An
IROS5
2025 Fine-grained recognition of citrus varieties via wavelet channel attention network
Fukai Zhang, Xiao-Bo Jin, Shan An, Qiang Lyu
Knowl. Based Syst.6
2025 scSPAF: Cell Similarity Purified Adaptive Fusion Network for Single-Cell Multi-Omics Clustering
abstract
The rapid advancement of single-cell sequencing technology has generated vast amounts of multi-omics data, presenting unprecedented opportunities for single-cell multi-omics clustering analysis. However, existing single-cell clustering algorithms focus on extracting shared representations, overlooking the interactions and correlations among cells. This oversight inevitably leads to biased or confounded cell clustering results. In this paper, we propose a cell similarity purified adaptive fusion network for single-cell multi-omics clustering, named scSPAF, which adopts a multi-level fusion approach to thoroughly explore the consistency and complementarity of omics data. Specifically, we design a cell similarity purification module to accurately incorporate neighborhood information among cells into cell features, thereby purifying the latent representation that reflects cell correlations. In addition, we align attribute features from different omics to extract the consistent representation across omics. Simultaneously, by employing an adaptive fusion mechanism, we integrate representations of omics-specific and the consistent representation across omics to generate more discriminative representations of omics, further enhancing the clustering performance. Experimental results obtained from six real-world datasets demonstrate the superiority of the scSPAF algorithm when compared with other state-of-the-art methods.
Shanghui Deng, Chang Tang, Xinwang Liu 0002, Yuanyuan Liu 0004, Shan An
IEEE Trans. Comput. Biol. Bioinform.6
2025 Debiased Estimation and Inference for Spatial-Temporal EEG/MEG Source Imaging
abstract
The development of accurate electroencephalography (EEG) and magnetoencephalography (MEG) source imaging algorithm is of great importance for functional brain research and non-invasive presurgical evaluation of epilepsy. In practice, the challenge arises from the fact that the number of measurement channels is far less than the number of candidate source locations, rendering the inverse problem ill-posed. A widely used approach is to introduce a regularization term into the objective function, which inevitably biased the estimated amplitudes towards zero, leading to an inaccurate estimation of the estimator's variance. This study proposes a novel debiased EEG/MEG source imaging (DeESI) algorithm for detecting sparse brain activities, which corrects the estimation bias in signal amplitude, dipole orientation and depth. The DeESI extends the idea of group Lasso by incorporating both the matrix Frobenius norm and the L1-norm, which guarantees the estimators are only sparse over sources while maintains smoothness in time and orientation. We also derived variance of the debiased estimators for standardization and hypothesis testing. A fast alternating direction method of multipliers (ADMM) algorithm is proposed for solving the matrix form optimization problem directly without the need for vectorization. The proposed algorithm is compared with eleven existing ESI methods using simulations and an open source EEG dataset whose stimulation locations are known precisely. The DeESI exhibits the best performance in peak localization and amplitude reconstruction.
Peifeng Tong, Xinru Ding, Yuchuan Ding, Xiaokun Geng, Shan An, Song Xi Chen
IEEE Trans. Medical Imaging6
2025 Multi-View Clustering via Multi-Stage Fusion
abstract
Multi-view clustering (MVC) exploits the information captured from diverse views to partition data into different groups and attracts much attention recently. Despite significant progress, most MVC methods fuse multi-view information via one-stage fusion while neglecting the merits of multi-stage fusion which causes insufficient in utilizing rich information within data and therefore degrades the clustering performance. To this end, designing a functional framework that can fully exploit multi-view information becomes a key challenge in multi-view clustering research. In this paper, we propose a novel multi-stage fusion method, which elegantly unifies the late and early fusion into one unified framework, to capture sufficient information underlying the multi-view data and to effectively reduce the effect of low-quality views. Specifically, we construct a low dimensional latent representation from multi-view data by learning proper correlation among multi-view data in the early fusion stage. The late fusion establishes a new optimal combinational data partition from base partitions constructed by spectral clustering, which suppresses the influence of low-quality basic partitions. Then we couple the low dimensional latent representation with the learned combinational data partition to share the same cluster structure by$k$-means and maximization alignment. As a result, we collaboratively learn an accurate and robust partition representation for the following clustering task. Besides, the late fusion and early fusion are jointly learned to achieve mutual collaboration for better performance. Finally, an alternating optimization algorithm is designed to solve the resultant optimization problem. Extensive experiments conducted on eight datasets show the superiority of our method in terms of effectiveness and efficiency.
Yu Gan 0004, Yunning You, Junjie Huang 0001, Sen Xiang, Chang Tang, Wei Hu 0001, Shan An
IEEE Trans. Multim.7
2024 A Bi-Pyramid Multimodal Fusion Method for the Diagnosis Of Bipolar Disorders
abstract
Previous research on the diagnosis of Bipolar disorder has mainly focused on resting-state functional magnetic resonance imaging. However, their accuracy can not meet the requirements of clinical diagnosis. Efficient multimodal fusion strategies have great potential for applications in multimodal data and can further improve the performance of medical diagnosis models. In this work, we utilize both sMRI and fMRI data and propose a novel multimodal diagnosis model for bipolar disorder. The proposed Patch Pyramid Feature Extraction Module extracts sMRI features, and the spatio-temporal pyramid structure extracts the fMRI features. Finally, they are fused by a fusion module to output diagnosis results with a classifier. Extensive experiments show that our proposed method outperforms others in balanced accuracy from 0.657 to 0.732 on the OpenfMRI dataset, and achieves the state of the art.
Sheng Shi, Shan An, Fengmei Fan, Wenshu Ge, Feng Yu 0003, Zhiren Wang
ICASSP3
2024 Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
abstract
Facial action units (AUs), as defined in the Facial Action Coding System (FACS), have received significant research interest owing to their diverse range of applications in facial state analysis. Current mainstream FAU recognition models have a notable limitation, i.e., focusing only on the accuracy of AU recognition and overlooking explanations of corresponding AU states. In this paper, we propose an end-to-end Vision-Language joint learning network for explainable FAU recognition (termed VL-FAU), which aims to reinforce AU representation capability and language interpretability through the integration of joint multimodal tasks. Specifically, VL-FAU brings together language models to generate fine-grained local muscle descriptions and distinguishable global face description when optimising FAU recognition. Through this, the global facial representation and its local AU representations will achieve higher distinguishability among different AUs and different subjects. In addition, multi-level AU representation learning is utilised to improve AU individual attention-aware representation capabilities based on multi-scale combined facial stem feature. Extensive experiments on DISFA and BP4D AU datasets show that the proposed approach achieves superior performance over the state-of-the-art methods on most of the metrics. In addition, compared with mainstream FAU recognition methods, VL-FAU can provide local- and global-level interpretability language descriptions with the AUs' predictions.
Xuri Ge, Junchen Fu, Fuhai Chen, Shan An, Nicu Sebe, Joemon M. Jose
ACM Multimedia4
2024 Heterogeneous Graph Guided Contrastive Learning for Spatially Resolved Transcriptomics Data
abstract
Spatial transcriptomics provides revolutionary insights into cellular interactions and disease development mechanisms by combining high-throughput gene sequencing and spatially resolved imaging technologies to analyze genes naturally associated with spatially variable tissue genes. However, existing methods typically map aggregated multi-view features into a unified representation, ignoring the heterogeneity and view independence of genes and spatial information. To this end, we construct a heterogeneous Graph guided Contrastive Learning (stGCL) for aggregating spatial transcriptomics data. The method is guided by the inherent heterogeneity of cellular molecules by dynamically coordinating triple-level node attributes through comparative learning loss distributed across view domains, thus maintaining view independence during the aggregation process. In addition, we introduce a cross-view hierarchical feature alignment module employing a parallel approach to decouple spatial and genetic views on molecular structures while aggregating multi-view features according to information theory, thereby enhancing the integrity of inter- and intra-views. Rigorous experiments demonstrate that stGCL outperforms existing methods in various tasks and related downstream applications.
Xiao He 0010, Chang Tang, Xinwang Liu 0002, Chuankun Li, Shan An, Zhenglai Li
ACM Multimedia5
2024 3SHNet: Boosting image-sentence retrieval via visual semantic-spatial self-highlighting
Xuri Ge, Songpei Xu, Fuhai Chen, Jie Wang 0072, Shan An, Joemon M. Jose
Inf. Process. Manag.6
2023 WCANet: Wavelet Channel Attention Network for Citrus Variety Identification
abstract
The effective fine-grained identification of citrus varieties plays a vital role in the differential production management of citrus orchards. To our knowledge, there are few studies and publicly available datasets on fine-grained identification of citrus varieties. In this study, we propose Wavelet Channel Attention Network (WCANet) to solve the problem of fine-grained visual classification of citrus varieties and create a Citrus Variety Dataset (CVD) consisting of tree canopy images. WCANet combines global average pooling to extract global features and wavelet transform to capture local features, which greatly improves the capability of channel attention modules for multi-scale feature extraction. Experimental results demonstrate that the WCANet outperforms the state-of-the-art confidence estimation approaches on various benchmarks. Our code and dataset will be open-sourced at https://github.com/fightero/WCANet.
Fukai Zhang, Xiao-Bo Jin, Shan An, Qiang Lyu
ICIP4
2023 An Open-Source Robotic Chinese Chess Player
abstract
Consumer robots can accompany children growing up, improving their abilities while playing and entertaining. This paper presents an open-source, practical, low-cost robotic Chinese chess player. The proposed system includes an elaborate mechanical structure, a simple kinematic solution, a novel robot operating system, real-time and accurate chess recognition. Regarding its mechanical design, it combines a magnetism structure and mechanical cam drive, while the overall system has just three servo motors. At the same time, its control strategy is simple and effective. Furthermore, a lightweight robot message communication mechanism, entitled TinyROS, is developed for computing resource-limited embedded chips. Concerning the recognition process, our CNNbased object detector determines chess and achieves accurate identification. As a result, our robotic Chinese chess player is exquisite and easy for large-scale promotion while improving users' chess skills. Aiming to facilitate future consumer robot research and popularize customer robots, the model's mechanical and software design and the TinyROS protocol are open-sourced at https://github.com/Star-Robot/chinese-chess-robot.
Shan An, Guangfu Che, Jinghao Guo, Konstantinos A. Tsintotas, Fukai Zhang, Junjie Ye 0004, Changhong Fu 0001, Haogang Zhu, Hong Zhang 0013
IROS1
2023 DPNET: Dynamic Poly-attention Network for Trustworthy Multi-modal Classification
abstract
With advances in sensing technology, multi-modal data collected from different sources are increasingly available. Multi-modal classification aims to integrate complementary information from multi-modal data to improve model classification performance. However, existing multi-modal classification methods are basically weak in integrating global structural information and providing trustworthy multi-modal fusion, especially in safety-sensitive practical applications (e.g., medical diagnosis). In this paper, we propose a novel Dynamic Poly-attention Network (DPNET) for trustworthy multi-modal classification. Specifically, DPNET has four merits: (i) To capture the intrinsic modality-specific structural information, we design a structure-aware feature aggregation module to learn the corresponding structure-preserved global compact feature representation. (ii) A transparent fusion strategy based on the modality confidence estimation strategy is induced to track information variation within different modalities for dynamical fusion. (iii) To facilitate more effective and efficient multi-modal fusion, we introduce a cross-modal low-rank fusion module to reduce the complexity of tensor-based fusion and activate the implication of different rank-wise features via a rank attention mechanism. (iv) A label confidence estimation module is devised to drive the network to generate more credible confidence. An intra-class attention loss is introduced to supervise the network training. Extensive experiments on four real-world multi-modal biomedical datasets demonstrate that the proposed method achieves competitive performance compared to other state-of-the-art ones.
Xin Zou 0001, Chang Tang, Zhenglai Li, Xiao He 0010, Shan An, Xinwang Liu 0002
ACM Multimedia6
2023 A meaningful learning method for zero-shot semantic segmentation
Xianglong Liu 0001, Shihao Bai, Shan An, Shuo Wang 0008, Wei Liu 0005, Yuqing Ma
Sci. China Inf. Sci.3
2023 ARCosmetics: a real-time augmented reality cosmetics try-on system
Shan An, Jianye Chen, Zhaoqi Zhu, Fangru Zhou, Yuxing Yang, Yuqing Ma, Xianglong Liu 0001, Haogang Zhu
Frontiers Comput. Sci.1
2023 SSA-ICL: Multi-domain adaptive attention with intra-dataset continual learning for Facial expression recognition
Hongxiang Gao, Min Wu 0008, Zhenghua Chen, Yuwen Li 0002, Xingyao Wang 0001, Shan An, Jianqing Li 0002, Chengyu Liu 0001
Neural Networks6
2023 Learning an insertion region for advertisement embedding on planes
Shan An, Fangru Zhou, Shuang Bai, Chang Tang, Haogang Zhu
Signal Process. Image Commun.1
2022 Attention-based Adversarial Partial Domain Adaptation
abstract
With the rapid development of vision-based deep learning (DL), it is an effective method to generate large-scale synthetic data to supplement real data to train the DL models for domain adaptation. However, previous vanilla domain adaptation methods generally assume the same label space, and such an assumption is no longer valid for a more realistic scenario where it requires adaptation from a larger and more diverse source domain to a smaller target domain with less number of classes. To handle this problem, we propose an attention-based adversarial partial domain adaptation (AADA). Specifically, we leverage adversarial domain adaptation to augment the target domain by using source domain, then we can readily turn this task into a vanilla domain adaptation. Meanwhile, to accurately focus on the transferable features, we apply attention-based method to train the adversarial networks to obtain better transferable semantic features. Experiments on four benchmarks demonstrate that the proposed method outperforms existing methods by a large margin, especially on the tough domain adaptation tasks, e.g. VisDA-2017.
Mengzhu Wang, Shan An, Xiao Luo 0001, Xiong Peng, Wei Yu 0029, Junyang Chen 0001, Zhigang Luo
ICASSP2
2022 Segmentation of ten fetal heart components with coarse-to-fine cascading and dynamic feature powering
abstract
Abstract Segmenting heart components in the apical four‐chamber view of fetal echocardiography is of critical significance in clinical practice. However, it is difficult to recognize these components due to small‐scale components and the imbalanced ventricular apex orientation. In this study, a novel segmentation framework is proposed to segment ten general fetal heart components for the first time. This framework consists of a multi‐directional fine‐density (MDFD) data augmentation method and a coarse‐to‐fine cascade network (CFCN). MDFD enhances the apex orientation diversity and balances the orientation distribution. CFCN has two stages including a coarse network and a fine network. These two stages have similar structures that consist of a feature extractor and a feature refined layer named as Element‐Wise Power with Dynamic Exponent layer (EWPDE). EWPDE which is a plug‐and‐play module for segmentation refines the features from the feature extractor to position small components accurately. By adopting EWPDE, the influence of each pixel is adjusted and hard pixels of small components are segmented precisely. Based on the dataset, the method is proved to be effective with the high mean intersection over union (mIoU) value and low missing ratio (MR). With MDFD and EWPDE, CFCN that adopts DeepLabV3+ as the feature extractor outperforms the best segmentation results (mIoU:0.480, MR:0.035). Compared to the original performance (mIoU:0.407, MR:0.085) of DeepLabV3+, the method improves the results significantly.
Tingyang Yang, Mengxiao Zhu 0002, Yan Wang 0076, Shan An, Jiancheng Han, Yihua He, Haogang Zhu
IET Image Process.5
2022 FastHand: Fast monocular hand pose estimation on embedded systems
Shan An, Xiajie Zhang, Haogang Zhu, Jianyu Yang 0002, Konstantinos A. Tsintotas
J. Syst. Archit.1
2021 Deep Balanced Learning for Long-tailed Facial Expressions Recognition
abstract
The analysis of facial expression is a very complex and challenging problem. Most researches for automated Facial Expression Recognition (FER) are mainly based on deep learning networks, rarely considering data imbalance. This paper commits to addressing the long-tail distribution problems among large-scale datasets in wild. Inspired by the continual learning method, we reconstruct multi-subsets first by randomly selecting from head classes and up-sampling tail classes. A pre-trained backbone is then introduced to learn general weights in a repeatedly train-prune fashion. Hereafter, our approach creatively trains a new classifier based on union parameters previously preserved and achieves an outperformance without extra parameters added in, using the gradual-prune technique. The results show that the independent training of classifiers has been a contributing factor. We successfully conduct this experiment with several classic networks, prove its effectiveness in training a deep network on imbalanced dataset. In the face of the poor performance in current FER, we find that domain knowledge is somehow affecting the accuracy of recognition by further exploring the obstacles from the image itself.Code available at https://github.com/Epicghx/FER
Hongxiang Gao, Shan An, Jianqing Li 0002, Chengyu Liu 0001
ICRA2
2021 Vanishing Point Aided LiDAR-Visual-Inertial Estimator
abstract
In this paper, we propose a vanishing point aided LiDAR-Visual-Inertial estimator to achieve real-time, low-drift and robust pose estimation. The proposed method is mainly composed of 3 sequential modules, namely IMU-aided vanishing point (VP) detection module, voxel-map based feature depth association module, and visual inertial fixed-lag smoother module. The IMU-aided VP detection module will detect feature points, line segments and vanishing points to establish robust correspondences in successive frames. In particular, we propose to use 1-line RANSAC method to provide stable VP hypotheses and polar grid to accelerate vanishing point hypothesis validation. After that, we propose a novel voxel-map based feature depth association method, to retrieve depth and assign depth to visual feature efficiently. Finally, the visual inertial fixed-lag smoother module is proposed to jointly minimize error terms. Experiments show that our method outperforms the state-of-the-art visual-inertial odometry and LiDAR-visual estimator in both indoor and outdoor environments.
Zheng Fang 0001, Shibo Zhao, Yongnan Chen, Shan An
ICRA6
2021 Real-Time Monocular Human Depth Estimation and Segmentation on Embedded Systems
abstract
Estimating a scene’s depth to achieve collision avoidance against moving pedestrians is a crucial and fundamental problem in the robotic field. This paper proposes a novel, low complexity network architecture for fast and accurate human depth estimation and segmentation in indoor environments, aiming to applications for resource-constrained platforms (including battery-powered aerial, micro-aerial, and ground vehicles) with a monocular camera being the primary perception module. Following the encoder-decoder structure, the proposed framework consists of two branches, one for depth prediction and another for semantic segmentation. Moreover, network structure optimization is employed to improve its forward inference speed. Exhaustive experiments on three self-generated datasets prove our pipeline’s capability to execute in real-time, achieving higher frame rates than contemporary state-of-the-art frameworks (114.6 frames per second on an NVIDIA Jetson Nano GPU with TensorRT) while maintaining comparable accuracy.
Shan An, Fangru Zhou, Haogang Zhu, Changhong Fu 0001, Konstantinos A. Tsintotas
IROS1
2021 ARShoe: Real-Time Augmented Reality Shoe Try-on System on Smartphones
abstract
Virtual try-on technology enables users to try various fashion items using augmented reality and provides a convenient online shopping experience. However, most previous works focus on the virtual try-on for clothes while neglecting that for shoes, which is also a promising task. To this concern, this work proposes a real-time augmented reality virtual shoe try-on system for smartphones, namely ARShoe. Specifically, ARShoe adopts a novel multi-branch network to realize pose estimation and segmentation simultaneously. A solution to generate realistic 3D shoe model occlusion during the try-on process is presented. To achieve a smooth and stable try-on effect, this work further develop a novel stabilization method. Moreover, for training and evaluation, we construct the very first large-scale foot benchmark with multiple virtual shoe try-on task-related labels annotated. Exhaustive experiments on our newly constructed benchmark demonstrate the satisfying performance of ARShoe. Practical tests on common smartphones validate the real-time performance and stabilization of the proposed approach.
Shan An, Guangfu Che, Jinghao Guo, Haogang Zhu, Junjie Ye 0004, Fangru Zhou, Zhaoqi Zhu, Aishan Liu, Wei Zhang 0031
ACM Multimedia1
2021 BR$^2$Net: Defocus Blur Detection Via a Bidirectional Channel Attention Residual Refining Network
abstract
Due to the remarkable potential applications, defocus blur detection, which aims to separate blurry regions from an image, has attracted much attention. Although significant progress has been made by many methods, there are still various challenges that hinder the results, e.g., confusing background areas, sensitivity to the scale and missing the boundary details of the defocus blur regions. To solve these issues, in this paper, we propose a deep convolutional neural network (CNN) for defocus blur detection via a Bi-directional Residual Refining network (BR2Net). Specifically, a residual learning and refining module (RLRM) is designed to correct the prediction errors in the intermediate defocus blur map. Then, we develop a bidirectional residual feature refining network with two branches by embedding multiple RLRMs into it for recurrently combining and refining the residual features. One branch of the network refines the residual features from the shallow layers to the deep layers, and the other branch refines the residual features from the deep layers to the shallow layers. In such a manner, both the low-level spatial details and highlevel semantic information can be encoded step by step in two directions to suppress background clutter and enhance the detected region details. The outputs of the two branches are fused to generate the final results. In addition, with the observation that different feature channels have different extents of discrimination for detecting blurred regions, we add a channel attention module to each feature extraction layer to select more discriminative features for residual learning. To promote further research on defocus blur detection, we create a new dataset with various challenging images and manually annotate their corresponding pixelwise ground truths. The proposed network is validated on two commonly used defocus blur detection datasets and our newly collected dataset by comparing it with 10 other state-of-the-art methods. Extensive experiments with ablation studies demonstrate that BR2Net consistently and significantly outperforms the competitors in terms of both the efficiency and accuracy.
Chang Tang, Xinwang Liu 0002, Shan An, Pichao Wang
IEEE Trans. Multim.3
2021 Partial tracking method based on siamese network
Chuanhao Li 0001, Shukuan Lin, Jianzhong Qiao, Shan An
Vis. Comput.4
2020 Transductive Relation-Propagation Network for Few-shot Learning
abstract
Few-shot learning, aiming to learn novel concepts from few labeled examples, is an interesting and very challenging problem with many practical advantages. To accomplish this task, one should concentrate on revealing the accurate relations of the support-query pairs. We propose a transductive relation-propagation graph neural network (TRPN) to explicitly model and propagate such relations across support-query pairs. Our TRPN treats the relation of each support-query pair as a graph node, named relational node, and resorts to the known relations between support samples, including both intra-class commonality and inter-class uniqueness, to guide the relation propagation in the graph, generating the discriminative relation embeddings for support-query pairs. A pseudo relational node is further introduced to propagate the query characteristics, and a fast, yet effective transductive learning strategy is devised to fully exploit the relation information among different queries. To the best of our knowledge, this is the first work that explicitly takes the relations of support-query pairs into consideration in few-shot learning, which might offer a new way to solve the few-shot learning problem. Extensive experiments conducted on several benchmark datasets demonstrate that our method can significantly outperform a variety of state-of-the-art few-shot learning methods.
Yuqing Ma, Shihao Bai, Shan An, Wei Liu 0005, Aishan Liu, Xiantong Zhen, Xianglong Liu 0001
IJCAI3
2019 Quarter-Point Codeword Expansion for Product Quantization
abstract
Due to its low storage cost and high query accuracy, Product Quantization (PQ) has been widely used for approximate nearest neighbor (ANN) search. However, almost all existing PQ-based methods use the nearest clustering center as the codeword, which might not fully utilize the information of distances from data points to clustering centers. In this paper, we propose a novel codeword expansion method for PQ-based methods, called Quarter-point Codeword Expansion (QCE), by estimating the distances from the query points to the database points using the quarter points instead of the clustering centers. The distances can be computed more precisely and it will result in a lower distortion using QCE, which is also a general method could be used to improve all PQ-based methods. Extensive experiments on approximate nearest neighbor search show that PQ-based methods with QCE can outperform the state-of-the-art.
Shan An, Zhibiao Huang, Guangfu Che, Xianglong Liu 0001, Xin Ma 0001
ICME1
2019 Fast and Incremental Loop Closure Detection Using Proximity Graphs
abstract
Visual loop closure detection, which can be considered as an image retrieval task, is an important problem in SLAM (Simultaneous Localization and Mapping) systems. The frequently used bag-of-words (BoW) models can achieve high precision and moderate recall. However, the requirement for lower time costs and fewer memory costs for mobile robot applications is not well satisfied. In this paper, we propose a novel loop closure detection framework titled FILD' (Fast and Incremental Loop closure Detection), which focuses on an on-line and incremental graph vocabulary construction for fast loop closure detection. The global and local features of frames are extracted using the Convolutional Neural Networks (CNN) and SURF on the GPU, which guarantee extremely fast extraction speeds. The graph vocabulary construction is based on one type of proximity graph, named Hierarchical Navigable Small World (HNSW) graphs, which is modified to adapt to this specific application. In addition, this process is coupled with a novel strategy for real-time geometrical verification, which only keeps binary hash codes and significantly saves on memory usage. Extensive experiments on several publicly available datasets show that the proposed approach can achieve fairly good recall at 100% precision compared to other state-of-the-art methods. The source code can be downloaded at https://github.comlAnshanTJU/FILD for further studies.
Shan An, Guangfu Che, Fangru Zhou, Xianglong Liu 0001, Xin Ma 0001
IROS1
2019 Coordinate CNNs and LSTMs to categorize scene images with multi-views and multi-levels of abstraction
Shuang Bai, Huadong Tang, Shan An
Expert Syst. Appl.3
2019 Quarter-Point Product Quantization for approximate nearest neighbor search
Shan An, Zhibiao Huang, Shuang Bai, Guangfu Che, Jie Luo 0004
Pattern Recognit. Lett.1
2019 RotateView: A Video Composition System for Interactive Product Display
abstract
Product display in e-commerce commonly uses static images or videos. In this paper, we design a novel video composition system “RotateView” which displays real-world products videos on the mobile phone in an interactive and 3D-like way. The system estimates the rotation direction of products using optical flow, segments products using an improved version of online adaptive video object segmentation, and adjusts motion and color for composition. The rotation direction estimation methods are evaluated on a video dataset of 1500 rotated videos that are uploaded by sellers, and on our collected real-world video object segmentation (RVOS) dataset. On our RVOS dataset, experimental results show an improvement in segmentation over the state of the art. Moreover, the proposed algorithms for motion and color adjustment are evaluated using extensive user studies. The RVOS dataset can be downloaded viahttps://realvos.github.io/for further studies and code will be made available.
Shan An, Si Liu 0001, Zhibiao Huang, Guangfu Che, Qian Bao, Zhaoqi Zhu, Dennis Z. Weng
IEEE Trans. Multim.1
2018 A survey on automatic image caption generation
Shuang Bai, Shan An
Neurocomputing2