Songlin Du

dblp:154/7687 · DBLP profile ↗
← Back
64ranked-venue papers
14as first author
54since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 28 since 2021Artificial intelligence and machine learning · 25 · 4 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Toward Free-Form Local Feature Matching
abstract
Existing feature matching methods are strongly coupled to their pre-defined position priors. For instance, sparse matchers are coupled to keypoints, and semi-dense matchers are coupled to grids. The coupled position prior dictates the distribution of matching points and imposes inherent limitations on the matcher. Consequently, sparse matchers suffer from a reliance on keypoint repeatability, while semi-dense matchers lack texture-based precision. Our preliminary work RCM leverages the keypoint prior in the source image and the grid prior in the target image, ensuring texture-based precision with keypoints while eliminating reliance on repeatability. However, RCM still relies heavily on keypoints in the source image, inheriting limitations such as sparsity and poor distribution in challenging scenes. To address these challenges, we introduce RCM+, which presents a novel free-form matching paradigm. By combining a position-agnostic encoder with a parameter-free decoder, we decouple the matcher from any position prior. As a result, the free-form matcher can match arbitrary input positions in a zero-shot manner, including detected keypoints, lines, edges, grids of any resolution, user-specified points, and more. This paradigm offers exceptional flexibility, allowing users to select position priors based on scene properties without retraining. Thus, RCM+ can leverage the advantages of various position priors without over-relying on any single prior, avoiding limitations in specific scenarios. To better match multiple position priors, we propose the Balancer, which reconciles all input position priors to achieve a more favorable point distribution for downstream tasks. Additionally, we enhance the view switcher and conflict-free matching layer introduced in RCM, further improving matching quality. Comprehensive experiments demonstrate the excellent performance, efficiency, and flexibility of RCM+, underscoring its promising potential for applications.
Xiaoyong Lu, Songlin Du, Yaping Yan, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Robust unsupervised visual tracking via image-to-video identity knowledge transferring
Bin Kang, Zongyu Wang, Dong Liang 0008, Tianyu Ding, Songlin Du
Pattern Recognit.5
2026 HSENet: Hierarchical semantic-enriched network for multi-modal image fusion
Rui Ming, Songlin Du, Lianghua He, Guobao Xiao
Pattern Recognit.3
2026 Parallel consensus transformer for local feature matching
Xiaoyong Lu, Bin Kang, Songlin Du
Pattern Recognit.4
2026 Single-domain generalization for fastener detection via sample reconstruction and class-wise domain contrast
Shixiang Su, Songlin Du, Xiaobo Lu
Pattern Recognit.2
2026 SceneGlue: Scene-Aware Transformer for Feature Matching Without Scene-Level Annotation
abstract
Local feature matching plays a critical role in understanding the correspondence between cross-view images. However, traditional methods are constrained by the inherent local nature of feature descriptors, limiting their ability to capture non-local scene information that is essential for accurate cross-view correspondence. In this paper, we introduce SceneGlue, a scene-aware feature matching framework designed to overcome these limitations. SceneGlue leverages a hybridizable matching paradigm that integrates implicit parallel attention and explicit cross-view visibility estimation. The parallel attention mechanism simultaneously exchanges information among local descriptors within and across images, enhancing the scene’s global context. To further enrich the scene awareness, we propose the Visibility Transformer, which explicitly categorizes features into visible and invisible regions, providing an understanding of cross-view scene visibility. By combining explicit and implicit scene-level awareness, SceneGlue effectively compensates for the local descriptor constraints. Notably, SceneGlue is trained using only local feature matches, without requiring scene-level groundtruth annotations. This scene-aware approach not only improves accuracy and robustness but also enhances interpretability compared to traditional methods. Extensive experiments on applications such as homography estimation, pose estimation, image matching, and visual localization validate SceneGlues superior performance. The source code is available at https://github.com/songlindu/ SceneGlue.
Songlin Du, Xiaoyong Lu, Yaping Yan, Guobao Xiao, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.1
2026 Spatially Aware Adaptive Diffusion: Unifying Low-Resolution Image Fusion and Super-Resolution
abstract
Low-resolution visible-infrared image fusion and super-resolution (LRVIF) are critical for enhancing image quality in low-resolution scenarios, yet limited information in the input images often constrains performance. To address these challenges, we propose SaDiff, a spatially-aware adaptive diffusion model that introduces diffusion processes into LRVIF for the first time, representing a major breakthrough in the field. Leveraging the generative capabilities of diffusion models, our approach unifies and enhances image fusion and super-resolution within a cohesive framework. A key component of SaDiff is the Spatial Residual Adaptation Block, which extends the diffusion process by dynamically adapting feature representations to spatial variations in the local regions of the input images. This module maximally preserves crucial information from the input images, such as texture details and contrast, while effectively suppressing noise, ensuring robust and context-aware feature refinement. Then we further propose Direct Diffusion Synthesis, a novel mechanism that utilizes noise predictions during diffusion to generate fused images, enabling joint training of the fusion and super-resolution networks. Additionally, a Cross-Feature Fusion Module integrates texture and contrast details, producing super-resolution fused images with improved clarity and structural integrity. Extensive experiments show that SaDiff achieves state-of-the-art performance, offering a robust and unified solution to infrared-visible image fusion and super-resolution. The code for the proposed method will be made available at https://github.com/guobaoxiao/SaDiff.
Jiajia Fu, Zhenni Yu, Haosheng Chen 0001, Songlin Du, Changcai Yang, Lianghua He, Guobao Xiao
IEEE Trans. Circuits Syst. Video Technol.4
2026 A Unified fNIRS Classification Framework Informed by Local Brain Activation Patterns
abstract
Functional near-infrared spectroscopy (fNIRS) is a promising noninvasive neuroimaging technique that detects cerebral hemodynamic responses in brain–computer interfaces. Recent studies have focused on task-specific and neuroscience-agnostic fNIRS classification models rather than a unified neuroscience-informed framework. We propose LoBrAFrame, a unified, neuroscience-informed fNIRS classification framework that leverages local brain activation patterns through a shared brain activation (SBA) module with a shared weight mechanism. SBA encodes shared activation dependencies into knowledge-aware hypersignals. A lightweight shared branch captures global spatio-temporal activation dependencies. Within this framework, researchers can easily enhance classification performance using simple or off-the-shelf methods on hypersignals, without redesigning complex models. To instantiate a concrete model, we introduce Mamba, a state space model, into the fNIRS domain and propose LoBrAMamba to learn activation patterns from hypersignals. Subject-specific and subject-independent experiments demonstrate the generalizability of LoBrAFrame and the superiority of LoBrAMamba on three open-access datasets. Our work will inspire interest in neuroscience-informed fNIRS frameworks.
Zenghui Wang 0009, Songlin Du
IEEE Trans. Ind. Informatics2
2026 Revisiting Semantic Correspondence: When Feature Aggregation Hurts Structural Integrity
abstract
Semantic correspondence seeks to establish matches between different instances of the same category. A common paradigm for this task leverages high-quality features from stable diffusion (SD) and DINOv2. However, we identify a widely overlooked yet critical issue: common feature aggregation disrupts the structural integrity of SD features, degrading semantic matching performance. We revisit and analyze this phenomenon and propose structure-aware aggregation (SAA) for SD features as a direct replacement for common feature aggregation methods. SAA uses filtering to decompose SD features into fine texture details and coarse contour structures. It aggregates only the texture components while preserving the contours. This divide-and-conquer mechanism enables SAA to significantly enhance the performance of state-of-the-art semantic correspondence models without increasing trainable parameters or computational overhead. Extensive qualitative and quantitative experiments confirm our analysis and validate the effectiveness of SAA. Moreover, SAA generalizes well to geometric, cross-species, and cross-family semantic correspondence tasks. Code is available at https://github.com/wzhlearning/SAA.
Zenghui Wang 0009, Songlin Du, Xiaobo Lu, Guobao Xiao
IEEE Trans. Image Process.2
2026 Simpler is Better: Feature Guard and Interaction for Semantic Correspondence
abstract
Semantic correspondence establishes keypoint correspondences between different instances of the same category. Fusing texture and semantic features from vision foundation models like stable diffusion (SD) and DINO significantly improves matching performance. However, we found an unnoticed yet essential problem: current feature fusion enhances the edge and semantic information in SD features with fine textures and DINOv2 features with fine semantics, but it destroys the semantic and structural information in SD features with weak and coarse semantics. We propose guard features (GuFT), a simple yet efficient method, to prevent feature degradation. Moreover, matching methods designed for traditional deep neural networks can be simplified based on two key insights: 1) vision foundation models provide rich visual knowledge; and 2) GuFT yields high-quality feature descriptors. We propose a bottleneck-style non-shared aggregation and backward interaction (NABI) module to efficiently capture intra- and inter-feature relationships, instead of common self- and cross-attention. The resulting framework, SimBetter, embodies a "simpler is better" design philosophy. It achieves state-of-the-art results with lower computation on SPair-71k, AP-10K, and PF-PASCAL, excelling in geometry-aware, cross-species, cross-family, and cross-dataset tasks. SimBetter also shows excellent potential in the applications of image-video semantic correspondence and sticker editing. Code is available at https://github.com/wzhlearning/SimBetter.
Zenghui Wang 0009, Songlin Du, Guobao Xiao
IEEE Trans. Image Process.2
2026 Progressive disentanglement for robust image anomaly detection
Yaping Yan, Yuanrui Zeng, Songlin Du
Vis. Comput.5
2025 JamMa: Ultra-lightweight Local Feature Matching with Joint Mamba
abstract
Existing state-of-the-art feature matchers capture long-range dependencies with Transformers but are hindered by high spatial complexity, leading to demanding training and high-latency inference. Striking a better balance between performance and efficiency remains a challenge in feature matching. Inspired by the linear complexity $\mathcal{O}(N)$ of Mamba, we propose an ultra-lightweight Mamba-based matcher, named JamMa, which converges on a single GPU and achieves an impressive performance-efficiency balance in inference. To unlock the potential of Mamba for feature matching, we propose Joint Mamba with a scan-merge strategy named JEGO, which enables: (1) Joint scan of two images to achieve high-frequency mutual interaction, (2) Efficient scan with skip steps to reduce sequence length, (3) Global receptive field, and (4) Omnidirectional feature representation. With the above properties, the JEGO strategy significantly outperforms the scan-merge strategies proposed in VMamba and EVMamba in the feature matching task. Compared to attention-based sparse and semi-dense matchers, JamMa demonstrates a superior balance between performance and efficiency, delivering better performance with less than 50% of the parameters and FLOPs. Project page: https://leoluxxx.github.io/JamMa-page/.
Xiaoyong Lu, Songlin Du
CVPR2
2025 Label distribution learning with structured manifold subspace
Yaping Yan, Yunlong Tang 0004, Yongxin Jiang, Songlin Du
Neurocomputing4
2025 SemMatcher: Semantic-aware feature matching with neighborhood consensus
Qimin Jiang, Xiaoyong Lu, Dong Liang 0008, Songlin Du
J. Vis. Commun. Image Represent.4
2025 Binary Banyan tree growth optimization: A practical approach to high-dimensional feature selection
abstract
High-dimensional feature spaces in Scientific and Technical Service Resources (STSR) classification present significant challenges, including increased computational costs and diminished accuracy. Identifying an optimal subset of features from raw text vectors is thus critical for effective data classification . This paper introduces a novel metaheuristic algorithm called Binary Banyan Tree Growth Optimization (BBTGO), specifically designed for high-dimensional feature selection (FS). Inspired by the unique growth patterns of the banyan tree , BBTGO leverages a combination of innovative Boolean vectors, including rooting, multi-trunk, and adjustment operator, along with a perturbation phase to enhance the search efficiency and reduce feature dimensionality. These operators enhance the search for promising regions and reduce features by utilizing the optimal solutions clustered within subgroups. Furthermore, BBTGO incorporates a dynamic adjustment mechanism that periodically activates different growth operators to meet the search demands of high-dimensional space. We rigorously evaluate the exploration and exploitation capabilities of BBTGO through comprehensive statistical analyses of various performance metrics. The proposed method demonstrates superior results on 12 high-dimensional benchmark datasets and is successfully applied to feature selection in STSR text classification tasks . Experimental results show that BBTGO significantly outperforms existing methods in terms of classification accuracy , selected features, convergence speed, and processing time. These results underscore the potential of BBTGO as a robust and versatile solution for high-dimensional FS, with broad applicability to real-world classification challenges.
Minrui Fei, Wenju Zhou, Songlin Du, Zixiang Fei, Huiyu Zhou 0001
Knowl. Based Syst.4
2025 Skeleton-Aware Representation of Spatio-Temporal Kinematics for 3D Human Motion Prediction
abstract
3D human motion prediction, which attempts to foresee the behaviors of human, is an issue of great significance in computer vision. Attention-based neural networks and graph convolution networks (GCNs) have recently shown great promise in 3D skeleton-based human motion prediction for their attractive performance in learning spatial and temporal kinematics. However, existing methods have several critical issues: 1) Spatial dependencies for distal joints in each independent frame are hard to learn; 2) The GCN ignores hierarchical structure and diverse motion patterns of different body parts; 3) Existing methods disregard the statistical interdependence inherent in time series data. To address these issues, this paper proposes a skeleton-aware representation of spatio-temporal kinematics for 3D human motion prediction. The proposed method makes three key contributions: a learnable temporal aggregation, a skeleton-aware spatio-temporal attention, and an upper/lower decoupling GCN. The learnable temporal aggregation selectively obtains past information by leveraging the dependencies between each time step and its historical moments. The skeleton-aware spatio-temporal attention method leverages the self-attention mechanism and a designed adjacency matrix to model the skeleton constraints of distal joints. The upper/lower decoupling GCN introduces a grouping strategy to learn the dynamics of various body parts separately. Experimental results on three publicly available datasets demonstrate that the proposed method achieves state-of-the-art performances for both short-term prediction and long-term prediction. Note to Practitioners—3D human motion prediction forms a fundamental component of human-centered automation systems by enabling safer, more efficient, and more natural interactions between humans and machines. This paper was motivated by the challenges of predicting human motion: 1) Explicitly capturing the complex spatial patterns of distal joints is challenging; 2) Neglecting the inter-part variations of motion dynamics is problematic; 3) Neglecting the statistical interdependence inherent in time series data of human motion leads to poor performance. This paper suggests a skeleton-aware representation of spatio-temporal kinematics for 3D human motion prediction through three innovations: a learnable temporal aggregation, a skeleton-aware spatio-temporal attention, and an upper/lower decoupling GCN. The three contributions overcome the weaknesses of existing works and made a pioneering attempt of skeleton-aware representation of spatio-temporal human kinematics. It will significatively advance the development of many automation systems relevant to human motion prediction such as human-robot interaction and teleoperation.
Songlin Du, Zhihan Zhuang, Zenghui Wang 0009, Yuan Li 0058, Takeshi Ikenaga
IEEE Trans Autom. Sci. Eng.1
2025 Tex2Sem: Learning From Textures to Semantics for Robust Semantic Correspondence
abstract
Recent advances in semantic correspondence have witnessed growing interest in vision foundation models, particularly stable diffusion (SD) and self-distillation with no labels (DINO). However, existing methods underutilize the matching potential of SD and DINOv2 features and show similar background interference patterns. They lack texture-to-semantic learning and intra- and inter-image feature interaction. This study proposes Tex2Sem, a framework learning from textures to semantics, to address the two problems. For the first problem, we propose a texture-to-semantic learning paradigm that achieves texture-semantic trade-offs on features and correlation maps, including progressive fusion and correlation map computation. The SD and DINOv2 features are aggregated from textures to semantics to produce multi-stage progressive fusion features. The resulting multi-stage progressive fusion correlation maps improve semantic correspondence significantly. For the second problem, MamFormer, a hybrid architecture of Mamba-2 and Transformer, is proposed to improve intra- and inter-image feature aggregation and interaction. It enhances foreground focus and background suppression. Given the high computational cost of processing all-stage progressive fusion features, the terminal-stage aggregation and interaction mechanism (TAIM) is proposed to enhance feature learning efficiency. Experiments demonstrate that Tex2Sem achieves state-of-the-art performance on SPair-71k, AP-10K, and PF-PASCAL. Furthermore, Tex2Sem shows remarkable generalization capabilities in cross-species, cross-family, and cross-dataset matching and demonstrates the potential for applications in video swap and human pose estimation. Code is available at https://github.com/wzhlearning/Tex2Sem.
Zenghui Wang 0009, Songlin Du, Yaping Yan, Guobao Xiao, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.2
2025 MambaMatch: Establishing Reliable Correspondences via Multi-Scale State Space Model
abstract
Correspondence pruning aims to identify inliers from correspondences severely disturbed by outliers. Although Transformers and graph neural networks have shown impressive results in this field, they are either limited by a narrow receptive field or encounter quadratic computational complexity. To tackle this challenge, this work pioneers the integration of state space model into correspondence pruning task, proposing a Mamba-based framework named MambaMatch. Specifically, to address the limitations of the Mamba architecture in local consensus modeling, we proposes a multi-scale scanning strategy. It first employs an adaptive clustering algorithm to map origin correspondences into spatially coherent feature clusters, constructing a dual-representation space encompassing both full-scale and clustered-scale features. Bidirectional scan operations are then performed at both scales: 1) full-scale scan preserves global structural context, and 2) clustered-scale scan enhances local consistency. Subsequently, a Multi-Scale Interaction layer is designed to dynamically fuse dual-scale features via a cross-attention mechanism, further integrated with a Gated Feed-Forward Network to significantly improve the network's feature discrimination capability. Extensive experiments validate that MambaMatch surpasses state-of-the-art approaches across multiple benchmarks for two-view geometry estimation. Furthermore, MambaMatch exhibits robust generalization across diverse scenarios, tasks, and feature extractors. The source code is available at: https://github.com/mxyttkx/MambaMatch.
Xiangyang Miao, Shunxing Chen, Shiping Wang, Songlin Du, Lianghua He, Guobao Xiao
IEEE Trans. Image Process.5
2025 Topology Learning for Two-View Correspondence Filtering
abstract
In this paper, we propose a novel neural network called Topology Learning Network (TL-Net), that exploits local and global geometric relation by topology graphs to handle the problem of correspondence filtering in complex scenes. Specifically, we first design a Multi-level Topology Encoder (MLTE), which fuses local and global topology graphs by a channel attention, to sufficiently extract the geometric relation among correspondences. MLTE not only includes local topology graphs by gathering the information of relative motion and multi-resolution group convolution, but also includes a global topology graph by aggregating the information of the similarity and the Graph Laplacian. In addition, inspired by Transformer, we design the backbone of TL-Net to generate enriched fdeature maps for correspondence filtering. Meanwhile, by simplifying the global context aggregation, we maintain the lightweight of the backbone, introducing the superiority of Transformer while avoiding extra parameters and calculations. Empirical experiments on several computer vision tasks show that the performance and generalization ability of TL-Net are significantly superior to the state of the art methods. Notably, on relative pose estimation, we achieve 5.63% and 5.03% mAP improvements under an error threshold of$5^{\circ }$outdoors and indoors, respectively.
Ziwei Shi, Xiangyang Miao, Guobao Xiao, Songlin Du, Zheng Wang 0044, Heng Tao Shen
IEEE Trans. Multim.4
2024 Contrastive Max-Correlation for Multi-view Clustering
Yanghao Deng, Zenghui Wang 0013, Songlin Du
ACCV (1)3
2024 Raising the Ceiling: Conflict-Free Local Feature Matching with Dynamic View Switching
Xiaoyong Lu, Songlin Du
ECCV (42)2
2024 Underlying-Complementarity and Surrounding-Correspondence for Multi-View Clustering
abstract
In this paper, we study the intrinsic connections between informative representation among different views and sufficient information among reconstructed views in multi-view clustering. To this end, we propose a novel method called underlying-complementarity and surrounding-correspondence for multi-view clustering (CCMC) including two goals: 1) The surrounding-correspondence is learned by domain correspondence of surrounding data points based on decoder regularization to capture the supplement structure information. 2) The underlying-complementarity is learned by pseudo-class-label with allocation matrix where the contrastive learning is applied to obtain the supervised information and the pyramid network auto-fuses multi-view information which is based on different latent layers of the different views. Compared to existing works, the advantages of the proposed method lie in enhancing the combinative as well as deep usage of complementary and domain correspondence information for better performance of clustering. CCMC is capable of guiding the network learning without intact and correct corresponding multi-view data. Extensive experiments demonstrate the promising performance even in the case of missing and unaligned data compared with state-of-the-art approaches on the multi-view clustering task.
Songlin Du
ICASSP2
2024 Beyond Global Cues: Unveiling the Power of Fine Details in Image Matching
abstract
To obtain feature descriptors, current detector-free feature matching algorithms usually leverage attention at a coarse level to model relationships between keypoints. Nevertheless, relying solely on global cues from other points would bring uncertainty in high-quality feature matching. Rethinking the matching process of humans, humans not only look back-and-forth but also reference the details surrounding keypoints for precise localization. Based on the above observations, a novel Transformer-based detector-free matcher, entitled FineFormer, is proposed. FineFormer not only aggregates global cues from keypoints but also takes local details around each keypoint into account. Extensive experiments across three datasets demonstrate the superiority of our method in both efficiency and effectiveness against existing state-of-the-arts.
Songlin Du
ICME2
2024 Single-Domain Generalization Combining Geometric Context Toward Instance Segmentation of Track Components
abstract
A stable track components segmentation model should have consistent performance across a broad spectrum of railroad conditions, particularly in unfamiliar locations. Despite this, satisfying this requirement proves challenging when working with a limited track dataset, as there is a substantial domain shift between the given dataset and unobserved distributions. The goal of this paper is to improve the generalization ability of track component segmentation in situations where single-domain training data is available. Toward this end, a novel track component instance segmentation method combining the geometric context is proposed. First, we design an initial mask prediction head (IMPH) that utilizes predicted box output from the object detector to generate initial masks by merging the geometric priors of track components. Meanwhile, a multiscale feature fusion structure is introduced to encourage IMPH to better capture the geometric context. Then, a final mask refinement head (FMRH) is introduced to get higher quality masks with a geometry-gated aggregation strategy. On the basis of experiments on track datasets, it was determined that the heads can be incorporated with several types of detection frameworks and have demonstrated consistent generalization enhancements across multiple object detectors. Furthermore, our method substantially enhances segmentation performance on various unobserved domains.
Shixiang Su, Songlin Du, Dezhou Wang, Shuzhen Tong, Xiaobo Lu
IJCNN2
2024 Kinematics-aware spatial-temporal feature transform for 3D human pose estimation
Songlin Du, Zhiwei Yuan, Takeshi Ikenaga
Pattern Recognit.1
2024 JoyPose: Jointly learning evolutionary data augmentation and anatomy-aware global-local representation for 3D human pose estimation
Songlin Du, Zhiwei Yuan, Peifu Lai, Takeshi Ikenaga
Pattern Recognit.1
2024 AnatPose: Bidirectionally learning anatomy-aware heatmaps for human pose estimation
Songlin Du, Takeshi Ikenaga
Pattern Recognit.1
2024 Fine-grained recognition via submodular optimization regulated progressive training
Bin Kang, Songlin Du, Dong Liang 0008, Xin Li 0086
Pattern Recognit.2
2024 Bi-Pose: Bidirectional 2D-3D Transformation for Human Pose Estimation From a Monocular Camera
abstract
Automatically estimating 3D human poses in video and inferring their meanings play an essential role in many human-centered automation systems. Existing researches made remarkable progresses by first estimating 2D human joints in video and then reconstructing 3D human pose from the 2D joints. However, mono-directionally reconstructing 3D pose from 2D joints ignores the interaction between information in 3D space and 2D space, losses rich information of original video, therefore limits the ceiling of estimation accuracy. To this end, this paper proposes a bidirectional 2D-3D transformation framework that bidirectionally exchanges 2D and 3D information and utilizes video information to estimate an offset for refining 3D human pose. In addition, a bone-length stability loss is utilized for the purpose of exploring human body structure to make the estimated 3D pose more natural and to further increase the overall accuracy. By evaluation, estimation error of the proposed method, measured by the mean per joint position error (MPJPE), is only 46.5 mm, which is much lower than state-of-the-art methods under the same experimental condition. The improvement on accuracy will make machines to better understand human poses for building superior human-centered automation systems.Note to Practitioners—This paper was motivated by the demand of human-centered automation systems needing to accurately understand human poses. Existing approaches mainly focus on inferring 3D human pose from 2D joints mono-directionally. Although they made remarkable contributions to estimating 3D human pose in such a mono-directional way, we found that they ignore the 2D-3D interaction and do not use original video when inferring 3D pose from 2D joints. This paper therefore suggests a bidirectional 2D-3D transformation that exchanges 2D and 3D information and utilizes video information to estimate more accurate 3D human pose for human-centered automation systems. This work is a pioneering attempt of interactively using 2D and 3D information for more accurate estimation of human pose. Benefited from the state-of-the-art accuracy, the proposed approach is expected to make significant contributions to many human-centered automation systems, such as human-machine interaction, biomimetic manipulation, and automatic surveillance systems.
Songlin Du, Zhiwei Yuan, Takeshi Ikenaga
IEEE Trans Autom. Sci. Eng.1
2024 ContextMatcher: Detector-Free Feature Matching With Cross-Modality Context
abstract
Existing feature matching methods tend to extract feature descriptors by relying on the visual appearance, leading to false matches which are obviously false from the geometric perspective. This paper proposes ContextMatcher, which goes beyond the visual appearance representation by introducing the geometric context to guild the feature matching. Specifically, our ContextMatcher includes visual descriptors generation, the neighborhood consensus module, and the geometric context encoder. To learn visual descriptors, Transformers situated in different branches are leveraged to obtain feature descriptors. In one branch, convolutions are integrated into self-attention layers elegantly to compensate for the lack of the local structure information. In another branch, a cross-scale Transformer is proposed through injecting heterogeneous receptive field sizes into tokens. To leverage and aggregate the geometric contextual information, a neighborhood consensus mechanism is proposed by re-ranking initial pixel-level matches to make a constraint of geometric consensus on neighborhood feature descriptors. Moreover, local feature descriptors are boosted through combining with the geometric properties of keypoints for refining matches to the sub-pixel level. Extensive experiments on relative pose estimations and image matching show that our proposed method outperforms existing state-of-the-art methods by a large margin.
Songlin Du
IEEE Trans. Circuits Syst. Video Technol.2
2023 ParaFormer: Parallel Attention Transformer for Efficient Feature Matching
abstract
Heavy computation is a bottleneck limiting deep-learning-based feature matching algorithms to be applied in many real-time applications. However, existing lightweight networks optimized for Euclidean data cannot address classical feature matching tasks, since sparse keypoint based descriptors are expected to be matched. This paper tackles this problem and proposes two concepts: 1) a novel parallel attention model entitled ParaFormer and 2) a graph based U-Net architecture with attentional pooling. First, ParaFormer fuses features and keypoint positions through the concept of amplitude and phase, and integrates self- and cross-attention in a parallel manner which achieves a win-win performance in terms of accuracy and efficiency. Second, with U-Net architecture and proposed attentional pooling, the ParaFormer-U variant significantly reduces computational complexity, and minimize performance loss caused by downsampling. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that ParaFormer achieves state-of-the-art performance while maintaining high efficiency. The efficient ParaFormer-U variant achieves comparable performance with less than 50% FLOPs of the existing attention-based models.
Xiaoyong Lu, Yaping Yan, Bin Kang, Songlin Du
AAAI4
2023 MSFORMER: Multi-Scale Transformer with Neighborhood Consensus for Feature Matching
abstract
Existing feature matching methods tend to extract feature descriptors by feeding down-sampled feature maps into a Transformer that is unable to extend feature scales, leading to false correspondences between small-size objects. This paper proposes MSFormer, which uses Transformers situated in different branches to obtain feature descriptors. In one branch, convolutions are integrated into self-attention layers elegantly to compensate for the lack of the local structure information. In another branch, a multi-scale Transformer is proposed through injecting heterogeneous receptive field sizes into tokens. Additionally, a neighborhood consensus mechanism is proposed by re-ranking initial matches to make a constraint of geometric consensus on neighborhood feature descriptors. Extensive experiments on indoor and outdoor pose estimations show that MSFormer outperforms existing state-of-the- art methods by a large margin.
Yaping Yan, Dong Liang 0008, Songlin Du
ICASSP4
2023 CFFMixer: Multi-Dimensional Feature Fusion for Object Detection
abstract
Object detection is a fundamental task in the field of computer vision, and one of its essential requirements is high-quality feature fusion. Previous works have made various efforts in this regard: CNN-based detectors use convolutional blocks to fuse local features and dense prior knowledge to predict objects, while query-based detectors fuse global features by self-attention then decode features with object queries. However, their feature fusion methods are relatively monotonous. Considering that different modules are applicable to different dimensions, we proposed an object detector named CFFMixer which used hybrid architecture to achieve multi-dimensional feature fusion. The sampling strategy to extract abundant local and global features was first introduced then the Comprehensive Feature Fusion Network (CFFN) was proposed to integrate them. CFFN not only achieved local and global features interaction in the spatial dimension, but also fused semantics in the channel dimension. Furthermore, we conducted experiments and made a comparison with competitive models, our model finally got 43.0 mAP on COCO 2017 dataset within 12 epochs. Experimental results showed that the model’s accuracy benefits from the powerful feature fusion capability of CFFN. Besides, we performed ablation studies on our modules to evaluate their effectiveness.
Weizhe Yuan, Bin Kang, Songlin Du
ICASSP4
2023 Scene-Aware Feature Matching
abstract
Current feature matching methods focus on point-level matching, pursuing better representation learning of individual features, but lacking further understanding of the scene. This results in significant performance degradation when handling challenging scenes such as scenes with large viewpoint and illumination changes. To tackle this problem, we propose a novel model named SAM, which applies attentional grouping to guide Scene-Aware feature Matching. SAM handles multi-level features, i.e., image tokens and group tokens, with attention layers, and groups the image tokens with the proposed token grouping module. Our model can be trained by ground-truth matches only and produce reasonable grouping results. With the sense-aware grouping guidance, SAM is not only more accurate and robust but also more interpretable than conventional feature matching models. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that our model achieves state-of-the-art performance.
Xiaoyong Lu, Yaping Yan, Songlin Du
ICCV4
2023 DNC-Net:Dual-neighbourhood Consensus Network for Feature Matching
abstract
Local feature matching is the core of many computer vision tasks. While dectetor-based methods lack repeatability and only consider small image regions, sparse feature matching methods is challenging in textureless scenes. This paper proposes a novel feature matching method based on dual-neighbourhood consensus, termed as DNC-Net. In DNC-Net, we build the connection between keypoints and feature maps, so that the matching can comprehensively consider the small and large image regions. Further, dual-neighbourhood consensus fully learn the consensus between keypoints and feature maps, which improve matching robustness especially in textureless scenes. Experiments on homography estimation, outdoor pose estimation and image matching show that the model is superior to other methods, and has achieved the most advanced results.
Qimin Jiang, Songlin Du
ICIP2
2023 Robust RGB-T Tracking via Consistency Regulated Scene Perception
abstract
RGB-T tracking has received increasing attention due to its significant advantage under severe weather conditions. Existing RGB-T tracking methods pay close attention to the representation of target appearance, ignoring the importance of scene information. In this paper, we propose a global reasoning-oriented method for RGB-T tracking. In particular, within a multi-task learning framework, our approach adopts a nested global reasoning model to regulate the consistency of scene perception (reasoning the relation between targets and the surrounding semantic regions) in different image domains. Moreover, a meta-unsupervised learning strategy is designed to enforce the nested global reasoning model to utilize partial multi-domain target information for the updating of scene perception. Extensive experiments on GTOT, RGBT210 and LasHeR datasets show the superior performance of our method when compared with related works.
Bin Kang, Songlin Du
ICIP4
2023 A knowledge-driven monarch butterfly optimization algorithm with self-learning mechanism
Tianpeng Xu, Fuqing Zhao, Jianxin Tang, Songlin Du, Jonrinaldi
Appl. Intell.4
2023 An effective discrete monarch butterfly optimization algorithm for distributed blocking flow shop scheduling with an assembly machine
Songlin Du, Wenju Zhou, Dakui Wu, Minrui Fei
Expert Syst. Appl.1
2023 UAV image stitching by estimating orthograph with RGB cameras
abstract
In the field of image stitching, cases with large camera optical center movement and large parallax have been the virgin territory of research. The goal of image stitching is to overcome the parallax and stitch a natural image. We look into this problem in the context of ultra-low altitude flight of a UAV . We model the 3D world in this scenario and quickly estimate orthographic projection by pairs of homography matrices. Our stitching method can achieve precise alignment since it takes parallax well into consideration. The stitching results are natural and the extra time consumed is short.
Wenxiao Cai, Songlin Du, Wankou Yang
J. Vis. Commun. Image Represent.2
2023 Enhanced Binary Black Hole algorithm for text feature selection on resources classification
Minrui Fei, Dakui Wu, Wenju Zhou, Songlin Du, Zixiang Fei
Knowl. Based Syst.5
2023 Bi-SCM: bidirectional spiking cortical model with adaptive unsharp masking for mammography image enhancement
Yaping Yan, Hongjuan Zhang, Songlin Du, Yide Ma
Multim. Tools Appl.3
2023 RFS-Net: Railway Track Fastener Segmentation Network With Shape Guidance
abstract
The fastener is one of the main components of a rail track system. In recent years, deep learning methods such as image segmentation have greatly boosted the fastener state detection process. However, there is still a need to improve the segmentation accuracy and speed, especially for the fasteners in complex environments. To handle this problem, a fast and accurate fastener semantic segmentation network named RFS-Net is proposed based on shape guidance, which can offer a better speed/accuracy trade-off performance via a very shallow architecture. Specifically, in the encoder, a two-stream structure (i.e., regular stream and shape stream) that processes the fastener and shape image in parallel is introduced. The shape image is created based on the geometric structure of the fastener, and it is served as input to the shape stream to guide the segmentation of the fastener. The decoder integrates deep features from the two-stream encoder and then recovers the shape information by the shape attention blocks with skipping connections. We provide two versions of RFS-Net: RFS-Net_S (1.0M, 1014FPS) and RFS-Net_L (12.01M, 453FPS) on the NVIDIA RTX 3060. Experimental results demonstrate the effectiveness of our method by achieving a promising trade-off between accuracy and inference speed. In particular, our method is faster and more accurate on a challenging dataset, from fast modes: 1014 FPS for RFS-Net_S versus 724 FPS for Segmenter, to high-quality segmentation: better performance than STDC with nearly one percent (92.36% versus 91.48% Mean IoU score).
Shixiang Su, Songlin Du, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Straight-Line Detection Within 1 Millisecond Per Frame for Ultrahigh-Speed Industrial Automation
abstract
Detecting straight lines in video plays a fundamental role in camera-based industrial automation. With the increasing demands on production efficiency, detection speed has become one of the bottlenecks for highly efficient industrial automation. Because of data dependence and hardware limitations, existing vision systems based on central processing unit/graphics processing unit are unable to detect straight lines at an ultrahigh speed. This article addresses this problem and proposes a hardware-friendly Hough transform that can be implemented in fully parallel for the ultrahigh-speed detection, because of the following two key features: it processes multiple pixels in parallel and directly calculates line parameters while capturing the current frame; and it simultaneously initializes the Hough parameter space and votes in the Hough parameter space without any delay. Based on the proposed hardware-friendly Hough transform, its chip-level implementation and system-level hardware design are presented. Experimental results show that the main benefits of the proposed architecture are in real-time performances at a high frame rate (784 frames/s) and an ultralow delay (0.7749 ms/frame).
Songlin Du, Ziwei Dong, Takeshi Ikenaga
IEEE Trans. Ind. Informatics1
2022 JointFusionNet: Parallel Learning Human Structural Local and Global Joint Features for 3D Human Pose Estimation
Zhiwei Yuan, Yaping Yan, Songlin Du, Takeshi Ikenaga
ICANN (4)3
2022 NCTR: Neighborhood Consensus Transformer for Feature Matching
abstract
This paper presents NCTR, a feature matcher that enhances input descriptors and finds the correspondences between them. In NCTR, Transformer is applied to aggregate global context for each descriptor. To solve the lack of neighborhood consensus that Transformer may bring, we propose a novel method to evaluate the neighborhood consensus of each key-point and integrate it into the attentional aggregation. The combination of global and local information greatly enhances the model capability and improves the match quality. The experiments on homography estimation and outdoor pose estimation show that NCTR outperforms other hand-designed or learning-based methods and achieves state-of-the-art results.
Xiaoyong Lu, Songlin Du
ICIP2
2022 More Than Accuracy: An Empirical Study of Consistency Between Performance and Interpretability
Dong Liang 0008, Rong Quan, Songlin Du, Yaping Yan
PRICAI (3)4
2022 Automatic Foreground Detection at 784 FPS for Ultra-High-Speed Human-Machine Interactions
abstract
Human-machine interactive systems show increasing demand for analysing fast moving objects in high-frame-rate videos. Robust foreground detection, which is able to reduce large amount of redundant background data from high-frame-rate video, becomes the essence to achieve ultra-high-speed human-machine interactions. This paper proposes a local spatial propagation based background model generation, a local linear illumination correction based background model update, and a regional central coordinates and edge keypoints constrained foreground region reselection. The three proposals make up a robust and hardware-friendly foreground detection method. Experimental results prove that the proposed hardware-friendly algorithm achieves high accuracy and robustness on various kinds of challenging cases. Meanwhile, the hardware implementation utilizes little hardware resources and achieves realtime processing of high-frame-rate (784 frame/second) video with the delay less than 1 ms/frame in image processing core. In addition, a practical system is implemented by combing a PC, a high-speed camera and a field programmable gate array (FPGA) for realworld applications. This work will significatively promote the development and application of high-speed human machine interaction. A demo of the proposed vision system working at 784 FPS is available athttps://wcms.waseda.jp/em/5f84f75136a6. Note to Practitioners—This paper was motivated by the problem of high-frame-rate video contains large amount of redundant background pixels which makes ultra-high-speed human-machine interactions inaccessible. Existing approaches are mainly focused on designing complex background models, but processing speed, which is the most important issue for ultra-high-speed human-machine interactions, has received relatively little attention. This paper suggests a robust and hardware-friendly foreground detection algorithm which has been implemented as a hardware system by using an FPGA, a high-frame-rate camera, and a PC. We show that the hardware implementation utilizes less hardware resources and achieves real-time processing speed of 784 FPS with the delay less than 1 ms/frame in the image processing core. This work is a pioneering attempt of ultra-high-speed foreground detection, which will significatively speed up the wide applications of ultra-high-speed human machine interactions.
Songlin Du, Peikun Cai, Takeshi Ikenaga
IEEE Trans Autom. Sci. Eng.1
2022 Highly-Parallel Hardwired Deep Convolutional Neural Network for 1-ms Dual-Hand Tracking
abstract
1-ms vision systems represent an extreme case of temporal development in video sensing techniques. Moreover, a 1-ms dual-hand tracking system leverages the dexterous functionality of hands and thus serves as a seamless and intuitive interface for Human-Computer Interaction. Deep CNN is promising for high tracking robustness, however, neither GPU-based nor FPGA-based implementation addresses the tracking task with ultra-high-speed. This paper proposes: (a) A paradigm to directly map a deep CNN as a hardwired circuit, so the entire network runs in parallel and high processing speed is obtained. The network is exempted from memory access since all intermediate neural values are implicitly represented in hardware states. And condensed binarization is used to reduce resource utilization; (b) Hardware design of the hardwired network on FPGA, inside which kernel-adapted convolutional trees are devised to maximize the parallelism. The speed bottleneck of the network is therefore removed by implementing convolutional layers as fine-grained pipelines with unified components; (c) FPGA-GPU hetero complementation, which utilizes an auxiliary GPU network to compensate for accuracy of the FPGA network without affecting its speed. The quick primary results on FPGA are intermittently refined using delayed but accurate hints from GPU. Implementation results show that the proposed method reaches 973fps and consumes merely 1.30ms to process on$640\times 480$images, while the accuracy is only 4.7% lower compared with the general method on test sequences. Video demonstrations are available athttps://wcms.waseda.jp/em/5f9d020f136e7.
Peiqi Zhang, Dingli Luo, Songlin Du, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.4
2022 Geometric Constraint and Image Inpainting-Based Railway Track Fastener Sample Generation for Improving Defect Inspection
abstract
Defective fastener images detection is an essential task in the vision-based railway track safety inspection. Although existing methods have achieved some level of success, the detection accuracy in this field suffers from the defective fasteners being far less common than normal fasteners. One way to tackle this problem is to expand the defect sample. However, current state-of-the-art defective fastener generation methods mainly rely on generative adversarial networks or simply augment the defect data through traditional image processing. These methods may not be ideal as it is difficult to produce images with high quality and rich diversity at the same time. This paper proposes a new method for fastener sample generation that actively divides the sample generation into two independent parts: defective foregrounds generation and complete backgrounds generation. The key to this method is to generate foregrounds and backgrounds based on geometric constraint and image inpainting, respectively. Specifically, we adopt a skeleton mapping algorithm to directionally control the generated types of defective foregrounds. Meanwhile, an image inpainting network is employed to expand the background. The experiments show that this enables us to generate better-quality and richer-diversity images by combining deep learning and image processing advantages. To the best of our knowledge, our method is the first to achieve state-of-the-art performance, i.e., the classification accuracy reaches 97.97%, without using real defective fastener images during the defect classification network training process.
Shixiang Su, Songlin Du, Xiaobo Lu
IEEE Trans. Intell. Transp. Syst.2
2021 An Algorithm Based on Monarch Butterfly Optimization with Learning Mechanism and Topological Structure
abstract
In the past decades, various attention has been paid to the global optimization problems. The Monarch Butterfly Optimization (MBO) algorithm is an effective meta-heuristic algorithm for the global optimization problems. However, in the MBO, the diversity of the population is lost in the late iteration. The MBO is easy to trap into the local optima. In this study, an algorithm based on MBO with learning mechanism and topological structure, named LTMBO, is proposed to enhance the ability of exploration and exploitation on the global optimization problems. The learning mechanism is present for the migration operator to increase the speed of the iteration. The topological structure is proposed for the butterfly adjusting operator to improve the diversity of the population. The experimental results demonstrated that the efficiency and significance of the proposed LTMBO algorithm.
Fuqing Zhao, Songlin Du, Jianxin Tang, Yi Zhang 0096, Weimin Ma
CSCWD2
2021 JointPose: Jointly Optimizing Evolutionary Data Augmentation and Prediction Neural Network for 3D Human Pose Estimation
Zhiwei Yuan, Songlin Du
ICANN (3)2
2021 A hybrid self-adaptive invasive weed algorithm with differential evolution
abstract
The invasive weed algorithm (IWO) is a meta-heuristic algorithm, which is an effective and promising optimiser to address the optimisation problems. In this study, a hybrid algorithm based on the self-adaptive invasive weed algorithm (IWO) and differential evolution algorithm (DE), named SIWODE, is proposed to address the continuous optimisation problems. In the proposed SIWODE, first, the two parameters are adaptively proposed to improve the convergence speed of the algorithm. Second, the crossover and mutation operations are introduced in SIWODE to improve the population diversity and increase the exploration capability during the iterative process. Furthermore, a local perturbation strategy is presented to improve exploitation ability during the late process. The exploration and exploitation ability of the algorithm is effectively balanced by cooperative mechanisms. The experiment results of SIWODE show that the SIWODE has the superior searching quality and stability than other mentioned approaches.
Fuqing Zhao, Songlin Du, Weimin Ma, Houbin Song
Connect. Sci.2
2021 STED-Net: Self-taught encoder-decoder network for unsupervised feature representation
Songlin Du, Takeshi Ikenaga
Multim. Tools Appl.1
2021 Multi-task neural network with physical constraint for real-time multi-person 3D pose estimation from monocular camera
Dingli Luo, Songlin Du, Takeshi Ikenaga
Multim. Tools Appl.2
2020 Resolution Irrelevant Encoding and Difficulty Balanced Loss Based Network Independent Supervision for Multi-Person Pose Estimation
abstract
Sustainable efforts are made to improve the accuracy performance in multi-person pose estimation, but the current accuracy is still not enough for real-world applications. Besides, most improvement approaches are designed for special basement networks and ignore the speed performance, which results in limited applicability and low cost-performance. This paper proposes two network independent supervision: Resolution Irrelevant Encoding and Difficulty Balanced Loss. The proposed methods reorganize task representatives, the loss calculation method, and the loss punishment ratio in one-stage pose estimation frameworks to improve the joints' location accuracy with general applicability and high computational efficiency. Resolution Irrelevant Encoding fuses heatmaps and proposed inner block offsets to fix pixel-level joints positions without resolution limitations. To improve network training efficiency, Difficulty Balanced Loss adjusts loss weight in spatial and sequential aspects. On the MS COCO keypoints detection benchmark, the mAP of OpenPose trained with our proposals outperforms the OpenPose baseline over 4.9%.
Dingli Luo, Songlin Du, Takeshi Ikenaga
HSI3
2020 Hetero Complementary Networks with Hard-Wired Condensing Binarization for High Frame Rate and Ultra-Low Delay Dual-Hand Tracking
abstract
High frame rate, ultra-low delay yet accurate hand tracking system provides a seamless and intuitive interface for Human Computer Interaction (HCI). Tracking multi-person's dual-hand from monocular RGB camera is challenging for hand's variant image feature. Although many CNN based trackers have been proposed on general hardware, they cannot address this challenge with ultra-high speed. This paper proposes: (A) Hetero complementary networks for ultra-high speed dual-hand tracking, where the quick primary result from an FPGA network is intermittently combined with delayed accurate result from a GPU network. (B) Hard-wired condensing binarization for ultrahigh speed network implementation on FPGA. The network is able to be directly mapped as hardware resource because complex computation is condensed into binary layers. The proposed method achieves 69.8% accuracy on test sequences, which is only 4.7% lower compared with the general method. Meanwhile, the estimated FPGA resource utilization is tremendously reduced to 54.7% on the target platform. This work shows the potential to track multi-person's dual-hand at millisecond-level speed.
Peiqi Zhang, Dingli Luo, Songlin Du, Takeshi Ikenaga
HSI3
2020 Local Spatio-Temporal Propagation Based Adaptive Model Generation and Update for High Frame Rate and Ultra-Low Delay Foreground Detection
abstract
High frame rate and ultra-low delay matching system plays an increasingly important role in human-machine interactive applications, which demands better experience and higher accuracy. Foreground detection is an indispensable preprocessing step to make the system suitable for complex scenes. Although many foreground detection algorithms have been proposed, few can achieve high speed in hardware due to their high complexity or high consumption. Based on the foreground detection algorithm ViBe, this paper proposes a local spatio-temporal propagation based adaptive model generation and update strategy for high frame rate and ultra-low delay foreground detection. Our algorithm predicts whether a region is a foreground by setting up detecting points, thereby adaptively adjusting the number of pixels that needs to be modeled. Secondly, the local linear illumination correlation is used to update models, which makes the algorithm more robust to illumination changes. The evaluation results show that the proposed algorithm successfully achieves real-time processing on the field-programmable gate array (FPGA) at a resolution of 640×480 pixels, with a delay of 0.908ms/frame.
Peikun Cai, Songlin Du, Takeshi Ikenaga
RTCSA2
2020 Hybrid biogeography-based optimization with enhanced mutation and CMA-ES for global optimization problem
Fuqing Zhao, Songlin Du, Yi Zhang 0096, Weimin Ma, Houbin Song
Serv. Oriented Comput. Appl.2
2019 Iterative Autoencoding and Clustering for Unsupervised Feature Representation
abstract
Unsupervised feature representation is a challenging problem in machine learning and computer vision. Since manual labels are unavailable for training, it is difficult to reduce the gap between learned features and image semantics. This paper proposes an iterative autoencoding and clustering approach, which consists of an autoencoding sub-network and a classification sub-network, for unsupervised feature representation. On one hand, the autoencoding sub-network maps images to features. On the other hand, using the features generated by the autoencoding sub-network, the classification sub-network maps the features to classes and estimates pseudo labels by clustering the features simultaneously. Through iterations between the feature representation and the pseudo-labels-supervised classification, the gap between features and image semantics is reduced. Experimental results on handwritten digits recognition and objects classification prove that the proposed approach achieves state-of-the-art performance compared with existing methods.
Songlin Du, Takeshi Ikenaga
ISCAS1
2019 Low-dimensional superpixel descriptor and its application in visual correspondence estimation
Songlin Du, Takeshi Ikenaga
Multim. Tools Appl.1
2018 Partial Descriptor Update and Isolated Point Avoidance Based Template Update for High Frame Rate and Ultra-Low Delay Deformation Matching
abstract
High frame rate and ultra-low delay matching system plays an important role in various human-machine interactive applications, which demands better performance in matching deformable and out-of-plane rotating objects. Although many algorithms have been proposed for deformation tracking and matching, few of them are suitable for hardware implementation due to complicated operations and large time consumption. This paper proposes a hardware-oriented template update method for high frame rate and ultra-low delay deformation matching system. In the proposed method, the new template is generated in real time by partially updating the template descriptor and adding new keypoints simultaneously with the matching process in pixels, and incorrect boundary points are avoided when judged as isolated with distance-reachability to solve the problem of template drift. Evaluation results indicate that the proposed method successfully supports the real-time processing of the 784fps and 640×480 resolution system on field-programmable gate array (FPGA), with a delay of 0.808ms/frame, as well as achieves satisfactory deformation matching results in comparison with other general methods.
Songlin Du, Takeshi Ikenaga
ICPR3
2017 LAP: a bio-inspired local image structure descriptor and its applications
Songlin Du, Yaping Yan, Yide Ma
Multim. Tools Appl.1
2016 When spatial distribution unites with spatial contrast: an effective blind image quality assessment model
abstract
Blind image quality assessment (BIQA), which aims to estimate the perceptual quality of images without any reference information, is a very important yet challenging task. Although human visual system is sensitive to degradations on both spatial contrast and spatial distribution, most of the existing structural degradation based BIQA models consider only one of them. This study introduces a novel BIQA model by taking into account degradations on both contrast and spatial distribution. First, the authors construct a multi‐threshold local tetra pattern (MTLTrP) instead of local binary pattern to measure the changes on spatial distribution. Second, Weber–Laplacian of Gaussian (WLOG) operator, which responds to intensity contrast in a small spatial neighbourhood, is proposed to extract local contrast features. Finally, the joint statistics of MTLTrP and WLOG are utilised for BIQA model learning. Experimental results on three large benchmark databases demonstrate that the proposed model outperforms state‐of‐the‐art BIQA models, as well as with several well‐known full reference quality assessment methods.
Yaping Yan, Songlin Du, Hongjuan Zhang, Yide Ma
IET Image Process.2
2015 Quantum-Accelerated Fractal Image Compression: An Interdisciplinary Approach
abstract
Fractal image compression (FIC) is one of the most widely approved image compression approaches for its high compression ratio and quality of retrieved images. However, FIC suffers from high computational cost in searching local self-similarities in natural image. Although many papers aiming at speeding up FIC have been published, they use pre-processing tools or approximation methods. Reducing the intrinsic computational complexity of FIC is still an open problem. Since quantum mechanics based Grover's quantum search algorithm (QSA) is able to achieve square-root speedup over classical algorithms in unsorted database searching, we propose an interdisciplinary approach by using Grover's QSA to reduce the intrinsic computational complexity of FIC in this letter. In particular, both domain blocks and range blocks are represented as quantum states, then Grover's QSA is employed to search the most similar domain block for each range block under the criterion of maximizing quantum fidelity between these two kinds of quantum states. Without sacrificing compression ratio, experimental results show that the execution time of the proposed method is 100 times shorter than that of the baseline FIC. Moreover, retrieved images from our proposal are also less distorted than those from other state-of-the-art FIC approaches.
Songlin Du, Yaping Yan, Yide Ma
IEEE Signal Process. Lett.1