Yuanlong Yu 0001

dblp:48/5332-1 · DBLP profile ↗
← Back
69ranked-venue papers
8as first author
48since 2021 · last 2026
0000-0002-2112-6214ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 5 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 20 since 2021Systems, architecture and hardware · 11 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Seeing in Double: Dual-Granularity BEV Segmentation via Mamba-Driven Alignment and Polar-Decoupled Experts
abstract
Bird's Eye View (BEV) representation has become pivotal for autonomous driving, yet existing polar coordinate-based approaches face two critical limitations: (1) distant semantic misprojection caused by radial resolution decay, and (2) region-specific geometric distortions from non-uniform polar discretization. To address these issues, we propose a novel framework addressing these challenges through three key innovations. First, we present a bilateral heterogeneous network constructs multi-granularity BEV spaces, efficiently exploiting dual-resolution visual information for distant detail preservation. Second, we employ an align-fusion strategy for multi-granularity feature aggregation. Specifically, the Mamba-Based Cross-Resolution Alignment module establishes semantic consistency for perspective features through shared state-space optimization. In the later stage, the Adaptive BEV Space Selector dynamically aggregates multi-granularity BEV features. Third, we introduce a Mixture of Radial-Angular Decoupled Experts, which employs polar-aware expert routing to disentangle radial compression and angular shear distortions through specialized geometric refinement. Comprehensive experiments on nuScenes and Lyft L5 demonstrate the state-of-the-art performance of our model across various resolution settings, visibility filtering, and perception ranges.
Jingze Su, Qi Li 0038, Wenjie Yang 0005, Yuanlong Yu 0001, Wenxi Liu
AAAI6
2026 Uncertainty-aware multi-instance partial-label learning via evidential deep model
Gaowen Jie, Fumiao Wang, Gaojie Song, Luojun Lin, Yuanlong Yu 0001, Qinghai Zheng
Neurocomputing5
2026 Partial multi-label feature selection via feature-label bidirectional association with adaptive label graph diffusion
Shuhao Shan, Yuanlong Yu 0001, Zhenzhen Sun
Neurocomputing4
2026 Multi-View Clustering via Cross-View Alignment and Anchor-Guided Representation
abstract
Due to its effectiveness and efficiency, anchor based multi-view clustering (MVC) has recently attracted much attention. However, existing anchor-based methods often ignore the balance of anchor distribution and fail to fully leverage the complementary nature of view-specific information. To address these issues, we propose Cross-view Alignment and Anchor-guided Representation(CAAR), a unified framework that jointly optimizes anchor learning, anchor graph construction, and clustering partition. To be specific, CAAR employs an explicit alignment mechanism to preserve both cross-view consistency and view-specific characteristics, and introduces a cluster-aware prior to encourage balanced anchor allocation across semantic clusters. The model is optimized via an alternating minimization algorithm and achieves linear computational complexity. Extensive experiments demonstrate the effectiveness and efficiency of CAAR on six real-life datasets.
Gaojie Song, Fumiao Wang, Gaowen Jie, Yuanlong Yu 0001, Qinghai Zheng
IEEE Signal Process. Lett.4
2026 Trust Under Attack: Environment Perturbation-Based Attacks on Trust-Aware Human-Robot Collaboration
abstract
In trust-aware human-robot collaboration (HRC), incorporating human trust into robot decision-making has shown promise in improving collaboration performance. However, most existing research overlooks the potential security vulnerabilities in trust-aware human-robot collaboration methods, which may compromise both collaboration efficiency and human trust. To directly reveal these vulnerabilities and evaluate their impact, this paper proposes an environment perturbation-based trust attack method for human-robot collaboration. The proposed method models the human-robot collaboration process under black-box assumptions and generates adversarial environmental perturbations using the Fast Gradient Sign Method (FGSM). Extensive simulation and real-world experiments show that both traditional and large language model (LLM)-based trust-aware human-robot collaboration methods are vulnerable to such attacks, leading to reductions in collaboration efficiency and human trust. Furthermore, the analysis reveals characteristic phenomena in trust dynamics and system performance under adversarial conditions. These results highlight the importance of considering security vulnerabilities in trust-aware human-robot collaboration.
Boyuan Du, Huaping Liu 0001, Yuanlong Yu 0001
IEEE Trans Autom. Sci. Eng.3
2026 COSOS-1k: A Benchmark Dataset and Occlusion-Aware Uncertainty Learning for Multi-View Video Object Detection
abstract
Confined spaces refer to partially or fully enclosed areas, e.g., sewage wells, where working conditions pose significant risks to the workers. The evaluation of COfined Space Operational Safety (COSOS) refers to verifying whether workers are properly equipped with safety equipment before entering a confined space, which is crucial for protecting their safety and health. Due to the crowded nature of such environments and the small size of certain safety equipment, existing methods face significant challenges. Moreover, there is a lack of dedicated datasets to support research in this domain. In this paper, in order to advance research in this challenging task, we present COSOS-1k, an extensive dataset constructed from diverse confined space scenarios. It comprises multi-view videos for each scenario, covers 10 essential safety protective equipments and 6 attributes of worker, and is annotated with expressive object locations, fine-grained attributes, and occlusion status. The COSOS-1k is the first dataset known to date, tailored explicitly for the real-world COSOS scenarios. In addition, we address the challenge of occlusion from three perspectives: instance, video, and view. Firstly, at the instance level, we propose Occlusion-aware Uncertainty Estimation (OUE) method, which leverages box-level occlusion annotations to enable part-level occlusion prediction for objects. Secondly, at the video level, we introduce Cross-Frame Cluster (CFC) attention, which integrates temporal context features from the same object category to mitigate the impact of occlusions in the current frame. Finally, we extend CFC to the view level and form Cross-View Cluster (CVC) attention, where complementary information is mined from another view. Extensive experiments demonstrate the effectiveness of the proposed methods and provide insights into the importance of dataset diversity and expressivity. The COSOS-1k dataset and code are available at https://github.com/deepalchemist/cosos-1k.
Wenjie Yang 0005, Yueying Kao, Yuanlong Yu 0001, Kaiqi Huang
IEEE Trans. Image Process.4
2026 From One Comes Two: A Tensorized Graph Learning Framework for Clustering
Qinghai Zheng, Jihua Zhu, Yuanlong Yu 0001, Haoyu Tang 0002
IEEE Trans. Knowl. Data Eng.3
2025 GSSN: A Graph Structural Similarity Network for Complex 2D Geometric Drawing Retrieval
Longlong Liao, Junyong Lu, Yuanlong Yu 0001
ICA3PP (5)5
2025 A Tiny Change, a Giant Leap: Long-Tailed Class-Incremental Learning via Geometric Prototype Alignment
Xinyi Lai, Luojun Lin, Weijie Chen 0006, Yuanlong Yu 0001
ICCV4
2025 Multi-Attention Guided Knowledge Distillation For High-Performance Object Detection
abstract
Knowledge distillation is beneficial for improving the performance of object detection models. However, the existing methods utilizing attention maps for feature weighting embrace limited flexibility, and may result in the loss of crucial channel and location information. To fully utilize these critical information, this paper introduce a novel Multi-Attention Guided Distillation framework that aims to enrich the channel of detail attention representations by enhancing local feature expressions, while patching feature maps concurrently. Meanwhile, a global spatial attention map is utilized to supplement global feature information. To alleviate the discrepancy in the feature attention map between the teacher and student, we use the original student’s features to mimic the weighted teacher’s features. Extensive ablation experiments have proven the effectiveness of our method. Compared with other distillation methods, the SOTA results demostrate the advanced nature of our model, and some student models even outperform the teacher model.
Zhihao Kong, Qifeng Lin, Qishen Shen, Jiayi Qiu, Gang Fu 0003, Yuanlong Yu 0001
ICME6
2025 Neural Collision Detection for Constrained Grasp Pose Optimization in Cluttered Environments
abstract
Robust robotic grasping in cluttered environments presents a significant challenge, as existing methods often neglect the complex interactions between the gripper, objects, and obstacles, leading to collisions and grasping failures. To address this, we propose a framework that integrates collision avoidance as a core constraint within the grasp pose optimization process. Central to this framework is a Neural Collision Detection (NCD) network that takes scene configurations and grasp poses as inputs, producing a collision score that approximates traditional collision detection functions. The NCD network provides critical feedback for refining grasp predictions and demonstrates strong generalization across diverse environments, facilitating efficient collision detection and constrained grasp pose optimization. Additionally, we incorporate frictional force closure, geometric symmetry, and surface alignment as regularization terms within the optimization function, enhancing the physical stability and geometric plausibility of the generated grasps. Extensive experiments conducted in real-world environments show a significant improvement in grasp success rates, with robust generalization to previously unseen objects and scenarios. These results validate the efficacy of our framework, highlighting its potential for enabling reliable robotic manipulation in complex and cluttered environments.
Longyuan Lin, Yixin Zhuang, Qinghai Zheng, Yuanlong Yu 0001
IROS5
2025 EEG super-resolution with Laplacian Regularized Coupled Matrix Decomposition: A case study of Autism Spectrum Disorder EEG enhancement
Yunbo Tang, Qifeng Lin, Yuanlong Yu 0001, Dan Chen 0001
Artif. Intell. Medicine3
2025 Geometry-aware triplane diffusion for single shape generation with feature alignment
Hongliang Weng, Qinghai Zheng, Yuanlong Yu 0001, Yixin Zhuang
Comput. Graph.3
2025 Trusted Cross-view Completion for incomplete multi-view classification
Peihuan Song, Qinghai Zheng, Yuanlong Yu 0001
Neurocomputing5
2025 Fault resilient on-device batched DNN inference for mobile devices with ARM TrustZone
Longlong Liao, Wenbin Zeng, Xinqi Liu, Yuanlong Yu 0001
J. Syst. Archit.5
2025 Cross-View Fusion for Multi-View Clustering
abstract
Multi-view clustering has attracted significant attention in recent years because it can leverage the consistent and complementary information of multiple views to improve clustering performance. However, effectively fuse the information and balance the consistent and complementary information of multiple views are common challenges faced by multi-view clustering. Most existing multi-view fusion works focus on weighted-sum fusion and concatenating fusion, which unable to fully fuse the underlying information, and not consider balancing the consistent and complementary information of multiple views. To this end, we propose Cross-view Fusion for Multi-view Clustering (CFMVC). Specifically, CFMVC combines deep neural network and graph convolutional network for cross-view information fusion, which fully fuses feature information and structural information of multiple views. In order to balance the consistent and complementary information of multiple views, CFMVC enhances the correlation among the same samples to maximize the consistent information while simultaneously reinforcing the independence among different samples to maximize the complementary information. Experimental results on several multi-view datasets demonstrate the effectiveness of CFMVC for multi-view clustering task.
Binqiang Huang, Qinghai Zheng, Yuanlong Yu 0001
IEEE Signal Process. Lett.4
2025 Supervised Contrastive Learning With Mixed Samples for Long-Tailed Recognition
abstract
In the domain of signal processing and deep learning, long-tailed data distributions present significant challenges due to the class imbalance in which a few classes contain a large number of samples, while most classes have far fewer. This imbalance hinders the ability of traditional models to effectively learn from minority classes. In this work, we focus on long-tailed supervised contrastive learning and introduce a novel approach termed Mixture-based Supervised Contrastive Learning (MixSCL), which integrates image mixing techniques into the supervised contrastive learning framework. By focusing on intra-class diversity and inter-class separability, our method aims to enhance the global uniformity of feature representations and improve model robustness. Specifically, MixSCL employs dual-stream projection heads designed to optimize separately for original and mixed samples, ensuring that the introduction of mixed samples does not distort the representations of original samples. We conduct extensive evaluations on bench mark datasets including CIFAR-100-LT and ImageNet-LT, which demonstrate that MixSCL achieves superior and more balanced performance in long-tailed scenarios.
Peihuan Song, Luojun Lin, Yuanlong Yu 0001, Wenjie Yang 0005, Qinghai Zheng
IEEE Signal Process. Lett.3
2025 Attention-Based Mean-Max Balance Assignment for Oriented Object Detection in Optical Remote Sensing Images
abstract
For objects with arbitrary angles in optical remote sensing (RS) images, the oriented bounding box regression task often faces the problem of ambiguous boundaries between positive and negative samples. The statistical analysis of existing label assignment strategies reveals that anchors with low Intersection over Union (IoU) between ground truth (GT) may also accurately surround the GT after decoding. Therefore, this article proposes an attention-based mean-max balance assignment (AMMBA) strategy, which consists of two parts: mean-max balance assignment (MMBA) strategy and balance feature pyramid with attention (BFPA). MMBA employs the mean-max assignment (MMA) and balance assignment (BA) to dynamically calculate a positive threshold and adaptively match better positive samples for each GT for training. Meanwhile, to meet the need of MMBA for more accurate feature maps, we construct a BFPA module that integrates spatial and scale attention mechanisms to promote global information propagation. Combined with S2ANet, our AMMBA method can effectively achieve state-of-the-art performance, with a precision of 80.91% on the DOTA dataset in a simple plug-and-play fashion. Extensive experiments on three challenging optical RS image datasets (DOTA-v1.0, HRSC, and DIOR-R) further demonstrate the balance between precision and speed in single-stage object detectors. Our AMMBA has enough potential to assist all existing RS models in a simple way to achieve better detection performance. The code is available athttps://github.com/promisekoloer/AMMBA.
Qifeng Lin, Daoye Zhu, Gang Fu 0003, Chuanxi Chen, Yuanlong Yu 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 Multiple Region Proposal Experts Network for Wide-Scale Remote Sensing Object Detection
abstract
Faced with the wide-scale characteristics of objects in optical remote sensing images, the current object detection models are always unable to provide satisfactory detection capabilities for remote sensing tasks. To achieve better wide-scale coverage for various remote sensing regions of interest, this article introduces a multiprediction mechanism to build a novel region generation model, namely, a multiple region proposal experts network (MRPENet). Meanwhile, to achieve both region proposal coverage and receptive field coverage of wide-scale objects, we constructed a prior design of an anchor (PDA) module and an adaptive features compensation (AFC) module to achieve the coverage of wide-scale remote sensing objects. To better utilize the multiexpert characteristics of our model, we customized a new training sample allocation strategy, dynamic scale-assigned expert learning (DSAEL), to cultivate the ability of experts to deal with objects at various scales. To the best of our knowledge, this is the first time that a multiple region proposal network (RPN) mechanism has been used in the object detection of optical remote sensing images. Extensive experiments have shown the generality and effectiveness of our MRPENet. Without bells and whistles, MRPENet achieves a new state-of-the-art (SOTA) on standard benchmarks, i.e., DOTA-v1.0 [82.02% mean average precision (mAP)], HRSC2016 (98.16% mAP), and FAIR1M-v1.0 (48.80% mAP).
Qifeng Lin, Daoye Zhu, Gang Fu 0003, Yuanlong Yu 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Incomplete Multi-View Clustering Via Inference and Evaluation
abstract
Multi-view clustering aims to improve the clustering performance by leveraging information from multiple views. Most existing works assume that all views are complete. However, samples in real-world scenarios cannot be always observed in all views, leading to the challenging problem of Incomplete Multi-View Clustering (IMVC). Although some attempts are made recently, they still suffer from the following two limitations: (1) they usually adopt shallow models, which are unable to sufficiently explore the consistency and complementary of multiple views; (2) they lack of a suitable measurement to evaluate the quality of the recovered data during the learning process. To address the aforementioned limitations, we introduce a novel Incomplete Multi-View Clustering via Inference and Evaluation (IMVC-IE). Specifically, IMVC-IE adopts the contrastive learning strategy on features of different views to excavate the underlying information from existing samples firstly. Subsequently, massive alternative simulated data are inferred for missing views and a novel evaluation strategy is presented to obtain the proper data for missing views completion. Extensive experiments are conducted and verify the effectiveness of our method.
Binqiang Huang, Shoujie Lan, Qinghai Zheng, Yuanlong Yu 0001
ICASSP5
2024 Learning feature alignment across attribute domains for improving facial beauty prediction
abstract
Facial beauty prediction (FBP) aims to develop a system to assess facial attractiveness automatically. Through prior research and our own observations, it has become evident that attribute information, such as gender and race, is a key factor leading to the distribution discrepancy in the FBP data. Such distribution discrepancy hinders current conventional FBP models from generalizing effectively to unseen attribute domain data, thereby discounting further performance improvement . To address this problem, in this paper, we exploit the attribute information to guide the training of convolutional neural networks (CNNs), with the final purpose of implicit feature alignment across various attribute domain data. To this end, we introduce the attribute information into convolution layer and batch normalization (BN) layer, respectively, as they are the most crucial parts for representation learning in CNNs. Specifically, our method includes: 1) Attribute-guided convolution (AgConv) that dynamically updates convolutional filters based on attributes by parameter tuning or parameter rebirth; 2) Attribute-guided batch normalization (AgBN) is developed to compute the attribute-specific statistics through an attribute guided batch sampling strategy; 3) To benefit from both approaches, we construct an integrated framework by combining AgConv and AgBN to achieve a more thorough feature alignment across different attribute domains. Extensive qualitative and quantitative experiments have been conducted on the SCUT-FBP, SCUT-FBP5500 and HotOrNot benchmark datasets. The results show that AgConv significantly improves the attribute-guided representation learning capacity and AgBN provides more stable optimization. Owing to the combination of AgConv and AgBN, the proposed framework (Ag-Net) achieves further performance improvement and is superior to other state-of-the-art approaches for FBP.
Zhishu Sun, Luojun Lin, Yuanlong Yu 0001
Expert Syst. Appl.3
2024 Multi-label feature selection via adaptive dual-graph optimization
Zhenzhen Sun, Yuanlong Yu 0001
Expert Syst. Appl.4
2024 Fault-tolerant deep learning inference on CPU-GPU integrated edge devices with TEEs
Hongjian Xu, Longlong Liao, Xinqi Liu, Shuguang Chen, Zhixuan Liang, Yuanlong Yu 0001
Future Gener. Comput. Syst.7
2024 You only label once: A self-adaptive clustering-based method for source-free active domain adaptation
abstract
Abstract With the growing significance of data privacy protection, Source‐Free Domain Adaptation (SFDA) has gained attention as a research topic that aims to transfer knowledge from a labeled source domain to an unlabeled target domain without accessing source data. However, the absence of source data often leads to model collapse or restricts the performance improvements of SFDA methods, as there is insufficient true‐labeled knowledge for each category. To tackle this, Source‐Free Active Domain Adaptation (SFADA) has emerged as a new task that aims to improve SFDA by selecting a small set of informative target samples labeled by experts. Nevertheless, existing SFADA methods impose a significant burden on human labelers, requiring them to continuously label a substantial number of samples throughout the training period. In this paper, a novel approach is proposed to alleviate the labeling burden in SFADA by only necessitating the labeling of an extremely small number of samples on a one‐time basis . Moreover, considering the inherent sparsity of these selected samples in the target domain, a Self‐adaptive Clustering‐based Active Learning (SCAL) method is proposed that propagates the labels of selected samples to other datapoints within the same cluster. To further enhance the accuracy of SCAL, a self‐adaptive scale search method is devised that automatically determines the optimal clustering scale, using the entropy of the entire target dataset as a guiding criterion. The experimental evaluation presents compelling evidence of our method's supremacy. Specifically, it outstrips previous SFDA methods, delivering state‐of‐the‐art (SOTA) results on standard benchmarks. Remarkably, it accomplishes this with less than 0.5% annotation cost, in stark contrast to the approximate 5% required by earlier techniques. The approach thus not only sets new performance benchmarks but also offers a markedly more practical and cost‐effective solution for SFADA, making it an attractive choice for real‐world applications where labeling resources are limited.
Zhishu Sun, Luojun Lin, Yuanlong Yu 0001
IET Image Process.3
2024 Ultra-High Resolution Image Segmentation via Locality-Aware Context Fusion and Alternating Local Enhancement
Wenxi Liu, Qi Li 0038, Xindai Lin, Weixiang Yang, Shengfeng He, Yuanlong Yu 0001
Int. J. Comput. Vis.6
2024 Double-level View-correlation Multi-view Subspace Clustering
Shoujie Lan, Qinghai Zheng, Yuanlong Yu 0001
Knowl. Based Syst.3
2024 Partial multi-label feature selection via low-rank and sparse factorization with manifold learning
Zhenzhen Sun, Zexiang Chen, Yewang Chen, Yuanlong Yu 0001
Knowl. Based Syst.5
2024 Monocular BEV Perception of Road Scenes via Front-to-Top View Projection
abstract
HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to expensive sensors and time-consuming computation. Camera-based methods usually need to perform road segmentation and view transformation separately, which often causes distortion and missing content. To push the limits of the technology, we present a novel framework that reconstructs a local map formed by road layout and vehicle occupancy in the bird's-eye view given a front-view monocular image only. We propose a front-to-top view projection (FTVP) module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. In addition, we apply multi-scale FTVP modules to propagate the rich spatial information of low-level features to mitigate spatial deviation of the predicted object location. Experiments on public benchmarks show that our method achieves various tasks on road layout estimation, vehicle occupancy estimation, and multi-class semantic estimation, at a performance level comparable to the state-of-the-arts, while maintaining superior efficiency.
Wenxi Liu, Qi Li 0038, Weixiang Yang, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Semi-supervised domain generalization with evolving intermediate domain
Luojun Lin, Zhishu Sun, Weijie Chen 0006, Wenxi Liu, Yuanlong Yu 0001, Lei Zhang 0038
Pattern Recognit.6
2024 Prototype learning based generic multiple object tracking via point-to-box supervision
Wenxi Liu, Qi Li 0038, Yinhua She, Yuanlong Yu 0001, Jia Pan 0001, Jason Gu
Pattern Recognit.5
2024 Learning Interpretable Brain Functional Connectivity via Self-Supervised Triplet Network With Depth-Wise Attention
abstract
Brain functional connectivity has been widely explored to reveal the functional interaction dynamics between the brain regions. However, conventional connectivity measures rely on deterministic models demanding application-specific empirical analysis, while deep learning approaches focus on finding discriminative features for state classification, having limited capability to capture the interpretable connectivity characteristics. To address the challenges, this study proposes a self-supervised triplet network with depth-wise attention (TripletNet-DA) to generate the functional connectivity: 1) TripletNet-DA firstly utilizes channel-wise transformations for temporal data augmentation, where the correlated & uncorrelated sample pairs are constructed for self-supervised training, 2) Channel encoder is designed with a convolution network to extract the deep features, while similarity estimator is employed to generate the similarity pairs and the functional connectivity representations, 3) TripletNet-DA applies Triplet loss with anchor-negative similarity penalty for model training, where the similarities of uncorrelated sample pairs are minimized to enhance model's learning capability. Experimental results on pathological EEG datasets (Autism Spectrum Disorder, Major Depressive Disorder) indicate that 1) TripletNet-DA demonstrates superiority in both ASD discrimination and MDD classification than the state-of-the-art counterparts, where the connectivity features in beta & gamma bands have respectively achieved the accuracy of 97.05%, 98.32% for ASD discrimination, 89.88%, 91.80% for MDD classification in the eyes-closed condition and 90.90%, 92.26% in the eyes-open condition, 2) TripletNet-DA enables to uncover significant differences of functional connectivity between ASD EEG and TD ones, and the prominent connectivity links are in accordance with the empirical findings, thus providing potential biomarkers for clinical ASD analysis.
Yunbo Tang, Weirong Huang, Rongchang Liu, Yuanlong Yu 0001
IEEE J. Biomed. Health Informatics4
2024 Learning Nighttime Semantic Segmentation the Hard Way
abstract
Nighttime semantic segmentation is an important but challenging research problem for autonomous driving. The major challenges lie in the small objects or regions from the under-/over-exposed areas or suffer from motion blur caused by the camera deployed on moving vehicles. To resolve this, we propose a novel hard-class-aware module that bridges the main network for full-class segmentation and the hard-class network for segmenting aforementioned hard-class objects. In specific, it exploits the shared focus of hard-class objects from the dual-stream network, enabling the contextual information flow to guide the model to concentrate on the pixels that are hard to classify. In the end, the estimated hard-class segmentation results will be utilized to infer the final results via an adaptive probabilistic fusion refinement scheme. Moreover, to overcome over-smoothing and noise caused by extreme exposures, our model is modulated by a carefully crafted pretext task of constructing an exposure-aware semantic gradient map, which guides the model to faithfully perceive the structural and semantic information of hard-class objects while mitigating the negative impact of noises and uneven exposures. In experiments, we demonstrate that our unique network design leads to superior segmentation performance over existing methods, featuring the strong ability of perceiving hard-class objects under adverse conditions.
Wenxi Liu, Qi Li 0038, Chenyang Liao, Jingjing Cao, Shengfeng He, Yuanlong Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Fuzzy Representation Learning on Dynamic Graphs
abstract
Exploring dynamic patterns from complex and large-scale networks is a significant and challenging task in graph analysis. One of the most advanced solutions is dynamic graph representation learning, which embeds structural and temporal correlations into a representative vector for each node or subgraph. Existing models have made some successes, such as overcoming the problems of induction for unseen nodes and scalability for large-scale evolving networks. However, these models usually rely on crisp representation learning that is incapable of modeling feature fuzziness and capturing uncertainties in dynamic graphs. While real-world dynamic networks as complex systems always contain non-negligible but inestimable uncertainties in node/link attributes and network topology. These uncertainties may cause the learned representations from crisp models hard to precisely reflect network evolution. To address the issues, we propose a new dynamic graph representation learning model, called FuzzyDGL, which first incorporates fuzzy representation learning to handle the uncertainties in dynamic graphs. Through combining acrlong CDGRL with fuzzy logic, the FuzzyDGL digests both of their advantages. On the one hand, it has flexible model scalability and brilliant inductive capability. On the other hand, it can model feature fuzziness to reduce the impact of uncertainties in dynamic graphs, improving the quality of learned representations. To demonstrate its effectiveness, we conduct two important tasks of network analysis, including link prediction and node classification, over eight real-world datasets. The experimental results show the strong competitiveness and generalization of the FuzzyDGL against a number of baseline models.
Yuanlong Yu 0001, Chun-Yang Zhang, Yue-Na Lin, Shang-Jia Li
IEEE Trans. Syst. Man Cybern. Syst.2
2023 FBPFormer: Dynamic Convolutional Transformer for Global-Local-Contexual Facial Beauty Prediction
Qipeng Liu 0004, Luojun Lin, Zhifeng Shen, Yuanlong Yu 0001
ICANN (10)4
2023 Learning High Frequency Surface Functions In Shells
abstract
Recently, coordinate-based MLPs have been shown to be powerful representations for 3D surfaces, where learning high-frequency details is facilitated by modulating surface functions with periodic functions [1], [2]. While shortening the periodicity helps in learning high frequencies, it leads to increasing ambiguity, i.e., more points along the axis directions become similar in the embedded space, so that many points on the surface and outside the surface have similar predictions. In addition, short periodicity increases local geometric variations, leading to unexpected noisy artifacts in untrained regions. Unlike existing methods that learn surface functions in a regular cube, we find surfaces within shells, a coarse form of the target surfaces constructed by a binary classifier. The advantage of build surfaces in shells is that MLPs focus on regions of interest, which inherently reduces ambiguity and also promotes training efficiency and test accuracy. We demonstrate the effectiveness of shells and show significant improvements over baseline methods in 3D surface reconstruction from raw point clouds.
Yuanlong Yu 0001, Xuelin Chen, Yixin Zhuang
ICME2
2023 Run and Chase: Towards Accurate Source-Free Domain Adaptive Object Detection
abstract
Recently, there has been increasing interest in the Source-Free Domain Adaptive Object Detection task, which involves training an object detector on the unlabeled target data using a pre-trained source model without accessing the source data. Most related methods are developed from the mean-teacher framework, which aims to train the student model closer to the teacher model via a pseudo labeling manner, where the teacher model is the exponential-moving-average of the student models at different time-steps. Following this line of works, we propose a Run-and-Chase Mutual-Learning method to strengthen the interactions between the student model and the teacher model in both feature and prediction levels. In our method, the student model is optimized to run away from the teacher model at the feature level, while chasing the teacher model at the prediction level. In this way, the student model is forced to be distinguishable at different time-steps, so that the teacher model can acquire more diverse task-related information and produce higher-accuracy pseudo labels. As the training goes, the student and teacher models are updated iteratively and promoted mutually, which can prevent the model collapse problem. Extensive experiments are conducted to validate the effectiveness of our method.
Luojun Lin, Zhifeng Yang, Qipeng Liu 0004, Yuanlong Yu 0001, Qifeng Lin
ICME4
2023 Deep Inversion Method for Attacking Lifelong Learning Neural Networks
abstract
Artificial neural networks suffer from catastrophic forgetting when knowledge needs to be learned from multi-batch or streaming data. In response to this problem, researchers have proposed a variety of lifelong learning methods to avoid catastrophic forgetting. However, current methods usually do not consider the possibility of malicious attacks. Meanwhile, in real lifelong learning scenarios, batch data or streaming data usually come from an incompletely trusted environment. Attackers can easily manipulate data or inject malicious samples into the training data set. As a result, the reliability of neural networks decreases. Recently, researches of lifelong learning attacks need to obtain real samples of the attacked classes, whether using backdoor attacks or data poisoning attacks. In this paper, we focus on an attack setting that is more suitable for lifelong learning scenario. This setting has two main features. The first is the setting does not require real samples of the attacked classes, and the second is it allows attacks to be performed on tasks that exclude the attacked classes. For this scenario, we propose a lifelong learning attack model based on deep inversion. In the scenario where EWC is used as the benchmark lifelong learning model, our experiments show that 1) in the data poisoning attack, the target accuracy can be significantly decreased by adding 0.5% of poisoned samples; 2) The backdoor attack with high accuracy can be achieved by adding 1% of backdoor samples.
Boyuan Du, Yuanlong Yu 0001, Huaping Liu 0001
IJCNN2
2023 A Multiple Prediction Mechanisms Ensemble for Complex Remote Sensing Scenes
abstract
Facing complex remote sensing scenes, detection models with single detection mechanisms cannot always provide satisfactory detection capabilities. In order to obtain better detection performance in various remote sensing scenes, this paper constructs a novel ensemble model, namely: the multiple prediction mechanisms ensemble (MPME). In order to improve the feature representation ability and region recognition ability of the ensemble model, we build the ensemble of feature pyramids (EFP) and the ensemble of detection heads (EDH) respectively. In order to further improve the detection accuracy of the ensemble model, we propose a training strategy (k-Nearest Loss Learning), so that each sub-detector does not need to learn a trade-off among all training samples, and also reduces the possibility of model over-fitting. The experimental results show that our MPME is a more efficient and effective ensemble model. Compared with other ensemble models, our MPME has a faster detection speed and better detection accuracy. Compared with other state-of-the-art detectors, our detector also achieves superior detection performance.
Qifeng Lin, Luojun Lin, Yuanlong Yu 0001, Gang Fu 0003
ACM Multimedia3
2023 Parameter Exchange for Robust Dynamic Domain Generalization
abstract
Agnostic domain shift is the main reason of model degradation on the unknown target domains, which brings an urgent need to develop Domain generalization (DG). Recent advances at DG use dynamic networks to achieve training-free adaptation on the unknown target domains, termed Dynamic Domain Generalization (DDG), which compensates for the lack of self-adaptability in static models with fixed weights. The parameters of dynamic networks can be decoupled into a static and a dynamic component, which are designed to learn domain-invariant and domain-specific features, respectively. Based on the existing arts, in this work, we try to push the limits of DDG by disentangling the static and dynamic components more thoroughly from an optimization perspective. Our main consideration is that we can enable the static component to learn domain-invariant features more comprehensively by augmenting the domain-specific information. As a result, the more comprehensive domain-invariant features learned by the static component can then enforce the dynamic component to focus more on learning adaptive domain-specific features. To this end, we propose a simple yet effective Parameter Exchange (PE) method to perturb the combination between the static and dynamic components. We optimize the model using the gradients from both the perturbed and non-perturbed feed-forward jointly to implicitly achieve the aforementioned disentanglement. In this way, the two components can be optimized in a mutually-beneficial manner, which can resist the agnostic domain shifts and improve the self-adaptability on the unknown target domain. Extensive experiments show that PE can be easily plugged into existing dynamic networks to improve their generalization ability without bells and whistles.
Luojun Lin, Zhifeng Shen, Zhishu Sun, Yuanlong Yu 0001, Lei Zhang 0038, Weijie Chen 0006
ACM Multimedia4
2023 MetaFBP: Learning to Learn High-Order Predictor for Personalized Facial Beauty Prediction
abstract
Predicting individual aesthetic preferences holds significant practical applications and academic implications for human society. However, existing studies mainly focus on learning and predicting the commonality of facial attractiveness, with little attention given to Personalized Facial Beauty Prediction (PFBP). PFBP aims to develop a machine that can adapt to individual aesthetic preferences with only a few images rated by each user. In this paper, we formulate this task from a meta-learning perspective that each user corresponds to a meta-task. To address such PFBP task, we draw inspiration from the human aesthetic mechanism that visual aesthetics in society follows a Gaussian distribution, which motivates us to disentangle user preferences into a commonality and an individuality part. To this end, we propose a novel MetaFBP framework, in which we devise a universal feature extractor to capture the aesthetic commonality and then optimize to adapt the aesthetic individuality by shifting the decision boundary of the predictor via a meta-learning mechanism. Unlike conventional meta-learning methods that may struggle with slow adaptation or overfitting to tiny support sets, we propose a novel approach that optimizes a high-order predictor for fast adaptation. In order to validate the performance of the proposed method, we build several PFBP benchmarks by using existing facial beauty prediction datasets rated by numerous users. Extensive experiments on these benchmarks demonstrate the effectiveness of the proposed MetaFBP method.
Luojun Lin, Zhifeng Shen, Jia-Li Yin, Qipeng Liu 0004, Yuanlong Yu 0001, Weijie Chen 0006
ACM Multimedia5
2023 Dual-graph with non-convex sparse regularization for multi-label feature selection
Zhenzhen Sun, Jin Gou, Yuanlong Yu 0001
Appl. Intell.5
2023 A novel decoder based on Bayesian rules for task-driven object segmentation
abstract
Abstract As a challenging problem in computer vision, salient object segmentation has attracted increasing attention in recent years. Though a lot of works based on encoder–decoder have been made, these methods can only recognize and segment one class of objects, but cannot segment the other classes of objects in the same image. To address this issue, this paper proposes a novel decoder based on Bayesian rules to perform task‐driven object segmentation, in which a control signal is added to the decoder to determine which class of objects need to be segmented. What's more, a Bayesian rule is established in the decoder, in which the control signal is set as the prior, and the latent features learned in encoder is transferred to the corresponding layer of decoder as observation, thus the posterior probability of each object with respect to the specific‐class can be calculated, and the objects belonging to this class can be segmented. This proposed method is evaluated for task‐driven salient object segmentation on several benchmark datasets, including MS COCO, DUT‐OMRON, ECSSD etc. Experimental results show that the approach tends to segment accurate, detailed, and complete objects, and improves the performance compared with the previous state‐of‐the‐art.
Yuanlong Yu 0001, Weijie Jiang 0001, Weitao Zheng, Renjie Su
IET Image Process.2
2022 Dynamic Domain Generalization
abstract
Domain generalization (DG) is a fundamental yet very challenging research topic in machine learning. The existing arts mainly focus on learning domain-invariant features with limited source domains in a static model. Unfortunately, there is a lack of training-free mechanism to adjust the model when generalized to the agnostic target domains. To tackle this problem, we develop a brand-new DG variant, namely Dynamic Domain Generalization (DDG), in which the model learns to twist the network parameters to adapt to the data from different domains. Specifically, we leverage a meta-adjuster to twist the network parameters based on the static model with respect to different data from different domains. In this way, the static model is optimized to learn domain-shared features, while the meta-adjuster is designed to learn domain-specific features. To enable this process, DomainMix is exploited to simulate data from diverse domains during teaching the meta-adjuster to adapt to the agnostic target domains. This learning mechanism urges the model to generalize to different agnostic target domains via adjusting the model without training. Extensive experiments demonstrate the effectiveness of our proposed method. Code is available: https://github.com/MetaVisionLab/DDG
Zhishu Sun, Zhifeng Shen, Luojun Lin, Yuanlong Yu 0001, Zhifeng Yang, Shicai Yang, Weijie Chen 0006
IJCAI4
2022 Robust multi-class feature selection via l2, 0-norm regularization minimization
abstract
Feature selection is an important data preprocessing in data mining and machine learning, that can reduce the number of features without deteriorating model’s performance. Recently, sparse regression has received considerable attention in feature selection task due to its good performance. However, because the l2,0-norm regularization term is non-convex, this problem is hard to solve, and most of the existing methods relaxed it by l2,1-norm. Unlike the existing methods, this paper proposes a novel method to solve the l2,0-norm regularized least squares problem directly based on iterative hard thresholding, which can produce exact row-sparsity solution for weights matrix, and features can be selected more precisely. Furthermore, two homotopy strategies are derived to reduce the computational time of the optimization method, which are more practical for real-world applications. The proposed method is verified on eight biological datasets, experimental results show that our method can achieve higher classification accuracy with fewer number of selected features than the approximate convex counterparts and other state-of-the-art feature selection methods.
Zhenzhen Sun, Yuanlong Yu 0001
Intell. Data Anal.2
2021 Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View Transformation
abstract
HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the absence of content. To push the limits of the technology, we present a novel framework that enables reconstructing a local map formed by road layout and vehicle occupancy in the bird’s-eye view given a front-view monocular image only. In particular, we propose a cross-view transformation module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. Considering the relationship between vehicles and roads, we also design a context-aware discriminator to further refine the results. Experiments on public benchmarks show that our method achieves the state-of-the-art performance in the tasks of road layout estimation and vehicle occupancy estimation. Especially for the latter task, our model outperforms all competitors by a large margin. Furthermore, our model runs at 35 FPS on a single GPU, which is efficient and applicable for real-time panorama HD map reconstruction.
Weixiang Yang, Qi Li 0038, Wenxi Liu, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001
CVPR4
2021 From Contexts to Locality: Ultra-high Resolution Image Segmentation via Locality-aware Contextual Correlation
abstract
Ultra-high resolution image segmentation has raised increasing interests in recent years due to its realistic applications. In this paper, we innovate the widely used high-resolution image segmentation pipeline, in which an ultrahigh resolution image is partitioned into regular patches for local segmentation and then the local results are merged into a high-resolution semantic mask. In particular, we introduce a novel locality-aware contextual correlation based segmentation model to process local patches, where the relevance between local patch and its various contexts are jointly and complementarily utilized to handle the semantic regions with large variations. Additionally, we present a contextual semantics refinement network that associates the local segmentation result with its contextual semantics, and thus is endowed with the ability of reducing boundary artifacts and refining mask contours during the generation of final high-resolution mask. Furthermore, in comprehensive experiments, we demonstrate that our model outperforms other state-of-the-art methods in public benchmarks. Our released codes are available at https://github.com/liqiokkk/FCtL.
Qi Li 0038, Weixiang Yang, Wenxi Liu, Yuanlong Yu 0001, Shengfeng He
ICCV4
2021 Deep Unsupervised Learning Based Visual Odometry with Multi-scale Matching and Latent Feature Constraint
abstract
A novel siamese autoencoder visual odometry system named SAEVO is proposed in this paper. SAEVO can jointly estimate the 6-DoF pose and the depth using deep neural networks trained with monocular clips only. The main idea of the proposed method is an unsupervised deep learning scheme that combines siamese networks with auto-encoder for multi-scale matching to estimate ego-motion. Also, two unsupervised losses are designed to align extracted features from the siamese autoencoder networks. A system overview is shown in Fig. 1. The experiments on KITTI and CityScapes datasets demonstrate the SAEVO achieves good performance in terms of pose and depth accuracy, and competitive performance to state-of-the-art methods.
Zhenzhen Liang, Yuanlong Yu 0001
IROS3
2021 Dynamic Domain Adaptation for Single-view 3D Reconstruction
abstract
Learning 3D object reconstruction from a single RGB image is a fundamental and extremely challenging problem for robots. As acquiring labeled 3D shape representations for real-world data is time-consuming and expensive, synthetic image-shape pairs are widely used for 3D reconstruction. However, the models trained on synthetic data set did not perform equally well on real-world images. The existing method used the domain adaptation to fill the domain gap between different data sets. Unlike the approach simply considered global distribution for domain adaptation, this paper presents a dynamic domain adaptation (DDA) network to extract domain-invariant image features for 3D reconstruction. The relative importance between global and local distributions are considered to reduce the discrepancy between synthetic and real-world data. In addition, graph convolution network (GCN) based mesh generation methods have achieved impressive results than voxel-based and point cloud-based methods. However, the global context in a graph is not effectively used due to the limited receptive field of GCN. In this paper, a multi-scale processing method for graph convolution network (GCN) is proposed to further improve the performance of GCN-based 3D reconstruction. The experiment results conducted on both synthetic and real-world data set have demonstrated the effectiveness of the proposed methods.
Housen Xie, Haihong Tian, Yuanlong Yu 0001
IROS4
2020 An Internal Covariate Shift Bounding Algorithm for Deep Neural Networks by Unitizing Layers' Outputs
abstract
Batch Normalization (BN) techniques have been proposed to reduce the so-called Internal Covariate Shift (ICS) by attempting to keep the distributions of layer outputs unchanged. Experiments have shown their effectiveness on training deep neural networks. However, since only the first two moments are controlled in these BN techniques, it seems that a weak constraint is imposed on layer distributions and furthermore whether such constraint can reduce ICS is unknown. Thus this paper proposes a measure for ICS by using the Earth Mover (EM) distance and then derives the upper and lower bounds for the measure to provide a theoretical analysis of BN. The upper bound has shown that BN techniques can control ICS only for the outputs with low dimensions and small noise whereas their control is not effective in other cases. This paper also proves that such control is just a bounding of ICS rather than a reduction of ICS. Meanwhile, the analysis shows that the high-order moments and noise, which BN cannot control, have great impact on the lower bound. Based on such analysis, this paper furthermore proposes an algorithm that unitizes the outputs with an adjustable parameter to further bound ICS in order to cope with the problems of BN. The upper bound for the proposed unitization is noise-free and only dominated by the parameter. Thus, the parameter can be trained to tune the bound and further to control ICS. Besides, the unitization is embedded into the framework of BN to reduce the information loss. The experiments show that this proposed algorithm outperforms existing BN techniques on CIFAR-10, CIFAR-100 and ImageNet datasets.
You Huang, Yuanlong Yu 0001
CVPR2
2019 Context-Aware Spatio-Recurrent Curvilinear Structure Segmentation
abstract
Curvilinear structures are frequently observed in various images in different forms, such as blood vessels or neuronal boundaries in biomedical images. In this paper, we propose a novel curvilinear structure segmentation approach using context-aware spatio-recurrent networks. Instead of directly segmenting the whole image or densely segmenting fixed-sized local patches, our method recurrently samples patches with varied scales from the target image with learned policy and processes them locally, which is similar to the behavior of changing retinal fixations in the human visual system and it is beneficial for capturing the multi-scale or hierarchical modality of the complex curvilinear structures. In specific, the policy of choosing local patches is attentively learned based on the contextual information of the image and the historical sampling experience. In this way, with more patches sampled and refined, the segmentation of the whole image can be progressively improved. To validate our approach, comparison experiments on different types of image data are conducted and the sampling procedures for exemplar images are illustrated. We demonstrate that our method achieves the state-of-the-art performance in public datasets.
Feigege Wang, Wenxi Liu, Yuanlong Yu 0001, Shengfeng He, Jia Pan 0001
CVPR4
2019 Visualizing the Invisible: Occluded Vehicle Segmentation and Recovery
abstract
In this paper, we propose a novel iterative multi-task framework to complete the segmentation mask of an occluded vehicle and recover the appearance of its invisible parts. In particular, firstly, to improve the quality of the segmentation completion, we present two coupled discriminators that introduce an auxiliary 3D model pool for sampling authentic silhouettes as adversarial samples. In addition, we propose a two-path structure with a shared network to enhance the appearance recovery capability. By iteratively performing the segmentation completion and the appearance recovery, the results will be progressively refined. To evaluate our method, we present a dataset, Occluded Vehicle dataset, containing synthetic and real-world occluded vehicle images. Based on this dataset, we conduct comparison experiments and demonstrate that our model outperforms the state-of-the-arts in both tasks of recovering segmentation mask and appearance for occluded vehicles. Moreover, we also demonstrate that our appearance recovery approach can benefit the occluded vehicle tracking in real-world videos.
Xiaosheng Yan, Yuanlong Yu 0001, Feigege Wang, Wenxi Liu, Shengfeng He, Jia Pan 0001
ICCV2
2019 Deformable Object Tracking With Gated Fusion
abstract
The tracking-by-detection framework receives growing attention through the integration with the convolutional neural networks (CNNs). Existing tracking-by-detection-based methods, however, fail to track objects with severe appearance variations. This is because the traditional convolutional operation is performed on fixed grids, and thus may not be able to find the correct response while the object is changing pose or under varying environmental conditions. In this paper, we propose a deformable convolution layer to enrich the target appearance representations in the tracking-by-detection framework. We aim to capture the target appearance variations via deformable convolution, which adaptively enhances its original features. In addition, we also propose a gated fusion scheme to control how the variations captured by the deformable convolution affect the original appearance. The enriched feature representation through deformable convolution facilitates the discrimination of the CNN classifier on the target object and background. The extensive experiments on the standard benchmarks show that the proposed tracker performs favorably against the state-of-the-art methods.
Wenxi Liu, Yibing Song, Dengsheng Chen, Shengfeng He, Yuanlong Yu 0001, Tao Yan 0001, Gerhard P. Hancke 0002, Rynson W. H. Lau
IEEE Trans. Image Process.5
2018 An efficient cascaded method for network intrusion detection based on extreme learning machines
Yuanlong Yu 0001, Zhifan Ye, Xianghan Zheng, Chunming Rong
J. Supercomput.1
2017 A multi-label classification algorithm based on kernel extreme learning machine
Fangfang Luo, Wenzhong Guo, Yuanlong Yu 0001
Neurocomputing3
2017 Sparse coding extreme learning machine for classification
Yuanlong Yu 0001, Zhenzhen Sun
Neurocomputing1
2017 Visual-Tactile Fusion for Object Recognition
abstract
The camera provides rich visual information regarding objects and becomes one of the most mainstream sensors in the automation community. However, it is often difficult to be applicable when the objects are not visually distinguished. On the other hand, tactile sensors can be used to capture multiple object properties, such as textures, roughness, spatial features, compliance, and friction, and therefore provide another important modality for the perception. Nevertheless, effective combination of the visual and tactile modalities is still a challenging problem. In this paper, we develop a visual–tactile fusion framework for object recognition tasks. This paper uses the multivariate-time-series model to represent the tactile sequence and the covariance descriptor to characterize the image. Further, we design a joint group kernel sparse coding (JGKSC) method to tackle the intrinsically weak pairing problem in visual–tactile data samples. Finally, we develop a visual–tactile data set, composed of 18 household objects for validation. The experimental results show that considering both visual and tactile inputs is beneficial and the proposed method indeed provides an effective strategy for fusion.
Huaping Liu 0001, Yuanlong Yu 0001, Fuchun Sun 0001, Jason Gu
IEEE Trans Autom. Sci. Eng.2
2017 An Efficient Method for Traffic Sign Recognition Based on Extreme Learning Machine
abstract
This paper proposes a computationally efficient method for traffic sign recognition (TSR). This proposed method consists of two modules: 1) extraction of histogram of oriented gradient variant (HOGv) feature and 2) a single classifier trained by extreme learning machine (ELM) algorithm. The presented HOGv feature keeps a good balance between redundancy and local details such that it can represent distinctive shapes better. The classifier is a single-hidden-layer feedforward network. Based on ELM algorithm, the connection between input and hidden layers realizes the random feature mapping while only the weights between hidden and output layers are trained. As a result, layer-by-layer tuning is not required. Meanwhile, the norm of output weights is included in the cost function. Therefore, the ELM-based classifier can achieve an optimal and generalized solution for multiclass TSR. Furthermore, it can balance the recognition accuracy and computational cost. Three datasets, including the German TSR benchmark dataset, the Belgium traffic sign classification dataset and the revised mapping and assessing the state of traffic infrastructure (revised MASTIF) dataset, are used to evaluate this proposed method. Experimental results have shown that this proposed method obtains not only high recognition accuracy but also extremely high computational efficiency in both training and recognition processes in these three datasets.
Zhiyong Huang 0005, Yuanlong Yu 0001, Jason Gu, Huaping Liu 0001
IEEE Trans. Cybern.2
2016 A pruning algorithm for extreme learning machine based on sparse coding
abstract
This paper presents a pruned sparse extreme learning machine (PS-ELM) algorithm, which can generate a compact single-hidden-layer neural network (SLNN) by automatically pruning the number of hidden nodes while keep high accuracy. In this PS-ELM algorithm, input connections between input and hidden layers are base vectors, which can sparsely map the input features into hidden layer by using gradient projection (GP) algorithm; The output weights between hidden and output layers can map the sparse features into class labels. This PS-ELM algorithm initializes the SLNN given superfluous number of hidden nodes. The subsequent training process consists of four iterative steps. The first one is to update base vectors by using Lagrange dual optimization. The second one is to prune zero base vectors which are considered to be insignificant. The third one is sparse coding which re-encodes the training samples given remained base vectors. The fourth one is to update output weight matrix by using ELM-like algorithm. This iterative process stops once the number of hidden nodes at current time is equal to the number at the last time. This PS-ELM algorithm can improve the sparsity and distinction of hidden layer feature representations. Meanwhile, the pruning is independent on hand-designed threshold. Experimental results on benchmark datasets have shown that the PS-ELM algorithm can automatically achieve a reasonable compact network structure while keep comparable or much higher accuracy in classification.
Yuanlong Yu 0001, Zhenzhen Sun
IJCNN1
2016 Protein Function Detection Based on Machine Learning: Survey and Possible Solutions
abstract
With the completion of the Human Genome Project, proteomics research has become one of the most important topics in the fields of life science and natural science. The project determined that proteins participate in life activities mainly in the form of complexes. At present, research on protein-protein interaction networks (PPINs) have mainly focused on detecting protein complexes or function modules. This problem has been transformed into a recognizable dense subgraph problem in a PPIN diagram.The situation in PPIN research in recent years is introduced in this study, including commonly used databases, traditional detection algorithms, recent solutions, and the application of the swarm intelligence algorithms in this field. We then propose a detection scheme based on particle swarm optimization (PSO) and gene ontology knowledge. This scheme combines PSO and biological gene ontology knowledge to identify complexes from PPINs. Simultaneously, network topology knowledge improves the detection accuracy of the protein module.
Xianghan Zheng, Chunming Rong, Yuanlong Yu 0001, Riqing Chen
ISPDC4
2016 ELM-based spammer detection in social networks
Xianghan Zheng, Yuanlong Yu 0001, M. Tahar Kechadi, Chunming Rong
J. Supercomput.3
2015 Detecting spammers on social networks
abstract
Social network has become a very popular way for internet users to communicate and interact online. Users spend plenty of time on famous social networks (e.g., Facebook, Twitter, Sina Weibo, etc.), reading news, discussing events and posting messages. Unfortunately, this popularity also attracts a significant amount of spammers who continuously expose malicious behavior (e.g., post messages containing commercial URLs, following a larger amount of users, etc.), leading to great misunderstanding and inconvenience on users׳ social activities. In this paper, a supervised machine learning based solution is proposed for an effective spammer detection. The main procedure of the work is: first, collect a dataset from Sina Weibo including 30,116 users and more than 16 million messages. Then, construct a labeled dataset of users and manually classify users into spammers and non-spammers. Afterwards, extract a set of feature from message content and users׳ social behavior, and apply into SVM (Support Vector Machines) based spammer detection algorithm. The experiment shows that the proposed solution is capable to provide excellent performance with true positive rate of spammers and non-spammers reaching 99.1% and 99.9% respectively.
Xianghan Zheng, Zhipeng Zeng, Zheyi Chen, Yuanlong Yu 0001, Chunming Rong
Neurocomputing4
2014 Spammer Detection on Weibo Social Network
abstract
Social network has become a very popular way for internet users to communicate and interact online. Users spend a great deal of time on famous social networks (e.g. Facebook, Twitter, Sina Weibo, etc.), reading news, discussing events and posting their messages. Unfortunately, this popularity also attracts a significant amount of spammers who continuously expose malicious behaviors (e.g. Post messages containing commercial topics or URLs, following a larger amount of users, etc.), leading to great inconvenience on normal users' social activities. In this paper, a supervised machine learning based spammer filtering method is proposed. We first collected a dataset from Sina Weibo that includes 30,116 users and more than 16 million messages, then, construct a labeled dataset of users and manually classify users into spammers and non-spammers, after that, abstract a set of novel features from message content and users' social behavior, and apply into SVM based spammer classifier. Our experiments show that true positive rate of spammers and non-spammers could reach 99.1% and 99.9%.
Zhipeng Zeng, Xianghan Zheng, Yuanlong Yu 0001
CloudCom4
2014 Simultaneous prototype selection and outlier isolation for traffic sign recognition: A collaborative sparse optimization method
abstract
Video-based traffic sign recognition is one of the most important task for unmanned autonomous vehicle. However, there always exists unavoidable outliers in the practical scenario. Therefore, robust prototype extraction from the noisy sample set is highly expected to help traffic sign recognition in video sequence. In this paper, we propose a novel approach for simultaneous prototype extraction and outlier isolation through collaborative sparse learning. The new model accounts for not only the reconstruction capability and the sparsity, but also the robustness. To solve the optimization problem, we adopt the Alternating Directional Method of Multiplier (ADMM) technology to design an iterative algorithm. Finally, the effectiveness of the approach is demonstrated by experiments on GTSRB dataset.
Huaping Liu 0001, Yuanlong Yu 0001, Fuchun Sun 0001
ICRA3
2014 Bhattacharyya distance-based irregular pyramid method for image segmentation
abstract
This paper proposes a new unsupervised image segmentation method by using Bhattacharyya distance‐based irregular pyramid, termed as ‘BDIP’ algorithm. The proposed BDIP algorithm obtains a suboptimal labelling solution under the condition that the number of segments is not manually given. It hierarchically builds each level of the irregular pyramid, with the result that the final segments emerge as they are represented by single nodes at certain levels. The BDIP algorithm employs Bhattacharyya distance to estimate the intra‐level similarity at higher pyramidal levels so as to improve the accuracy and robustness to noise. Furthermore, an adaptive neighbour search method is proposed such that the BDIP algorithm can self‐determine the number of segments. This method considers not only the graphic constraint, but also the similarity constraint in the sense that a candidate node is selected as a neighbour of the centre node if there is no boundary evidence between these two nodes. With the pyramidal accumulation, this evaluation is aggregated into the approximately global evidence, based on which the number of segments can be self‐determined. Experimental results have shown that this proposed BDIP algorithm outperforms other benchmark segmentation algorithms in terms of segmentation accuracy, labelling cost and robustness to noise.
Yuanlong Yu 0001, Jason Gu
IET Comput. Vis.1
2014 Diversified Key-Frame Selection Using Structured ${L_{2, 1}}$ Optimization
abstract
In this paper, a structured L2,1optimization model, which simultaneously characterizes the reconstruction capability and diversity, is proposed to provide a semantically meaningful representation of a short video clip acquired from digital cameras or a mobile robot. In this model, a mutual inhabitation penalty term is imposed to prevent similar samples from being selected simultaneously. The proposed model is highly flexible to incorporate different mutual inhabitation terms and the temporal redundancy in video is exploited to encourage the diversity. The constructed objective function is nonconvex and an iterative algorithm is developed to solve the optimization problem. The performance is evaluated using various video clips from YouTube and also based on practical video captured by an indoor mobile robot. The results clearly indicate that the proposed strategy helps the optimization model to achieve more diversified key frames than the other existing work method.
Huaping Liu 0001, Yunhui Liu 0003, Yuanlong Yu 0001, Fuchun Sun 0001
IEEE Trans. Ind. Informatics3
2013 Development and Evaluation of Object-Based Visual Attention for Automatic Perception of Robots
abstract
Bottom-up visual attention is an automatic behavior to guide visual perception to a conspicuous object in a scene. This paper develops a new object-based bottom-up attention (OBA) model for robots. This model includes four modules: Extraction of preattentive features, preattentive segmentation, estimation of space-based saliency, and estimation of proto-object-based saliency. In terms of computation, preattentive segmentation serves as a bridge to connect the space-based saliency and object-based saliency. This paper therefore proposes a preattentive segmentation algorithm, which is able to self-determine the number of proto-objects, has low computational cost, and is robust in a variety of conditions such as noise and spatial transformations. Experimental results have shown that the proposed OBA model outperforms space-based attention model and other object-based attention methods in terms of accuracy of attentional selection, consistency under a series of noise settings and object completion.
Yuanlong Yu 0001, Jason Gu, George K. I. Mann, Ray G. Gosine
IEEE Trans Autom. Sci. Eng.1
2010 Target tracking for moving robots using object-based visual attention
abstract
Visual tracking is a quite challenging issue for a moving robot due to the appearance changes of both the background and targets, large variation of motion, partial or full occlusion and so on. However, humans are capable to cope with those difficulties to achieve satisfactory tracking performance. Thus this paper presents a biologically-inspired method of visual tracking for moving robots by using object-based visual attention mechanism. This tracking method consists of four modules: pre-attentive segmentation, top-down attentional biasing, post-attentive completion processing and online learning of the target model. Experimental results in natural and cluttered scenes are shown to validate this general and robust tracking method.
Yuanlong Yu 0001, George K. I. Mann, Ray G. Gosine
IROS1
2010 An Object-Based Visual Attention Model for Robotic Applications
abstract
By extending integrated competition hypothesis, this paper presents an object-based visual attention model, which selects one object of interest using low-dimensional features, resulting that visual perception starts from a fast attentional selection procedure. The proposed attention model involves seven modules: learning of object representations stored in a long-term memory (LTM), preattentive processing, top-down biasing, bottom-up competition, mediation between top-down and bottom-up ways, generation of saliency maps, and perceptual completion processing. It works in two phases: learning phase and attending phase. In the learning phase, the corresponding object representation is trained statistically when one object is attended. A dual-coding object representation consisting of local and global codings is proposed. Intensity, color, and orientation features are used to build the local coding, and a contour feature is employed to constitute the global coding. In the attending phase, the model preattentively segments the visual field into discrete proto-objects using Gestalt rules at first. If a task-specific object is given, the model recalls the corresponding representation from LTM and deduces the task-relevant feature(s) to evaluate top-down biases. The mediation between automatic bottom-up competition and conscious top-down biasing is then performed to yield a location-based saliency map. By combination of location-based saliency within each proto-object, the proto-object-based saliency is evaluated. The most salient proto-object is selected for attention, and it is finally put into the perceptual completion processing module to yield a complete object region. This model has been applied into distinct tasks of robots: detection of task-specific stationary and moving objects. Experimental results under different conditions are shown to validate this model.
Yuanlong Yu 0001, George K. I. Mann, Ray G. Gosine
IEEE Trans. Syst. Man Cybern. Part B1
2008 An object-based visual attention model for robots
abstract
In this paper an object-based visual attention model extending Duncan’s integrated competition hypothesis is presented for robots. Based on Gestalt rules the model segments the visual field into primitive groupings by evaluating both edge continuity and color similarity. An object representation is also built in long-term memory by using contour and color features. Dependent on the task and object representation, top-down modulation performs on pre-attentive features, followed by bottom-up competition. The object-based salience is evaluated by combination of pixel-wise salience within each pre-attentive grouping. The attended object is finally refined to reach an accurate representation in working memory. This model has been applied into two tasks of mobile robots: task-specific still and moving object detection. Experimental results in cluttered scenes are shown to validate this model.
Yuanlong Yu 0001, George K. I. Mann, Ray G. Gosine
ICRA1