Jin Wu 0002

dblp:18/1539-2 · DBLP profile ↗
← Back
55ranked-venue papers
11as first author
42since 2021 · last 2026
0000-0001-5930-4170ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 6 first-author · 12 since 2021Systems, architecture and hardware · 10 · 2 first-author · 10 since 2021Computer networks · 5 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Explainable multiscale representation learning for anticancer peptide prediction
Yongqing Zhang 0001, Zhigan Zhou, Yugui Xu, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047
Eng. Appl. Artif. Intell.7
2026 Multiple minds are better than one: Enhancing temporal knowledge graph forecasting with mixture of diverse graph experts
Yichen Xin, Hanting Shen, Shichong Li, Zhangtao Cheng, Xueting Liu 0005, Jin Wu 0002, Fan Zhou 0002
Inf. Process. Manag.6
2026 GNN-Based Spatio-Temporal Manifold Learning: An Application of Landslide Prediction
Liu Yu 0001, Rongfan Li, Kunpeng Zhang 0001, Siyuan Liu 0001, Goce Trajcevski, Jin Wu 0002, Fan Zhou 0002
Mach. Learn.6
2026 MCFusion-DDI: Multimodal cross-attention fusion of local-global features and latent drug associations for explainable DDI prediction
Yongqing Zhang 0001, Yugui Xu, Zhigan Zhou, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047
Neural Networks5
2026 CapSRA: LMM-guided caption semantic retrieval aggregation for hateful meme detection
Yili Li, Yong Wang 0046, Jin Wu 0002, Kunpeng Zhang 0001, Fan Zhou 0002
Pattern Recognit.5
2026 Guiding multimodal LLMs for efficient visual place recognition
Zhijian He, Jintao Cheng, Yipu Zhang 0002, Chi-Man Vong, Jin Wu 0002, Xieyuanli Chen
Pattern Recognit. Lett.6
2026 DOSE+: A Timestep-Aware Dropout Strategy for Diffusion Models in Speech Enhancement
abstract
Diffusion-based speech enhancement (SE) models have recently demonstrated superior performance compared to traditional single-step models. In this work, we revisit the advantages of diffusion models from a multi-source learning perspective, highlighting that their ability to jointly leverage data likelihood and conditional mapping makes them theoretically superior to deterministic models when controllability is ensured. From this standpoint, we identify a key limitation in DOSE, a recent diffusion-based SE model that enhances controllability by applying fixed dropout ratio to non-conditional inputs, leading to unnecessary information loss at every timestep. To address this, we propose a timestep-aware dropout mechanism that dynamically adjusts the dropout intensity at each denoising step. Extensive experiments across matched and cross-dataset benchmarks show that our method consistently outperforms DOSE and other state-of-the-art diffusion-based SE methods, achieving superior speech enhancement with high efficiency. The code and audio samples are publicly available athttps://github.com/ICDM-UESTC/DOSE.
Siqi Yang 0008, Jin Wu 0002, Yue Lei, Wenxin Tai, Fan Zhou 0002
IEEE Signal Process. Lett.2
2026 Greedy Priority Inheritance With Backtracking for Multi-Agent Pathfinding Problem
abstract
Some real-world transportation systems require moving robots from their start points to goal points without collision, which can be formalized as the multi-agent pathfinding (MAPF) problem. The Priority Inheritance with Backtracking (PIBT) algorithm is an efficient approach for MAPF with low-degree polynomial time complexity and proven completeness. PIBT employs a dynamic prioritization scheme to determine the order of sequential agent planning. However, this scheme may lead to low solution quality, which is measured by the sum of time steps each agent takes to reach its goal for the first time. In this work, we propose two greedy variants of the PIBT algorithm that improve solution quality while preserving completeness: Backflow-based Greedy PIBT (GPIBT-B) and Reduction-based Greedy PIBT (GPIBT-R). GPIBT-B introduces a backflow mechanism that allows a determined agent to adjust its plan during the planning of another agent, achieving higher solution quality while maintaining the same low time complexity as PIBT. GPIBT-R formulates the planning problem as a Mixed-Integer Linear Programming (MILP), enabling the use of efficient MILP solvers to find high-quality solutions. Experimental results show that GPIBT-B substantially improves solution quality with minimal additional computation time, while GPIBT-R achieves even better solution quality at the cost of increased computational time.
Mingkai Tang 0002, Yuanhang Li, Lu Gan 0001, Chengxi Zhang, Yuxiang Sun 0002, Jin Wu 0002
IEEE Trans Autom. Sci. Eng.6
2026 MATCH: Multi-Agentic Evidence Grounding for Explainable Hate Video Detection
abstract
The growing prevalence of hate videos promoting intolerance, bigotry, and discrimination presents significant psychosocial threats to both individuals and society. Current detection methods often rely on black-box models, which lack interpretability — a crucial factor for fostering more reliable content moderation and trustworthy AI. To bridge this gap, we propose MATCH, the first attempt to achieve interpretable hate video detection via multiple Large Multimodal Model (LMM) agent collaboration. Our method facilitates a noveldiversely generate-then-verifyparadigm, where LMM agents work in tandem to generate diverse clues and verify them to yield more faithful explanations. MATCH begins by proposing a new Dual-Perspective Proposing paradigm, where two LMM agents are regarded asProposersto independently identify evidential clues from opposing angles – hate and non-hate. Leveraging these comprehensive clues, we introduce an innovative Spatiotemporal Evidence-Grounded Verification mechanism. In this mechanism, a third LMM agent acts as aVerifier, rigorously validating and reconciling the proposed clues against spatiotemporal evidences directly extracted from video content, yielding coherent and faithful explanations. Finally, these explanations are integrated with video features, enabling accurate identification of complex and ambiguous hateful content. Extensive experiments conducted on three benchmark datasets demonstrate that MATCH not only achieves state-of-the-art performance, but also provides reliable and trustworthy rationales for the predictions. Code is available at https://anonymous.4open.science/r/MATCH-HVD.
Kaiju Li, Rongpei Hong, Jian Lang, Jin Wu 0002, Fan Zhou 0002, Jingkuan Song
IEEE Trans. Circuits Syst. Video Technol.4
2025 GS-LIVM: Real-Time Photo-Realistic LiDAR-Inertial-Visual Mapping with Gaussian Splatting
abstract
In this paper, we introduce GS-LIVM, a real-time photo-realistic LiDAR-Inertial-Visual mapping framework with Gaussian Splatting tailored for outdoor scenes. Compared to existing methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), our approach enables real-time photo-realistic mapping while ensuring high-quality image rendering in large-scale unbounded outdoor environments. In this work, Gaussian Process Regression (GPR) is employed to mitigate the issues resulting from sparse and unevenly distributed LiDAR observations. The voxel-based 3D Gaussians map representation facilitates real-time dense mapping in large outdoor environments with acceleration governed by custom CUDA kernels. Moreover, the overall framework is designed in a covariance-centered manner, where the estimated covariance is used to initialize the scale and rotation of 3D Gaussians, as well as update the parameters of the GPR. We evaluate our algorithm on several outdoor datasets, and the results demonstrate that our method achieves state-of-the-art performance in terms of mapping efficiency and rendering quality. The source code is available on GitHub.
Yusen Xie, Zhenmin Huang, Jin Wu 0002, Jun Ma 0008
ICCV3
2025 From Satellite to Street: Semantic and Depth Information for Enhanced Geo-Localization
abstract
Accurate positioning is essential for autonomous driving, but localization using 2D maps is challenging due to the domain gap between perspective view and 2D map. While GNSS accuracy is often limited by atmospheric effects, multipath, and signal blockages. We propose a novel positioning method that combines perspective view images with satellite images retrieved based on rough GNSS positions to achieve precise three-degree-of-freedom (3-DoF) pose estimation. Our method leverages the Swin Transformer for satellite image processing and semantic completion for monocular image analysis. By extracting depth and semantic information from monocular images, we convert these to overhead projections, effectively bridging the gap between different viewpoints. This cross-view transformation allows for precise alignment of features from monocular images onto semantically enriched satellite images. Additionally, we integrate a robust global position estimator using the semantic information from satellite images to further enhance accuracy and robustness. The experimental results demonstrate that our method excels in various complex scenarios; we successfully improved the positioning accuracy within 1 m to 80.67% and the heading in 1° to 33.78%. However, longitudinal localization remains more challenging, with higher errors than lateral positioning.
Yilong Zhu, Jianhao Jiao, Hexiang Wei, Jin Wu 0002, Bohuan Xue, Shaojie Shen
IROS4
2025 Real-Time AIoT for AAV Antenna Interference Detection via Edge-Cloud Collaboration
abstract
In the fifth-generation (5G) era, eliminating communication interference sources is crucial for maintaining network performance. Interference often originates from unauthorized or malfunctioning antennas, and radio monitoring agencies must address numerous sources of such antennas annually. Autonomous aerial vehicles (AAVs) can improve inspection efficiency. However, the data transmission delay in the existing cloud-only (CO) artificial intelligence (AI) mode fails to meet the low latency requirements for real-time performance. Therefore, we propose a computer vision-based AI of Things (AIoT) system to detect antenna interference sources for AAVs. The system adopts an optimized edge-cloud collaboration (ECC+) mode, combining a keyframe selection algorithm (KSA), focusing on reducing end-to-end latency (E2EL) and ensuring reliable data transmission, which aligns with the core principles of ultrareliable low-latency communication (URLLC). At the core of our approach is an end-to-end antenna localization scheme based on the tracking-by-detection (TBD) paradigm, including a detector (EdgeAnt) and a tracker (AntSort). EdgeAnt achieves state-of-the-art (SOTA) performance with a mean average precision (mAP) of 42.1% on our custom antenna interference source dataset, requiring only three million parameters and 14.7 GFLOPs. On the COCO dataset, EdgeAnt achieves 38.9% mAP with 5.4 GFLOPs. We deployed EdgeAnt on Jetson Xavier NX (TRT) and Raspberry Pi 4B (NCNN), achieving real-time inference speeds of 21.1 (1088) and 4.8 (640) frames/s (FPS), respectively. Compared with CO mode, the ECC+ mode reduces E2EL by 88.9%, increases accuracy by 28.2%. Additionally, the system offers excellent scalability for coordinated multiple AAVs inspections. The detector code is publicly available athttps://github.com/SCNU-RISLAB/EdgeAnt.
Jintao Cheng, Jin Wu 0002, Chengxi Zhang, Shunyi Zhao
IEEE Internet Things J.3
2025 Real-Time Metric-Semantic Mapping for Autonomous Navigation in Outdoor Environments
abstract
The creation of a metric-semantic map, which encodes human-prior knowledge, represents a high-level abstraction of environments. However, constructing such a map poses challenges related to the fusion of multi-modal sensor data, the attainment of real-time mapping performance, and the preservation of structural and semantic information consistency. In this paper, we introduce an online metric-semantic mapping system that utilizes LiDAR-Visual-Inertial sensing to generate a global metric-semantic mesh map of large-scale outdoor environments. Leveraging GPU acceleration, our mapping process achieves exceptional speed, with frame processing taking less than$7ms$, regardless of scenario scale. Furthermore, we seamlessly integrate the resultant map into a real-world navigation system, enabling metric-semantic-based terrain assessment and autonomous point-to-point navigation within a campus environment. Through extensive experiments conducted on both publicly available and self-collected datasets comprising 24 sequences, we demonstrate the effectiveness of our mapping and navigation methodologies. Note to Practitioners—This paper tackles the challenge of autonomous navigation for mobile robots in complex, unstructured environments with rich semantic elements. Traditional navigation relies on geometric analysis and manual annotations, struggling to differentiate similar structures like roads and sidewalks. We propose an online mapping system that creates a global metric-semantic mesh map for large-scale outdoor environments, utilizing GPU acceleration for speed and overcoming the limitations of existing real-time semantic mapping methods, which are generally confined to indoor settings. Our map integrates into a real-world navigation system, proven effective in localization and terrain assessment through experiments with both public and proprietary datasets. Future work will focus on integrating kernel-based methods to improve the map’s semantic accuracy.
Jianhao Jiao, Ruoyu Geng, Yuanhang Li, Ren Xin, Jin Wu 0002, Lujia Wang 0001, Ming Liu 0001, Rui Fan 0001, Dimitrios Kanoulas
IEEE Trans Autom. Sci. Eng.6
2025 Enhancing Attitude Tracking With Self-Learning Control Using Tanh-Type Learning Intensity
abstract
This paper investigates the attitude tracking control problem for spacecraft. A tanh-type self-learning control (TSLC) approach with variable learning intensity (VLI) is proposed, which avoids saturation while overcoming previous algorithms’ long response time disadvantage. Unlike the previously introduced VLI method, the enhanced TSLC does not tweak the learning intensity based on the previous controller output. Instead, it relates learning intensity to an intermediate variable directly related to the system state and tunes the learning intensity using a tanh-type function. Since the system state reflects the tracking error in real-time, the transformed tanh-type function has a higher decay rate than the exponential function, which not only significantly reduces the saturation response but also improves the response speed and achieves higher steady-state accuracy. Simulation proved TSLC’s superiority, considering adverse actuator factors such as dead zone, bias torque, and saturation. The proposed approach has also been validated on the Quanser helicopter platform, confirming its better performance.
Chengxi Zhang, Weijia Lu, Shunyi Zhao, Jin Wu 0002, Zhijie Liu 0001, Wei He 0001
IEEE Trans Autom. Sci. Eng.4
2025 Factor Graph Optimization for Flexibly Modeled INS/GPS Navigation in Graphical State-Space
abstract
This article investigates loosely coupled inertial navigation system/global positioning system (INS/GPS) integration for land vehicle navigation. To achieve navigation with higher accuracy and lower computational complexity, we present an integration solution using factor graph optimization (FGO) based on the graphical state-space model (GSSM). This solution is referred to as GSSM-FGO. Compared with traditional methods, the unique specialty of our work lies in both modeling and problem-solving aspects under the assumption of calibration parameter invariance. Specifically, we suggest that the time-series state-space model is not always suitable for widely existing constant calibration parameters. Thus, we propose GSSM as a more flexible and accurate state description by extracting the constant states as singular nodes. The FGO is adopted to manage this novel graphical model, while traditional filter-based algorithms fail when faced with the cyclic model structure. The universality of our approach is validated through a real-world land vehicle navigation dataset, featuring four distinct-grade inertial measurement units. Compared to the methods based on extended Kalman filter and FGO with the traditional state-space model, our approach demonstrates a substantial enhancement in estimation accuracy and computational speed.
Shunyi Zhao, Chengxi Zhang, Jin Wu 0002, Biao Huang 0001
IEEE Trans. Ind. Informatics4
2024 RELEAD: Resilient Localization with Enhanced LiDAR Odometry in Adverse Environments
abstract
LiDAR-based localization is valuable for applications like mining surveys and underground facility maintenance. However, existing methods can struggle when dealing with uninformative geometric structures in challenging scenarios. This paper presents RELEAD, a LiDAR-centric solution designed to address scan-matching degradation. Our method enables degeneracy-free point cloud registration by solving constrained ESIKF updates in the front end and incorporates multisensor constraints, even when dealing with outlier measurements, through graph optimization based on Graduated Non-Convexity (GNC). Additionally, we propose a robust Incremental Fixed Lag Smoother (rIFL) for efficient GNC-based optimization. RELEAD has undergone extensive evaluation in degenerate scenarios and has outperformed existing state-of-the-art LiDAR-Inertial odometry and LiDAR-Visual-Inertial odometry methods.
Yuhua Qi, Shipeng Zhong, Dapeng Feng, Jin Wu 0002, Weisong Wen, Ming Liu 0001
ICRA6
2024 MF-MOS: A Motion-Focused Model for Moving Object Segmentation
abstract
Moving object segmentation (MOS) provides a reliable solution for detecting traffic participants and thus is of great interest in the autonomous driving field. Dynamic capture is always critical in the MOS problem. Previous methods capture motion features from the range images directly. Differently, we argue that the residual maps provide greater potential for motion information, while range images contain rich semantic guidance. Based on this intuition, we propose MF-MOS, a novel motion-focused model with a dual-branch structure for LiDAR moving object segmentation. Novelly, we decouple the spatial-temporal information by capturing the motion from residual maps and generating semantic features from range images, which are used as movable object guidance for the motion branch. Our straightforward yet distinctive solution can make the most use of both range images and residual maps, thus greatly improving the performance of the LiDAR-based MOS task. Remarkably, our MF-MOS achieved a leading IoU of 76.7% on the MOS leaderboard of the SemanticKITTI dataset upon submission, demonstrating the current state-of-the-art performance. The implementation of our MF-MOS has been released at https://github.com/SCNU-RISLAB/MF-MOS.
Jintao Cheng, Kang Zeng, Zhuoxu Huang, Jin Wu 0002, Chengxi Zhang, Xieyuanli Chen, Rui Fan 0001
ICRA5
2024 Generalized Correspondence Matching via Flexible Hierarchical Refinement and Patch Descriptor Distillation
abstract
Correspondence matching plays a crucial role in numerous robotics applications. In comparison to conventional hand-crafted methods and recent data-driven approaches, there is significant interest in plug-and-play algorithms that make full use of pre-trained backbone networks for multi-scale feature extraction and leverage hierarchical refinement strategies to generate matched correspondences. The primary focus of this paper is to address the limitations of deep feature matching (DFM), a state-of-the-art (SoTA) plug-and-play correspondence matching approach. First, we eliminate the pre-defined threshold employed in the hierarchical refinement process of DFM by leveraging a more flexible nearest neighbor search strategy, thereby preventing the exclusion of repetitive yet valid matches during the early stages. Our second technical contribution is the integration of a patch descriptor, which extends the applicability of DFM to accommodate a wide range of backbone networks pre-trained across diverse computer vision tasks, including image classification, semantic segmentation, and stereo matching. Taking into account the practical applicability of our method in real-world robotics applications, we also propose a novel patch descriptor distillation strategy to further reduce the computational complexity of correspondence matching. Extensive experiments conducted on three public datasets demonstrate the superior performance of our proposed method. Specifically, it achieves an overall performance in terms of mean matching accuracy of 0.68, 0.92, and 0.95 with respect to the tolerances of 1, 3, and 5 pixels, respectively, on the HPatches dataset, outperforming all other SoTA algorithms. Our source code, demo video, and supplement are publicly available at mias.group/GCM.
Ziwei Long, Yanting Zhang 0001, Jin Wu 0002, Zhijun Fang 0001, Rui Fan 0001
ICRA4
2024 An Image Acquisition Scheme for Visual Odometry based on Image Bracketing and Online Attribute Control
abstract
Visual odometry (VO) system is challenged by complex illumination environments. Image quality and its consistency in the time domain directly determine feature detection and tracking performance, which further affect the robustness and accuracy of the entire system. In this paper, an image acquisition scheme with image bracketing patterns is proposed. Images with different exposure levels are continuously captured to sufficiently explore the scene under varying illumination. An attribute control method is designed to adjust image exposures within the brackets online. Gaussian process regression fits the relationship between image quality metric and exposure via image synthesis technique. The optimal exposures for the next bracket are obtained directly without attempts to ensure a quick response. Experiments show our acquisition system’s effectiveness and performance improvement for VO tasks in complex illumination scenes.
Jinhao He, Bohuan Xue, Jin Wu 0002, Pengyu Yin, Jianhao Jiao, Ming Liu 0001
ICRA4
2024 CoLRIO: LiDAR-Ranging-Inertial Centralized State Estimation for Robotic Swarms
abstract
Collaborative state estimation using different heterogeneous sensors is a fundamental prerequisite for robotic swarms operating in GPS-denied environments, posing a significant research challenge. In this paper, we introduce a centralized system to facilitate collaborative LiDAR-ranging-inertial state estimation, enabling robotic swarms to operate without the need for anchor deployment. The system efficiently distributes computationally intensive tasks to a central server, thereby reducing the computational burden on individual robots for local odometry calculations. The server back-end establishes a global reference by leveraging shared data and refining joint pose graph optimization through place recognition, global optimization techniques, and removal of outlier data to ensure precise and robust collaborative state estimation. Extensive evaluations of our system, utilizing both publicly available datasets and our custom datasets, demonstrate significant enhancements in the accuracy of collaborative SLAM estimates. Moreover, our system exhibits remarkable proficiency in large-scale missions, seamlessly enabling ten robots to collaborate effectively in performing SLAM tasks. In order to contribute to the research community, we will make our code open-source and accessible at https://github.com/PengYu-team/Co-LRIO.
Shipeng Zhong, Yuhua Qi, Dapeng Feng, Jin Wu 0002, Weisong Wen, Ming Liu 0001
ICRA6
2024 PGSL: A probabilistic graph diffusion model for source localization
Xovee Xu, Tangjiang Qian, Jin Wu 0002, Fan Zhou 0002
Expert Syst. Appl.5
2023 Trajectory-User Linking via Trajectory Convolution
abstract
Trajectory-user linking is the basis for many applications such as personalized recommendation and urban planning. Although plenty of efforts have been devoted to this topic, the results achieved are still not good enough. Existing methods mainly employ Recurrent Neural Networks (RNNs) to model trajectories semantically due to the inherent sequential attribute of trajectories. However, these approaches are weak at Point of Interest (POI) representation learning and trajectory feature detection. Thus, the performance of existing solutions is far from the requirements of practical applications. In this paper, we propose a novel Trajectory Convolution-based Trajectory-User Linking (TCTUL) method. Firstly, we connect all POI according to trajectories from all users. The result is a connected graph that can be used to generate more informative POI sequences than other approaches. Secondly, we employ the Node2Vec algorithm to encode each POI into a low-dimensional real value vector. Then, we transform each trajectory into an image-like matrix with fixed dimensions. Finally, a CNN is designed to detect features and predict the user of a given trajectory. The CNN can extract informative features from the matrix representations of trajectories by convolutional operations, Batch normalization, and$K$-max pooling operations. Extensive experiments on real datasets demonstrate that TCTUL substantially outperforms existing solutions in terms of macro-Precision, macro-Recall, macro-F1, and accuracy.
Xucheng Luo, Jin Wu 0002, Fan Zhou 0002
ICC3
2023 Completely Rational $\text{SO}(n)$ Orthonormalization
abstract
The rotation orthonormalization on the special orthogonal group$\text{SO}(n)$, also known as the high dimensional nearest rotation problem, has been revisited. A new generalized simple iterative formula has been proposed that solves this problem in a completely rational manner. Rational operations allow for efficient implementation on various platforms and also significantly simplify the synthesis of large-scale circuitization. The developed scheme is also capable of designing efficient fundamental rational algorithms, for example, quaternion normalization, which outperforms long-exisiting solvers. Furthermore, an$\text{SO}(n)$neural network has been developed for further learning purpose on the rotation group. Simulation results verify the effectiveness of the proposed scheme and show the superiority against existing representatives. Applications show that the proposed orthonormalizer is of potential in robotic pose estimation problems, e.g., hand-eye calibration.
Jin Wu 0002, Soheil Sarabandi, Jianhao Jiao, Huaiyang Huang, Bohuan Xue, Ruoyu Geng, Lujia Wang 0001, Ming Liu 0001
ICRA1
2023 Self-Supervised Drivable Area Segmentation Using LiDAR's Depth Information for Autonomous Driving
abstract
Drivable area segmentation is an essential component of the visual perception system for autonomous driving vehicles. Recent efforts in deep neural networks have sig-nificantly improved semantic segmentation performance for autonomous driving. However, most DNN-based methods need a large amount of data to train the models, and collecting large-scale datasets with manually labeled ground truth is costly, tedious, time consuming and requires the availability of experts, making DNN-based methods often difficult to implement in real world applications. Hence, in this paper, we introduce a novel module named automatic data labeler (ADL), which leverages a deterministic LiDAR-based method for ground plane segmentation and road boundary detection to create large datasets suitable for training DNNs. Furthermore, since the data generated by our ADL module is not as accurate as the manually annotated data, we introduce uncertainty estimation to compensate for the gap between the human labeler and our ADL. Finally, we train the semantic segmentation neural networks using our automatically generated labels on the KITTI dataset [10] and KITTI-CARLA dataset [7]. The experimental results demonstrate that our proposed ADL method not only achieves impressive performance compared to manual labeling but also exhibits more robust and accurate results than both traditional methods and state-of-the-art self-supervised methods.
Fulong Ma, Yang Liu 0477, Sheng Wang 0017, Jin Wu 0002, Weiqing Qi, Ming Liu 0001
IROS4
2023 Mixup-based Unified Framework to Overcome Gender Bias Resurgence
abstract
Unwanted social biases are usually encoded in pretrained language models (PLMs). Recent efforts are devoted to mitigating intrinsic bias encoded in PLMs. However, the separate fine-tuning on applications is detrimental to intrinsic debiasing. A bias resurgence issue arises when fine-tuning the debiased PLMs on downstream tasks. To eliminate undesired stereotyped associations in PLMs during fine-tuning, we present a mixup-based framework Mix-Debias from a new unified perspective, which directly combines debiasing PLMs with fine-tuning applications. The key to Mix-Debias is applying mixup-based linear interpolation on counterfactually augmented downstream datasets, with expanded pairs from external corpora. Besides, we devised an alignment regularizer to ensure original augmented pairs and gender-balanced counterparts are spatially closer. Experimental results show that Mix-Debias can reduce biases in PLMs while maintaining a promising performance in applications.
Liu Yu 0001, Yuzhou Mao, Jin Wu 0002, Fan Zhou 0002
SIGIR3
2023 Generalized n-Dimensional Rigid Registration: Theory and Applications
abstract
The generalized rigid registration problem in high-dimensional Euclidean spaces is studied. The loss function is minimized with an equivalent error formulation by the Cayley formula. The closed-form linear least-square solution to such a problem is derived which generates the registration covariances, i.e., uncertainty information of rotation and translation, providing quite accurate probabilistic descriptions. Simulation results indicate the correctness of the proposed method and also present its efficiency on computation-time consumption, compared with previous algorithms using singular value decomposition (SVD) and linear matrix inequality (LMI). The proposed scheme is then applied to an interpolation problem on the special Euclidean group SE(n) with covariance-preserving functionality. Finally, experiments on covariance-aided Lidar mapping show practical superiority in robotic navigation.
Jin Wu 0002, Miaomiao Wang 0001, Hassen Fourati, Hui Li 0037, Yilong Zhu, Chengxi Zhang, Yi Jiang 0007, Xiangcheng Hu, Ming Liu 0001
IEEE Trans. Cybern.1
2023 A VT-HMM-Based Framework for Countdown Timer Traffic Light State Estimation
abstract
Traffic lights are important components of traffic systems, and perceptual tasks on traffic lights are crucial for intelligent agents on the road. Auxiliary countdown timers, providing the remaining time of the current traffic phase, improve the safety and smoothness of the entire traffic system. This work proposes a state estimation framework for countdown timer traffic lights. Time-domain information is adequately integrated into a variable transition Hidden Markov Model (VT-HMM), and our system provides optimal estimates of traffic light colors and countdown numbers based on noisy detection inputs. A dynamic state transition matrix is designed based on a 1-step transition logic and a probability of the number of transitions related to the current state sojourn duration. A recursive decoding method based on the Viterbi algorithm is proposed to update all the state candidates and select the optimal state chain. Extensive experiments evaluate the robustness and effectiveness of the proposed work. The performance boundaries of this system are also found under various input noise levels. The source code is available here:https://github.com/ShuyangUni/countdown-timer-traffic-light-estimation
Qingwen Zhang, Feiyi Chen, Jin Wu 0002, Jianhao Jiao, Lujia Wang 0001
IEEE Trans. Intell. Transp. Syst.4
2023 Uncertainty-Aware Heterogeneous Representation Learning in POI Recommender Systems
abstract
Discovering interesting yet unvisited point-of-interests (POIs) is among the most practical applications but challenging problems in location-based social networks (LBSNs). Popular approaches face several issues, such as data sparsity and difficulties in modeling latent nonlinearity between users and POIs. Furthermore, the uncertainty in LBSNs poses additional obstacles to learning good representations of users’ general and current interests. To effectively address these issues, we postulate that fusing multiple sources of information is paramount. Toward that, we propose a novel deep generative recommender system—Wasserstein autoencoder for POI recommendation (WaPOIR). It unifies the information from users’ personal preference, social influence, and geographical data, and captures users’ general interests from historical check-ins, while modeling users’ current interests from recently visited POIs. Unlike previous methods, WaPOIR learns the latent distribution of data in the Wasserstein space as a potential representation for each POI and each user in LBSNs. This enables simultaneous maintenance of social and POI interactions and modeling the uncertainty of their relationships. WaPOIR is a stochastic recommendation approach that allows Bayesian inference and approximation of variational posterior distribution. Extensive experiments conducted on real-world LBSN datasets demonstrate that WaPOIR achieves better performance over the state-of-the-art approaches.
Fan Zhou 0002, Tangjiang Qian, Yuhua Mo, Zhangtao Cheng, Chunjing Xiao, Jin Wu 0002, Goce Trajcevski
IEEE Trans. Syst. Man Cybern. Syst.6
2022 FusionPortable: A Multi-Sensor Campus-Scene Dataset for Evaluation of Localization and Mapping Accuracy on Diverse Platforms
abstract
Combining multiple sensors enables a robot to maximize its perceptual awareness of environments and enhance its robustness to external disturbance, crucial to robotic navigation. This paper proposes the FusionPortable benchmark, a complete multi-sensor dataset with a diverse set of sequences for mobile robots. This paper presents three contributions. We first advance a portable and versatile multi-sensor suite that offers rich sensory measurements: 10Hz LiDAR point clouds, 20Hz stereo frame images, high-rate and asynchronous events from stereo event cameras, 200Hz inertial readings from an IMU, and 10Hz GPS signal. Sensors are already temporally synchronized in hardware. This device is lightweight, self-contained, and has plug-and-play support for mobile robots. Second, we construct a dataset by collecting 17 sequences that cover a variety of environments on the campus by exploiting multiple robot platforms for data collection. Some sequences are challenging to existing SLAM algorithms. Third, we provide ground truth for the decouple localization and mapping performance evaluation. We additionally evaluate state-of-the-art SLAM approaches and identify their limitations. The dataset, consisting of raw sensor measurements, ground truth, calibration data, and evaluated algorithms, will be released.
Jianhao Jiao, Hexiang Wei, Tianshuai Hu, Xiangcheng Hu, Yilong Zhu, Zhijian He, Jin Wu 0002, Jingwen Yu, Xupeng Xie, Huaiyang Huang, Ruoyu Geng, Lujia Wang 0001, Ming Liu 0001
IROS7
2022 Mining Spatio-Temporal Relations via Self-Paced Graph Contrastive Learning
abstract
Modeling complex spatial and temporal dependencies are indispensable for location-bound time series learning. Existing methods, typically relying on graph neural networks (GNNs) and temporal learning modules based on recurrent neural networks, have achieved significant performance improvements. However, their representation capabilities and prediction results are limited when pre-defined graphs are unavailable. Unlike spatio-temporal GNNs focusing on designing complex architectures, we propose a novel adaptive graph construction strategy: Self-Paced Graph Contrastive Learning (SPGCL). It learns informative relations by maximizing the distinguishing margin between positive and negative neighbors and generates an optimal graph with a self-paced strategy. Specifically, the existing neighborhoods iteratively absorb more reliable nodes with the highest affinity scores as new neighbors to generate the next-round neighborhoods, and augmentations are applied to improve the transferability and robustness. As the adaptively self-paced graph approaches the optimized graph for prediction, the mutual information between nodes and the corresponding neighbors is maximized. Our work provides a new perspective of addressing spatio-temporal learning problems beyond information aggregation in Euclidean space and can be generalized to different tasks. Extensive experiments conducted on two typical spatio-temporal learning tasks (traffic forecasting and land displacement prediction) demonstrate the superior performance of SPGCL against the state-of-the-art.
Rongfan Li, Ting Zhong, Xinke Jiang, Goce Trajcevski, Jin Wu 0002, Fan Zhou 0002
KDD5
2022 Identification and classification of promoters using the attention mechanism based on long short-term memory
Lei Xu 0047, Quan Zou 0001, Jin Wu 0002
Frontiers Comput. Sci.5
2022 MesoGRU: Deep Learning Framework for Mesoscale Eddy Trajectory Prediction
abstract
This work presents a novel framework, named MesoGRU, to accurately predict trajectories of mesoscale eddies (MEs), which is beneficial to study the natural ocean phenomena and perform statistical analysis of oceanic data. MesoGRU first extracts ME trajectory data and establishes two datasets according to the information of location and date in the South China Sea (SCS). Then it effectively processes SCS trajectory data and integrates SLA data and AVISO data into our combined ME dataset (CDME). Furthermore, a designated neural network with a new loss function [(weighted mean square estimation (WMSE)] is designed to learn the characteristics of ME trajectories. After thoroughly analyzing correlations of every two features, the MesoGRU network iteratively extracts key ME features and renormalizes the future time sequence of eddy trajectories. Our verification results demonstrate that MesoGRU has conducted a significant prediction improvement. The mean daily center error of ME trajectory prediction with our scheme is about 8 km and its center error for seven-day forecasting is approximately 18.2 km. Moreover, MesoGRU achieves a higher matching ratio of ME trajectory forecasting compared to other deep learning methods with a single dataset, indicating MesoGRU can predict trajectories as a state-of-the-art method.
Xuegong Wang, Chong Li 0006, Dalei Song, Peng Ren 0001, Jin Wu 0002
IEEE Geosci. Remote. Sens. Lett.7
2022 CRCF: A Method of Identifying Secretory Proteins of Malaria Parasites
abstract
Malaria is a mosquito-borne disease that results in millions of cases and deaths annually. The development of a fast computational method that identifies secretory proteins of the malaria parasite is important for research on antimalarial drugs and vaccines. Thus, a method was developed to identify the secretory proteins of malaria parasites. In this method, a reduced alphabet was selected to recode the original protein sequence. A feature synthesis method was used to synthesise three different types of feature information. Finally, the random forest method was used as a classifier to identify the secretory proteins. In addition, a web server was developed to share the proposed algorithm. Experiments using the benchmark dataset demonstrated that the overall accuracy achieved by the proposed method was greater than 97.8 percent using the 10-fold cross-validation method. Furthermore, the reduced schemes and characteristic performance analyses are discussed.
Changli Feng, Jin Wu 0002, Haiyan Wei, Lei Xu 0047, Quan Zou 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 H∞-Based Minimal Energy Adaptive Control With Preset Convergence Rate
abstract
This work studies the${H}_{\infty }$-based minimal energy control with a preset convergence rate (PCR) problem for a class of disturbed linear time-invariant continuous-time systems with matched external disturbance. This problem aims to design an optimal controller so that the energy of the control input satisfies a predetermined requirement. Moreover, the closed-loop system asymptotic stability with PCR is ensured simultaneously. To deal with this problem, a modified game algebraic Riccati equation (MGARE) is proposed, which is different from the game algebraic Riccati equation in the traditional${H}_{\infty } $control problem due to the state cost being lost. Therefore, a unique positive-definite solution of the MGARE is theoretically analyzed with its existing conditions. In addition, based on this formulation, a novel approach is proposed to solve the actuator magnitude saturation problem with the system dynamics being exactly known. To relax the requirement of the knowledge of system dynamics, a model-free policy iteration approach is proposed to compute the solution of this problem. Finally, the effectiveness of the proposed approaches is verified through two simulation examples.
Yi Jiang 0007, Kai Zhang 0004, Jin Wu 0002, Chengxi Zhang, Wenqian Xue, Tianyou Chai, Frank L. Lewis
IEEE Trans. Cybern.3
2022 SE(n)++: An Efficient Solution to Multiple Pose Estimation Problems
abstract
In robotic applications, many pose problems involve solving the homogeneous transformation based on the special Euclidean group SE(n) . However, due to the nonconvexity of SE(n) , many of these solvers treat rotation and translation separately, and the computational efficiency is still unsatisfactory. A new technique called the SE(n)++ is proposed in this article that exploits a novel mapping from SE(n) to SO(n + 1) . The mapping transforms the coupling between rotation and translation into a unified formulation on the Lie group and gives better analytical results and computational performances. Specifically, three major pose problems are considered in this article, that is, the point-cloud registration, the hand-eye calibration, and the SE(n) synchronization. Experimental validations have confirmed the effectiveness of the proposed SE(n)++ method in open datasets.
Jin Wu 0002, Ming Liu 0001, Yulong Huang 0003, Yuanxin Wu, Changbin Yu
IEEE Trans. Cybern.1
2022 Quadratic Pose Estimation Problems: Globally Optimal Solutions, Solvability/Observability Analysis, and Uncertainty Description
abstract
Pose estimation problems are fundamental in robotics. Most of these problems are challenging due to the nonconvex nature. This also sets up an obstacle for uncertainty description that is essential for pose integration and quality control. In this article, we show that a large class of related problems can be categorized as the quadratic pose estimation problems (QPEPs) and we propose a general quaternion-based mathematical model to unify these problems. To solve the nonconvex QPEPs, a Gröbner-basis method is investigated to derive their globally optimal and robust solutions. Furthermore, we develop the rules for characterizing the solvability and observability of these solutions. In addition, the uncertainty description, i.e., covariance matrix, as an important piece of information in robotic state estimation frameworks, is analyzed in detail. Theoretical results show that the covariance can be estimated via online optimization, in an efficient and unbiased manner. In this way, both the solution and covariance are guaranteed to be globally optimal. Through simulations and experiments, we show that the proposed QPEP-based solver is not only accurate, robust, and efficient but outperforms the representatives for covariance estimation. The designed algorithms are also assembled as a C++/MATLAB/Octave/ROS library, while these developed interfaces are built for main stream platforms and simultaneous localization and mapping schemes.
Jin Wu 0002, Yu Zheng 0001, Zhi Gao 0005, Yi Jiang 0007, Xiangcheng Hu, Yilong Zhu, Jianhao Jiao, Ming Liu 0001
IEEE Trans. Robotics1
2021 Trajectory-User Linking via Graph Neural Network
abstract
Trajectory-User Linking (TUL) refers to classifying trajectories into the corresponding generated users and has emerged as an essential spatio-temporal data mining task with a broad spectrum of applications, ranging from personalized location recommendation and trip planning to criminal behavior detection and object tracking. Despite the progress made by recent deep learning-based human mobility learning models, some critical factors related to personal context and user-location interactions have not yet been fully explored. Besides, existing works suffer from high computational cost issues due to the increased complexity of trajectory learning and contextual location embedding. In this work, we propose a novel end-to-end model called GNNTUL, composed of a graph neural network (GNN) module and a classifier, to effectively and efficiently learn human mobility and associate the traces to the users in online social networks. GNNTUL is the first GNN-based human mobility learning model exploiting implicit transition patterns behind sparse user traces in online social networks while extracting users' unique motion features and discriminating the motion traces. Extensive experiments conducted on two real- world datasets demonstrate the superiority of GNNTUL over several state-of-the-art baselines in terms of both linking accuracy and learning efficiency.
Fan Zhou 0002, Shupei Chen, Jin Wu 0002, Chengtai Cao
ICC3
2021 Differential Information Aided 3-D Registration for Accurate Navigation and Scene Reconstruction
abstract
A novel 3-dimensional (3-D) alignment method for point-cloud registration is proposed where the time-differential information of the measured points is employed. The new problem turns out to be a novel multi-dimensional optimization. Analytical solution to this optimization is then obtained, which sets the ground of further correspondence matching using k-D trees. Finally, via many examples, we show that the new method owns better registration accuracy in real-world experiments.
Jin Wu 0002, Yilong Zhu, Ruoyu Geng, Zhongtao Fu, Fulong Ma, Ming Liu 0001
ICRA1
2021 A spectral clustering with self-weighted multiple kernel learning method for single-cell RNA-seq data
abstract
Single-cell RNA-sequencing (scRNA-seq) data widely exist in bioinformatics. It is crucial to devise a distance metric for scRNA-seq data. Almost all existing clustering methods based on spectral clustering algorithms work in three separate steps: similarity graph construction; continuous labels learning; discretization of the learned labels by k-means clustering. However, this common practice has potential flaws that may lead to severe information loss and degradation of performance. Furthermore, the performance of a kernel method is largely determined by the selected kernel; a self-weighted multiple kernel learning model can help choose the most suitable kernel for scRNA-seq data. To this end, we propose to automatically learn similarity information from data. We present a new clustering method in the form of a multiple kernel combination that can directly discover groupings in scRNA-seq data. The main proposition is that automatically learned similarity information from scRNA-seq data is used to transform the candidate solution into a new solution that better approximates the discrete one. The proposed model can be efficiently solved by the standard support vector machine (SVM) solvers. Experiments on benchmark scRNA-Seq data validate the superior performance of the proposed model. Spectral clustering with multiple kernels is implemented in Matlab, licensed under Massachusetts Institute of Technology (MIT) and freely available from the Github website, https://github.com/Cuteu/SMSC/.
Ren Qi, Jin Wu 0002, Fei Guo 0001, Lei Xu 0047, Quan Zou 0001
Briefings Bioinform.2
2021 An in silico approach to identification, categorization and prediction of nucleic acid binding proteins
abstract
The interaction between proteins and nucleic acid plays an important role in many processes, such as transcription, translation and DNA repair. The mechanisms of related biological events can be understood by exploring the function of proteins in these interactions. The number of known protein sequences has increased rapidly in recent years, but the databases for describing the structure and function of protein have unfortunately grown quite slowly. Thus, improving such databases is meaningful for predicting protein-nucleic acid interactions. Furthermore, the mechanism of related biological events, such as viral infection or designing novel drug targets, can be further understood by understanding the function of proteins in these interactions. The information for each sequence, including its function and interaction sites, were collected and identified, and a database called PNIDB was built. The proteins in PNIDB were grouped into 27 classes, such as transcription, immune system, and structural protein, etc. The function of each protein was then predicted using a machine learning method. Using our method, the predictor was trained on labeled sequences, and then the function of a protein was predicted based on the trained classifier. The prediction accuracy achieved a score of 77.43% by 10-fold cross validation.
Lei Xu 0047, Jin Wu 0002, Quan Zou 0001
Briefings Bioinform.3
2021 Improving human mobility identification with trajectory augmentation
Fan Zhou 0002, Ruiyang Yin, Goce Trajcevski, Kunpeng Zhang 0001, Jin Wu 0002, Ashfaq Khokhar 0001
GeoInformatica5
2021 rBPDL: Predicting RNA-Binding Proteins Using Deep Learning
abstract
RNA-binding protein (RBP) is a powerful and wide-ranging regulator that plays an important role in cell development, differentiation, metabolism, health and disease. The prediction of RBPs provides valuable guidance for biologists. Although experimental methods have made great progress in predicting RBP, they are time-consuming and not flexible. Therefore, we developed a network model, rBPDL, by combining a convolutional neural network and long short-term memory for multilabel classification of RBPs. Moreover, to achieve better prediction results, we used a voting algorithm for ensemble learning of the model. We compared rBPDL with state-of-the-art methods and found that rBPDL significantly improved identification performance for the RBP68 dataset, with a macro-Area Under Curve (AUC), micro-AUC, and weighted AUC of 0.936, 0.962, and 0.946, respectively. Furthermore, through AUC statistical analysis of the RBP domain, we analyzed the performance of rBPDL and found that the RBP identification performance in the same domain was similar. In addition, we analyzed the performance preferences and physicochemical properties of the binding protein amino acids and explored the characteristics that affect the binding by using the RBP86 dataset.
Mengting Niu, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047
IEEE J. Biomed. Health Informatics2
2020 Continual Information Cascade Learning
abstract
Modeling the information diffusion process is an essential step towards understanding the mechanisms driving the success of information. Existing methods either exploit various features associated with cascades to study the underlying factors governing information propagation, or leverage graph representation techniques to model the diffusion process in an end-to-end manner. Current solutions are only valid for a static and fixed observation scenario and fail to handle increasing observations due to the challenge of catastrophic forgetting problems inherent in the machine learning approaches used for modeling and predicting cascades. To remedy this issue, we propose a novel dynamic information diffusion model CICP (Continual Information Cascades Prediction). CICP employs graph neural networks for modeling information diffusion and continually adapts to increasing observations. It is capable of capturing the correlations between successive observations while preserving the important parameters regarding cascade evolution and transition. Experiments conducted on real-world cascade datasets demonstrate that our method not only improves the prediction performance with accumulated data but also prevents the model from forgetting previously trained tasks.
Fan Zhou 0002, Xin Jing 0003, Xovee Xu, Ting Zhong, Goce Trajcevski, Jin Wu 0002
GLOBECOM6
2020 Recommendation via Collaborative Autoregressive Flows
Fan Zhou 0002, Yuhua Mo, Goce Trajcevski, Kunpeng Zhang 0001, Jin Wu 0002, Ting Zhong
Neural Networks5
2020 MARG Attitude Estimation Using Gradient-Descent Linear Kalman Filter
abstract
The magnetic, angular rate, and gravity (MARG) sensor array has been widely used for attitude estimation tasks. In this article, we report our new advances on related fusion algorithm based on the gradient-descent algorithm (GDA). Combining with complementary filters, GDA has been very popular for attitude estimation in industrial applications. The integration of the Kalman filter introduces covariance information and significantly improves in-run quality control. Some useful results are derived to build up the framework of a novel linear Kalman filter called GDA-LKF. A new simplified linear measurement quaternion model is proposed. This article also deals with the analytical adaptive problem of determining gradient-descent step length. We, for the first time, give the solution in the sense of least square based on aided vector measurements. Simulations are carried out to validate the noise sensitivity and covariance characteristics. The proposed schemes are also evaluated by real-world experiments that provide the audience with comparisons on accuracy, convergence, and execution time consumption performances between the proposed GDA-LKFs and recent representative methods. The results show that the proposed GDA-LKF can accurately estimate attitude with fast convergence and high accuracy. Note to Practitioners-Attitude estimation using magnetic, angular rate, and gravity (MARG) sensors is very common for ground, air, and underwater vehicles. The novel findings in this article aim to give the engineers a brand new perspective on the related filter design according to the widely employed gradient-descent algorithm (GDA) method. Some specific techniques, e.g., optimal adaptive law of the designed scheme, will also bring flexibility to the implementation with multiple considerations of object motions.
Jin Wu 0002
IEEE Trans Autom. Sci. Eng.1
2020 Fast Symbolic 3-D Registration Solution
abstract
3-D registration has always been performed invoking singular value decomposition (SVD) or eigenvalue decomposition (EIG) in real engineering practices. However, these numerical algorithms suffer from uncertainty of convergence in many cases. A novel fast symbolic solution is proposed in this article by following our recent publication in this journal. The equivalence analysis shows that our previous solver can be converted to deal with the 3-D registration problem. Rather, the computation procedure is studied for further simplification of computing without complex-number support. Experimental results show that the proposed solver does not loose accuracy and robustness but improves the execution speed to a large extent by almost 50%-80%, on both a personal computer (PC) and an embedded processor. Note to Practitioners-3-D registration usually has a large computational burden in engineering tasks. The proposed symbolic solution can directly solve the eigenvalue and its associated eigenvector. A lot of computation resources can then be saved for better overall system performance. The deterministic behavior of the proposed solver also ensures long-endurance stability and can help an engineer better design thread timing logic.
Jin Wu 0002, Ming Liu 0001, Zebo Zhou, Rui Li 0037
IEEE Trans Autom. Sci. Eng.1
2020 Hybrid graph convolutional networks with multi-head attention for location recommendation
Ting Zhong, Fan Zhou 0002, Kunpeng Zhang 0001, Goce Trajcevski, Jin Wu 0002
World Wide Web6
2019 Convexity Analysis of Optimization Framework of Attitude Determination from Vector Observations
abstract
In the past several years, there have been several representative attitude determination methods developed using derivative-based optimization algorithms. Optimization techniques e.g. gradient-descent algorithm (GDA), Gauss-Newton algorithm (GNA), Levenberg-Marquadt algorithm (LMA) suffer from local optimum in real engineering practices. A brief discussion on the convexity of this problem is presented recently [1] stating that the problem is neither convex nor concave. In this paper, we give analytic proofs on this problem. The results reveal that the target loss function is convex in the common practice of quaternion normalization, which leads to non-existence of local optimum.
Jin Wu 0002, Zebo Zhou, Hassen Fourati, Ming Liu 0001
CoDIT1
2019 Reference-Free Adaptive Attitude Determination Method Using Low-Cost MARG Sensors
Jin Wu 0002, Mingsen Deng, Ming Liu 0001
ICVS2
2019 Robust Rotation Interpolation Based on SO(n) Geodesic Distance
Jin Wu 0002, Ming Liu 0001, Mingsen Deng
ICVS1
2019 Adversarial Point-of-Interest Recommendation
abstract
Point-of-interest (POI) recommendation is essential to a variety of services for both users and business. An extensive number of models have been developed to improve the recommendation performance by exploiting various characteristics and relations among POIs (e.g., spatio-temporal, social, etc.). However, very few studies closely look into the underlying mechanism accounting for why users prefer certain POIs to others. In this work, we initiate the first attempt to learn the distribution of user latent preference by proposing an Adversarial POI Recommendation (APOIR) model, consisting of two major components: (1) the recommender (R) which suggests POIs based on the learned distribution by maximizing the probabilities that these POIs are predicted as unvisited and potentially interested; and (2) the discriminator (D) which distinguishes the recommended POIs from the true check-ins and provides gradients as the guidance to improve R in a rewarding framework. Two components are co-trained by playing a minimax game towards improving itself while pushing the other to the boundary. By further integrating geographical and social relations among POIs into the reward function as well as optimizing R in a reinforcement learning manner, APOIR obtains significant performance improvement in four standard metrics compared to the state of the art methods.
Fan Zhou 0002, Ruiyang Yin, Kunpeng Zhang 0001, Goce Trajcevski, Ting Zhong, Jin Wu 0002
WWW6
2019 Generalized Linear Quaternion Complementary Filter for Attitude Estimation From Multisensor Observations: An Optimization Approach
abstract
Focusing on generalized sensor combinations, this paper deals with the attitude estimation problem using a linear complementary filter (CF). The quaternion observation model is obtained via a gradient descent algorithm. An additive measurement model is then established according to derived results. The filter is named as the generalized CF where the observation model is simplified as a linear one that is quite different from previous-reported brute-force nonlinear results. Moreover, we prove that representative derivative-based optimization algorithms are essentially equivalent to each other. Derivations are given to establish the state model based on the quaternion kinematic equation. The proposed algorithm is validated under several experimental conditions involving the free-living environment, harsh external field disturbances, and aerial flight test aided by robotic vision. Using the specially designed experimental devices, data acquisition and algorithm computations are performed to give comparisons on accuracy, robustness, time-consumption, and so on with representative methods. The results show that not only the proposed filter can give fast, accurate, and stable estimates in terms of various sensor combinations but also produces robust attitude estimation in the scenario of harsh situations, e.g., irregular magnetic distortion. Note to Practitioners-Multisensor attitude estimation is a crucial technique in robotic devices. Many existing methods focus on the orientation fusion of specific sensor combinations. In this paper, we make the problem more concise. The results given in this paper are very general and can significantly decrease the space consumption and computation burden without losing the original estimation accuracy. Such performance will be of benefit to robotic platforms requiring flexible and easy-to-tune attitude estimation in the future.
Jin Wu 0002, Zebo Zhou, Hassen Fourati, Rui Li 0037, Ming Liu 0001
IEEE Trans Autom. Sci. Eng.1
2018 DeepLink: A Deep Learning Approach for User Identity Linkage
abstract
The typical aim of User Identity Linkage (UIL) is to detect when users from across different social platforms are actually one and the same individual. Existing efforts to address this problem of practical relevance span from user-profile-based, through user-generated-content-based, user-behavior-based approaches to supervised or unsupervised learning frameworks, to subspace learning-based models. Most of them often require extraction of relevant features (e.g., profile, location, biography, networks, behavior, etc.) to model the user consistently across different social networks. However, these features are mainly derived based on prior knowledge and may vary for different platforms and applications. Inspired by the recent successes of deep learning in different tasks, especially in automatic feature extraction and representation, we propose a deep neural network based algorithm for UIL, called DeepLink. It is a novel end-to-end approach in a semi-supervised learning manner, without involving any hand-crafting features. Specifically, DeepLink samples the networks and learns to encode network nodes into vector representation to capture local and global network structures which, in turn, can be used to align anchor nodes through deep neural networks. A dual learning based paradigm is exploited to learn how to transfer knowledge and update the linkage using the policy gradient method. Experiments conducted on several public datasets show that DeepLink outperforms the state-of-the-art methods in terms of both linking precision and identity-match ranking.
Fan Zhou 0002, Kunpeng Zhang 0001, Goce Trajcevski, Jin Wu 0002, Ting Zhong
INFOCOM5
2018 Fast Linear Attitude Estimation and Angular Rate Genseration
abstract
This paper focuses on the approach to attitude estimation and virtual gyroscope compensation. First, the rotation transformation is built between the body frame and the reference frame by means of an accelerometer-magnetometer triad. Fast linear attitude estimation method is used in the process to improve the computation efficiency. The attitude quaternion is invoked for parameterization of orientation in Kalman filtering. Furthermore, to generate the virtual-gyro output in the case of gyroscope failures, virtual-gyro Kalman filter is established for angular rate estimation. Experiments are carried out to illustrate the validity and efficiency of the proposed attitude and angular rate estimation approaches.
Zebo Zhou, Jin Wu 0002, Shuang Du, Hassen Fourati
IPIN3
2018 Fast Linear Quaternion Attitude Estimator Using Vector Observations
abstract
As a key problem for multisensor attitude determination, Wahba's problem has been studied for almost 50 years. Different from existing methods, this paper presents a novel linear approach to solve this problem. We name the proposed method the fast linear attitude estimator (FLAE) because it is faster than known representative algorithms. The original Wahba's problem is extracted to several 1-D equations based on quaternions. They are then investigated with pseudoinverse matrices establishing a linear solution to n-D equations, which are equivalent to the conventional Wahba's problem. To obtain the attitude quaternion in a robust manner, an eigenvalue-based solution is proposed. Symbolic solutions to the corresponding characteristic polynomial are derived, showing higher computation speed. Simulations are designed and conducted using test cases evaluated by several classical methods, e.g., Shuster's quaternion estimator, Markley's singular value decomposition method, Mortari's second estimator of the optimal quaternion, and some recent representative methods, e.g., Yang's analytical method and Riemannian manifold method. The results show that FLAE generates attitude estimates as accurate as that of several existing methods, but consumes much less computation time (about 50% of the known fastest algorithm). Also, to verify the feasibility in embedded application, an experiment on the accelerometer-magnetometer combination is carried out where the algorithms are compared via C++ programming language. An extreme case is finally studied, revealing a minor improvement that adds robustness to FLAE, inspired by Cheng et al.
Jin Wu 0002, Zebo Zhou, Bin Gao 0003, Rui Li 0037, Yuhua Cheng 0001, Hassen Fourati
IEEE Trans Autom. Sci. Eng.1