Yunji Liang

dblp:22/10796 · DBLP profile ↗
← Back
40ranked-venue papers
15as first author
32since 2021 · last 2026
0000-0002-8381-8187ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 5 first-author · 10 since 2021Computer networks · 10 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Security and privacy · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 VoiceFormer: Fusing Non-Acoustic Motion Sensors for High-Fidelity Voice Synthesis in Mobile Devices
abstract
With the popularity of mobile devices, a variety of motion sensors are integrated to enhance the user experience. Although existing studies demonstrated that non-acoustic motion sensors can be attacked by adversaries, they overlook the limited sampling frequencies of motion sensors (e.g., < 500 Hz) in mobile devices and are evaluated in the controlled laboratory settings. In this article, we explore a new attack model on non-acoustic motion sensors based on the off-the-shelf mobile devices. We propose a general framework named VoiceFormer to synthesize high-fidelity speeches based on the vibrations of accelerometers and gyroscopes with a low sampling frequency. Specifically, in VoiceFormer , we introduce a signal alignment approach to remove the time offsets between two nonsynchronous signals, and leverage Time Interleaved Analog-Digital-Conversion (TI-ADC) to generate a high-frequency synthetic signal (e.g., > 8 KHz) based on the vibration signals of accelerometers and gyroscopes on the same motherboard. To synthesize the high-fidelity acoustic waveforms, we propose a wavelet-based generative adversarial network to learn the spatiotemporal latent mapping between vibrations and original speech signals. Extensive experimental results demonstrate the feasibility of voice synthesis by spying the low-frequency non-acoustic motion sensors in off-the-shelf mobile devices. VoiceFormer shows impressive performance in the synthesized acoustical signals with a Mean Opinion Score of 3.38. Although there are significant differences of mobile devices in hardware settings, VoiceFormer shows robust performance in synthesizing intelligible voice signals. Our results suggest that eavesdropping an off-the-shelf mobile device remotely by fusing non-acoustic sensors is feasible.
Xiaokai Yan, Yunji Liang, Lei Liu 0073, Sagar Samtani, Bin Guo 0001, Zhiwen Yu 0001
ACM Trans. Priv. Secur.2
2026 Corner Case Detection and Generation for Autonomous Driving: An Overview
abstract
Safety concerns remain one of the most significant obstacles to the large-scale deployment and continued advancement of autonomous driving (AD) systems. A major underlying cause of many safety-related incidents in AD systems is suboptimal or erroneous decision-making when the vehicle encounters corner cases (CCs)—rare, unexpected, or extreme situations that fall outside typical operating conditions. Although recent advances in artificial intelligence have driven substantial progress in both autonomous driving and corner-case research, the field still lacks a coherent conceptual foundation and a systematic, widely accepted categorization of CCs. In this survey, we address this gap by offering a structured review of the existing corner-case literature along three key dimensions: understanding, detection, and generation. The key contribution is a three-level classification of corner cases—spanning data-level, model-level, and semantic-level CCs—that clarifies and disambiguates competing definitions and perspectives on AD corner cases. Building on this framework, we examine simulator-based methods for corner-case selection and data generation, and we identify open challenges, promising directions, and potential solutions to more effectively handle corner cases in autonomous driving systems.
Yunji Liang, Junteng Liu, Xiaokai Yan, Xiaolong Zheng 0001, Lei Tang 0002, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001
IEEE Trans. Intell. Transp. Syst.1
2026 Cooperative Air-Ground Instant Delivery by UAVs and Crowdsourced Taxis: Joint UAV Station Deployment and Delivery Scheduling
abstract
Instant delivery has become an essential service in daily life, requiring strict delivery timelines. However, traditional delivery methods that employ human couriers struggle to meet the soaring delivery demands due to labor shortages. While researchers have explored alternative solutions using ground vehicles (e.g., crowdsourced taxis) and Unmanned Aerial Vehicles (UAVs), their inherent limitations, such as constrained delivery detour for crowdsourced taxis and limited battery capacity of UAVs, greatly constrain their effectiveness. To address these challenges, this paper proposes a novel air-ground delivery paradigm that cooperatively integrates UAVs and crowdsourced taxis. First, UAV stations are strategically deployed based on delivery gaps between the delivery demands and taxis' delivery capacity, instead of delivery demands only; Then, a predictive UAV repositioning strategy is designed to bridge instantaneously dynamic delivery gaps. Thereafter, a transfer learning-based (TL-based) algorithm that mines the delivery knowledge of human couriers is designed to optimize the cooperative performance. This algorithm extracts behavioral insights from human couriers and transfers them to enhance the delivery capabilities of UAVs and taxis. Finally, parcel assignment is formulated as optimization problems aimed at maximizing total preferences of UAVs and taxis, and maximizing delivery number while minimizing cost, respectively. Evaluations on real-world datasets demonstrate that the proposed method delivers 27.4% more parcels, saves 19.2% delivery cost, and preserves 36.3% more of the travel experience of taxi passengers than the state-of-the-art (SOTA) air-ground cooperative approach for instant delivery.
Qianru Wang, Xin Zhang 0018, Xiang Zhao 0002, Yunji Liang, Bin Guo 0001, Qingye Han, Yan Pan 0003
IEEE Trans. Mob. Comput.6
2026 Hard Sample Mining: A New Paradigm of Efficient and Robust Model Training
abstract
Over the past two decades, deep learning (DL) has achieved unprecedented breakthroughs across diverse application domains spanning computer vision (CV) to natural language processing (NLP). However, despite significant advances in computational resources and algorithmic frameworks, the training of deep neural networks continues to present formidable challenges due to persistent issues of training inefficiency and inherent data distribution biases. Recent years have witnessed the emergence of hard sample mining (HSM) as a promising paradigm to mitigate training inefficiencies and enhance model robustness through representative sample selection. Although HSM is reshaping contemporary AI research, its critical role in enabling efficient and robust model training has not yet been systematically explored. This article presents a comprehensive survey of HSM methodologies by: 1) establishing unified definitions of hard samples through rigorous sample complexity quantification criteria; 2) proposing a systematic taxonomy of HSM approaches with in-depth technical analysis; and 3) identifying pivotal research frontiers in this evolving field. This survey not only consolidates the foundations of HSM but also provides a roadmap for advancing efficient, robust, and generalizable deep learning models.
Lei Liu 0073, Yunji Liang, Xiaokai Yan, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001, Yanyong Zhang, Daniel Dajun Zeng
IEEE Trans. Neural Networks Learn. Syst.2
2026 PersuHSG: Adaptive Persuasion Strategy Planning for Dialogue Agents Based on Hierarchical Strategy Graph
abstract
Persuasion, a vital social skill, influences beliefs, attitudes, and behaviors through conversation. Yet, current dialogue agents either rely on scenario-specific strategies, restricting their cross-context adaptability, or neglect persuasion’s logical structure. They focus on isolated strategy classification, overlooking the significance of fine-grained sequential planning for real-world scenarios. To address these limitations, inspired by basic human mental activities, we present PersuHSG, an adaptive persuasion strategy planning framework. The core idea is to conceptualize persuasion as a tripartite framework comprising cognition, affection, and volition, with each stage represented as a graph layer and principle-based strategies for efficient multi-stage persuasion. Specifically, we first develop PersuInstruct, a fine-tuning dataset to improve dialogue agents’ strategic planning and response generation. Then, we propose a graph-aware planning algorithm for stage-strategy-response reasoning to generate persuasive responses for diverse scenarios. Extensive experiments confirm that PersuHSG significantly enhances the persuasiveness of Large Language Models (LLMs), allows smaller models (e.g., 9B, 13B) to achieve competitive performance, and demonstrates the efficacy of structured strategy planning in improving model efficiency and adaptability.
Bin Guo 0001, Hao Wang 0182, Jingqi Liu, Yan Liu 0045, Yunji Liang, Yan Pan 0003, Zhiwen Yu 0001
ACM Trans. Inf. Syst.6
2026 A Hierarchical Hard Negative Sampling Strategy for Robust Out-of-Distribution Object Detection
abstract
Out-of-distribution (OOD) detection is crucial for deploying models in open-world environments. This process aims to mitigate the issue of overconfident predictions, which is a common problem for models designed for closed-domain tasks when they encounter OOD data. Recent progress in OOD detection has shown that integrating auxiliary datasets during model training can greatly enhance OOD detection performance. However, existing methods tend to heavily depend on these auxiliary datasets to establish the decision boundary for in-distribution (ID) data, while not adequately addressing OOD object detection in safety–critical applications. In this article, we propose object-level OOD detection and introduce a hierarchical hard negative sampling (HNS) strategy that does not necessitate auxiliary data. Specifically, we offer a new metric that strategically considers difficult negatives near the decision boundary between inter-class and intra-class instances. Inspired by adversarial thinking, we sample outliers for each class, ensuring that the negative samples capture both diversity and informative traits. We conducted comprehensive experiments on three public datasets. The results demonstrate that HNS performs superiorly in object-level OOD detection, even without auxiliary datasets. The source code can be accessed at: https://github.com/Aurevior1/HNS .
Junteng Liu, Zizhe Wang, Yunji Liang, Sagar Samtani, Lei Tang 0002, Zhiwen Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 DiffSQL: Leveraging Diffusion Model for Zero-Shot Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation has attracted significant attention due to its broad applications in autonomous driving and robotics. Although significant performance improvements have been achieved by learning the relative distance of objects with the introduction of Self Query Layer (SQL), it struggles with zero-shot generalization due to the lack of geometric features and the fixed number of query sizes. To address these problems, we propose a diffusion-augmented self-supervised depth estimation framework, named DiffSQL, to learn geometric priors for feature augmentation. Additionally, we introduce a dynamic self-query layer that implicitly computes the relative distances between objects by adjusting the query size according to the feature distribution. Experimental results on the KITTI dataset show that DiffSQL outperforms SQLdepth by 1.03% in terms of AbsRel and 2.79% in terms of SqRel. Furthermore, our experiments demonstrate that DiffSQL is superior in zero-shot generalization.
Heyuan Zheng, Yunji Liang, Lei Liu 0073, Zhiwen Yu 0001
IJCAI2
2025 Scenario-Based Accelerated Testing for SOTIF in Autonomous Driving: A Review
abstract
The development of intelligent driving systems has drawn significant attention to enhancing the safety of autonomous vehicles and their intended functionality. Despite this, current accelerated testing approaches remain inadequate in assessing system reliability, as they fail to simulate scenarios involving collisions between vehicles and pedestrians and identify unknown risks. To address these limitations, scenario-based testing methods have been proposed, which seek to identify critical scenarios with a high frequency of exposure to safety risks. A comprehensive review of these methods is thus of paramount significance. In this article, we provide a timely and systematic literature review of existing accelerated testing for autonomous vehicles. We propose a taxonomy of these methods, discuss each subfield, and highlight open problems and future directions. Our objective is to provide a clear and concise overview of the state of the art in this field and to offer insights into the effectiveness of scenario-based testing approaches. By doing so, we aim to facilitate the identification of critical scenarios and the assessment of risk exposure frequencies, which are essential for enhancing the safety and reliability of autonomous vehicles.
Lei Tang 0002, Zhanwen Liu, Yunji Liang, Yuanyuan Niu, Wei Zhu 0004, Zongtao Duan
IEEE Internet Things J.4
2025 Upper bound on the predictability of rating prediction in recommender systems
En Xu, Zhiwen Yu 0001, Hui Wang 0011, Helei Cui, Yunji Liang, Bin Guo 0001
Inf. Process. Manag.7
2025 $\gamma$-Razor: Hardness-Aware Dataset Pruning for Efficient Neural Network Training
abstract
Training deep neural networks (DNNs) on large-scale datasets is often inefficient with large computational needs and significant energy consumption. Although great efforts have been taken to optimize DNNs, few studies focused on the inefficiency caused by the data samples with less value for model training. In this article, we empirically demonstrate that sample complexity is important for model efficiency and selecting representative samples is constructive to the model efficiency. In particular, we propose hardness-aware dataset pruning method ($\gamma$-Razor) to select representative samples from large-scale datasets to remove the less valuable data samples for model training.$\gamma$-Razor is a two-stage framework that includes interclass sampling and intraclass sampling. First, we introduce the inverse self-paced learning strategy to learn hard samples and adjust their weights adaptively according to the inverse frequency of effective samples of each class. For intraclass sampling, hardness-aware cluster sampling algorithm is proposed to downsample easy samples within each class. To evaluate the performance of$\gamma$-Razor, we conducted extensive experiments on three large-scale datasets for image classification tasks. The experimental results show that models trained with the pruned datasets show competitive performances against their counterparts trained with the original large-scale datasets in terms of robustness and efficiency. Furthermore, models trained with the pruned datasets converge faster with lower energy consumption.
Lei Liu 0073, Peng Zhang 0139, Yunji Liang, Lia Morra, Bin Guo 0001, Zhiwen Yu 0001, Yanyong Zhang, Daniel Dajun Zeng
IEEE Trans. Comput. Soc. Syst.3
2025 Optimizing Matching for On-Demand Ride-Pooling with Stochastic Day-to-Day Dynamics
abstract
Ride-pooling significantly reduces traffic congestion by enhancing fleet utilization through effective ride-matching. Real-world ride-pooling systems are dynamic, with fluctuations in driver availability and demand throughout the day. This necessitates adaptive ride-matching strategies that can quickly adjust to changing proximities and identify new carpooling opportunities by recalculating driver-rider correlations. However, most current methods primarily focus on static demand-supply scenarios and short-term accessibility, falling short in dynamic environment. In this study, we introduce a dynamic heterogeneous network model that captures the evolving nature of ride-pooling systems, where new requests and carpooling arrangements continuously emerge. We propose an embedding model-based matching decision process that operates online, adjusting to changes in the network’s structure. This process involves constructing a dynamic heterogeneous ride-pooling network that encompasses diverse node attributes and driver-rider connections, updating these representations to reflect the network’s evolution, and quickly identifying and ranking candidate riders for efficient online matching. Our approach demonstrates improved performance in offline evaluations using datasets from Austin, TX (RideAustin) and Chengdu, China (DiDi Chuxing). We observe a reduction in the necessary fleet size as new orders are placed, and an improvement in drivers’ matching probability compared to existing methods (e.g., an increase of 5.4–31.1% in the assignment rate on DiDi dataset), showcasing the advantage of employing dynamic network embedding to cut down on matching time (e.g., a decrease of 3.7–228.8 seconds in running time on DiDi dataset). Furthermore, we develop a simulated ride-pooling system (SRPool) that mimics dynamic demand-supply fluctuations and supports vehicle routing, providing a robust platform for evaluating ride-matching strategies. Our strategy not only excels in the SRPool environment but also effectively minimizes the total trip distance and rider waiting times.
Yaling Zhao, Lei Tang 0002, Yunji Liang, Junchi Ma
ACM Trans. Knowl. Discov. Data3
2024 SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth Estimation
abstract
Recently, self-supervised monocular depth estimation has gained popularity with numerous applications in autonomous driving and robotics. However, existing solutions primarily seek to estimate depth from immediate visual features, and struggle to recover fine-grained scene details. In this paper, we introduce SQLdepth, a novel approach that can effectively learn fine-grained scene structure priors from ego-motion. In SQLdepth, we propose a novel Self Query Layer (SQL) to build a self-cost volume and infer depth from it, rather than inferring depth from feature maps. We show that, the self-cost volume is an effective inductive bias for geometry learning, which implicitly models the single-frame scene geometry, with each slice of it indicating a relative distance map between points and objects in a latent space. Experimental results on KITTI and Cityscapes show that our method attains remarkable state-of-the-art performance, and showcases computational efficiency, reduced training complexity, and the ability to recover fine-grained scene details. Moreover, the self-matching-oriented relative distance querying in SQL improves the robustness and zero-shot generalization capability of SQLdepth. Code is available at https://github.com/hisfog/SfMNeXt-Impl.
Youhong Wang, Yunji Liang, Shaohui Jiao, Hongkai Yu
AAAI2
2024 MELPD-Detector: Multi-level ensemble learning method based on adaptive data augmentation for Parkinson disease detection via free-KD
Yafang Yang, Bin Guo 0001, Kaixing Zhao, Yunji Liang, Zhiwen Yu 0001
CCF Trans. Pervasive Comput. Interact.4
2024 Cross-Modality 3D Multiobject Tracking Under Adverse Weather via Adaptive Hard Sample Mining
abstract
3-D multiobject tracking (MOT) is an important task in numerous applications, including robotics and autonomous driving. Nevertheless, existing 3-D MOT solutions suffer from significant performance degradation under adverse weather conditions. Inspired by the fact that hard objects (e.g., missed detections or wrongly-associated objects) are more constructive to performance improvement, in this article, we leverage hard samples for robust 3-D MOT in adverse weather conditions. Specifically, we implement a cross-modality 3-D MOT framework to learn the 3-D region proposals from point clouds and RGB images, respectively. To minimize the risk of missed detection and wrong association, we introduce an adaptive hard sample mining scheme to align the 3-D region proposals provided by two modalities. We quantify the hard level by comparing the confidence values of the same object in the two branches and their distance in the embedding space. Meanwhile, we dynamically adjust the weights of hard samples during training to enhance the representation learning for robust 3-D MOT. Extensive experimental results showcase that our proposed solution effectively mitigates the missed detection and reduces wrong association with good generalization.
Lifeng Qiao, Peng Zhang 0139, Yunji Liang, Xiaokai Yan, Luwen Huangfu, Xiaolong Zheng 0001, Zhiwen Yu 0001
IEEE Internet Things J.3
2024 APGVAE: Adaptive disentangled representation learning with the graph-based structure information
Qiao Ke, Xinhui Jing, Marcin Wozniak, Yunji Liang, Jiangbin Zheng 0001
Inf. Sci.5
2024 Cross-device free-text keystroke dynamics authentication using federated learning
Yafang Yang, Bin Guo 0001, Yunji Liang, Kaixing Zhao, Zhiwen Yu 0001
Pers. Ubiquitous Comput.3
2024 Learning Cross-modality Interaction for Robust Depth Perception of Autonomous Driving
abstract
As one of the fundamental tasks of autonomous driving, depth perception aims to perceive physical objects in three dimensions and to judge their distances away from the ego vehicle. Although great efforts have been made for depth perception, LiDAR-based and camera-based solutions have limitations with low accuracy and poor robustness for noise input. With the integration of monocular cameras and LiDAR sensors in autonomous vehicles, in this article, we introduce a two-stream architecture to learn the modality interaction representation under the guidance of an image reconstruction task to compensate for the deficiencies of each modality in a parallel manner. Specifically, in the two-stream architecture, the multi-scale cross-modality interactions are preserved via a cascading interaction network under the guidance of the reconstruction task. Next, the shared representation of modality interaction is integrated to infer the dense depth map due to the complementarity and heterogeneity of the two modalities. We evaluated the proposed solution on the KITTI dataset and CALAR synthetic dataset. Our experimental results show that learning the coupled interaction of modalities under the guidance of an auxiliary task can lead to significant performance improvements. Furthermore, our approach is competitive against the state-of-the-art models and robust against the noisy input. The source code is available at https://github.com/tonyFengye/Code/tree/master .
Yunji Liang, Nengzhen Chen, Zhiwen Yu 0001, Lei Tang 0002, Hongkai Yu, Bin Guo 0001, Daniel Dajun Zeng
ACM Trans. Intell. Syst. Technol.1
2024 Learning Entangled Interactions of Complex Causality via Self-Paced Contrastive Learning
abstract
Learning causality from large-scale text corpora is an important task with numerous applications—for example, in finance, biology, medicine, and scientific discovery. Prior studies have focused mainly on simple causality, which only includes one cause-effect pair. However, causality is notoriously difficult to understand and analyze because of multiple cause spans and their entangled interactions. To detect complex causality, we propose a self-paced contrastive learning model, namely N2NCause, to learn entangled interactions between multiple spans. Specifically, N2NCause introduces data enhancement operations to convert implicit expressions into explicit expressions with the most rational causal connectives for the synthesis of positive samples and to invert the directed connection between a cause-effect pair for the synthesis of negative samples. To learn the semantic dependency and causal direction of positive and negative samples, self-paced contrastive learning is proposed to learn the entangled interactions among spans, including the interaction direction and interaction field. We evaluated the performance of N2NCause in three cause-effect detection tasks. The experimental results show that, with the least data annotation efforts, N2NCause demonstrates competitive performance in detecting simple cause-effect relations, and it is superior to existing solutions for the detection of complex causality.
Yunji Liang, Lei Liu 0073, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001, Daniel Dajun Zeng
ACM Trans. Knowl. Discov. Data1
2023 Identifying emotional causes of mental disorders from social media for effective intervention
abstract
Identifying the emotional causes of mental illnesses is key to effective intervention. Existing emotion-cause analysis approaches can effectively detect simple emotion-cause expressions where only one cause and one emotion exist. However, emotions may often result from multiple causes, implicitly or explicitly, with complex interactions among these causes. Moreover, the same causes may result in multiple emotions. How to model the complex interactions between multiple emotion spans and cause spans remains under-explored. To tackle this problem, a contrastive learning-based framework is presented to detect the complex emotion-cause pairs with the introduction of negative samples and positive samples. Additionally, we developed a large-scale emotion-cause dataset with complex emotion-cause instances based on subreddits associated with mental health. Our proposed approach was compared to prevailing CNN-based, LSTM-based, Transformer-based and GNN-based methods. Extensive experiments have been conducted and the quantifiable outcomes indicate that our proposed solution achieves competitive performance on simple emotion-cause pairs and significantly outperformed baseline methods in extracting complex emotion-cause pairs. Empirical studies further demonstrated that our proposed approach can be used to reveal the emotional causes of mental disorders for effective intervention.
Yunji Liang, Lei Liu 0073, Yapeng Ji, Luwen Huangfu, Daniel Dajun Zeng
Inf. Process. Manag.1
2023 An Escalated Eavesdropping Attack on Mobile Devices via Low-Resolution Vibration Signals
abstract
With the global prevalence of mobile devices, concerns about mobile devices regarding privacy breaches and data leakage are rising. Although sensor permissions are required for mobile applications to access outputs of built-in sensors, motion sensors (e.g., accelerometer and gyroscope) can be visited directly without permission requirement. Extant studies have shown that motion sensors may cause breaches of confidential information, such as passwords, digits, and voice-based commands, but whether it is possible to synthesize intelligible speech waveforms from low-resolution motion sensors has been understudied. In this article, we present an escalated side-channel attack of built-in speakers by synthesizing intelligible speech waveforms from low-resolution vibration signals. Opposite to traditional classification problems, we formulate this task as a generative problem and introduce an end-to-end synthesis framework dubbed asAccMyrinxto eavesdrop on the speaker via the low-resolution vibration signals. InAccMyrinx, we introduce the data alignment solution to provide the pair-wise voice-vibration sequences and present wavelet-based MelGAN (WMelGAN) with multi-scale time-frequency domain discriminators to generate intelligible acoustic waveforms. We conducted intensive experiments and demonstrated the feasibility of synthesizing the intelligible acoustic signals from low-resolution solid-borne vibration signals. Compared with existing synthesis solutions, our proposed solution outperforms the baselines in both subject and object metrics with the smoothed word error rate of 42.67% and the Mel-Cepstral distortion of 0.298. In addition, the quality of synthetic speeches could be impacted by several factors, including gender, speech rate, volume, and sampling frequency.
Yunji Liang, Yuchen Qin, Qi Li 0048, Xiaokai Yan, Luwen Huangfu, Sagar Samtani, Bin Guo 0001, Zhiwen Yu 0001
IEEE Trans. Dependable Secur. Comput.1
2023 Learning Dynamic App Usage Graph for Next Mobile App Recommendation
abstract
Next mobile app recommendation aims to recommend the next app that a user is most likely to use based on the user’s app usage behaviors, which is beneficial for improving user experience, app pre-loading, and system optimization. However, existing works ignore the complex correlations between apps in the app usage sessions. In addition, they do not consider the dynamics of user interests over time. To address these concerns, we propose a novel model named dynamic usage graph network (DUGN) to recommend the next app that a user is most likely to use. To model the complex correlations among apps explicitly, we adopt the dynamic graph structure to learn the dynamics of user interests. Firstly, we extract user interests in each app usage graph by using the hierarchical graph attention mechanism. Secondly, we capture user interests evolving over time, and generate the dynamic user embeddings by modeling the temporal dependencies among multiple app usage graphs. Finally, we obtain the current user interests in the current app usage graph, fuse multiple user interests and generate comprehensive user embeddings for next mobile app recommendation. We conduct experiments on real-world datasets. The results show that our model outperforms the state-of-art recommendation methods.
Yi Ouyang 0003, Bin Guo 0001, Qianru Wang, Yunji Liang, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.4
2023 DeepApp: characterizing dynamic user interests for mobile application recommendation
Yunji Liang, Lei Liu 0073, Luwen Huangfu, Zhu Wang 0001, Bin Guo 0001
World Wide Web (WWW)1
2022 Robust Detection of Malicious URLs With Self-Paced Wide & Deep Learning
abstract
As cybercrimes grow in scale with devastating economic costs, it is important to protect potential victims against diverse attacks. It is the uniform resource locators (URLs) that connect vulnerable users with potential attacks. Although numerous solutions (e.g., rule-based solutions and machine learning-based methods) are proposed for malicious URL detection, they can not provide robust performance due to the diversity of cybercrimes and can not cope with the explosive growth of malicious URLs with the evolution of obfuscation strategies. In this paper, we propose a deep learning-based system, dubbed as CyberLen, to detect malicious URLs robustly and effectively. Specifically, we use factorization machine (FM) to learn the latent interaction among lexical features. For the deep structural features, position embedding is introduced for token vectorization to reduce the ambiguity of URL tokens. Meanwhile, temporal convolution network (TCN) is utilized to learn the long-distance dependency among URL tokens. To fuse heterogeneous features, self-paced wide & deep learning strategy is proposed to train a robust model effectively. The proposed solution is evaluated on a large-scale URL dataset. Our experimental results show that position embedding is constructive to reducing the ambiguity of URL tokens, and the self-paced wide & deep learning strategy shows superior performance in terms of F1 score and convergence speed.
Yunji Liang, Kang Xiong, Xiaolong Zheng 0001, Zhiwen Yu 0001, Daniel Dajun Zeng
IEEE Trans. Dependable Secur. Comput.1
2022 MetaDetector: Meta Event Knowledge Transfer for Fake News Detection
abstract
The blooming of fake news on social networks has devastating impacts on society, the economy, and public security. Although numerous studies are conducted for the automatic detection of fake news, the majority tend to utilize deep neural networks to learn event-specific features for superior detection performance on specific datasets. However, the trained models heavily rely on the training datasets and are infeasible to apply to upcoming events due to the discrepancy between event distributions. Inspired by domain adaptation theories, we propose an end-to-end adversarial adaptation network, dubbed as MetaDetector , to transfer meta knowledge (event-shared features) between different events. Specifically, MetaDetector pushes the feature extractor and event discriminator to eliminate event-specific features and preserve required meta knowledge by adversarial training. Furthermore, the pseudo-event discriminator is utilized to evaluate the importance of news records in historical events to obtain partial knowledge that are discriminative for detecting fake news. Under the coordinated optimization among all the submodules, MetaDetector accurately transfers the meta knowledge of historical events to the upcoming event for fact checking. We conduct extensive experiments on two real-world datasets collected from Sina Weibo and Twitter. The experimental results demonstrate that MetaDetector outperforms the state-of-the-art methods, especially when the distribution discrepancy between events is significant.
Yasan Ding, Bin Guo 0001, Yan Liu 0045, Yunji Liang, Haocheng Shen, Zhiwen Yu 0001
ACM Trans. Intell. Syst. Technol.4
2022 Who Will Travel With Me? Personalized Ranking Using Attributed Network Embedding for Pooling
abstract
In ride matching, the search results can be personalized for a particular driver. Given a query with trip plans, it is advantageous to rank potential riders in terms of who are most appealing to the driver for increasing occupancy rates. While personalized ranking approaches such as collaborative filtering and factorization are available, they are not suitable for pooling because candidate riders are associated with different preferences, and their travel is sparsely distributed with a long tail of users for a few popular destinations. The user embedding method is a good candidate in terms of alleviating data sparsity, but it has issues such as difficulty encoding user preferences from rich information. In this study, we explore user embedding techniques for the purposes of short-term personalized rider ranking, where the aim is to present to drivers a set of potential riders who share similar itineraries with them and can be picked up on their current route. Considering trip requests, along with the preferences issued in advance, this study uses attribute representations to rank the riders based on the higher-order similarities in the participants’ itineraries in a three-step manner: (i) start with a distributed representation of the riders’ preference regarding the cost of extra distance, (ii) generate user embeddings in a heterogeneous network with the meeting points and associated waiting times, and (iii) match and rank riders for drivers depending on an attribute fusion operation by adopting a personal route and schedule. Our proposed method performs well in an offline estimation on a huge dataset from DiDi in Chengdu, China. Experimental results indicate that with the learned embeddings, we can obtain statistically significant advancements (e.g., 4.6–29.5% increase in mean reciprocal rank (MRR); 2.8–17.4% in normalized discounted cumulative gain (nDCG)) over current methods for pooling ranking. Furthermore, we implement the proposed method on our simulated pooling system. These results validate that personalized ranking can undoubtedly boost the number of trips served, and reduce the total trip distance and waiting time.
Lei Tang 0002, Rongguo Zhang, Zongtao Duan, Yunji Liang
IEEE Trans. Intell. Transp. Syst.5
2022 App Popularity Prediction by Incorporating Time-Varying Hierarchical Interactions
abstract
App popularity prediction is a significant task in mobile service development, which predicts an app's future popularity based on its current behaviors. It provides benefits from app development to targeted investment. Popularity is affected by two factors, i.e., internal ones like reviews and external ones like interaction among apps. However, most related studies only explore internal factors but neglect external ones. In fact, external factor plays an important role in popularity prediction modelling since it is the promoting and/or inhibiting influence resulted by app interaction. The app interaction has two major characteristics, i.e., interactivity and dynamicity, which brings challenges to app popularity prediction due to two reasons: 1) interactivity—it is hard to evaluate the existence and influence intensity of interactions; 2) dynamicity—the nature of interaction influence, e.g., promoting or inhibiting, and its intensity on popularity change with time. In this paper, we propose DeePOP, a popularity prediction model that innovatively leverages time-varying hierarchical interactions. First, we propose Hierarchical Interaction Graph, which is first studied in this work, to organically characterize the relationship and influence among apps. Second, DeePOP integrates internal factors and time-varying hierarchical interactions as inputs to build the prediction model. It develops multi-level modules based on Recurrent Neural Network with attention mechanism and generates multi-step time series predictions by fusing the outputs of modules. Experiments on a real-world dataset show that DeePOP outperforms state-of-the-art methods in prediction accuracy, effectively reducing the Root Mean Square Error (RMSE) to 0.088.
Jiaqi Liu 0002, Bin Guo 0001, Zhu Wang 0001, Yunji Liang, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.5
2021 JointCS: Joint Search for Deep Model Compression and Segmentation on Heterogeneous IoT Devices
abstract
Deep neural networks (DNNs) play an important role in a variety of intelligent applications (e.g. image classification and target recognition), yet at the cost of heavy computation burden, that makes DNNs difficult to deploy on resource-constrained IoT devices. To solve this problem, there are two categories of model computation adjustment methods: model compression and model segmentation. However, model compression mainly reduces resource consumption at the cost of accuracy while model segmentation reduces resource consumption according to the cost of communication latency. In this paper, we propose Joint Search for Model Compression and Segmentation (JointCS) that highlights the following aspects: 1) we integrate both model compression and model segmentation under an automatic and progressive framework, it simplifies model to fit the different IoT resource requirements. JointCS achieves a series slim models that outperform better both in accuracy and latency. 2) we train a network architecture-aware latency predictor to fast measure the latency of the slimed model on heterogeneous IoT devices. 3) we introduce a search algorithm to select the optimal state in progressively joint search. Finally, we evaluate the performance of our proposed method for image classification on CIFAR datasets comparing with the state-of-the-art approach, the inference time of the proposed method has inference speedup of 12.2 % −30.9 % under the same accuracy.
Bin Guo 0001, Sicong Liu 0005, Chen Qiu 0002, Yunji Liang, Zhiwen Yu 0001
ICPADS5
2021 Detection of Behavior Aging from Keystroke Dynamics
abstract
Keystroke dynamics-based authentication (KDA) is one of human behavioral-based authentication methods based on the unique typing rhythm of an individual. Nevertheless, the typing characteristics gradually change over time. Various solutions have been suggested to remedy the concept drift problem, including multimodal and unimodal adaptive methods. However, these solutions don't consider that temporal concept drift has a negative impact on performance and update frequency increases computation cost. The paper proposes weighted EDDM to detect concept drift and capture permanent concept drift (behavioral natural aging). Experimental results show that our method can accurately capture behavioral natural aging and filter temporal concept drift. Our proposed method has better performance and less computation.
Yafang Yang, Bin Guo 0001, Yunji Liang, Zhiwen Yu 0001
ICPADS3
2021 Fusion of heterogeneous attention mechanisms in multi-view convolutional neural network for text classification
Yunji Liang, Bin Guo 0001, Zhiwen Yu 0001, Xiaolong Zheng 0001, Sagar Samtani, Daniel Dajun Zeng
Inf. Sci.1
2021 DeepDepict: Enabling Information Rich, Personalized Product Description Generation With the Deep Multiple Pointer Generator Network
abstract
In e-commerce platforms, the online descriptive information of products shows significant impacts on the purchase behaviors. To attract potential buyers for product promotion, numerous workers are employed to write the impressive product descriptions. The hand-crafted product descriptions are less-efficient with great labor costs and huge time consumption. Meanwhile, the generated product descriptions do not take consideration into the customization and the diversity to meet users’ interests. To address these problems, we propose one generic framework, namely DeepDepict, to automatically generate the information-rich and personalized product descriptive information. Specifically, DeepDepict leverages the graph attention to retrieve the product-related knowledge from external knowledge base to enrich the diversity of products, constructs the personalized lexicon to capture the linguistic traits of individuals for the personalization of product descriptions, and utilizes multiple pointer-generator network to fuse heterogeneous data from multi-sources to generate informative and personalized product descriptions. We conduct intensive experiments on one public dataset. The experimental results show that DeepDepict outperforms existing solutions in terms of description diversity, BLEU, and personalized degree with significant margin gain, and is able to generate product descriptions with comprehensive knowledge and personalized linguistic traits.
Shaoyang Hao, Bin Guo 0001, Hao Wang 0182, Yunji Liang, Lina Yao 0001, Qianru Wang, Zhiwen Yu 0001
ACM Trans. Knowl. Discov. Data4
2021 Energy-efficient Collaborative Sensing: Learning the Latent Correlations of Heterogeneous Sensors
abstract
With the proliferation of Internet of Things (IoT) devices in the consumer market, the unprecedented sensing capability of IoT devices makes it possible to develop advanced sensing and complex inference tasks by leveraging heterogeneous sensors embedded in IoT devices. However, the limited power supply and the restricted computation capability make it challenging to conduct seamless sensing and continuous inference tasks on resource-constrained devices. How to conduct energy-efficient sensing and perform rich-sensor inference tasks on IoT devices is crucial for the success of IoT applications. Therefore, we propose a novel energy-efficient collaborative sensing framework to optimize the energy consumption of IoT devices. Specifically, we explore the latent correlations among heterogeneous sensors via an attention mechanism in temporal convolutional network to quantify the dependency among sensors, and characterize the heterogeneous sensors in terms of energy consumption to categorize them into low-power sensors and energy-intensive sensors . Finally, to decrease the sampling frequency of energy-intensive sensors , we propose a multi-task learning strategy to predict the statuses of energy-intensive sensors based on the low-power sensors . To evaluate the performance of the proposed collaborative sensing framework, we develop a mobile application to collect concurrent heterogeneous data streams from all sensors embedded in Huawei Mate 8. The experimental results show that latent correlation learning is greatly helpful to understand the latent correlations among heterogeneous streams, and it is feasible to predict the statuses of energy-intensive sensors by low-power sensors with high accuracy and fast convergence. In terms of energy consumption, the proposed collaborative sensing framework is able to preserve the energy consumption of IoT devices by nearly 50% for continuous data acquisition tasks.
Yunji Liang, Zhiwen Yu 0001, Bin Guo 0001, Xiaolong Zheng 0001, Sagar Samtani
ACM Trans. Sens. Networks1
2021 A multi-view attention-based deep learning system for online deviant content detection
Yunji Liang, Bin Guo 0001, Zhiwen Yu 0001, Xiaolong Zheng 0001, Zhu Wang 0001, Lei Tang 0002
World Wide Web1
2020 Behavioral Biometrics for Continuous Authentication in the Internet-of-Things Era: An Artificial Intelligence Perspective
abstract
In the Internet-of-Things (IoT) era, user authentication is essential to ensure the security of connected devices and the customization of passive services. However, conventional knowledge-based and physiological biometric-based authentication systems (e.g., password, face recognition, and fingerprints) are susceptible to shoulder surfing attacks, smudge attacks, and heat attacks. The powerful sensing capabilities of IoT devices, including smartphones, wearables, robots, and autonomous vehicles enable continuous authentication (CA) based on behavioral biometrics. The artificial intelligence (AI) approaches hold significant promise in sifting through large volumes of heterogeneous biometrics data to offer unprecedented user authentication and user identification capabilities. In this survey article, we outline the nature of CA in IoT applications, highlight the key behavioral signals, and summarize the extant solutions from an AI perspective. Based on our systematic and comprehensive analysis, we discuss the challenges and promising future directions to guide the next generation of AI-based CA research.
Yunji Liang, Sagar Samtani, Bin Guo 0001, Zhiwen Yu 0001
IEEE Internet Things J.1
2019 Leverage Temporal Convolutional Network for the Representation Learning of URLs
abstract
Cyber crimes including computer virus/malwares, spam, illegal sales, and phishing websites are proliferated aggressively via the disguised Uniform Resource Locators (URL). Although numerous studies were conducted for the URL classification task, the traditional URL classification solutions retreated due to the hand-crafted feature engineering and the boom of newly generated URLs. In this paper, we study the representation learning of URLs, and explore the URL classification using deep learning. Specifically, we propose URL2vec to extract both the structural and lexical features of URLs, and apply temporal convolutional network (TCN) for the URL classification task. The experimental results show that URL2vec outperforms both word2vec and character-level embedding for URL representation, and TCN achieves the best performance than baselines with the precision up to 95.97%.
Yunji Liang, Zhiwen Yu 0001, Bin Guo 0001, Xiaolong Zheng 0001, Saike He
ISI1
2016 From Mobile Phone Sensing to Human Geo-Social Behavior Understanding
abstract
The study of geo‐social behaviors has long been a scientific problem. In contrast to traditional social science, which suffers from the problems such as high data collection cost and imported user subjectivity, a new approach is presented to study social behaviors based on mobile phone sensing data. Different from other similar studies on mobile social sensing, three different types of geo‐social behaviors, including online interaction, offline interaction, and mobility patterns, are characterized based on a newly released Nokia mobile phone data set. We further discuss the impact factors to these behaviors as well as the correlation among them. The findings in this article are crucial for many different fields, ranging from urban planning, location‐based services, to social recommendation.
Bin Guo 0001, Yunji Liang, Zhiwen Yu 0001, Minshu Li, Xingshe Zhou 0001
Comput. Intell.2
2015 Facilitating medication adherence in elderly care using ubiquitous sensors and mobile social networks
Zhiwen Yu 0001, Yunji Liang, Bin Guo 0001, Xingshe Zhou 0001, Hongbo Ni
Comput. Commun.2
2014 Energy-Efficient Motion Related Activity Recognition on Mobile Devices for Pervasive Healthcare
Yunji Liang, Xingshe Zhou 0001, Zhiwen Yu 0001, Bin Guo 0001
Mob. Networks Appl.1
2013 From the internet of things to embedded intelligence
Bin Guo 0001, Daqing Zhang 0001, Zhiwen Yu 0001, Yunji Liang, Zhu Wang 0001, Xingshe Zhou 0001
World Wide Web4
2012 Energy Efficient Activity Recognition Based on Low Resolution Accelerometer in Smart Phones
Yunji Liang, Xingshe Zhou 0001, Zhiwen Yu 0001, Bin Guo 0001
GPC1
2012 Understanding the Regularity and Variability of Human Mobility from Geo-trajectory
abstract
Over the last few years, many efforts have been devoted to revealing human mobility patterns. However, the regularity and variability of human mobility from a microscopic view, i.e., what factors affect human mobility patterns, has yet not been investigated. In this paper, we aim to study the impact factors that may affect the regularity and variability of human mobility patterns using social network analysis. Specifically, we introduce the spatial interaction matrix to represent the interaction strength and interaction semantics among spatial regions. Based on the spatial interaction matrix, we investigate the factors that impact the mobility patterns, including temporal factors, occupational factors and age factors. Our experimental results demonstrate that lots of factors such as environmental, temporal and age factors contribute to the shape of human mobility patterns.
Yunji Liang, Xingshe Zhou 0001, Bin Guo 0001, Zhiwen Yu 0001
Web Intelligence1