Yao Liu 0017

dblp:64/424-17 · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-6393-4295ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Learning Causality-Aware Exploration with Transformers for Goal-Oriented Navigation
abstract
Navigation is a fundamental task in the research of Embodied AI, and recent advances in machine learning algorithms have garnered growing interest in developing versatile Embodied AI systems. However, current research in this domain reveals opportunities for improvement. First, the direct application of RNNs and Transformers often overlooks the distinct characteristics of navigation tasks compared to traditional sequential data modeling. These methods are inherently designed to capture long-term dependencies, which are relatively weak in navigation scenarios, potentially limiting their performance in such tasks. Second, the reliance on task-specific configurations, such as pre-trained modules and dataset-specific logic, compromises the generalizability of these methods. We address these constraints by initially exploring the unique differences between Navigation tasks and other sequential data tasks through the lens of Causality, presenting a causal framework to elucidate the inadequacies of conventional sequential methods for Navigation. By leveraging this causal perspective, we propose Causality-Aware Transformer (CAT) Networks for Navigation, featuring a Causal Understanding Module to enhance the model’s Environmental Understanding capability. Meanwhile, our method is devoid of task-specific inductive biases and can be trained in an End-to-End manner, which enhances the method’s generalizability across various contexts. Empirical evaluations demonstrate that our methodology consistently surpasses benchmark performances across a spectrum of settings, tasks, and simulation environments, specifically, in Object Navigation within RoboTHOR, Objective Navigation, Point Navigation in Habitat, and R2R Navigation. Extensive ablation studies reveal that the performance gains can be attributed to the Causal Understanding Module, which demonstrates effectiveness and efficiency in both Reinforcement Learning and Supervised Learning settings. Additionally, further analysis highlights the robustness of our method, demonstrating its capacity to consistently perform well across diverse experimental settings and varying conditions. This robustness underscores the adaptability and generalizability of our approach, reinforcing its potential for application across a wide range of tasks.
Ruoyu Wang 0038, Tong Yu 0001, Mingjie Li 0006, Yuanjiang Cao, Yao Liu 0017, Lina Yao 0001
ACM Trans. Intell. Syst. Technol.5
2025 Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
abstract
Visual Language Navigation (VLN) is a fundamental task within the field of Embodied AI, focusing on the ability of agents to navigate complex environments based on natural language instructions. Despite the progress made by existing methods, these methods often present some common challenges. First, they rely on pre-trained backbone models for visual perception, which struggle with the dynamic viewpoints in VLN scenarios. Second, the performance is limited when using pre-trained LLMs or VLMs without fine-tuning, due to the absence of VLN domain knowledge. Third, while fine-tuning LLMs and VLMs can improve results, their computational costs are higher than those without fine-tuning. To address these limitations, we propose Weakly-supervised Partial Contrastive Learning (WPCL), a method that enhances an agent’s ability to identify objects from dynamic viewpoints in VLN scenarios by effectively integrating pre-trained VLM knowledge into the perception process, without requiring VLM fine-tuning. Our method enhances the agent’s ability to interpret and respond to environmental cues while ensuring computational efficiency. Experimental results have shown that our method outperforms the baseline methods on multiple benchmarks, which validates the effectiveness, robustness, and generalizability of our method.
Ruoyu Wang 0038, Tong Yu 0001, Junda Wu, Yao Liu 0017, Julian J. McAuley, Lina Yao 0001
IROS4
2024 Multi-agent Traffic Prediction via Denoised Endpoint Distribution
abstract
The exploration of high-speed movement by robots or road traffic agents is crucial for autonomous driving and navigation. Trajectory prediction at high speeds requires considering historical features and interactions with surrounding entities, a complexity not as pronounced in lower-speed environments. Prior methods have assessed the spatiotemporal dynamics of agents but often neglected intrinsic intent and uncertainty, thereby limiting their effectiveness. We present the Denoised Endpoint Distribution model for trajectory prediction, which distinctively models agents’ spatio-temporal features alongside their intrinsic intentions and un-certainties. By employing Diffusion and Transformer models to focus on agent endpoints rather than entire trajectories, our approach significantly reduces model complexity and enhances performance through endpoint information. Our experiments on open datasets, coupled with comparison and ablation studies, demonstrate our model’s efficacy and the importance of its components. This approach advances trajectory prediction in high-speed scenarios and lays groundwork for future developments.
Yao Liu 0017, Ruoyu Wang 0038, Yuanjiang Cao, Quan Z. Sheng, Lina Yao 0001
IROS1
2024 Regularized Multi-LLMs Collaboration for Enhanced Score-Based Causal Discovery
Xiaoxuan Li 0003, Yao Liu 0017, Ruoyu Wang 0038, Lina Yao 0001
WISE (4)2
2024 Uncertainty-aware pedestrian trajectory prediction via distributional diffusion
abstract
Tremendous efforts have been put forth on predicting pedestrian trajectory with generative models to accommodate uncertainty and multi-modality in human behaviors. An individual’s inherent uncertainty, e.g., change of destination, can be masked by complex patterns resulting from the movements of interacting pedestrians. However, latent variable-based generative models often entangle such uncertainty with complexity, leading to limited either latent expressivity or predictive diversity. In this work, we propose to separately model these two factors by implicitly deriving a flexible latent representation to capture intricate pedestrian movements, while integrating predictive uncertainty of individuals with explicit bivariate Gaussian mixture densities over their future locations. More specifically, we present a model-agnostic uncertainty-aware pedestrian trajectory prediction framework, parameterizing sufficient statistics for the mixture of Gaussians that jointly comprise the multi-modal trajectories. We further estimate these parameters of interest by approximating a denoising process that progressively recovers pedestrian movements from noise. Unlike previous studies, we translate the predictive stochasticity to explicit distributions, allowing it to readily generate plausible future trajectories indicating individuals’ self-uncertainty. Moreover, our framework is compatible with different neural net architectures. We empirically show the performance gains over state-of-the-art even with lighter backbones, across most scenes on two public benchmarks.
Yao Liu 0017, Zesheng Ye, Rui Wang 0088, Binghao Li, Quan Z. Sheng, Lina Yao 0001
Knowl. Based Syst.1
2024 Attention-Aware Social Graph Transformer Networks for Stochastic Trajectory Prediction
abstract
Trajectory prediction is fundamental to various intelligent technologies, such as autonomous driving and robotics. The motion prediction of pedestrians and vehicles helps emergency braking, reduces collisions, and improves traffic safety. Current trajectory prediction research faces problems of complex social interactions, high dynamics and multi-modality. Especially, it still has limitations in long-time prediction. We propose Attention-aware Social Graph Transformer Networks for multi-modal trajectory prediction. We combine Graph Convolutional Networks and Transformer Networks by generating stable resolution pseudo-images from Spatio-temporal graphs through a designed stacking and interception method. Furthermore, we design the attention-aware module to handle social interaction information in scenarios involving mixed pedestrian-vehicle traffic. Thus, we maintain the advantages of the Graph and Transformer, i.e., the ability to aggregate information over an arbitrary number of neighbors and the ability to perform complex time-dependent data processing. We conduct experiments on datasets involving pedestrian, vehicle, and mixed trajectories, respectively. Our results demonstrate that our model minimizes displacement errors across various metrics and significantly reduces the likelihood of collisions. It is worth noting that our model effectively reduces the final displacement error, illustrating the ability of our model to predict for a long time.
Yao Liu 0017, Binghao Li, Xianzhi Wang 0001, Claude Sammut, Lina Yao 0001
IEEE Trans. Knowl. Data Eng.1
2024 Two-stream Multi-level Dynamic Point Transformer for Two-person Interaction Recognition
abstract
As a fundamental aspect of human life, two-person interactions contain meaningful information about people’s activities, relationships, and social settings. Human action recognition serves as the foundation for many smart applications, with a strong focus on personal privacy. However, recognizing two-person interactions poses more challenges due to increased body occlusion and overlap compared to single-person actions. In this article, we propose a point cloud-based network named Two-stream Multi-level Dynamic Point Transformer for two-person interaction recognition. Our model addresses the challenge of recognizing two-person interactions by incorporating local-region spatial information, appearance information, and motion information. To achieve this, we introduce a designed frame selection method named Interval Frame Sampling (IFS), which efficiently samples frames from videos, capturing more discriminative information in a relatively short processing time. Subsequently, a frame features learning module and a two-stream multi-level feature aggregation module extract global and partial features from the sampled frames, effectively representing the local-region spatial information, appearance information, and motion information related to the interactions. Finally, we apply a transformer to perform self-attention on the learned features for the final classification. Extensive experiments are conducted on two large-scale datasets, the interaction subsets of NTU RGB+D 60 and NTU RGB+D 120. The results show that our network outperforms state-of-the-art approaches in most standard evaluation settings.
Yao Liu 0017, Gangfeng Cui, Xiaojun Chang, Lina Yao 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Mixed Precision Based Parallel Optimization of Tensor Mathematical Operations on a New-generation Sunway Processor
abstract
As an important part of high-performance computing (HPC) applications, tensor mathematical operations have a wide and significant impact on application performance. However, due to the unique heterogeneous architecture and software environment of the new-generation Sunway processors, it is critical to utilize the computing capacities of the processor for tensor mathematical operations. The existing research has not fully considered the computing characteristics of tensor mathematical operations and the hardware features of the new-generation Sunway processor. In this paper, we propose an optimization method for tensor mathematical operations on the new-generation Sunway processor. Firstly, an optimization method for elementary functions is proposed, which implements high-performance vector elementary functions with variable precision. Then, an mixed precision optimization method is proposed, which realizes expression computation with variable precision according to precision requirements of users. Finally, a multi-level parallel optimization method is proposed, which realizes asynchronous parallelism of the master core and the slave cores. The experimental results show that, compared with the native implementation, optimized tensor mathematical operations can achieve an average speedup of 112.19× on 64 cores, which exceeds the theoretical speedup.
Shuwei Fan, Yao Liu 0017, Juliang Su, Xianyou Wu, Qiong Jiang
CCGrid2
2023 NP-SSL: A Modular and Extensible Self-supervised Learning Library with Neural Processes
abstract
Neural Processes (NPs) are a family of supervised density estimators devoted to probabilistic function approximation with meta-learning. Despite extensive research on the subject, the absence of a unified framework for NPs leads to varied architectural solutions across diverse studies. This non-consensus poses challenges to reproducing and benchmarking different NPs. Moreover, existing codebases mainly prioritize generative density estimation, yet rarely consider expanding the capability of NPs to self-supervised representation learning, which however has gained growing importance in data mining applications. To this end, we present NP-SSL, a modular and configurable framework with built-in support, requiring minimal effort to 1) implement classical NPs architectures; 2) customize specific components; 3) integrate hybrid training scheme (e.g., contrastive); and 4) extend NPs to act as a self-supervised learning toolkit, producing latent representations of data, and facilitating diverse downstream predictive tasks. To illustrate, we discuss a case study that applies NP-SSL to model time-series data. We interpret that NP-SSL can handle different predictive tasks such as imputation and forecasting, by a simple switch in data samplings, without significant change to the underlying structure. We hope this study can reduce the workload of future research on leveraging NPs to tackle more a broader range of real-world data mining applications. Code and documentation are at https://github.com/zyecs/NP-SSL.
Zesheng Ye, Jing Du 0003, Yao Liu 0017, Yihong Zhang 0001, Lina Yao 0001
CIKM3
2023 MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model Evaluation
abstract
Curated datasets for healthcare are often limited due to the need of human annotations from experts.In this paper, we present MEDEVAL, a multi-level, multi-task, and multi-domain medical benchmark to facilitate the development of language models for healthcare.MEDEVAL is comprehensive and consists of data from several healthcare systems and spans 35 human body regions from 8 examination modalities.With 22,779 collected sentences and 21,228 reports, we provide expert annotations at multiple levels, offering a granular potential usage of the data and supporting a wide range of tasks.Moreover, we systematically evaluated 10 generic and domain-specific language models under zero-shot and finetuning settings, from domain-adapted baselines in healthcare to general-purposed state-of-the-art large language models (e.g., ChatGPT).Our evaluations reveal varying effectiveness of the two categories of language models across different tasks, from which we notice the importance of instruction tuning for few-shot usage of large language models.Our investigation paves the way toward benchmarking language models for healthcare and provides valuable insights into the strengths and limitations of adopting large language models in medical domains, informing their practical applications and future advancements 1 . Chest
Zexue He, Yu Wang 0170, An Yan 0003, Yao Liu 0017, Eric Y. Chang, Amilcare Gentili, Julian J. McAuley, Chun-Nan Hsu
EMNLP4
2023 Multi-level Attention Network with Weather Suppression for All-Weather Action Detection in UAV Rescue Scenarios
Yao Liu 0017, Binghao Li, Claude Sammut, Lina Yao 0001
ICONIP (9)1
2023 Automatic Multi-Parameter Performance Modeling of HPC Applications on a New Sunway Supercomputer
abstract
As the successor to Sunway TaihuLight, the new Sunway supercomputer has ultra-high computing capacity, but the unique heterogeneous architecture presents performance optimization challenges for High Performance Computing (HPC) applications. Performance modeling is an effective way to discover the performance bottlenecks and then improve the performance of HPC applications. Existing performance modeling techniques do not work well on large-scale HPC applications due to high overhead and low accuracy, and are not suitable for the heterogeneous architecture due to a lack of support for multi-resource parameters. To address the above challenges, we propose an automatic multi-parameter performance modeling method for HPC applications on the new Sunway supercomputer. First, a lightweight performance profiling method is proposed to achieve low overhead performance profiling. Then, performance models with multiple resource parameters based on the Fourier neural operator are built, achieving high prediction accuracy and generalization ability. Finally, the Fourier neural operator is extended on the new Sunway supercomputer to realize the performance modeling automatically. Experimental results show that the average prediction error is less than 10% and the average overhead is less than 4%, and the results are superior to the baselines.
Yilian Zhang, Yao Liu 0017, Penglong Jiao, Yiping Zhou, Tongquan Wei
IEEE Trans. Parallel Distributed Syst.2
2022 Social Graph Transformer Networks for Pedestrian Trajectory Prediction in Complex Social Scenarios
abstract
Pedestrian trajectory prediction is essential for many modern applications, such as abnormal motion analysis and collision avoidance for improved traffic safety. Previous studies still face challenges in embracing high social interaction, dynamics, and multi-modality for achieving high accuracy with long-time predictions. We propose Social Graph Transformer Networks for multi-modal prediction of pedestrian trajectories, where we combine Graph Convolutional Network and Transformer Network by generating stable resolution pseudo-images from Spatio-temporal graphs through a designed stacking and interception method. Specifically, we adopt adjacency matrices to obtain Spatio-temporal features and Transformer for long-time trajectory predictions. As such, we retrain the advantages of both, i.e., the ability to aggregate information over an arbitrary number of neighbors and to conduct complex time-dependent data processing. Our experimental results show that our model reduces the final displacement error and achieves state-of-the-art in multiple metrics. The module's effectiveness is demonstrated through ablation experiments.
Yao Liu 0017, Lina Yao 0001, Binghao Li, Xianzhi Wang 0001, Claude Sammut
CIKM1
2022 Interpolation graph convolutional network for 3D point cloud analysis
abstract
The feature analysis of point clouds, a popular representation of three-dimensional (3D) objects, is rising as a hot research topic nowadays. Point cloud data bear a sparse and unordered nature, making many commonly used feature extraction methods, for example, Convolutional Neural Networks (CNNs) inapplicable, while previous models suitable for the task are usually complex. We aim to reduce model complexity by reducing the number of parameters while achieving better (or at least comparable) performance. We propose an Interpolation Graph Convolutional Network (IGCN) for extracting features of point clouds. IGCN uses the point cloud graph structure and a specially designed Interpolation Convolution Kernel to mimic the operations of CNN for feature extraction. On the basis of weight postfusion and multilevel-resolution aggregation, IGCN not only reduces the cost of calculating the interpolation operation but also improves the model's performance. We validate the performance of IGCN on both point cloud classification and segmentation tasks and explore the contribution of each module of our model through ablation experiments. Furthermore, we embed the IGCN point cloud feature extraction module as a plug-and-play module into other frameworks and perform point cloud registration experiments.
Yao Liu 0017, Lina Yao 0001, Binghao Li, Claude Sammut, Xiaojun Chang
Int. J. Intell. Syst.1
2022 Customer Adaptive Resource Provisioning for Long-Term Cloud Profit Maximization under Constrained Budget
abstract
As an efficient commercial information technology, cloud computing has attracted more and more users and enterprises to use it. Faced with such a large number and variety of customers, it is necessary for cloud providers (CPs) with limited budget to provide satisfactory customized pricing services, profitable customer and system investments, and flexible system resource provisioning strategies to improve both customer experience and long-term profit. Existing profit optimization research rarely considers customer diversity and dynamics, which may have a negative impact on long-term profit growth due to poor management of customer relations. In this article, we implement customer relationship management by considering both customer diversity and dynamics, and propose a customer adaptive resource provisioning scheme to maximize long-term profit under constrained budget. We consider four customer types (i.e., loyal, old, new, and lost) that can transition to each other during the customer's lifetime of interaction with the CP. The CP builds multiple cloud service sub-platforms, each of which contains multiple multiserver systems and serves the same type of customers. For the cloud service platform, we first analyze single multiserver system using an analytical method to obtain its optimal profit, invested funding, and system configuration. In particular, for systems serving new and lost customers, we develop a novel customer lifetime value (CLV)-based customer investment scheme that selects valuable customers for investment under limited marketing budget. Based on the above analysis, we then present a customer retention rate (CRR)-driven three-stage heuristic scheme that prioritizes investment in multiserver systems with endangered customers under limited infrastructure budget for reducing customer churn and promoting long-term profit growth. We conduct extensive simulation experiments to validate the effectiveness of our method. Simulation results show that compared with the benchmark algorithms, our method can improve the long-term profit and CRR by up to 3.4x and 7.8x, respectively.
Peijin Cong, Junlong Zhou, Xin Liu 0081, Yao Liu 0017, Tongquan Wei
IEEE Trans. Parallel Distributed Syst.5
2021 Parallelization and Optimization of NSGA-II on Sunway TaihuLight System
abstract
Sunway TaihuLight system is the first supercomputer offering a peak performance over 100 PFlops, which can be utilized to parallelize Non-dominated Sorting Genetic Algorithm II (NSGA-II), a standard approach to multi-objective optimization. However, insufficient off-chip memory bandwidth and limited scratchpad memory capacity of the supercomputer hinder the performance improvement of parallellizing NSGA-II. In this article, we propose an optimized parallel NSGA-II on Sunway TaihuLight system, called swNSGA-II, by utilizing process- and thread-level parallelism of the system based on an improved island/master-slave model. To overcome the hurdles of low memory bandwidth and capacity, we propose a data sharing scheme based on register-level communication that can efficiently parallelize non-dominated sorting and crowding-distance computation of NSGA-II. Several optimization techniques including vectorization, direct memory accessing, and double buffering are also adopted to further accelerate swNSGA-II. Experiment results show that the proposed swNSGA-II can achieve a speedup of 41284 on a use case of path planning, and a speedup of 62692 on ZDT1 as compared to conventional NSGA-II.
Xin Liu 0081, Su Wang 0005, Yao Liu 0017, Tongquan Wei
IEEE Trans. Parallel Distributed Syst.5
2020 Performance Modeling of Stencil Computation on SW26010 Processors
Yao Liu 0017, Mengtao Hu, Wei Wang 0033, Wei Xue 0003, Qingting Zhu
ICA3PP (1)1
2018 A Scalable Heterogeneous Parallel SOM Based on MPI/CUDA
abstract
Self-Organizing Map (SOM) is a kind of artificial neural network used in unsupervised machine learning, which is widely applied to clustering, dimension reduction and visualization for high-dimensional data, etc. There are two major versions of the training algorithm: original algorithm and batch algorithm. Compared with the original, the batch algorithm has some advantages including faster convergence and less computation, and is suitable for parallelization. However, it is still confronted with the challenge of eficiency in the case of massive data, high-dimensional data or a large-scale map. In this paper, a scalable heterogeneous parallel SOM based on the batch algorithm is proposed which combines process-level and thread-level parallelism by MPI and CUDA. To boost the parallel efficiency on GPUs and make full use of the high floating-point computing capability, we design matrix operations for the the most time-consuming steps, the computation of best match units and weights update, making the steps available for the implementation by cuBLAS. In addition, the memory optimization methods are adopted. The experiments show that the proposed heterogeneous parallel SOM is effective, efficient and scalable.
Yao Liu 0017, Su Wang 0005, Kai Zheng 0008
ACML1
2014 A DSCP-Based Method of QoS Class Mapping between WLAN and EPS Network
Yao Liu 0017, Fengling Cai, Qian Kong
ICA3PP (1)1