VLDB 2026 Research / reviewers in the wild / expert
Yang Guan
dblp:22/10238
· DBLP profile ↗
18ranked-venue papers
10as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Computer networks · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhanced Integrated Decision and Control for High-Level Automated Vehicles and Its Experiment VerificationabstractLearning through experience is essential for high-level autonomous driving systems, as it has the potential to enhance driving performance in corner cases. However, current decision and control modules adopt an empirical design paradigm for engineering efficiency, relying heavily on expert rules or real-vehicle data, failing to fully cover and optimize edge scenarios. To address this gap, we propose an enhanced integrated decision and control method that leverages reinforcement learning as the optimal control problem solver, endowing high-level automated vehicles with experience data usage. Specifically, a constrained mixed policy gradient algorithm is developed, which combines and dynamically adjusts the application ratio of experience data and the environmental model during training. This approach achieves fast convergence while maintaining high performance even with inaccurate analytic models. Furthermore, an attention based encoding network is designed to accommodate diverse driving states in urban traffic, integrating an embedding network for feature extraction and a weighting network for feature fusion, realizing order-insensitive encoding and importance differentiation of road users. The trained policy is deployed on a fully functional autonomous vehicle. Experiments at a signalized intersection show that the proposed method can accurately identify critical surrounding obstacles and execute safe, efficient, and intelligent driving behaviors across 32 scenarios. Yang Guan, Liye Tang, Yao Lyu, Shengbo Eben Li, Kehua Sheng, Keqiang Li 0002 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | A novel diagnostic framework based on vibration image encoding and multi-scale neural network
Yang Guan, Zong Meng, Jimeng Li, Dengyun Sun, Fengjie Fan |
Expert Syst. Appl. | 1 |
| 2024 | Model-Based Chance-Constrained Reinforcement Learning via Separated Proportional-Integral LagrangianabstractSafety is essential for reinforcement learning (RL) applied in the real world. Adding chance constraints (or probabilistic constraints) is a suitable way to enhance RL safety under uncertainty. Existing chance-constrained RL methods, such as the penalty methods and the Lagrangian methods, either exhibit periodic oscillations or learn an overconservative or unsafe policy. In this article, we address these shortcomings by proposing a separated proportional-integral Lagrangian (SPIL) algorithm. We first review the constrained policy optimization process from a feedback control perspective, which regards the penalty weight as the control input and the safe probability as the control output. Based on this, the penalty method is formulated as a proportional controller, and the Lagrangian method is formulated as an integral controller. We then unify them and present a proportional-integral Lagrangian method to get both their merits with an integral separation technique to limit the integral value to a reasonable range. To accelerate training, the gradient of safe probability is computed in a model-based manner. The convergence of the overall algorithm is analyzed. We demonstrate that our method can reduce the oscillations and conservatism of RL policy in a car-following simulation. To prove its practicality, we also apply our method to a real-world mobile robot navigation task, where our robot successfully avoids a moving obstacle with highly uncertain or even aggressive behaviors. Baiyu Peng, Jingliang Duan, Jianyu Chen 0002, Shengbo Eben Li, Genjin Xie, Congsheng Zhang, Yang Guan, Yao Mu 0001, Enxin Sun |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Integrated Decision and Control: Toward Interpretable and Computationally Efficient Driving IntelligenceabstractDecision and control are core functionalities of high-level automated vehicles. Current mainstream methods, such as functional decomposition and end-to-end reinforcement learning (RL), suffer high time complexity or poor interpretability and adaptability on real-world autonomous driving tasks. In this article, we present an interpretable and computationally efficient framework called integrated decision and control (IDC) for automated vehicles, which decomposes the driving task into static path planning and dynamic optimal tracking that are structured hierarchically. First, the static path planning generates several candidate paths only considering static traffic elements. Then, the dynamic optimal tracking is designed to track the optimal path while considering the dynamic obstacles. To that end, we formulate a constrained optimal control problem (OCP) for each candidate path, optimize them separately, and follow the one with the best tracking performance. To unload the heavy online computation, we propose a model-based RL algorithm that can be served as an approximate-constrained OCP solver. Specifically, the OCPs for all paths are considered together to construct a single complete RL problem and then solved offline in the form of value and policy networks for real-time online path selecting and tracking, respectively. We verify our framework in both simulations and the real world. Results show that compared with baseline methods, IDC has an order of magnitude higher online computing efficiency, as well as better driving performance, including traffic efficiency and safety. In addition, it yields great interpretability and adaptability among different driving scenarios and tasks. Yang Guan, Yangang Ren, Qi Sun 0004, Shengbo Eben Li, Haitong Ma, Jingliang Duan, Bo Cheng 0003 |
IEEE Trans. Cybern. | 1 |
| 2022 | Beyond backpropagate through time: Efficient model-based training through time-splittingabstractModel-based policy gradient (MBPG) has been employed to seek an approximate solution to the optimal control problem. However, there is coupling between adjacent states due to temporal dependencies, making the training time grow linearly with the time horizon. This paper reshapes the training process of MBPG with the time-splitting technique to establish a time-independent algorithm called Training Through Time-Splitting (T3S). First, copy the coupled variables to obtain two independent variables. Meanwhile, an extra variable together with an equivalence constraint is introduced for problem consistency. Then, the transformed problem divides into subproblems with carefully derived loss functions. Subproblems own decoupled variables and shared policy networks, which means they can be optimized concurrently. Guided by the algorithm design, this paper further proposes an asynchronous parallel training scheme to accelerate training efficiency. Numerical simulation shows that the T3S algorithm outperforms the MBPG algorithm by 83.6% in wall-clock time with a trajectory tracking task. Jiaxin Gao 0002, Yang Guan, Shengbo Eben Li, Junqing Wei, Keqiang Li 0002 |
Int. J. Intell. Syst. | 2 |
| 2022 | Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation ErrorsabstractIn reinforcement learning (RL), function approximation errors are known to easily lead to the Q -value overestimations, thus greatly reducing policy performance. This article presents a distributional soft actor-critic (DSAC) algorithm, which is an off-policy RL method for continuous control setting, to improve the policy performance by mitigating Q -value overestimations. We first discover in theory that learning a distribution function of state-action returns can effectively mitigate Q -value overestimations because it is capable of adaptively adjusting the update step size of the Q -value function. Then, a distributional soft policy iteration (DSPI) framework is developed by embedding the return distribution function into maximum entropy RL. Finally, we present a deep off-policy actor-critic variant of DSPI, called DSAC, which directly learns a continuous return distribution by keeping the variance of the state-action returns within a reasonable range to address exploding and vanishing gradient problems. We evaluate DSAC on the suite of MuJoCo continuous control tasks, achieving the state-of-the-art performance. Jingliang Duan, Yang Guan, Shengbo Eben Li, Yangang Ren, Qi Sun 0004, Bo Cheng 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Model-based Constrained Reinforcement Learning using Generalized Control Barrier FunctionabstractModel information can be used to predict future trajectories, so it has huge potential to avoid dangerous regions when applying reinforcement learning (RL) on real-world tasks, like autonomous driving. However, existing studies mostly use model-free constrained RL, which causes inevitable constraint violations. This paper proposes a model-based feasibility enhancement technique of constrained RL, which enhances the feasibility of policy using generalized control barrier function (GCBF) defined on the distance to constraint boundary. By using the model information, the policy can be optimized safely without violating actual safety constraints, and the sample efficiency is increased. The infeasibility in solving the constrained policy gradient is handled by an adaptive coefficient mechanism. We evaluate the proposed method in both simulations and real vehicle experiments in a complex autonomous driving collision avoidance task. The proposed method achieves up to four times fewer constraint violations and converges 3.36 times faster than baseline constrained RL approaches. Haitong Ma, Jianyu Chen 0002, Shengbo Eben Li, Ziyu Lin, Yang Guan, Yangang Ren, Sifa Zheng |
IROS | 5 |
| 2021 | Separated Proportional-Integral Lagrangian for Chance Constrained Reinforcement LearningabstractSafety is essential for reinforcement learning (RL) applied in real-world tasks like autonomous driving. Imposing chance constraints (or probabilistic constraints) is a suitable way to enhance RL safety under model uncertainty. Existing chance constrained RL methods like the penalty methods and the Lagrangian methods either exhibit periodic oscillations or learn an over-conservative or unsafe policy. In this paper, we address these shortcomings by elegantly combining these two methods and propose a separated proportional-integral Lagrangian (SPIL) algorithm. We first rewrite penalty methods as optimizing safe probability according to the proportional value of constraint violation, and Lagrangian methods as optimizing according to the integral value of the violation. Then we propose to add up both the integral and proportion values to optimize the policy, with an integral separation technique to limit the integral value within a reasonable range. Besides, the gradient of policy is computed in a model-based paradigm to accelerate training. The proposed method is proved to reduce oscillations and conservatism while ensuring safety by a car-following experiment. Baiyu Peng, Yao Mu 0001, Jingliang Duan, Yang Guan, Shengbo Eben Li, Jianyu Chen 0002 |
IV | 4 |
| 2021 | Cover: International Journal of Intelligent Systems, Volume 36 Issue 8 August 2021abstractCover Caption: The cover image is based on the Research Article Direct and indirect reinforcement learning by Yang Guan et al., https://doi.org/10.1002/int.22466. Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003 |
Int. J. Intell. Syst. | 1 |
| 2021 | Direct and indirect reinforcement learningabstractReinforcement learning (RL) algorithms have been successfully applied to a range of challenging sequential decision-making and control tasks. In this paper, we classify RL into direct and indirect RL according to how they seek the optimal policy of the Markov decision process problem. The former solves the optimal policy by directly maximizing an objective function using gradient descent methods, in which the objective function is usually the expectation of accumulative future rewards. The latter indirectly finds the optimal policy by solving the Bellman equation, which is the sufficient and necessary condition from Bellman's principle of optimality. We study policy gradient (PG) forms of direct and indirect RL and show that both of them can derive the actor–critic architecture and can be unified into a PG with the approximate value function and the stationary state distribution, revealing the equivalence of direct and indirect RL. We employ a Gridworld task to verify the influence of different forms of PG, suggesting their differences and relationships experimentally. Finally, we classify current mainstream RL algorithms using the direct and indirect taxonomy, together with other ones, including value-based and policy-based, model-based and model-free. Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003 |
Int. J. Intell. Syst. | 1 |
| 2020 | Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalabstractHashing-based cross-modal search which aims to map multiple modality features into binary codes has attracted increasingly attention due to its storage and search efficiency especially in large-scale database retrieval. Recent unsupervised deep cross-modal hashing methods have shown promising results. However, existing approaches typically suffer from two limitations: (1) They usually learn cross-modal similarity information separately or in a redundant fusion manner, which may fail to capture semantic correlations among instances from different modalities sufficiently and effectively. (2) They seldom consider the sampling and weighting schemes for unsupervised cross-modal hashing, resulting in the lack of satisfactory discriminative ability in hash codes. Shengsheng Qian, Yang Guan, Jiawei Zhan, Long Ying |
SIGIR | 3 |
| 2017 | A robust and synthesized-unseen watermarking for the DRM of DIBR-based 3D video
Xiyao Liu 0001, Fangfang Li 0004, Jingyu Du, Yang Guan, Yuesheng Zhu, Beiji Zou 0001 |
Neurocomputing | 4 |
| 2014 | Power efficient peer-to-peer streaming to co-located mobile usersabstractThe increasing popularity of Internet-capable mobile devices leads to the explosion of mobile data traffic. Although the emerging mobile broadband systems such as 4G LTE promise higher bandwidth and lower latency for video traffic, it is not power efficient to deliver video traffic over cellular networks. This paper studies how peer-to-peer communications via wireless local networks could complement the cellular networks so as to minimize the power consumption of mobile devices. Specifically, we consider the scenario in which a group of nearby mobile devices share the streaming contents, originated from the cellular networks, over the wireless local area networks in a peer-to-peer manner. We focus on investigating how the mobile devices should cooperate to minimize their own power consumption. The problem has been formulated into a linear programming (LP) model, and numerical results show that at least 31% of the power consumption can be saved on mobile devices through the cooperation among them. Yang Guan, Leonard J. Cimini Jr., Chien-Chung Shen |
CCNC | 1 |
| 2014 | MobiCacher: Mobility-aware content caching in small-cell networksabstractSmall-cell networks have been proposed to meet the demand of ever growing mobile data traffic. One of the prominent challenges faced by small-cell networks is the lack of sufficient backhaul capacity to connect small-cell base stations (small-BSs) to the core network. We exploit the effective application layer semantics of both spatial and temporal locality to reduce the backhaul traffic. Specifically, we envision that small-BSs are equipped with storage facility to cache contents requested by users. As the cache hit ratio increases, most of the users' requests can be satisfied locally without incurring traffic over the backhaul. To make informed caching decisions, the mobility patterns of users must be carefully considered as users might frequently migrate from one small cell to another. We study the issue of mobility-aware content caching, which is formulated into an optimization problem with the objective to maximize the caching utility. As the problem is NP-complete, we develop a polynomial-time heuristic solution termed MobiCacher with bounded approximation ratio. We conduct trace-based simulations to evaluate the performance of MobiCacher, which show that MobiCacher yields better caching utility than existing solutions. Yang Guan, Hao Feng 0006, Chien-Chung Shen, Leonard J. Cimini Jr. |
GLOBECOM | 1 |
| 2014 | A digital blind watermarking scheme based on quantization index modulation in depth map for 3D videoabstract3D video provides an immersive experience to viewers and is getting more and more popular. The solution to create 3D video from 2D video is low-cost compared with that captures 3D video directly, and the generation of depth map from 2D video is a key in the 2D-3D video conversion systems. Therefore, protection of depth map is vital for 3D video. In this paper, a digital blind watermarking scheme based on Quantization Index Modulation (QIM) algorithm is proposed in which the copyright information is embedded in the DCT coefficients of depth map imperceptibly. The experimental results show that the proposed scheme has good robustness against video attacks such as salt noise, median filtering, wiener filtering, and scaling. In the meanwhile, the stereo video embedded watermarking can accomplish zero distortion in comparison with the original one. Yang Guan, Yuesheng Zhu, Xiyao Liu 0001, Guibo Luo, Ziqiang Sun, Liming Zhang 0002 |
ICARCV | 1 |
| 2011 | MAC scheduling for high throughput underwater acoustic networksabstractUnderwater acoustic networks (UWANs) have emerged as the primary tool to monitor and act upon the well-being of marine environment. However, the significantly slower propagation speed of acoustic signals, in contrast to RF signals, introduces the spatio-temporal uncertainty, which makes existing medium access control (MAC) solutions for terrestrial RF wireless networks unsuitable for UWANs. In this paper, we investigate transmission scheduling for time-based MAC protocols and design scheduling algorithms that take advantage of the long propagation delay of acoustic signals to facilitate concurrent transmissions and receptions of acoustic communications. Specifically, we specify the constraints that MAC protocols need to satisfy to avoid conflicts and model these constrains into a Mixed Integer Linear Programming model. We also design heuristics that compute conflict-free transmission schedules, and demonstrate via simulation that the heuristics significantly improve network throughput. Yang Guan, Chien-Chung Shen, Justin Yackoski |
WCNC | 1 |
| 2011 | CSR: Cooperative source routing using virtual MISO in wireless ad hoc networksabstractCooperative transmissions combat various fading effects in wireless communications by employing multiple antennas from different nodes to achieve spatial diversity. Virtual Multiple-Input Single-Output (MISO) is one instance of cooperative transmissions capable of achieving higher receiving SNR, which can either extend the transmission range or increase the transmission rate. While the physical layer performance of Virtual MISO has been well studied, this paper describes the Cooperative Source Routing (CSR) protocol to convert physical layer gain into network level performance improvement. With both route request and route reply control packets being transmitted cooperatively, CSR can explore routes with high cooperative diversity. We demonstrate CSR's performance through simulation and compare CSR with other protocols. Yang Guan, Chien-Chung Shen, Leonard J. Cimini Jr. |
WCNC | 1 |
| 2011 | Location-Aware cooperative routing in multihop wireless networksabstractGeographic routing is a scalable routing scheme for wireless networks, where nodes make local routing decisions using position information. Cooperative transmissions utilize spatial diversity to combat fading in wireless channels. In this paper, we incorporate cooperative transmissions into geographic routing, and propose the Location-Aware Cooperative Routing (LACR). In LACR, a node that receives Route Request (RREQ) makes an individual decision on whether and how to rebroadcast RREQ based on its position and capability to cooperate. A theoretical analysis of the impact of cooperative transmissions on the transmission range extension is presented to guide the measurement of the potential performance of each node. Through simulation, we show that LACR performs better in terms of higher probability to find a route and higher throughput in comparison to Single-Input-Single-Output (SISO) based routing protocol. Yang Guan, Wei Chen 0002, Chien-Chung Shen, Leonard J. Cimini Jr. |
WCNC | 2 |