Shiyu Liang

dblp:158/4797 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021Computer networks · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
abstract
Supervised fine-tuning (SFT) on long chainof-thought (CoT) trajectories has emerged as a crucial technique for enhancing the reasoning abilities of large language models (LLMs).However, the standard cross-entropy loss treats all tokens equally, ignoring their heterogeneous contributions across a reasoning trajectory.This uniform treatment leads to misallocated supervision and weak generalization, especially in complex, long-form reasoning tasks.To address this, we introduce Variance-Controlled Optimization-based REweighting (VCORE), a principled framework that reformulates CoT supervision as a constrained optimization problem.By adopting an optimization-theoretic perspective, VCORE enables a principled and adaptive allocation of supervision across tokens, thereby aligning the training objective more closely with the goal of robust reasoning generalization.Empirical evaluations demonstrate that VCORE achieves the strongest overall average performance, with especially clear gains on lower-capacity models.Across both in-domain and out-of-domain settings, VCORE achieves substantial performance gains on mathematical and coding benchmarks, using models from the Qwen3 series (4B, 8B, 32B) and LLaMA-3.1-8B-Instruct.Moreover, we show that VCORE serves as a more effective initialization for subsequent reinforcement learning, establishing a stronger foundation for advancing the reasoning capabilities of LLMs. 1
Senmiao Wang, Hanbo Huang, Ruoyu Sun 0001, Shiyu Liang
ACL (1)5
2026 Deep-Saliency Foveated Ray Tracing For Real-time VR Rendering
abstract
Immersive VR applications demand high resolutions and refresh rates, posing significant challenges for real-time rendering. Foveated rendering mitigates this cost by exploiting properties of the Human Visual System (HVS), but conventional approaches often rely on oversimplified heuristic models that neglect high-level attentional cues, resulting in artifacts in peripheral regions. To this end, we present a neural saliency-driven foveated ray tracing framework that overcomes these limitations. Our method introduces a motion-aware foveation model to capture temporal dynamics and employs a lightweight convolutional neural network to predict saliency maps that reflect complex attentional patterns derived from eye-gaze data. The combination of these guides adaptive path tracing and filtering, enabling perceptually optimized rendering with minimal artifacts. Experimental results show that our approach improves perceptual quality over prior methods while sustaining real-time performance.
Yang Gao 0032, Wencan Li, Shiyu Liang, Weizichuan Feng, Qing Xia 0002, Shuai Li 0001, Aimin Hao
IEEE Trans. Vis. Comput. Graph.3
2025 MM4flow: A Pre-trained Multi-modal Model for Versatile Network Traffic Analysis
abstract
Network traffic analysis is a critical research area, playing an essential role in enhancing network security and ensuring high-quality network services. Existing methods, which primarily rely on a single modality, face two significant limitations. First, while existing approaches may achieve strong performance in specific tasks, they often lack sufficient adaptability for diverse tasks. Second, existing pre-trained models are only trained with GB-scale traffic, with which increases the risk of over-fitting and limiting the models' overall performance. To address these challenges, we propose MM4flow, a pre-trained multi-modal model designed for versatile network traffic analysis. We divide network flows into two modalities: raw byte streams and transmission patterns, which encapsulate the content and behavior information, respectively. MM4flow is composed of two key stages: uni-modal pre-training and multi-modal fine-tuning. We develop an efficient data collection scheme enabling TB-scale traffic pre-training. Leveraging a real-world traffic that exceeds 70 TB, MM4flow conducts uni-modal pre-training on each modality with a modified BERT architecture tailored for network flows. For specific downstream tasks, we introduce a modal fusion module based on cross-attention mechanisms. The fusion module facilitates effective integration of multi-modal information, enabling MM4flow to fully utilize both content and behavior cues during fine-tuning with minimal labeled dataset. We evaluate MM4flow on six public datasets covering six various tasks. Extensive experiments demonstrate that MM4flow achieves superior accuracy than baselines. Especially, compared to existing pre-trained models, MM4flow achieves an 84% improvement in accuracy for website identification under encrypted tunnels. Moreover, the pre-trained MM4flow significantly reduces the reliance on high-quality labeled training data for downstream tasks.
Luming Yang, Lin Liu 0018, Junjie Huang 0001, Zhuotao Liu, Shiyu Liang, Shaojing Fu
CCS5
2025 A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
abstract
Privacy-sensitive users require deploying large language models (LLMs) within their own infrastructure (on-premises) to safeguard private data and enable customization.However, vulnerabilities in local environments can lead to unauthorized access and potential model theft.To address this, prior research on small models has explored securing only the output layer within hardware-secured devices to balance model confidentiality and customization.Yet this approach fails to protect LLMs effectively.In this paper, we discover that (1) query-based distillation attacks targeting the secured top layer can produce a functionally equivalent replica of the victim model; (2) securing the same number of layers, bottom layers before a transition layer provide stronger protection against distillation attacks than top layers, with comparable effects on customization performance; and (3) the number of secured layers creates a trade-off between protection and customization flexibility.Based on these insights, we propose SOLID, a novel deployment framework that secures a few bottom layers in a secure environment and introduces an efficient metric to optimize the trade-off by determining the ideal number of hidden layers.Extensive experiments on five models (1.3B to 70B parameters) demonstrate that SOLID outperforms baselines, achieving a better balance between protection and downstream customization.Our code can be found at: https://github.com/ OTTO-OTO/SOLID-OnPremiseDeployment.
Hanbo Huang, Lin Liu 0018, Zhuotao Liu, Ruoyu Sun 0001, Shiyu Liang
EMNLP8
2025 $\text{CO}_{2}$-Net: A Physics-Informed Spatio-Temporal Model for Global Surface $\text{CO}_{{2}}$ Reconstruction
Hanbo Huang, Chaofan Sun, Enhui Liao, Shiyu Liang
ICCV9
2025 NUTS: Eddy-Robust Reconstruction of Surface Ocean Nutrients via Two-Scale Modeling
abstract
Reconstructing ocean surface nutrients from sparse observations is critical for understanding long-term biogeochemical cycles. Most prior work focuses on reconstructing atmospheric fields and treats the reconstruction problem as image inpainting, assuming smooth, single-scale dynamics. In contrast, nutrient transport follows advection–diffusion dynamics under nonstationary, multiscale ocean flow. This mismatch leads to instability, as small errors in unresolved eddies can propagate through time and distort nutrient predictions. To address this, we introduce NUTS, a two-scale reconstruction model that decouples large-scale transport and mesoscale variability. The homogenized solver captures stable, coarse-scale advection under filtered flow. A refinement module then restores mesoscale detail conditioned on the residual eddy field. NUTS is stable, interpretable, and robust to mesoscale perturbations, with theoretical guarantees from homogenization theory. NUTS outperforms all data-driven baselines in global reconstruction and achieves site-wise accuracy comparable to numerical models. On real observations, NUTS reduces NRMSE by 79.9% for phosphate and 19.3% for nitrate over the best baseline. Ablation studies validate the effectiveness of each module.
Shiyu Liang, Chaofan Sun, Lei Bai 0001, Enhui Liao
NeurIPS2
2025 CertTA: Certified Robustness Made Practical for Learning-Based Traffic Analysis
Jinzhu Yan, Zhuotao Liu, Shiyu Liang, Lin Liu 0018, Ke Xu 0002
USENIX Security Symposium4
2025 AI-agent communication network for 6G: vision, architecture, and key technologies
abstract
The booming of artificial intelligence (AI) agents has brought about promising business scenarios for sixth-generation (6G) mobile networks, while simultaneously posing significant challenges to network functionalities and infrastructure. These AI agents can be deployed on end devices (e.g., intelligent robots and intelligent cars) or as digital entities (e.g., personal AI assistants). As novel service entities with autonomous decision-making and task execution capabilities, AI agents introduce potential risks of uncontrollable actions and privacy disclosures. AI agents also require new 6G capabilities beyond traditional communication, including multimodality information interaction (e.g., AI models and tokens) and support for service requirements (e.g., computing and sensing of data). In this article, we introduce the concept of AI-agent communication network (ACN), a new paradigm to enable global information interaction and on-demand capability provisioning for single or multiple AI agents. We first introduce the vision and architectural framework of ACN. Then, key technologies and future research directions related to ACN are discussed. Furthermore, we provide potential use cases to elaborate on how ACN can expand the service capabilities of 6G networks.
Xiaodong Duan, Zhenglei Huang, Shiyu Liang, Shaowen Zheng, Lu Lu 0016, Tao Sun 0010
Frontiers Inf. Technol. Electron. Eng.3
2025 Saliency-Aware Foveated Path Tracing for Virtual Reality Rendering
abstract
Foveated rendering reduces computational load by distributing resources based on the human visual system. This enables the implementation of ray tracing in virtual reality applications, where a high frame rate is essential to achieve visual immersion. However, traditional foveation methods based solely on eccentricity cannot adequately account for the complex behavior of visual attention. This is one of the main reasons that leads to lower perceived quality compared to non-foveated techniques. In this study, we introduce a novel rendering pipeline that incorporates ocular attention through the use of visual saliency. Based on foveation saliency, our approach facilitates the real-time production of high-quality images utilizing path tracing by distributing samples according to saliency metrics derived from geometric and historical data. To further augment image quality, an adaptive filtering process, aligned with the saliency metrics, is employed to reduce visible artifacts in non-foveal regions. Our experiments prove that this novel approach can demonstrate superior performance compared to previous methods, both in terms of quantitative metrics and perceived visual quality.
Yang Gao 0032, Wencan Li, Shiyu Liang, Aimin Hao, Xiaohui Tan
IEEE Trans. Vis. Comput. Graph.3
2025 Efficient Photon Beam Diffusion for Directional Subsurface Scattering
abstract
Real-time subsurface scattering techniques are widely used in translucent material rendering. Among advanced methods that rely on the bidirectional scattering-surface reflectance distribution function (BSSRDF), screen space algorithms exhibit limited translucency, while existing large-distance methods are inefficient and yield poor illumination details. To address these limitations for better large-distance scattering, we develop a novel algorithm by extending the photon beam diffusion (PBD) model within the light view and screen space. Unlike surface irradiance in prior methods, we incorporate the refracted beam in the medium into real-time scattering estimation, presenting a new consideration for photon beam utilization. Concretely, we store all photon beam samples in light view textures and utilize an adaptive sampling pattern for beam sample selection in large filtering kernel sizes. This can reduce the sample count based on surface attributes. In screen space, virtual sources are derived from samples to estimate PBD contributions, with an approximation that preserves boundary conditions. To avoid possible overestimation, we implement correction factors that scale contributions, effectively aligning our results with path-tracing references. Through these reformulations, our efficient PBD generates results closest to references among existing methods. The experiments accurately represent better front-face illumination details and backlit translucency effects, while significantly accelerating performance compared to previous large-distance methods.
Shiyu Liang, Yang Gao 0032, Chonghao Hu, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.1
2024 Temporal Generalization Estimation in Evolving Graphs
abstract
Graph Neural Networks (GNNs) are widely deployed in vast fields, but they often struggle to maintain accurate representations as graphs evolve. We theoretically establish a lower bound, proving that under mild conditions, representation distortion inevitably occurs over time. To estimate the temporal distortion without human annotation after deployment, one naive approach is to pre-train a recurrent model (e.g., RNN) before deployment and use this model afterwards, but the estimation is far from satisfactory. In this paper, we analyze the representation distortion from an information theory perspective, and attribute it primarily to inaccurate feature extraction during evolution. Consequently, we introduce Smart, a straightforward and effective baseline enhanced by an adaptive feature extractor through self-supervised graph reconstruction. In synthetic random graphs, we further refine the former lower bound to show the inevitable distortion over time and empirically observe that Smart achieves good estimation performance. Moreover, we observe that Smart consistently shows outstanding generalization estimation on four real-world evolving graphs. The ablation studies underscore the necessity of graph reconstruction. For example, on OGB-arXiv dataset, the estimation metric MAPE deteriorates from 2.19% to 8.00% without reconstruction.
Bin Lu 0005, Tingyan Ma, Xiaoying Gan, Xinbing Wang, Yunqiang Zhu, Chenghu Zhou, Shiyu Liang
ICLR7
2024 State of the Art in Efficient Translucent Material Rendering with BSSRDF
abstract
Abstract Sub‐surface scattering is always an important feature in translucent material rendering. When light travels through optically thick media, its transport within the medium can be approximated using diffusion theory, and is appropriately described by the bidirectional scattering‐surface reflectance distribution function (BSSRDF). BSSRDF methods rely on assumptions about object geometry and light distribution in the medium, which limits their applicability to general participating media problems. However, despite the high computational cost of path tracing, BSSRDF methods are often favoured due to their suitability for real‐time applications. We review these methods and discuss the most recent breakthroughs in this field. We begin by summarizing various BSSRDF models and then implement most of them in a 2D searchlight problem to demonstrate their differences. We focus on acceleration methods using BSSRDF, which we categorize into two primary groups: pre‐computation and texture methods. Then we go through some related topics, including applications and advanced areas where BSSRDF is used, as well as problems that are sometimes important yet are ignored in sub‐surface scattering estimation. In the end of this survey, we point out remaining constraints and challenges, which may motivate future work to facilitate sub‐surface scattering.
Shiyu Liang, Yang Gao 0032, Chonghao Hu, Aimin Hao, Lili Wang 0006, Hong Qin 0001
Comput. Graph. Forum1
2024 Hi-PART: Going Beyond Graph Pooling with Hierarchical Partition Tree for Graph-Level Representation Learning
abstract
Graph pooling refers to the operation that maps a set of node representations into a compact form for graph-level representation learning. However, existing graph pooling methods are limited by the power of the Weisfeiler–Lehman (WL) test in the performance of graph discrimination. In addition, these methods often suffer from hard adaptability to hyper-parameters and training instability. To address these issues, we propose Hi-PART, a simple yet effective graph neural network (GNN) framework with Hi erarchical Par tition T ree (HPT). In HPT, each layer is a partition of the graph with different levels of granularities that are going toward a finer grain from top to bottom. Such an exquisite structure allows us to quantify the graph structure information contained in HPT with the aid of structural information theory. Algorithmically, by employing GNNs to summarize node features into the graph feature based on HPT’s hierarchical structure, Hi-PART is able to adequately leverage the graph structure information and provably goes beyond the power of the WL test. Due to the separation of HPT optimization from graph representation learning, Hi-PART involves the height of HPT as the only extra hyper-parameter and enjoys higher training stability. Empirical results on graph classification benchmarks validate the superior expressive power and generalization ability of Hi-PART compared with state-of-the-art graph pooling approaches.
Yuyang Ren, Haonan Zhang 0004, Luoyi Fu, Shiyu Liang, Lei Zhou 0016, Xinbing Wang, Xinde Cao, Chenghu Zhou
ACM Trans. Knowl. Discov. Data4
2024 Graph Out-of-Distribution Generalization With Controllable Data Augmentation
abstract
Graph Neural Network (GNN) has demonstrated extraordinary performance in classifying graph properties. However, due to the selection bias of training and testing data (e.g., training on small graphs and testing on large graphs, or training on dense graphs and testing on sparse graphs), distribution deviation is widespread. More importantly, we often observehybrid structure distribution shiftof both scale and density, despite of one-sided biased data partition. The spurious correlations over hybrid distribution deviation degrade the performance of previous GNN methods and show large instability among different datasets. To alleviate this problem, we proposeOOD-GMixupto jointly manipulate the training distribution withcontrollable data augmentationin metric space. Specifically, we first extract the graph rationales to eliminate the spurious correlations due to irrelevant information. Secondly, we generate virtual samples with perturbation on graph rationale representation domain to obtain potential OOD training samples. Finally, we propose OOD calibration to measure the distribution deviation of virtual samples by leveraging Extreme Value Theory, and further actively control the training distribution by emphasizing the impact of virtual OOD samples. Extensive studies on several real-world datasets on graph classification demonstrate the superiority of our proposed method over state-of-the-art baselines.
Bin Lu 0005, Ze Zhao, Xiaoying Gan, Shiyu Liang, Luoyi Fu, Xinbing Wang, Chenghu Zhou
IEEE Trans. Knowl. Data Eng.4
2024 FlowerCast: Efficient Time-sensitive Multicast in Wireless Sensor Networks with Link Uncertainty
abstract
This article studies time-sensitive multicast in wireless sensor networks (WSNs) with link uncertainty, where information from the source needs to be delivered to multiple receivers within an imposed delay constraint. Prior art on static WSNs minimizes the multicast delay via the construction of a multicast tree that approximates the Steiner tree in length, which, however, may be invalidated by the time-varying network topology of WSNs with uncertain link states. Moreover, for multicast in WSNs with link uncertainty, the possible link failure necessitates a suitable measurement of the uncertain communication distance and calls for the performance guarantee in both delay and delivery ratio. In this work, by modeling a WSN as a random graph with each link associated with a transmission probability, we propose FlowerCast, an efficient multicast scheme, to jointly minimize the expected multicast delay and to maximize the expected delivery ratio of multicast under delay constraint. The core of FlowerCast is to quantify the uncertain communication distance by the expected transmission delay of a time-varying path, based on which a delay-optimal multicast tree is constructed in accordance with the directionality of delay. Candidate paths with a high expected delivery ratio and low expected delay are then selected in a distributed manner to conditionally connect adjacent multicast members and thus transform the multicast tree into a multicast flower. Despite the NP-hardness of optimal candidate paths’ addition, the transformation with the highest expected delivery ratio of multicast under delay constraint can be guaranteed through a pseudo-polynomial time derandomization-based greedy approach. We further demonstrate the time and energy efficiency of FlowerCast through asymptotic analysis. To make full use of the possible overlapping links in a multicast flower, a hybrid routing strategy is presented to wisely switch between sequential routing and synchronous routing for extra enhancement of the multicast performance. Extensive experiments on various datasets verify the superiority of FlowerCast and hybrid routing over baselines and indicate their wide applicability to practical scenarios.
Jianzhi Tang, Luoyi Fu, Shiyu Liang, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou
ACM Trans. Sens. Networks3
2023 Cluster-Specific Dictionary Learning Based Active User Detection for mMTC With Massive MIMO
abstract
Massive machine-type communication (mMTC) is an important scenario for 5G and future 6G networks, as it can provide massive connectivity for internet of things (IoT) devices. However, the large number of supported devices raises challenges to random access with limited spectrum resources. In this paper, we propose a dictionary learning based method for active user detection (AUD) in massive MIMO systems, which leverages the potential spatial channel characteristics of users. Our approach separates users into clusters and reuses the same pilot pool among different clusters, which greatly saves the pilot resource. To resolve collisions caused by the reuse of pilots, we propose a cluster-specific dictionary to differentiate multiple active users of different clusters. Numerical experiments demonstrate the improved performance of the proposed AUD algorithm in comparison to the existing methods.
Shiyu Liang, Wei Chen 0016, Ning Wang 0004, Bo Ai 0001
GLOBECOM1
2018 Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networks
Shiyu Liang, Yixuan Li 0001, R. Srikant 0001
ICLR (Poster)1
2018 Understanding the Loss Surface of Neural Networks for Binary Classification
abstract
It is widely conjectured that training algorithms for neural networks are successful because all local minima lead to similar performance; for example, see (LeCun et al., 2015; Choromanska et al., 2015; Dauphin et al., 2014). Performance is typically measured in terms of two metrics: training performance and generalization performance. Here we focus on the training performance of neural networks for binary classification, and provide conditions under which the training error is zero at all local minima of appropriately chosen surrogate loss functions. Our conditions are roughly in the following form: the neurons have to be increasing and strictly convex, the neural network should either be single-layered or is multi-layered with a shortcut-like connection, and the surrogate loss function should be a smooth version of hinge loss. We also provide counterexamples to show that, when these conditions are relaxed, the result may not hold.
Shiyu Liang, Ruoyu Sun 0001, Yixuan Li 0001, R. Srikant 0001
ICML1
2018 Adding One Neuron Can Eliminate All Bad Local Minima
abstract
One of the main difficulties in analyzing neural networks is the non-convexity of the loss function which may have many bad local minima. In this paper, we study the landscape of neural networks for binary classification tasks. Under mild assumptions, we prove that after adding one special neuron with a skip connection to the output, or one special neuron per layer, every local minimum is a global minimum.
Shiyu Liang, Ruoyu Sun 0001, Jason D. Lee, R. Srikant 0001
NeurIPS1
2018 B4 and after: managing hierarchy, partitioning, and asymmetry for availability and scale in google's software-defined WAN
abstract
Private WANs are increasingly important to the operation of enterprises, telecoms, and cloud providers. For example, B4, Google's private software-defined WAN, is larger and growing faster than our connectivity to the public Internet. In this paper, we present the five-year evolution of B4. We describe the techniques we employed to incrementally move from offering best-effort content-copy services to carrier-grade availability, while concurrently scaling B4 to accommodate 100x more traffic. Our key challenge is balancing the tension introduced by hierarchy required for scalability, the partitioning required for availability, and the capacity asymmetry inherent to the construction and operation of any large-scale network. We discuss our approach to managing this tension: i) we design a custom hierarchical network topology for both horizontal and vertical software scaling, ii) we manage inherent capacity asymmetry in hierarchical topologies using a novel traffic engineering algorithm without packet encapsulation, and iii) we re-architect switch forwarding rules via two-stage matching/hashing to deal with asymmetric network failures at scale.
Chi-Yao Hong, Subhasree Mandal, Mohammad Al-Fares, Richard Alimi, Kondapa Naidu Bollineni, Chandan Bhagat, Sourabh Jain, Jay Kaimal, Shiyu Liang, Kirill Mendelev, Steve Padgett, Faro Rabe, Saikat Ray, Malveeka Tewari, Matt Tierney, Monika Zahn, Jonathan Zolla, Joon Ong, Amin Vahdat
SIGCOMM10
2018 FINE: A Framework for Distributed Learning on Incomplete Observations for Heterogeneous Crowdsensing Networks
Luoyi Fu, Songjun Ma, Lingkun Kong, Shiyu Liang, Xinbing Wang
IEEE/ACM Trans. Netw.4
2017 Why Deep Neural Networks for Function Approximation?
Shiyu Liang, R. Srikant 0001
ICLR (Poster)1
2015 Prioritization of potential candidate disease genes by topological similarity of protein-protein interaction network and phenotype data
Jiawei Luo 0001, Shiyu Liang
J. Biomed. Informatics2
2014 Are we still friends: Kernel multivariate survival analysis
abstract
Online Social Network becomes the most prevalent platform for exchanging information between users, maintaining friendships online. As is well-known to us, however, some friendships even those intimate ones might vanish. Therefore, precisely modeling and predicting state of each online relationship is worthwhile in many respects. For social communication services such modeling permits new and novel online services. In addition, constructing this model might enlighten us in exploiting information spreading pattern in online social network. In this paper, we propose a model in determining a probability distribution which describes the ‘surviving time’ of each friendships by applying one commonly used method in sociology, survival analysis. We discuss a series of social explanatory variables that highly affect this probability distribution. Moreover, methods in the moving average process are devoted to determining the appropriate parameter in survival model. Furthermore, to avoid the high computational complexity in kernel learning we impose sparsity in our model. Finally, with the experiments on real data, the proposed survival model is proven to be of high accuracy, and thus of great potential for further applications.
Shiyu Liang, Ruotian Luo, Songjun Ma, Weijie Wu, Li Song 0001, Xiaohua Tian, Xinbing Wang
GLOBECOM1