Biyu Zhou

dblp:144/7425 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Computer networks · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
abstract
Large Language Models (LLMs) demonstrate impressive capabilities across diverse tasks, yet their safety mechanisms remain susceptible to adversarial exploitation of cognitive biases---systematic deviations from rational judgment. Unlike prior studies focusing on isolated biases, this work highlights the overlooked power of multi-bias interactions in undermining LLM safeguards. Specifically, we propose CognitiveAttack, a novel red-teaming framework that adaptively selects optimal ensembles from 154 human social psychology-defined cognitive biases, engineering them into adversarial prompts to effectively compromise LLM safety mechanisms. Experimental results reveal systemic vulnerabilities across 30 mainstream LLMs, particularly open-source variants. CognitiveAttack achieves a substantially higher attack success rate than the SOTA black-box method PAP (60.1% vs. 31.6%), exposing critical limitations in current defenses. Through quantitative analysis of successful jailbreaks, we further identify vulnerability patterns in safety-aligned LLMs under synergistic cognitive biases, validating multi-bias interactions as a potent yet underexplored attack vector. This work introduces a novel interdisciplinary perspective by bridging cognitive science and LLM safety, paving the way for more robust and human-aligned AI systems.
Xikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu 0001
AAAI2
2026 More Thinking, Less Talking: Internalizing Deliberative Safety into LLM Parameters
abstract
Prevailing safety alignment methods still leave Large Language Models (LLMs) vulnerable to sophisticated jailbreak attacks.To bolster defenses, explicit reasoning mechanisms like Safety-oriented Chain-of-Thought (SCoT) have emerged, significantly enhancing robustness.However, this transparency introduces a critical trade-off: the exposed reasoning process itself becomes a new attack surface, risking the leakage of harmful information and revealing the model's safety logic to adversaries.This paper directly confronts this dilemma, asking: Can we achieve the full benefits of deliberative safety without the costs of explicit reasoning generation?We propose Safety Reasoning Internalization to make the deliberative process in SCoT "available but not visible".This approach is grounded in a key theoretical insight: the corrective influence of an SCoT can be effectively approximated by a targeted, low-rank update to the model's Feed-Forward Network (FFN) layers.We operationalize this through Hierarchical Internalization of Adversarially-Guided Reasoning (HIAR), a layer-wise safety alignment framework that internalizes safety reasoning into an implicit computational pathway using Low-Rank Adaptation (LoRA).HIAR enables the model to reach a safe conclusion within a single forward pass, entirely eliminating the need to generate vulnerable SCoT text.Extensive experiments on various LLMs demonstrate that HIAR achieves a 43% lower Attack Success Rate (ASR) against distinct jailbreak attacks compared to strong baselines.
Xuehai Tang, Biyu Zhou, Jizhong Han, Songlin Hu 0001
ACL (1)3
2026 Resolving the Security-Auditability Dilemma with Auditable Latent Chain-of-Thought Alignment
abstract
To address the increasingly severe safety risk of large language models (LLMs), reasoningbased safety alignment methods have emerged.These methods overcome the limitations of 'shallow alignment' by exposing the model's Chain-of-Thought (CoT), enabling auditability of safety reasoning process through both training-phase supervision and post-generation verification.However, this transparency creates a critical vulnerability, a tension we define as the Security Auditability Dilemma: while explicit reasoning is a prerequisite for safety, its textual Auditable paradoxically transforms it into an optimization target for adaptive attackers and induces the model to unintentionally copy harmful content from its own reasoning context.To address this, we propose Auditable Latent CoT Alignment (ALCA), a framework that decouples internal reasoning from external output.ALCA shifts the safety deliberation process into a continuous latent space.This allows the safety reasoning process to guide the generation of harmless outputs, while eliminates the discrete textual surface that facilitates internal copying and adaptive attack.Yet, this process is not a black box.we introduce a restricted Self-Decoding mechanism that allows the model to reconstruct its latent reasoning into human-readable text for supervision under specific guidance.Extensive experiments show that ALCA achieves robustness alignment, reducing the success rate of adaptive jailbreak attacks by over 40% compared to strong baselines, while preserving performance.Our framework presents a path toward building LLMs that are both robustly secure and auditable.
Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu 0001
ACL (1)2
2025 LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editing
abstract
Large Language Models often contain factually incorrect or outdated knowledge, giving rise to model editing methods for precise knowledge updates.However, current mainstream locate-then-edit approaches exhibit a progressive performance decline during sequential editing, due to inadequate mechanisms for long-term knowledge preservation.To tackle this, we model the sequential editing as a constrained stochastic programming.Given the challenges posed by the cumulative preservation error constraint and the gradually revealed editing tasks, LyapLock is proposed.It integrates queuing theory and Lyapunov optimization to decompose the long-term constrained programming into tractable stepwise subproblems for efficient solving.This is the first model editing framework with rigorous theoretical guarantees, achieving asymptotic optimal editing performance while meeting the constraints of long-term knowledge preservation.Experimental results show that our framework scales sequential editing capacity to over 10,000 edits while stabilizing general capabilities and boosting average editing efficacy by 11.89% over SOTA baselines.Furthermore, it can be leveraged to enhance the performance of baseline methods.Our code is released on https://github.com/caskcsg/LyapLock.
Peng Wang 0028, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu 0001
EMNLP2
2025 Segment-Recurrent Transformer with Multi-Scale Fusion for Long-Term Time Series Forecasting
abstract
Long-term time series forecasting (LTSF) seeks to make accurate long-term predictions by leveraging extensive historical data, which is crucial for solving scientific and engineering challenges. Traditional transformer-based methods process historical segments individually, leading to a limited view that overlooks distant dependencies within the entire time series. In this paper, we introduce the Segment-Recurrent Transformer (SRTrans), designed to provide a more comprehensive understanding of historical time series dynamics. By incorporating segment-level recurrence into the Transformer, our model enhances inter-segment information flow, capturing longer-term and global dependencies. We also propose a multi-scale adaptive fusion module that efficiently integrates diverse patterns using a variable-scale chunking mechanism and a weight-mixing strategy. Additionally, our spectrum purge operation improves data preprocessing by extracting significant long-term patterns from the frequency domain. Extensive experiments on eight real-world datasets demonstrate SRTrans’s effectiveness in accuracy and efficiency, offering a promising new solution for LTSF tasks.
Ziang Yang, Lingwei Wei, Biyu Zhou, Xuehai Tang, Ruixuan Li 0001, Songlin Hu 0001
ICASSP3
2024 Breaking the Weak Semantics Bottleneck of Transformers in Time Series Forecasting
abstract
Transformer with self-attention was initially crafted to model language sequences, where discrete tokens (i.e., words) showcase high semantic density. However, when applied to time series token inputs (i.e., datapoints) with weak-density semantics and temporal redundancy, it faces challenges as these time-domain tokens impede its ability to capture the intricate latent properties of time series dynamics. While time-frequency transformation presents a viable solution by bringing forth a new space with heightened expressive power, existing approaches fall short of fully exploiting its potential. In response to these limitations, we propose a general-purpose transformer-based model, named Scattering Transformer, for multivariate time series forecasting and self-supervised representation learning. It is based on two innovative components: i) scattering self-attention mechanism incorporating wavelet key/value and standard query to unify the learning of cross-domain relationships between the time and wavelet domains; and ii) stochastic scaling positional encoding scheme that relies solely on order information, emulating longer sequence positions to generalize up to ultra-long horizon case. Extensive experiments on eight real-world benchmarks show the potential of our Scattering Transformer as a robust and versatile solution, showcasing its quadruple efficacy of non-stationary forecasting, ultra-long horizons forecasting, representation learning, and reduction in time and space complexity.
Ziang Yang, Biyu Zhou, Xuehai Tang, Ruixuan Li 0001, Songlin Hu 0001
ECAI2
2024 Quartet: A Holistic Hybrid Parallel Framework for Training Large Language Models
Weigang Zhang, Biyu Zhou, Xing Wu 0002, Chaochen Gao, Xuehai Tang, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001
Euro-Par (2)2
2024 ProFetch: Accelerate Deep Recommendation System Training with Proactively Designed Data Layout and Dynamic Prefetching
Biyu Zhou, Weigang Zhang, Xuehai Tang, Ruixuan Li 0001, Songlin Hu 0001
ICONIP (5)2
2024 Parquet-Based CTR Model Training in Production Environment
abstract
CTR(click through rate) model has played an important role in modern recommendation systems. Most of the recommendation models in industrial scenario are trained by TensorFlow. However, we observed that, TFRecord, the native k-v data format in TensorFlow, is not the best choice for CTR training. Those keys take up to 54% of the storage space in TFRecord formatted training data. To overcome this, we introduce Apache Parquet, a column-oriented data format, into CTR tasks to improve spatial efficiency. Besides, to use in production environment, we further give some high performance implementations of Parquet training scheme. Firstly, GPU data preprocessing method is adopted in replace of original Spark based solution to generate Parquet training data and accelerate data preprocessing. Secondly, we modify data loader in TensorFlow to consume the Parquet training data with high efficiency. Experimental results show that, the size of preprocessed Criteo dataset is 95.29% smaller in comparison to TFRecord and the data preprocessing time also reduces 99.6%. Without any model performance damage, we speed up the training process by 1.45x. Our scheme has applied to our internal business and has obtained similar performance benefits.
Jinrong Guo, Biyu Zhou, Xiaokun Zhu, Yongjun Bao, Jizhong Han, Songlin Hu 0001
SMC3
2023 MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale Models
abstract
The rapid development of large-scale deep neural networks has put forward an urgent demand for the efficiency of parallel training. Recently, bidirectional pipeline parallelism has been recognized as an effective approach for improving training throughput. This paper proposes MixPipe, a novel bidirectional pipeline parallelism for efficiently training large-scale models in synchronous scenarios. Compared with previous proposals, MixPipe achieves a better balance between pipeline utilization and device utilization, which benefits from the flexible regulating for the number of micro-batches injected into the bidirectional pipelines at the beginning. MixPipe also features a mixed schedule to balance memory usage and further reduce the bubble ratio. Evaluation results show that: for Transformer based language models (i.e., Bert and GPT-2 models), MixPipe improves the training throughput by up to 2.39× over the state-of-the-art synchronous pipeline approaches.
Weigang Zhang, Biyu Zhou, Xuehai Tang, Zhaoxing Wang, Songlin Hu 0001
DAC2
2023 Orthrus: A Dual-Branch Model for Time Series Forecasting with Multiple Exogenous Series
Ziang Yang, Biyu Zhou, Xuehai Tang, Ruixuan Li 0001, Songlin Hu 0001
DASFAA (1)2
2023 A Multi-source Domain Adaption Approach to Minority Disk Failure Prediction
Wang Wang, Xuehai Tang, Biyu Zhou, Yangchen Dong, Yuanhang Feng, Jizhong Han, Songlin Hu 0001
ICA3PP (2)3
2022 Improving disk failure detection accuracy via data augmentation
abstract
Frequently happening of disk failures seriously affects the dependability and service quality of cloud data centers. Recently, machine learning (ML) based methods are popularly adopted to proactively predict forthcoming disk failures via supervised learning. However, the high imbalance of failure samples and healthy samples is a huge obstacle for existing detection methods to establish high performance detection model. This paper presents a data augmentation method MSGMD, which can efficiently generate high quality failure samples to alleviate the data imbalance of the training set, so as to effectively improve the performance of any supervised failure detection models. First, MSGMD converts failure samples (multivariate time series) into multiple univariate time series via decomposing the spatial relations among features. Then it learns the temporal correlation of each feature via a policy-based reinforcement learning model trained in an adversarial way. After that, it generates failure samples by combining feature series sampled from learned distribution. Finally, it filters out low quality generated samples with a confidence-based method. Experimental results on real-world datasets show that, through data augmentation, MSGMD can improve the FDR and F1-Score of the state-of-the-art disk failure detection model by 31.59% and 30.74% respectively on average.
Wang Wang, Xuehai Tang, Biyu Zhou, Wenjie Xiao, Jizhong Han, Songlin Hu 0001
IWQoS3
2021 Fed-Tra: Improving Accuracy of Deep Learning Model on Non-iid in Federated Learning
Wenjie Xiao, Xuehai Tang, Biyu Zhou, Wang Wang, Yangchen Dong, Liangjun Zang, Jizhong Han, Songlin Hu 0001
ICA3PP (1)3
2021 Workload Prediction and VM Clustering Based Server Energy Optimization in Enterprise Cloud Data Center
Wantao Liu, Biyu Zhou, Congfeng Jiang, Ruixuan Li 0001, Songlin Hu 0001
ICA3PP (3)3
2020 PStream: Priority-Based Stream Scheduling for Heterogeneous Paths in Multipath-QUIC
abstract
Web latency remains the main obstacle to improving user experience with the continuous development of the web. A lot of works have been made in this course. Quick UDP Internet Connection (QUIC) embeds stream multiplexing to solve the head-of-line blocking caused by the in-order requirement of TCP. Multipath-QUIC (MPQUIC) brings further improvements by utilizing multiple paths, as is done in MultiPath TCP (MPTCP). Different from MPTCP schedulers, MPQUIC schedulers are stream-aware and thus can provide finer granularity of multipath scheduling. As streams are with different features based on their contents, the resource preferences of a stream are highly related to its feature. We find that scheduling without the recognition of the stream features can aggravate inter-stream blocking when sharing paths. We fill this gap and propose PStream - a priority-based online stream scheduling mechanism for MPQUIC, which performs path scheduling based on the stream features. We examine the effectiveness of PStream under different path heterogeneity comparing to the original and the latest scheduler of MPQUIC. Our evaluation shows that our scheduler can reduce up to 25.4% of page load time in high path heterogeneity.
Lin Wang 0015, Fa Zhang 0001, Biyu Zhou, Zhiyong Liu 0002
ICCCN4
2019 Bridging Server and Cooling: Toward Effective Energy Management in Data Centers
abstract
This paper studies the energy management problem in data centers. As the most energy-consuming sub-systems, server and cooling have been considered as the key of energy saving for data center. However, due to the fact that servers are not directly connected to cooling devices, it is not trivial to optimize them jointly. In this paper, we formulate the joint server and cooling energy management problem and analyze its challenges, and design a Learning-based Energy management Scheme (LEES) to solve the joint energy management problem. Firstly, machine learning methods are investigated in the cooling power modeling and cabinet temperature modeling. Based on the models, a simple yet efficient energy management scheme is proposed to find an optimized configuration of server load distribution, server status and air supply temperature setting for the data center to reduce the total energy consumption. Extensive simulation results demonstrate that our method is able to reduce the total energy consumption of data centers significantly.
Biyu Zhou, Songlin Hu 0001
ICPADS1
2017 Joint Optimization of Server and Network Resource Utilization in Cloud Data Centers
abstract
Virtual machine placement is a key component of cloud resource management, which may affect network bandwidth allocation. In this paper, we revisit the virtual machine placement problem in cloud data centers and aim to maximize the overall resource utilization in multiple dimensions, while ensuring that the resource constraints on both the server such as CPU capacity and the network such as bandwidth are not violated. We model the bandwidth-guaranteed virtual machine placement problem and prove its NP-hardness, and design offline and online algorithms to solve the problem. We first consider the offline version and develop approximation algorithms with bounded performance ratios for both the homogeneous and the heterogeneous cases. Then, for the online version, we propose simple and efficient heuristics based on the insights from the offline algorithm design. Comprehensive experimental results verify that the overall resource utilization can be significantly improved by applying our proposals.
Biyu Zhou, Jie Wu 0001, Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002
GLOBECOM1
2017 Online Flow Scheduling with Deadline for Energy Conservation in Data Center Networks
abstract
We study the problem of flow scheduling in data center networks. Using speed scaling, our aim is to find an online scheduling algorithm that minimizes the total energy consumption of the network by determining both the transmission order and rates of the arriving flows while providing a strict flow deadline guarantee. Observing the superlinear property of link power consumption, the key challenge is in constantly determining the minimum transmission rate for “delay-tolerable” flows without any priori knowledge. To leverage the flow arrival pattern, we propose a probability-based flow prediction model to capture the uncertainty of the network flows. Based on the prediction model, we propose a tunable online flow scheduling algorithm to solve the online flow scheduling problem effectively. By introducing a scaling factor on bandwidth allocation, this algorithm allows us to conduct arbitrary trade-offs between the conservative and aggressive behaviors in terms of energy conser- vation. The effectiveness of the proposed algorithm is validated through rigorous theoretical analysis and further confirmed by extensive numerical simulations.
Biyu Zhou, Jie Wu 0001, Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002
ICPADS1
2017 Resource optimization for survivable embedding of virtual clusters in cloud data centers
abstract
With the popularity of cloud computing, optimizing cloud resource consumption while providing predictable cloud service has become one of the focuses of research in recent years. In order to ensure a predictable performance, the requests from tenants are abstracted as Virtual Clusters, which not only specify the computing demands, but also establish the communication requirements among virtual machines. While much work has been done on virtual cluster embedding under a variety of goals, very few people have studied this issue in consideration of service survivability, which also plays a vital role in ensuring the performance in cloud data centers. In this paper, we study the resource optimization for survivable embedding of virtual clusters and aim to minimize the consumption of cloud resources in terms of server and bandwidth, while ensuring that both the resource constraints and the survivability constraints are not violated. We formally define this problem and analyze its complexity, and design efficient algorithms to solve the problem. Comprehensive experimental results verify that the overall resource consumption can be significantly reduced by applying our proposals.
Biyu Zhou, Jie Wu 0001, Fa Zhang 0001, Zhiyong Liu 0002
IPCCC1
2017 Joint Optimization of Operational Cost and Performance Interference in Cloud Data Centers
abstract
Virtual machine (VM) scheduling is an important technique for the efficient operation of the computing resources in a data center. Previous work has mainly focused on consolidating VMs to improve resource utilization and to optimize energy consumption. However, the interference between collocated VMs is usually ignored, which can result in much worse performance degradation of the applications running on the VMs due to the contention of the shared resources. Based on this observation, we aim at designing efficient VM assignment and scheduling strategies in which we consider optimizing both the operational cost of the data center and the performance degradation of the running applications. We then propose a general model that captures the tradeoff between the two contradictory objectives. We present offline and online solutions for this problem by exploiting the spatial and temporal information of performance interference of VM collocation, where VM scheduling is performed by jointly considering the combinations and the life-cycle overlap of the VMs. Evaluation results show that the proposed methods can generate efficient schedules for VMs, achieving low operational cost while significantly reducing the performance degradation of applications in cloud data centers.
Xibo Jin, Fa Zhang 0001, Lin Wang 0015, Songlin Hu 0001, Biyu Zhou, Zhiyong Liu 0002
IEEE Trans. Cloud Comput.5
2016 HDEER: A Distributed Routing Scheme for Energy-Efficient Networking
abstract
The proliferation of new online Internet services has substantially increased the energy consumption in wired networks, which has become a critical issue for Internet service providers. In this paper, we target the network-wide energy-saving problem by leveraging speed scaling as the energy-saving strategy. We propose a distributed routing scheme-HDEER-to improve network energy efficiency in a distributed manner without significantly compromising traffic delay. HDEER is a two-stage routing scheme where a simple distributed multipath finding algorithm is firstly performed to guarantee loop-free routing, and then a distributed routing algorithm is executed for energy-efficient routing in each node among the multiple loop-free paths. We conduct extensive experiments on the NS3 simulator and simulations with real network topologies in different scales under different traffic scenarios. Experiment results show that HDEER can reduce network energy consumption with a fair tradeoff between network energy consumption and traffic delay.
Biyu Zhou, Fa Zhang 0001, Lin Wang 0015, Chenying Hou, Antonio Fernández 0001, Athanasios V. Vasilakos, Youshi Wang, Jie Wu 0001, Zhiyong Liu 0002
IEEE J. Sel. Areas Commun.1
2014 DEER: A distributed routing scheme for achieving network energy efficiency
abstract
The rapid growth of Internet services has brought emergent concerns over network energy efficiency. This study aims to improve network energy efficiency using power-down technique. We propose DEER, a fully distributed routing scheme. The main concept of DEER is to dynamically allocate the traffic demands in the nodes so that some links connected to the nodes can be put into sleep mode, thus reducing the energy consumption. The special features of DEER include that it does not need global traffic matrix of the network and that it uses only the local information of link loads, making DEER be able to be implemented in a distributed manner without centralized control. We develop algorithms in DEER to dynamically change the link state (into active or sleep mode) according to link utilization and to balance the loads of links by adjusting the link weights. With the traffic load varying over time, the link state transformation is triggered when any of the pre-defined thresholds is violated. Extensive simulations with the network topology and real traffic traces from the GÉANT network confirm that by involving DEER, up to 50% of the links can be put into sleep while the frequency of chaining the state of a link stays fairly low.
Biyu Zhou, Lin Wang 0015, Fa Zhang 0001, Xibo Jin, Zhiyong Liu 0002
LANMAN1