Qi Huang 0001

dblp:46/4397-1 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-8637-0269ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
abstract
We present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware.
Chao Jin 0007, Ziheng Jiang, Zhihao Bai, Juncai Liu, Xiang Li 0067, Ningxin Zheng, Qi Huang 0001, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng 0001, Xuegui Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086
EuroSys10
2026 Complementary Online Learning Network for Probabilistic Load Forecasting Against Extreme Weather
abstract
Extreme weather events, such as heatwaves, cold snaps, and storms, frequently cause sudden and unpredictable shifts in electricity consumption, significantly complicating accurate load forecasting. Existing forecasting methods, predominantly offline-trained deep learning models, struggle to rapidly adapt to these abrupt changes due to limitations in real-time processing and the issue of catastrophic forgetting, and they rarely capture the uncertainties inherent in load predictions under extreme weather conditions. To overcome these challenges, this study proposes a novel complementary online learning network (COLNet) explicitly designed for probabilistic load forecasting during extreme weather events. The key innovations of COLNet include the following: first, a fast adaptation mechanism to rapidly assimilate new load patterns; second, an associative memory module to preserve historical load information and mitigate catastrophic forgetting; finally, a weather-aware gating mechanism that dynamically incorporates real-time meteorological variables, enhancing the model's sensitivity and forecasting robustness. Extensive comparative evaluations using real-world hourly datasets from the 2022 Australian floods, covering three affected regions over about 14 months, and from the 2021 Texas cold snap, comprising statewide load over about 24 months with a test set covering the mid-February event, are conducted. These evaluations confirm that COLNet substantially outperforms state-of-the-art methods, achieving mean absolute percentage error of 3.205%, 2.457%, and 4.882% during the flood period and 0.84% on the Texas event, corresponding to an average reduction of about 21% in point error relative to the strongest baselines and about 19% in probabilistic error, thereby improving both accuracy and uncertainty quantifications.
Pengfei Zhao 0003, Weihao Hu, Qi Huang 0001, Zhe Chen 0007
IEEE Trans. Ind. Informatics5
2026 Topology Change Aware Distributed State Estimation Based on Unsupervised Bipartite Graph-Enabled Causality-Inspired Sparse Learning
abstract
Topology changes in a distribution network are common due to planned reconfigurations and unintentional switching events during practical operations. Topology changes make it challenging for existing optimization- and learning-based distributed system state estimation methods to maintain accuracy. This difficulty arises from the lack of accurate structural information for the new topology and the absence of labeled data (recorded state variables) for model retraining. To this end, this article proposes an unsupervised-on-target learning-based state estimation method for the distribution network after topology changes without relying on the topology information and labeled data. In particular, a bipartite graph learning (BGL) method with rank constraints is first designed to learn the representation of each topology with a restricted set of measurements. Then, the Euclidean distance is employed to select the best-matched source domain historical topology according to the representation learned by the BGL. To extract invariant causal structures across the two topologies, a causality-inspired sparse structure learning for domain adaptation network is further designed. It relaxes the correlations between the selected historical and new topologies into an associative structure, represented by attention scores derived from the proposed inter- and intravariable attention networks. This allows the leverage of the causality to enhance the state estimation performance of the distribution network after topology changes without relying on accurate topology information and recorded labels used for training. The comparison results on two standard IEEE test systems validate the efficacy of the proposed method.
Zhiping Lin 0003, Weihao Hu, Pengfei Zhao 0003, Sayed Abulanwar, Qi Huang 0001, Zhe Chen 0007
IEEE Trans. Ind. Informatics6
2025 ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUs
abstract
Scaling long-context ability is essential for Large Language Models (LLMs). To amortize the memory consumption across multiple devices in long-context training, inter-data partitioning (a.k.a. Data Parallelism) and intra-data partitioning (a.k.a. Context Parallelism) are commonly used. Current training frameworks predominantly treat the two techniques as orthogonal, and establish static communication groups to organize the devices as a static mesh (e.g., a 2D mesh). However, the sequences for LLM training typically vary in lengths, no matter for texts, multi-modalities or reinforcement learning. The mismatch between data heterogeneity and static mesh causes redundant communication and imbalanced computation, degrading the training efficiency.
Junda Feng, Qi Huang 0001, Fangcheng Fu, Xiaonan Nie, Lei Zuo 0004, Haibin Lin, Bin Cui 0001, Xin Liu 0086
SIGCOMM3
2025 A Novel Spatiotemporal Pyramidal Graph Modeling Approach for Short-Term Residential Load Forecasting
abstract
Precise short-term residential load forecasting (STRLF) is essential for maintaining stable and cost-effective operations on the demand side. Both spatial and temporal information are important for the STRLF tasks, but effectively extracting them remains a significant challenge due to the highly volatile and stochastic nature of residential consumption patterns. To this end, this article proposes pyramidal attention (PATNet), a low-complexity transformer network with spatiotemporal PATNet, to explore multiresolution spatiotemporal representations of residential load series and forecast multiple residential loads several steps ahead. Specifically, the temporal and spatial patterns of residential load series are, first, formulated as a temporal pyramidal graph and a spatial pyramidal graph according to the periodic characteristics of load time series and the spatial correlations of different residential units, respectively. Two types of low-complexity attention mechanisms—temporal and spatial PATNet—are, then, specifically designed for the temporal and spatial pyramidal graphs such that the short- and long-range temporal dependencies and dynamic spatial correlations among various groups of residents can be captured. Moreover, to enhance multistep forecast performance, we design a gated fusion unit that is capable of adaptively fusing extracted spatiotemporal information and a transform attention block that can translate historical loads into future forecasts. Numerical simulations using several real-world residential load datasets demonstrate that the proposed framework outperforms state-of-the-art load prediction methods by 7.98% at least in single-step forecasting and 11.38% at least in multistep forecasting.
Pengfei Zhao 0003, Weihao Hu, Xingtao Bai, Qi Huang 0001, Zhe Chen 0007
IEEE Trans. Ind. Informatics6
2024 MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 0001, Yangrui Chen, Zhi Zhang 0005, Yanghua Peng, Xiang Li 0067, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Zhang Zhang 0003, Pengfei Nie, Leqi Zou, Sida Zhao, Zherui Liu, Xiaoying Jia 0001, Jianxi Ye, Xin Jin 0008, Xin Liu 0086
NSDI4
2023 Identification and classification for multiple cyber attacks in power grids based on the deep capsule CNN
Guangdou Zhang, Jian Li 0056, Olusola Bamisile, Yankai Xing, Qi Huang 0001
Eng. Appl. Artif. Intell.6
2022 A Multiagent Deep Reinforcement Learning Based Approach for the Optimization of Transformer Life Using Coordinated Electric Vehicles
abstract
The uncertainties of charging behavior of electric vehicle (EV) owners have a negative impact on the loss of life (LOL) of distribution transformer. This article proposes a decentralized EV charging framework for optimization of the LOL of distribution transformer considering the dissatisfactions of EV owners. Specifically, long-short-term memory (LSTM) neural network is first utilized to capture the uncertainties caused by the load demand and electricity price. After that, each EV is modeled as an intelligent agent and a multiagent deep reinforcement learning approach is applied to solve the coordinated charging problem based on the forecasting information by the LSTM network. All the agents are trained in a centralized manner to develop coordinated control strategies while informing decisions based on local information when finishing the training process. The proposed approach can achieve coordinated charging management of EVs based on local information, which helps preserve the privacy of EV owners, reduce the cost induced by the deployment of communication devices, and avoid single-point failure. In addition, the parameter space noise and deep dense architecture in reinforcement learning are introduced to overcome premature convergence, training instability, and inefficiency due to the large action space of multiagent scenario. Comparative tests are carried out among several benchmarks utilizing real-world data to illustrate the effectiveness of the proposed approach.
Sichen Li, Weihao Hu, Zhenyuan Zhang 0004, Qi Huang 0001, Zhe Chen 0007, Frede Blaabjerg
IEEE Trans. Ind. Informatics5
2022 A Dynamic Bayesian Network Control Strategy for Modeling Grid-Connected Inverter Stability
abstract
The dynamic performance of a grid-connected inverter in a distributed generation system brings new challenges by affecting the power quality and dynamic stability. The traditional linear control method for the inverter requires complex processes including decoupling and control parameter tuning and relies on a pulsewidth modulation module. The nonlinear control method, e.g., the model predictive control (MPC), can address some of the aforementioned challenges but is incapable of dealing with system parameter changes. A data-driven method using the dynamic Bayesian network-based model predictive control (DBN-MPC) is proposed, which takes advantage of the predictive ability of the DBNs to implement the MPC strategy. A predictive model is built, and the prediction signals are generated based on the DBNs. Then, by constructing the cost function and optimization criteria, the controller generates the optimal switching state combination. Using the proposed DBN-MPC method, the predictive model implements parameter learning online, providing more accurate prediction signals for further optimization and the feedback correction. In addition, the control law of the proposed controller can update over time, which enables the grid-connected inverter system to achieve the optimal control. The case studies using the grid-connected inverter system and IEEE 39-bus benchmark power system integrated with the battery energy storage system demonstrate and verify the superiority of the proposed DBN-MPC method.
Qi Huang 0001, Zhaojun Li 0001
IEEE Trans. Reliab.2
2021 A Novel Hybrid Short-Term Load Forecasting Method of Smart Grid Using MLR and LSTM Neural Network
abstract
The short-term load forecasting is crucial in the power system operation and control. However, due to its nonstationary and complicated random features, an accurate forecast of the load behavior is challenging. An improved short-term load forecasting method is proposed in this article. At first, the load is decomposed into different frequency components varying from the low to high levels realized by the ensemble empirical-mode decomposition algorithm. Then, the smooth and periodic low-frequency components are predicted by the multivariable linear regression method while maintaining the efficient computation capacity, while the high-frequency components with strong randomness are forecasted by the long short-term memory neural network algorithms. Thus, the actual load behavior is obtained by combining these two methods. Finally, the proposed method is validated by experiments, in which the tested data from the west area of China, Uzbekistan, and PJM Interconnection (USA) are used. The prediction of the load behavior is accurate globally along with the local details, as presented in the experiments, which verify the effectiveness of the proposed method.
Jian Li 0056, Daiyu Deng, Junbo Zhao 0001, Dongsheng Cai, Weihao Hu, Qi Huang 0001
IEEE Trans. Ind. Informatics7
2021 A Novel Belief Function Based Framework for UOPF With Multiprobability-Characterized and Knowledge Deficient Power Sources
abstract
A probabilistic model for the predicted power sources is typically assumed for the existing uncertain optimal power flow (UOPF). However, obtaining accurate information of that is very challenging in practice due to the limited available data of renewable energy and loads with complex correlations. To address that, this article proposes a belief function based framework for UOPF with a large number of uncertain power sources. Multiple imperfect models with q-least committed joint basic belief density are developed to characterize the knowledge deficient power sources. This yields the integration of the generalized Bayesian theorem with the traditional evidence theory into a unified manner. The former allows estimating the uncertain model of the predicted power sources, whereas the latter is to obtain the probability box of the UOPF variables. Comparison results with the Monte Carlo simulations and the three-point estimation approach show that the proposed method is able to get accurate UOPF results while achieving high computational efficiency for large-scale systems with a large number of knowledge deficient power sources.
Bi Liu, Qi Huang 0001, Junbo Zhao 0001, Weihao Hu
IEEE Trans. Ind. Informatics2
2021 Design for Reliability Through Text Mining and Optimal Product Verification and Validation Planning
abstract
Failure mode and effects analysis (FMEA) has been widely used in product design process as a reliability analysis technique. Design FMEA (DFMEA), which is used in the product development phase to identify and mitigate product risk, is one of the three application scenarios of FMEA. During the DFMEA process, verification and validation (V&V) activities are proposed to mitigate the risk of the identified potential failure modes. The V&V activities can be further planned and implemented to improve the product reliability under development. However, the DFMEA report usually contains rich text descriptions of potential failure modes and causes, and it is difficult and nonintuitive to fully understand these information for design improvements. In addition, it is also very challenging to optimize the planning of V&V activities by selecting a set of V&V activities to achieve expected reliability improvement effectiveness with the minimum resource consumption requirement. To address these two challenges, this article first proposes a method of applying text mining to the DFMEA report to obtain two types of hidden reliability information, including the classifications of failure modes/causes and the correlation between keywords. Then, a mathematical model is proposed to optimize the product V&V planning by selecting an optimal set of V&V activities. The application of the above proposed methods is illustrated through the product development of a diesel engine power generation system.
Qi Huang 0001, Gongyu Wu, Zhaojun Li 0001
IEEE Trans. Reliab.1
2013 Stability analysis of trilateral haptic collaboration
abstract
This paper presents a criterion for absolute stability of a general class of three-port networks. Trilateral haptic systems, which have recently found many interesting applications, can be modeled as three-port networks. Traditionally, existing criteria (Llewellyn‘s criterion) have facilitated the stability analysis of bilateral haptic systems modeled as two-port networks. If the same criteria were to be used for stability analysis of a three-port network, its third port would need to be assumed known for it to reduce to a two-port network. However, this is restrictive because, according to the definition of absolute stability, all three terminations of the three-port network must be allowed to be arbitrary (while passive). In this paper, extending Llewellyn's criterion, we present closed-form necessary and sufficient conditions for absolute stability of a general class of three-port networks — the three terminations need to be passive but are otherwise arbitrary. To this end, we first find a symmetrization condition under which a general asymmetric impedance (or admittance) matrix Z3×3has an equivalent symmetric counterpart Zeq; this Zeq models a reciprocal three-port network with the same stability characterization as the general nonreciprocal three-port network modeled by Z. Then, based on the equivalence of passivity and absolute stability for the equivalent reciprocal network, an absolute stability condition for the original nonreciprocal network is derived. To show how the resulting absolute stability criterion can be utilized at the system design stage, we have applied it to the problem of designing controllers for triple-user collaborative haptic virtual environment systems. The validity of the resulting absolute stability conditions have been verified via simulations.
Jian Li 0056, Mahdi Tavakoli, Qi Huang 0001
World Haptics3
2013 Conservatism of passivity criteria for stability analysis of trilateral haptic systems
abstract
Trilateral haptic systems can be modeled as three-port networks. Analysis of coupled stability of a three-port network can be accomplished in either the passivity or the absolute stability frameworks assuming all three ports are connected to passive but otherwise unknown terminations. This paper first reviews our recent results in terms of passivity and absolute stability criteria for general three-port networks - both criteria are founded on the properties of a positive-real Hermitian matrix. Next, we show that the absolute stability criterion is less conservative than the passivity criterion and that the two criteria become the same when the trilateral system is represented by a reciprocal immitance matrix. Then, to show how the two criteria may be utilized at the system design stage, we apply them to the problem of designing controllers for a dual-user haptic teleoperation system. Using the two criteria, controllers are then designed and compared in terms of conservatism and performance in simulations.
Jian Li 0056, Mahdi Tavakoli, Victor Mendez, Qi Huang 0001
World Haptics4
2010 Nonlinear adaptive bilateral control of teleoperation systems with uncertain dynamics and kinematics
abstract
Research so far on adaptive bilateral control of master-slave teleoperation systems considers dynamic uncertainties but stops short of considering kinematic uncertainties. However, when picking up objects of unknown lengths, orientations and gripping points, the overall kinematics of a robot in the teleoperation system becomes uncertain. Therefore, new controllers are required that can guarantee the stability and motion tracking performance of the system in the presence of both dynamic and kinematic uncertainties in the master and the slave robots. In this paper, first the uncertain dynamics of the human operator and the environment are incorporated into the dynamics of the master and the slave, respectively. Then, for a teleoperation system with uncertain dynamics and kinematics, nonlinear adaptive controllers are designed for both the master and the slave. The controllers do not need exact knowledge of the dynamics of the master, the slave, the operator, or the environment, or of the kinematics of the master or the slave. The stability and position tracking convergence of the entire teleoperation system are studied. The validity of the theoretical results is verified by simulations.
Mahdi Tavakoli, Qi Huang 0001
IROS3