EDBT 2026 Demo / reviewers in the wild / expert
Xiaowei Jiang
dblp:31/2484
· DBLP profile ↗
67ranked-venue papers
29as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 17 · 11 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 9 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Finite-time hybrid cooperative control of nonlinear time-delay multiagent systems and its applications in Chua's systems
Xiaowei Jiang, Ranran Jiao, Bo Li 0124, Yan-Wu Wang |
Sci. China Inf. Sci. | 1 |
| 2026 | Asynchronous Saturation-Constrained Impulsive Consensus of Nonlinear Multi-Agent Systems and Its Applications in Chua's Circuit
Xiaowei Jiang, Feixue Chen, Xiaofan Ma, Qiang Lai |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Output Tracking Control Performance of Communication Networked Systems Over Uncertain Delay Channel and Its ApplicationsabstractDue to the various constraints in the communication process, such as bandwidth, quantization, time delay, etc., it will inevitably have some impacts on the performance of communication networked systems. More seriously, it will cause the system to lose its stability. This study investigates the output tracking performance (OTP) of communication networked systems under the combined effects of uncertain time delays, packet losses, and channel noise. By using the methods of first order Pade approximation, all-pass decomposition and Youla parameterization, the explicit expressions of OTP are derived by the design of the optimal controller, which mainly contain single-degree-of-freedom (SDOF) and two-degree-of-freedom(TDOF), respectively. The results show that the OTP of NCSs has strong connections with the inherent properties of the plant and the communication parameters of channel. Finally, the correctness of the theoretical results is verified by the simulations for the inverted pendulum system. Xiaowei Jiang, Yan-Wu Wang |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Robust Complete Synchronization of Coupled Boolean Networks With Stuck-At FaultabstractThis paper investigates the robust complete synchronization of coupled Boolean networks (BNs) subject to stuck-at fault. When stuck-at fault occurs, certain nodes become permanently fixed and the original synchronization conditions may no longer be applicable. To address this issue, we propose a new concept of fault-preserving subset to characterize the admissible invariant state evolution induced by stuck-at fault. Based on this concept, necessary and sufficient conditions are derived to determine whether complete synchronization remains valid without reconstructing the faulty network model. Furthermore, for coupled Boolean control networks (BCNs), a robust feedback complete synchronization scheme is developed by designing state feedback controllers based on maximal control fault-preserving subset. Finally, several numerical examples are provided to demonstrate the effectiveness of the proposed results. Xiaofan Ma, Xiaowei Jiang, Zhi-Wei Liu 0002 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | iFuzz-Meta: An Interpretable Fuzzy Learning Framework Bridging Top-Down and Bottom-Up Knowledge IntegrationabstractInterpretable representation learning remains a key challenge in modern neural computation, particularly when models are expected not only to perform but also to explain their reasoning. This paper introduces iFuzz-Meta, an inter pretable fuzzy rule-based learning framework that preserves human-understandable reasoning structures within modern neural architectures. Each fuzzy rule corresponds to a semantic and spatial prototype defined in the original feature space, enabling transparent inference and direct interpretability. Meta-learning is employed as an analytical paradigm to examine how these interpretable rules reorganize across tasks and domains, providing a principled means to link algorithmic adaptation with cognitive representation. A knowledge-guided regularization mechanism further enables a top-down–bottom-up integration, in which theoretical priors act as soft inductive biases while data-driven learning refines and extends them. This dual process ensures that adaptation proceeds along semantically and physiologically meaningful trajectories, rather than arbitrary parameter shifts. Evaluations demonstrate that iFuzz-Meta achieves interpretable reasoning and stable cross-domain generalization, establishing a potential general pathway toward explainable and knowledge aware fuzzy systems. Xiaowei Jiang, Daniel Leong, Beining Cao, Yingtao Ren, Thomas Do, Chin-Teng Lin |
IEEE Trans. Fuzzy Syst. | 1 |
| 2026 | Best Achievable Control Performance of Networked Systems Over Bandwidth-Constrained ChannelsabstractDuring data transmission, networked control systems (NCSs) are inevitably subject to constraints including bandwidth limitations, quantization effects, and time delays factors that exert a significant impact on the systems’ stability and control performance. This research explores the output tracking performance (OTP) of NCSs under the combined effects of uncertain time delays, bandwidth restrictions, and channel noise constraints. Through the adoption of first-order Pade approximation, all-pass factorization, and Youla parameterization techniques, the OTP bounds for single-degree-of-freedom systems and the modified performance limits for two-degree-of-freedom systems are derived analytically, yielding explicit formulations of the optimal performance. The results demonstrate that the OTP of NCSs is inherently associated with the intrinsic properties of the controlled plant and the parameters of the communication channel. Numerical simulations conducted on a metal rolling system further verify the theoretical predictions. Ruyan Li, Houjun Liang, Xiaowei Jiang |
IEEE Trans. Ind. Informatics | 3 |
| 2026 | High-Precision Camera Distortion Correction: A Decoupled Approach With Rational FunctionsabstractThis paper presents a robust, decoupled approach to camera distortion correction using a rational function model (RFM), designed to address challenges in accuracy and flexibility within precision-critical applications. Camera distortion is a pervasive issue in fields such as medical imaging, robotics, and 3D reconstruction, where high fidelity and geometric accuracy are crucial. Traditional distortion correction methods rely on radial-symmetry-based models, which have limited precision under tangential distortion and require nonlinear optimization. In contrast, general models do not rely on radial symmetry geometry and are theoretically generalizable to various sources of distortion. There exists a gap between the theoretical precision advantage of the Rational Function Model (RFM) and its practical applicability in real-world scenarios. This gap arises from uncertainties regarding the model's robustness to noise, the impact of sparse sample distributions, and its generalizability out of the training sample range. In this paper, we provide a mathematical interpretation of how RFM is suitable for the distortion correction problem through sensitivity analysis. The precision and robustness of RFM are evaluated through synthetic and real-world experiments, considering distortion level, noise level, and sample distribution. Moreover, a practical and accurate decoupled distortion correction method is proposed using just a single captured image of a chessboard pattern. The correction performance is compared with the current state-of-the-art using camera calibration, and experimental results indicate that more precise distortion correction can enhance the overall accuracy of camera calibration. In summary, this decoupled RFM-based distortion correction approach provides a flexible, high-precision solution for applications requiring minimal calibration steps and reliable geometric accuracy, establishing a foundation for distortion-free imaging and simplified camera models in precision-driven computer vision tasks. Jiachuan Yu, Yuankai Zhou, Xiaowei Jiang |
IEEE Trans. Image Process. | 4 |
| 2025 | Pretraining Large Brain Language Model for Active BCI: Silent SpeechabstractThis paper explores silent speech decoding in active brain-computer interface (BCI) systems, which offer more natural and flexible communication than traditional BCI applications. We collected a new silent speech dataset of over 120 hours of electroencephalogram (EEG) recordings from 12 subjects, capturing 24 commonly used English words for language model pretraining and decoding. Following the recent success of pretraining large models with self-supervised paradigms to enhance EEG classification performance, we propose Large Brain Language Model (LBLM) pretrained to decode silent speech for active BCI. To pretrain LBLM, we propose Future Spectro-Temporal Prediction (FSTP) pretraining paradigm to learn effective representations from unlabeled EEG data. Unlike existing EEG pretraining methods that mainly follow a masked-reconstruction paradigm, our proposed FSTP method employs autoregressive modeling in temporal and frequency domains to capture both temporal and spectral dependencies from EEG signals. After pretraining, we finetune our LBLM on downstream tasks, including word-level and semantic-level classification. Extensive experiments demonstrate significant performance gains of the LBLM over fully-supervised and pretrained baseline models. For instance, in the difficult cross-session setting, our model achieves 47.2% accuracy on semantic-level classification and 42.3% in word-level classification, outperforming baseline methods substantially. Our research advances silent speech decoding in active BCI systems, offering an innovative solution for EEG language model pretraining and a new dataset for fundamental research. Jinzhao Zhou, Zehong Cao, Yiqun Duan, Connor Barkley, Daniel Leong, Xiaowei Jiang, Quoc-Toan Nguyen, Thomas Do, Sheng-Fu Liang, Chin-Teng Lin |
ACM Multimedia | 6 |
| 2025 | Finite-Time Consensus of Second-Order Multiagent Systems With Input Saturation via Hybrid Sliding-Mode ControlabstractThis paper addresses the finite-time consensus (FTC) issue for second-order multi-agent systems (MASs) with nonlinear disturbances. To tackle the challenges posed by increasingly complex communication environments, an innovative integral sliding-mode surface is designed using event-triggered control. Furthermore, the study explores communication constraints between the leader and followers, employing impulsive control to facilitate communication only at specific intervals. A novel hybrid integral sliding-mode control protocol is advanced for the second-order MASs, which effectively eliminates the “Zeno phenomenon" and notably reduces both communication frequency and energy consumption, thereby enhancing overall communication efficiency. Ultimately, the efficacy of the put forward protocol is verified through simulation examples including comparison experiment. Xiaowei Jiang, Ranran Jiao, Bo Li 0124, Huaicheng Yan 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Modified Tracking Performance of NCSs Over Erasure Channel and Its Applications in Vehicle ControlabstractThe data communication among sensors, actuators, and controllers in network systems poses various challenges, including packet loss, network-induced delay, and channel noise interference, all of which adversely affect performance of control systems. Based on single-degree-of-freedom (SDOF) controller and two-degree-of-freedom (TDOF) controller, respectively, this article discuss networked control systems (NCSs) with packet loss, network-induced delay, and logarithmic quantization constraints in the communication channel, and proposes a discrete-time modified performance index. By the frequency domain analysis methods and Youla parameterization techniques of controllers, explicit expressions of the modified performance limitations are derived, which quantitatively reveal the impact of the fundamental characteristics of the plant and communication constraints on systems' performance. In addition, a linear dual-freedom vehicle system is modeled and the obtained theorems are applied to verify the correctness of the results. The results show that communication constraints have an adverse impacts on systems' performance, and modified performance index can better measure systems' performance. Xiaowei Jiang, Chuan-Ke Zhang, Yan-Wu Wang |
IEEE Trans. Cybern. | 1 |
| 2025 | Limited Impulsive Control of Time-Delay Multiagent Systems With Packet Loss and Parameter MismatchabstractThis article investigates the leader-following consensus of nonlinear time-delay multiagent system under impulsive control with simultaneous consideration of packet loss and parameter mismatch. Specifically, the inherent parameter mismatch between the leader's dynamics and followers' dynamics is explicitly addressed. To mitigate communication frequency, two novel impulsive control protocols are developed: 1) a pure impulsive scheme for theoretical analysis and 2) a limited impulsive strategy for practical implementation. Furthermore, an auxiliary function is introduced to characterize packet loss phenomena during information transmission, ensuring alignment with real-world communication constraints. By integrating impulsive control theory with reverse average dwell-time analysis, sufficient consensus criteria are rigorously derived for multiagent system with time delays and heterogeneous parameters. Finally some numerical simulations validate the effectiveness of the proposed control framework, demonstrating its capability to achieve consensus under practical communication imperfections. Le You, Xiaowei Jiang, Chuan-Ke Zhang, Yan-Wu Wang, Huaicheng Yan 0001 |
IEEE Trans. Cybern. | 2 |
| 2025 | A Fuzzy Logic-Based Approach to Predict Human Interaction by Functional Near-Infrared SpectroscopyabstractIn this article, we introduce the Fuzzy logic-based attention (Fuzzy Attention Layer) mechanism, a novel computational approach designed to enhance the interpretability and efficacy of neural models in psychological research. The fuzzy attention layer integrated into the transformer encoder model to analyze complex psychological phenomena from neural signals captured by functional near-infrared spectroscopy (fNIRS). By leveraging fuzzy logic, the fuzzy attention layer learns and identifies interpretable patterns of neural activity. This addresses a significant challenge in using transformers: the lack of transparency in determining which specific brain activities most contribute to particular predictions. Our experimental results, obtained from fNIRS data engaged in social interactions involving handholding, reveal that the fuzzy attention layer not only learns interpretable patterns of neural activity but also enhances model performance. In addition, these patterns provide deeper insights into the neural correlates of interpersonal touch and emotional exchange. The application of our model shows promising potential in understanding the complex aspects of human social behavior, verify psychological theory with machine learning algorithms, thereby contributing significantly to the fields of social neuroscience and AI. Xiaowei Jiang, Liang Ou, Na Ao, Thomas Do, Chin-Teng Lin |
IEEE Trans. Fuzzy Syst. | 1 |
| 2025 | Fully Distributed Dynamic Event-Triggered Consensus of Heterogeneous Multiagent Systems With Applications to Voltage ControlabstractThis article investigates a heterogeneous multiagent system composed of first-order agents and second-order agents. Control strategies for achieving average consensus and bipartite consensus are proposed, along with the design of a fully distributed dynamic event-triggered control algorithm. Under the proposed event-triggered mechanism, agent updates the control input and exchanges information with neighboring agents solely when its own triggering times arrive, requiring only local information without relying on global system knowledge or the real-time states of neighboring agents. It is rigorously proven that “Zeno behavior” is excluded. Simulation results are given to demonstrate that the proposed control protocols are effective and feasible. Xiaowei Jiang, Weichao Liu, Bo Li 0124 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | iFuzzyTL: Interpretable Fuzzy Transfer Learning for Steady-State Visual Evoked Potentials Brain-Computer Interfaces System
Xiaowei Jiang, Beining Cao, Liang Ou, Thomas Do, Chin-Teng Lin |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2025 | Optimal Performance of Discrete Networked Systems With Cyber-Attack and Packet DropoutsabstractIn this study, the limitations of the tradeoff performance of multiple-input–multiple-output (MIMO) discrete networked control systems (NCSs) with forward channel subject to cyber-attack, additive white Gaussian noise and packet dropouts were analyzed. The performance of intrusion detection systems under cyber-attack with incomplete information was analyzed using game theory. Then, explicit expressions for the optimal tradeoff performance between tracking error and control input are derived based on the two-degree-of-freedom (TDOF) controller using frequency domain analysis, coprime factorization technique and Youla parameterization method. Results show that the tradeoff performance of the system is affected by their fundamental properties, such as the direction and position of the nonminimum phase (NMP) zeros and unstable poles (UPs) in the plant as well as communication constraints, such as cyber-attack, additive white Gaussian noise and packet dropouts. Finally, an illustrative simulation is discussed to verify the aforementioned conclusions. Xiaowei Jiang, Xinyu Ren, Bo Li 0124, Feng Liu 0042, Wu-Hua Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Contrastive Masked Autoencoders for Character-Level Open-Set Writer IdentificationabstractIn the realm of digital forensics and document authentication, writer identification plays a crucial role in determining the authors of documents based on handwriting styles. The primary challenge in writer-id is the “open-set scenario”, where the goal is accurately recognizing writers unseen during the model training. To overcome this challenge, representation learning is the key. This method can capture unique handwriting features, enabling it to recognize styles not previously encountered during training. Building on this concept, this paper introduces the Contrastive Masked Auto-Encoders (CMAE) for Character-level Open-Set Writer Identification. We merge Masked Auto-Encoders (MAE) with Contrastive Learning (CL) to simultaneously and respectively capture sequential information and distinguish diverse handwriting styles. Demonstrating its effectiveness, our model achieves state-of-the-art (SOTA) results on the CASIA online handwriting dataset, reaching an impressive precision rate of 89.7%. Our study advances universal writer-id with a sophisticated representation learning approach, contributing substantially to the ever-evolving landscape of digital handwriting analysis, and catering to the demands of an increasingly interconnected world. Xiaowei Jiang, Yiqun Duan, Thomas Do, Chin-Teng Lin |
SMC | 1 |
| 2024 | Enhancing Marine Navigation Performance Using the Head-Up InterfaceabstractModern marine navigation places significant physical and mental demands on officers stationed on ship bridges, primarily due to the continuous observation and evaluation of real-time navigational information displayed on scattered electronic equipment. To alleviate the high cognitive load experienced by marine officers and allow them to focus on essential tasks during complex situations, the integration of head-up displays (HUDs) in marine applications has emerged as a promising solution. HUDs offer the potential to provide crucial information, enhancing the accessibility and organization of previously disordered data. However, there is limited information on the impact of HUDs on marine officers, which has been explored in the presented work. In this work, a novel immersive navigation experiment with three conditions: traditional display (NonAR), augmented reality (AR) based information presentation, and a variant of AR with essential information only (AR-Indicator) has been conducted. The objective is to explore the effects of these three conditions on navigation performance and mental workload. Our findings indicate that the AR-based information presentation, specifically the variant that includes only essential information, is preferred by participants and showed performance improvements measured by time to complete tasks, gaze duration, and pupil dilation compared to the traditional display and full AR condition. These results have an impact on the design and development of HUDs in marine-related tasks. This pilot research sheds light on the potential benefits of HUDs in improving maritime navigation and paves the way for further advancements in this field. Jinzhao Zhou, Chin-Teng Lin, Sara Lal, Ami Eidels, Xiaowei Jiang, Scott D. Brown |
SMC | 6 |
| 2024 | Distributed Control for Nonlinear Time-Delay Multiagent Systems: Hybrid Saturation-Constraint Impulsive ApproachabstractThis article mainly studies the problem of impulse consensus of multiagent systems under communication constraints and time delay. Considering the limited communication bandwidth of the agent, global and partial saturation constraints are considered. In addition, so as to further improve communication efficiency by reducing communication frequency, the novel control protocol combining event-triggered strategy and general impulse control protocol is proposed. Under this kind of novel control protocol, the communication frequency of multiagent systems can be reduced while avoiding "Zeno behavior." Through theoretical analysis, sufficient conditions for the systems to achieve consensus are obtained for the above two saturation constraint cases. In the end, the effectiveness of the novel protocols is proved by providing two different simulation instances. Xiaowei Jiang, Le You, Bo Li 0124, Huaicheng Yan 0001 |
IEEE Trans. Cybern. | 1 |
| 2024 | Tracking Performance of Feedback Systems Over a Fading Channel With Limited BandwidthabstractIn this study, we investigated the optimal tracking performance (OTP) of feedback control systems with limited bandwidth and colored noise in a fading channel. For the steady state of the feedback control systems, an equivalent average channel (EAC) model was developed by retaining the effects of the first and second moments of the multiplicative channel output, and on the basis of the coprime decomposition, all-pass factorization, and Youla parameterization of controllers, exact expressions for the OTP were derived by designing two compensators. The expressions quantitatively show the relationship between the OTP and inherent features of the plant. Specifically, the directions and locations of unstable poles (UPs) and nonminimum phase (NMP) zeros adversely affect the tracking performance. Furthermore, the bandwidth limitation and the presence of colored noise also degrade the tracking performance. Finally, our conclusions were verified by considering a numerical arithmetic example. Xiaowei Jiang, Bin Zhang 0040, Choon Ki Ahn, Shiqi Zheng, Huaicheng Yan 0001 |
IEEE Trans. Cybern. | 1 |
| 2024 | Impulsive Formation Tracking of Nonlinear Fuzzy Multiagent Systems With Input Saturation ConstraintsabstractThis article investigates leader-following formation problems for second-order fuzzy multiagent systems (MASs) with input saturation constraints, by using an impulsive control strategy. Traditional communication methods generate large amounts of unnecessary information, resulting in a waste of resources. Thus, fuzzy logic systems are established to estimate unknown nonlinear functions. Subsequently, impulsive formation control strategies are proposed to reduce the costs of continuous communication, in which followers only communicate with the leader at a fixed impulsive time. In addition, considering the limited nature of the actual physical actuator, input saturation constraints are introduced to control the formation of MASs. The conclusions are extended to the case in which the interaction topology switches with time. Finally, two simulation examples are provided to validate the derived results and the feasibility of the proposed control protocol. Xiaowei Jiang, Le You, Huaicheng Yan 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | Communication Limited Hybrid Impulsive Control of Fuzzy Time-Delay Multiagent NetworkabstractIn this article, the consensus problem of time-delay fuzzy multiagent systems (TDFMASs) with saturation-constraint impulsive control is considered. To better meet the complex requirements of real-world models, a nonlinear model of TDFMASs and the corresponding fuzzy rules are designed. Then, in view of the limited communication channel, all agents communication channel constraint and partial agents communication channel constraint are all considered in impulsive control, respectively. Furthermore, for the sake of increasing the antiinterference of the agents, a hybrid impulse control protocol is designed. In addition, with the help of Lyapunov stability theory, a suitable Lyapunov auxiliary function is constructed to obtain sufficient conditions for subsystem stability and solve the problem. Finally, numerical simulations are conducted to verify the feasibility of the theoretical results. Le You, Xiaowei Jiang, Shiqi Zheng, Huaicheng Yan 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Impulsive Layered Control of Heterogeneous Multi-Agent Systems Under Limited CommunicationabstractIn this article, the consensus problem of heterogeneous multi-agent systems (HMASs) with communication limited impulsive control (IC) is considered. For better meet the complex requirements of real-world models, a HMASs model with different dimensions of leader and followers is designed. Due to the different dimensions of the agents in the MASs, the followers and the leader cannot communicate directly, so a virtual layer as an intermediary is designed to complete the communication between the followers and the leader. Then, in view of the limited communication channel, all agents communication channel constraint have been considered in IC. Furthermore, for the sake of comparing the influence of the proportion of information constraints on the speed of the system to achieve consensus, the global information saturation constraints and local information saturation constraints are considered, respectively. Finally, numerical simulations are conducted to verify the feasibility of the theoretical results. Le You, Xiaowei Jiang, Bo Li 0124, Huaicheng Yan 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Optimal Tracking Performance of Networked Control Systems Under Communication Channel NoiseabstractIn networked control systems (NCSs), the network-induced delay, packet dropouts, noise and other constraints in the communication network will affect the optimal tracking performance (OTP) and even the system’s stability, which is also a problem that the NCSs approach must solve in practical applications. In this study, we primarily examine the OTP of NCSs that consider packet dropouts and nonzero mean additive white noise (AWN) constraints in communication networks. Based on the single-degree-of-freedom (SDOF) controller and the two-degree-of-freedom (TDOF) controller, respectively, using the coprime factorization and Youla parameterization approach, the explicit expressions of the OTP limitation of the NCSs under the constraints of nonzero mean noise and packet dropouts are obtained. The results reveal that the intrinsic features of the plant and the communication parameters of the network channel will affect the OTP of the NCSs. Finally, the correctness of the theoretical results is verified by the simulation of a multi-input and multioutput plant and an inverted pendulum system. Xiaowei Jiang, Bo Li 0124, Xiangyong Chen, Huaicheng Yan 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Tracking performance limitations of MIMO discrete-time networked control systems with multiple constraints
Xiaowei Jiang, Bin Zhang 0040, Xiangyong Chen, Huaicheng Yan 0001 |
Sci. China Inf. Sci. | 1 |
| 2023 | Control for Nonlinear Fuzzy Time-Delay Multiagent Systems: Two Kinds of Distributed Saturation-Constraint Impulsive ApproachabstractThe impulsive consensus problems of nonlinear time-delay multiagent system under communication constraints are analyzed and investigated in this article. In view of the limited communication channel, global input saturation constraint is integrated into the design of impulse control protocol. Furthermore, in order to reduce energy consumption and communication cost, two kinds of impulsive control protocols are designed that conclude general impulsive control protocols and event-triggered impulsive control protocols. For event-triggered impulsive control protocols, by setting event decision time at impulse time, this novel consensus control protocols not only decrease communication frequency and energy consumption but also avoid “Zeno behavior.” Through theoretical derivation, some sufficient conditions that guarantee the leader-following consensus of multiagent system are presented. In the end, for the sake of explaining the superiority of impulse control combined with event-triggered control strategy in more detail, the simulations of two control methods are given, respectively, for the same multiagent system, which can be better compared and explained. Le You, Xiaowei Jiang, Bo Li 0124, Huaicheng Yan 0001, Tingwen Huang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | Order batching and sequencing for minimising the total order completion time in pick-and-sort warehouses
Xiaowei Jiang, Yuankai Zhang 0003, Xiangpei Hu |
Expert Syst. Appl. | 1 |
| 2022 | Output Tracking Control Performance of Discrete Networked Systems Over Erasure Channel With Model UncertaintyabstractFor networked control systems, it is known that various communication parameters in the channel will pose some fundamental limitations on output tracking control (OTC) performance. In this study, we mainly discuss the limitations resulting from model uncertainties, involving channel and plant. Through using the bivariate stochastic process to model packet loss, and the assumption that channel noise is additive white Gaussian noise (AWGN), two explicit expressions of output tracking performance limitations are derived with the single-degree-of-freedom (SDOF) and two-degree-of-freedom (TDOF) control structure, which shows that the performance of OTC is closely related to the inherent characteristics of the plant, as well as the packet loss rate and power spectral density (PSD) of AWGN. Finally, by considering an illustrative example, the simulation results are verified and analyzed to ensure the effectiveness of treatment methods and results. Xiaowei Jiang, Xiangyong Chen, Huaicheng Yan 0001, Tingwen Huang |
IEEE Trans. Cybern. | 1 |
| 2022 | Output Tracking Control of Single-Input-Multioutput Systems Over an Erasure ChannelabstractThe output tracking control problem is investigated in this article. First, a new tradeoff performance index is presented for single-input-multioutput (SIMO) systems. Based on the frequency-domain method, the tracking performance limitations under time delay, packet loss, and channel noise effects are derived. We use a bivariate stochastic process to model the packet loss, and assume that channel noise is additive white Gaussian noise (AWGN). Two explicit expressions of the best tradeoff performance are given with the single-degree-of-freedom (SDOF) and two-degree-of-freedom (TDOF) control structures. It is shown that the tracking control performance has a close relation with the intrinsic characteristic of the plant, as well as the time delay, packet-dropouts rate, and power spectral density of AWGN. We also demonstrate that compared with the SDOF control structure, the TDOF control structure can improve the systems' attainable performance. A simulation example is finally discussed to validate the conclusions. Xiaowei Jiang, Xiangyong Chen, Tingwen Huang, Huaicheng Yan 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Distributed Edge Event-Triggered Control of Nonlinear Fuzzy Multiagent Systems With Saturation Constraint Hybrid Impulsive ProtocolsabstractIn this article, we devote to solve the consensus problem for nonlinear multiagent system under energy consumption constraint. A novel control protocol which combines impulse control with event-triggered mechanism is proposed. By setting the event decision time at pulse time, it can not only reduce triggered times but also avoid the occurrence of “Zeno behavior.” In order to reduce the conservatism of control protocol, input saturation constraint and state saturation constraint are both discussed, for the limited receiving channel of the agent and limited sampling range of the sampler, respectively. Further, the dynamic model of multiagent systems with fuzzy information is proposed for the uncertain factors in actual environment. Finally, for the sake of proving the effectiveness about the control protocols, some simulation examples are presented. Le You, Xiaowei Jiang, Huaicheng Yan 0001, Tingwen Huang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | Brain Decoding Using fNIRSabstractBrain activation can reflect semantic information elicited by natural words and concepts. Increasing research has been conducted on decoding such neural activation patterns using representational semantic models. However, prior work decoding semantic meaning from neurophysiological responses has been largely limited to ECoG, fMRI, MEG, and EEG techniques, each having its own advantages and limitations. More recently, the functional near infrared spectroscopy (fNIRS) has emerged as an alternative hemodynamic-based approach and possesses a number of strengths. We investigate brain decoding tasks under the help of fNIRS and empirically compare fNIRS with fMRI. Primarily, we find that: 1) like fMRI scans, activation patterns recorded from fNIRS encode rich information for discriminating concepts, but show limits on the possibility of decoding fine-grained semantic clues; 2) fNIRS decoding shows robustness across different brain regions, semantic categories and even subjects; 3) fNIRS has higher accuracy being decoded based on multi-channel patterns as compared to single-channel ones, which is in line with our intuition of the working mechanism of human brain. Our findings prove that fNIRS has the potential to promote a deep integration of NLP and cognitive neuroscience from the perspective of language understanding. We release the largest fNIRS dataset by far to facilitate future research. Dandan Huang, Yue Zhang 0004, Xiaowei Jiang |
AAAI | 4 |
| 2021 | CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersabstractAs the computation density and memory capacity continues to grow, DRAM errors have become the leading cause of server crashes and/or system failures in modern data centers. While myriads of techniques have been proposed to mitigate their impact on system reliability, these solutions either incur significant overhead on performance, power and memory capacity or require modifying multiple system components; hence, they are impractical to implement or deploy. This paper proposes CARE, a novel error tolerance framework for efficient and elastic resilience on DRAM errors. It introduces a cache-like structure in the memory controller for dynamic error tracking and proactive resilience enhancement to achieve high error tolerance economically and practically. Experiment results show that with around 58KB area overhead in the memory controller, CARE achieves near Chipkill reliability without any memory capacity penalty and incurs negligible performance overhead compared with the baseline SEC-DED systems. CARE provides an attractive alternative to enhance the reliability in data centers. Xiaowei Jiang, Liyin Liu, Huifeng Xu |
HPCA | 2 |
| 2021 | LIBRA: Clearing the Cloud Through Dynamic Memory Bandwidth ManagementabstractModern Cloud Service Providers (CSP) heavily co-schedule tasks with different priorities on the same computing node to increase server utilization. To ensure the performance of high priority jobs, CSPs usually employ Quality-of-Service (QoS) mechanisms to manage or regulate the usage of shared hardware resources. Among the critical shared hardware resources, there has been very limited analysis on effective sharing of memory bandwidth among co-scheduled jobs, mainly for two reasons: (1) The correlation between application performance and its memory bandwidth allocation is complicated. (2) An effective hardware throttling mechanism for precise memory bandwidth control is unavailable. These limitations drive CSPs to design conservative policies to ensure the performance of the high priority tasks, which significantly degrades the throughput of batch jobs and reduces the overall benefits of workload co-scheduling. This paper proposes LIBRA, a holistic framework for dynamic memory bandwidth management in production data centers. LIBRA incorporates a novel hardware throttling mechanism, Dynamic Resource Control, to support self-adaptive memory bandwidth regulation. It also employs a lightweight control policy to further enhance the bandwidth scalability for the throttled tasks. Our evaluation results on a cluster demonstrate that LIBRA is capable of increasing the performance of batch jobs by up to 52.8% compared to existing QoS schemes. Ying Zhang 0016, Xiaowei Jiang, Ian M. Steiner, Andrew Herdrich, Kevin Shu, Ripan Das, Long Cui, Litrin Jiang |
HPCA | 3 |
| 2021 | Swift: Reliable and Low-Latency Data Processing at Cloud ScaleabstractNowadays, it is a rapidly rising demand yet challenging issue to run large-scale applications on shared infrastructures such as data centers and clouds with low execution latency and high resource utilization. This paper reports our experience with Swift, a system capable of efficiently running real-time and interactive data processing jobs at cloud scale. Taking directed acyclic graph DAG as the job model, Swift achieves the design goal by three new mechanisms: 1 fine-grained scheduling that can efficiently partition a job into graphlets i.e., sub-graphs based on new shuffle heuristics and that does scheduling in the unit of graphlet, thus avoiding resource fragmentation and waste, 2 adaptive memory-based in-network shuffling that reduces IO overhead and data transfer time by doing shuffle in memory and allowing jobs to select the most efficient way to fulfill shuffling, and 3 lightweight fault tolerance and recovery that only prolong the whole job execution time slightly with the help of timely failure detection and fine-grained failure recovery. Experimental results show that Swift can achieve an average speedup of 2.11× on TPC-H, and 14.18× on Terasort when compared with Spark. Swift has been deployed in production, supporting as many as 140,000 executors and processing millions of jobs per day. Experiments with production traces show that Swift outperforms JetScope and Bubble Execution by 2.44× and 1.23× respectively. Yangyu Tao, Yifeng Lu, Xiaowei Jiang, Jinlei Jiang |
ICDE | 7 |
| 2021 | A novel wideband DOA estimation method based on a fast sparse frameabstractAbstract In this study, a novel fast wideband direction of arrival (DOA) estimation algorithm is proposed to reduce the computational complexity. First, a multiple measurement vector (MMV)‐based compact structure for a wideband signal is established. Combined with the focus operation, the array manifolds of different frequency bins are transformed into the dictionary of the reference frequency. Then, two efficient novel methods named adaptive step‐size‐based null space tuning with hard thresholding and feedback (ASNHF) and MMV‐ASNHF are proposed to process single measurement vector and MMV problem, respectively. Finally, wideband DOA estimation can be achieved by MMV‐ASNHF algorithm. Compared with other algorithms, the proposed algorithm has higher accuracy at a low signal‐to‐noise ratio and lower computational complexity. Simulation results show that the proposed estimator is effective and feasible. Haihong Tao, Jian Xie 0001, Xiaowei Jiang |
IET Commun. | 4 |
| 2021 | FlashP: An Analytical Pipeline for Real-time Forecasting of Time-Series Relational DataabstractInteractive response time is important in analytical pipelines for users to explore a sufficient number of possibilities and make informed business decisions. We consider a forecasting pipeline with large volumes of high-dimensional time series data. Real-time forecasting can be conducted in two steps. First, we specify the part of data to be focused on and the measure to be predicted by slicing, dicing, and aggregating the data. Second, a forecasting model is trained on the aggregated results to predict the trend of the specified measure. While there are a number of forecasting models available, the first step is the performance bottleneck. A natural idea is to utilize sampling to obtain approximate aggregations in real time as the input to train the forecasting model. Our scalable real-time forecasting system FlashP (Flash Prediction) is built based on this idea, with two major challenges to be resolved in this paper: first, we need to figure out how approximate aggregations affect the fitting of forecasting models, and forecasting results; and second, accordingly, what sampling algorithms we should use to obtain these approximate aggregations and how large the samples are. We introduce a new sampling scheme, called GSW sampling, and analyze error bounds for estimating aggregations using GSW samples. We introduce how to construct compact GSW samples with the existence of multiple measures to be analyzed. We conduct experiments to evaluate our solution its alternatives on real data. Shuyuan Yan, Bolin Ding, Jingren Zhou 0001, Zhewei Wei, Xiaowei Jiang, Sheng Xu 0007 |
Proc. VLDB Endow. | 6 |
| 2021 | An Elastic Task Scheduling Scheme on Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) are increasingly employed as domain-specific accelerators due to their efficiency and flexibility. A CGRA typically relies on compilers to perform task scheduling. The longstanding problem of static scheduling is that it suffers from insufficient parallelism in handling irregularities due to over-serialization and workload imbalance, which leads to severe resource underutilization and performance loss. To counteract the limitations of static scheduling in CGRAs, it is essential to exploit dynamic parallelism automatically and manage hardware resources adaptively. However, existing dynamic scheduling mechanisms, e.g., work stealing, often reschedule aggressively for instant performance but sacrifice efficiency, which is unfavorable to CGRAs that emphasize efficiency and fewer reconfigurations. This article proposes an elastic task scheduling scheme that enables lightweight dynamic scheduling in CGRAs. Tasks are rescheduled at runtime according to the classic tagged-token dataflow paradigm to enable dynamic task-level parallelism. Meanwhile, tasks are dynamically resized according to run-time throughputs via duplication, combination, and substitution operators for balanced multitask execution. We implement the elastic task scheduling scheme on a well-known reconfigurable architecture - triggered instruction architecture (TIA). Evaluation on the MachSuite benchmarks shows that the proposed scheme is effective in improving performance and energy efficiency. The average speedup is 2× over the baseline. Also, our design attains a 57 percent improvement in the area-normalized performance and a 49 percent better energy efficiency. Compared with a state-of-the-art dynamic scheduling method, our scheme achieves 1.6× speedup and 1.6× energy efficiency than work-stealing mechanism on the same substrate. Longlong Chen, Jianfeng Zhu 0001, Yangdong Deng, Zhaoshi Li, Xiaowei Jiang, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | DWT: Decoupled Workload Tracing for Data CentersabstractWorkload tracing is the foundational technology that many applications hinge upon. However, recent paradigm shift to-ward cloud computing has caused tremendous challenges to traditional workload tracing. Existing solutions either require a dedicated offline cluster or fail to capture the full-spectrum workload characteristics. This paper proposes DWT, a novel framework that leverages fast online instruction tracing, and uses synthetic data offline for memory access pattern reconstruction, thereby capturing the full workload characteristics while obviating the need of dedicated clusters. Experiment results show that the stack distance profiles generated from synthetic address traces match well with the original ones across all SPEC CPU 2017 programs and representative cloud applications, with correlation coefficient R^2 no less than 0.9. The page-level access frequencies also match well with those of the original programs. This decoupled tracing approach not only removes the roadblocks on workload characterization for data centers, but also enables new applications such as efficient online resource management. Xiaowei Jiang, Zheng Cao 0003 |
HPCA | 3 |
| 2020 | EFLOPS: Algorithm and System Co-Design for a High Performance Distributed Training PlatformabstractDeep neural networks (DNNs) have gained tremendous attractions as compelling solutions for applications such as image classification, object detection, speech recognition, and so forth. Its great success comes with excessive trainings to make sure the model accuracy is good enough for those applications. Nowadays, it becomes challenging to train a DNN model because of 1) the model size and data size keep increasing, which usually needs more iterations to train; 2) DNN algorithms evolve rapidly, which requires the training phase to be short for a quick deployment. To address those challenges, distributed training platforms have been proposed to leverage massive server nodes for training, with the hope of significant training time reduction. Therefore, scalability is a critical performance metric to evaluate a distributed training platform. Nevertheless, our analysis reveals that traditional server clusters have poor scalability for training due to the traffic congestions within the server and beyond. The intra-server traffic on the I/O fabric can result in severe congestions and skewed quality of service as high performance devices are competing with each other. Moreover, the traffic congestions on the Ethernet for inter-server communication could also incur significant performance degradation. In this work, we devise a novel distributed training platform, EFLOPS, that adopts an algorithm and system co-design methodology to achieve good scalability. A new server architecture is proposed to alleviate the intra-server congestions. Moreover, a new network topology, BiGraph, is proposed to divide the network into two separate parts, so that there is always a direct connection between any nodes from different parts. Finally, accompany with BiGraph, a topology-aware allreduce algorithm is proposed to eliminate the traffic congestion on the direct connection. The experimental results show that eliminating the congestions on network interface can gain up to 11.3xcommunication speedup. The proposed algorithm and topology can provide further improvement up to 6.08x. The overall performance of ResNet-50 training achieves near-linear scalability, and is competitive to the top-rankings of MLPerf results. Jianbo Dong, Zheng Cao 0003, Jianxi Ye, Shaochuang Wang, Liuyihan Song, Liwei Peng, Yiqun Guo, Xiaowei Jiang, Lingbo Tang, Yin Du, Yingya Zhang |
HPCA | 12 |
| 2020 | CETUS: Towards Proportional Capacity Provisioning and Cost-Effectiveness in Frontend ServersabstractHyper-scale data centers emerged in the last decade largely adopt a multi-tiered architecture with the frontend clusters dedicated to serve high-demand user facing web traffic. In order to mitigate the overhead caused by Transport Layer Security (TLS) that protects the communications between users and the data center, the frontend clusters usually apply hardware TLS acceleration. In this paper, we analyze the inefficiencies that lie in today's frontend clusters, and propose CETUS, an improved data center frontend system architecture. CETUS improves the cost effectiveness of frontend clusters through cluster consolidation that fully offloads the TLS and network stack to a CETUS SoC; it enables proportional capacity provisioning through pooling of resources and dynamic division of frontend tasks with a flow diverter. Compared to existing frontend clusters that are equipped with commercial TLS acceleration solutions, CETUS balances out the frontend resource utilization, and provides up to 86.2% in cost reduction while maintaining at the same level of throughput and latency of the frontend. Xiaowei Jiang, Zheng Cao 0003 |
ISPASS | 3 |
| 2020 | TFE: Energy-efficient Transferred Filter-based Engine to Compress and Accelerate Convolutional Neural NetworksabstractAlthough convolutional neural network (CNN) models have greatly enhanced the development of many fields, the untenable number of parameters and computations in these models yield significant performance and energy challenges in hardware implementations. Transferred filter-based methods, as very promising techniques that have not yet been explored in the architecture domain, can substantially compress CNN models. However, their straightforward hardware implementation inherently incurs massive redundant computations, causing significant energy and time consumption. In this work, a highly efficient transferred filter-based engine (TFE) is developed to alleviate this deficiency, with CNN models compressed and accelerated. First, the filters of CNN models are flexibly transferred according to specific tasks to reduce the model size. Then, two hardware-friendly mechanisms are proposed in the TFE to remove duplicate computations caused by transferred filters, which can further accelerate transferred CNN models. The first mechanism exploits the shared weights hidden in each row of transferred filters and reuses the corresponding same partial sums, reducing at least 25% of repetitive computations in each row. The second mechanism can intelligently schedule and access the memory system to reuse the repetitive partial sums among different rows of the transferred filters with at least 25% of computations eliminated. Furthermore, an efficient hardware architecture is proposed in the TFE to fully reap the benefits of the two proposed mechanisms such that different types of networks are flexibly supported. To achieve high energy efficiency, the sub-array-based filter mapping method (SAFM) is proposed, where the process element (PE) subarray is used as the elementary computational unit to support various filters. Therein, input data can be efficiently broadcast in each PE sub-array and the load can be stripped from each PE and intensively alleviated, which can dramatically reduce the area and power consumption. Excluding MobileNet-like networks that adopt depth-wise convolution, most mainstream networks can be compressed and accelerated by the proposed TFE. Two state-of-the-art transferred filter-based methods, i.e., doubly CNN and symmetry CNN are implemented by exploiting the TFE. Compared with Eyeriss, average speedup improvements of 2.93× and 3.17× are achieved in the convolutional layers of various modern CNNs. The overall energy efficiency can be improved by 12.66× and 13.31× on average. Compared with other state-of-the-art related works, the TFE can maximally achieve a parameter reduction of 4.0×, a speedup of 2.72× and an energy efficiency improvement of 10.74× on VGGNet. Huiyu Mo, Leibo Liu, Wenjing Hu, Wenping Zhu, Eric Q. Li, Ang Li 0033, Shouyi Yin, Xiaowei Jiang, Shaojun Wei |
MICRO | 9 |
| 2020 | Optimal performance of LTI systems over power constrained erasure channels
Xiaowei Jiang, Xiangyong Chen, Ming-Feng Ge |
Inf. Sci. | 1 |
| 2020 | Alibaba Hologres: A Cloud-Native Service for Hybrid Serving/Analytical ProcessingabstractIn existing big data stacks, the processes of analytical processing and knowledge serving are usually separated in different systems. In Alibaba, we observed a new trend where these two processes are fused: knowledge serving incurs generation of new data, and these data are fed into the process of analytical processing which further fine tunes the knowledge base used in the serving process. Splitting this fused processing paradigm into separate systems incurs overhead such as extra data duplication, discrepant application development and expensive system maintenance. In this work, we propose Hologres, which is a cloud native service for hybrid serving and analytical processing (HSAP). Hologres decouples the computation and storage layers, allowing flexible scaling in each layer. Tables are partitioned into self-managed shards. Each shard processes its read and write requests concurrently independent of each other. Hologres leverages hybrid row/column storage to optimize operations such as point lookup, column scan and data ingestion used in HSAP. We propose Execution Context as a resource abstraction between system threads and user tasks. Execution contexts can be cooperatively scheduled with little context switching overhead. Queries are parallelized and mapped to execution contexts for concurrent execution. The scheduling framework enforces resource isolation among different queries and supports customizable schedule policy. We conducted experiments comparing Hologres with existing systems specifically designed for analytical processing and serving workloads. The results show that Hologres consistently outperforms other systems in both system throughput and end-to-end query latency. Xiaowei Jiang, Yuejun Hu, Guangran Jiang, Chen Xia, Weihua Jiang, Jihong Ma, Li Su 0005, Kai Zeng 0002 |
Proc. VLDB Endow. | 1 |
| 2019 | Analysis and Optimization of the Memory Hierarchy for Graph Processing WorkloadsabstractGraph processing is an important analysis technique for a wide range of big data applications. The ability to explicitly represent relationships between entities gives graph analytics a significant performance advantage over traditional relational databases. However, at the microarchitecture level, performance is bounded by the inefficiencies in the memory subsystem for single-machine in-memory graph analytics. This paper consists of two contributions in which we analyze and optimize the memory hierarchy for graph processing workloads. First, we perform an in-depth data-type-aware characterization of graph processing workloads on a simulated multi-core architecture. We analyze 1) the memory-level parallelism in an out-of-order core and 2) the request reuse distance in the cache hierarchy. We find that the load-load dependency chains involving different application data types form the primary bottleneck in achieving a high memory-level parallelism. We also observe that different graph data types exhibit heterogeneous reuse distances. As a result, the private L2 cache has negligible contribution to performance, whereas the shared L3 cache shows higher performance sensitivity. Abanti Basak, Shuangchen Li, Xing Hu 0001, Sang Min Oh, Xinfeng Xie, Xiaowei Jiang, Yuan Xie 0001 |
HPCA | 7 |
| 2019 | H∞ Output Tracking Control for Networked Systems With Adaptively Adjusted Event-Triggered SchemeabstractThe H∞output tracking control problem of the networked control systems (NCSs) under an adaptively adjusted event-triggered scheme is investigated in this paper. Firstly, a novel adaptively adjusted event-triggered scheme developed in the NCSs with stochastic sensor faults is proposed to choose the necessary packets of sampled data to be transmitted through the networks. Then, the considered system is described as a time-delay system with a delay-distribution for investigation. Based on the established model, a novel stability criterion and the state-feedback controller design with a desired performance of H∞output tracking control for the systems are both derived by using Lyapunov functional. An algorithm is also presented to explain the process of adaptively adjusted event-triggered scheme method. Finally, a satellite tracking case is provided to demonstrate the effectiveness of the proposed approach. Huaicheng Yan 0001, Chenyang Hu, Hao Zhang 0008, Hamid Reza Karimi, Xiaowei Jiang, Ming Liu 0014 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2018 | Order batching and sequencing problem under the pick-and-sort strategy in online supermarketsabstractThe order picking and sorting-packing processes are important warehouses operations in online supermarkets under the pick-and-sort batch picking strategy. The buffer areas in the middle of these two processes play a buffering role, but the buffer size is limited because of storage facilities and finite room. Due to the differences between different batches, there may be too many batches blocking in the limited buffer areas which causes the stagnation of the picking process, or no batch in the buffer areas which causes the idleness of the sorting-packing process. This paper studies the order batching and sequencing problem with limited buffers with the objective of minimizing the total time of two processes for a given set of orders. We solve the problem with a modified seed algorithm. Xiaowei Jiang, Yaxian Zhou, Yuankai Zhang 0003, Xiangpei Hu |
KES | 1 |
| 2016 | Event-driven multi-consensus of multi-agent networks with repulsive links
Bin Hu 0008, Zhi-Hong Guan, Xiaowei Jiang, Rui-Quan Liao, Chaoyang Chen 0001 |
Inf. Sci. | 3 |
| 2016 | The minimal signal-to-noise ratio required for stability of control systems over a noisy channel in the presence of packet dropouts
Xiaowei Jiang, Bin Hu 0008, Zhi-Hong Guan, Li Yu 0003 |
Inf. Sci. | 1 |
| 2015 | Best achievable tracking performance for networked control systems with encoder-decoder
Xiaowei Jiang, Bin Hu 0008, Zhi-Hong Guan, Li Yu 0003 |
Inf. Sci. | 1 |
| 2013 | Flexible Capacity Partitioning in Many-Core Tiled CMPsabstractChip Multi-Processors (CMP) have become a mainstream computing platform. As transistor density shrinks and the number of cores increases, more scalable CMP architectures will emerge. Recently, tiled architectures have shown such scalable characteristics and been used in many industry chips. The memory hierarchy in tiled architectures presents interesting design challenges. One major challenge is the organization of the Last Level Cache (LLC). Shared but distributed LLCs are preferred over private LLCs due to better utilization of the aggregate cache capacity. However, such architectures suffer from high on-chip hit latency. Breaking down the the shared LLC into smaller domains called clusters where each cluster is associated with one processor VM can reduce the on-chip hit latency significantly. However, having static cluster sizes may not be the best option as some processes may need more cache capacity than others. In this paper, we propose a novel inter-cluster capacity partitioning scheme called Flexible TiledCMP Capacity Partitioning (FlexTCP). FlexTCP maintains the small hit latency of cluster caches while at the same time enables flexible capacity partitioning across clusters such that clusters with high cache demand can steal capacity from underutilized clusters. FlexTCP proposes multiple ways of shrinking/expanding the cluster size. When applied to a 64-coretiled-CMP running a mix of SPEC CPU2006 and Parsec 2.1 workloads, FlexTCP achieves an average of 21% and 18% improvement in Weighted Speedup over two rival schemes. Ahmad Samih, Xiaowei Jiang, Yan Solihin |
CCGRID | 2 |
| 2013 | Reducing cache and TLB power by exploiting memory region and privilege level semantics
Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Shih-Lien Lu, Ravi R. Iyer 0001, Tong Li 0003 |
J. Syst. Archit. | 3 |
| 2012 | Summarizing Semantic Associations Based on Focused Association Graph
Xiaowei Jiang, Wei Gui, Feifei Gao 0001, Peng Wang 0004, Fengbo Zhou |
ADMA | 1 |
| 2012 | Exploiting Semantics of Virtual Memory to Improve the Efficiency of the On-Chip Memory System
Bin Li 0018, Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Andrew Herdrich, Ravi R. Iyer 0001, Srihari Makineni |
Euro-Par | 4 |
| 2012 | QuickIA: Exploring heterogeneous architectures on real prototypesabstractOver the last decade, homogeneous multi-core processors emerged and became the de-facto approach for offering high parallelism, high performance and scalability for a wide range of platforms. We are now at an interesting juncture where several critical factors (smaller form factor devices, power challenges, need for specialization, etc) are guiding architects to consider heterogeneous chips and platforms for the next decade and beyond. Exploring heterogeneous architectures is challenging since it involves re-evaluating architecture options, OS implications and application development. In this paper, we describe these research challenges and then introduce a heterogeneous prototype platform called QuickIA that enables rapid exploration of heterogeneous architectures employing multiple generations of Intel processors for evaluating the implications of asymmetry and FPGAs to experiment with specialized processors or accelerators. We also show example case studies using the QuickIA research prototype to highlight its value in conducting heterogeneous architecture, OS and applications research. Bhushan Chitlur, Ganapati Srinivasa, Scott Hahn, Dheeraj Reddy, David A. Koufaty, Paul Brett, Abirami Prabhakaran, Li Zhao 0002, Nelson Ijih, Suchit Subhaschandra, Sabina Grover, Xiaowei Jiang, Ravi R. Iyer 0001 |
HPCA | 13 |
| 2012 | HiRe: using hint & release to improve synchronization of speculative threadsabstractThread-Level Speculation (TLS) is a promising technique for improving performance of serial codes on multi-cores by automatically extracting threads and running them in parallel. However, the speculation efficiency as well as the performance gain of TLS systems are reduced by cross-thread data dependence violations. Reducing the cost and frequency of violations are key to improving the efficiency of TLS. One method to keep a dependence from violating is to predict it and communicate the value via synchronization. However, prior work in this field still cannot handle enough violating dependences, especially hard-to-predict ones and those in non-loop TLS tasks. Also, they suffer from over-synchronization and/or introduce complicated hardware. The major reason is that these techniques are highly sensitive to the accuracy of the dependence prediction, which is hard to improve in the face of irregular dependence and task patterns. Xiaowei Jiang, Wei Liu 0014, Youfeng Wu, James Tuck 0001 |
ICS | 2 |
| 2012 | Reducing L1 caches power by exploiting software semanticsabstractTo access a set-associative L1 cache in a high-performance processor, all ways of the selected set are searched and fetched in parallel using physical address bits. Such a cache is oblivious of memory references' software semantics such as stack-heap bifurcation of the memory space, and user-kernel ring levels. This constitutes a waste of energy since e.g., a user-mode instruction fetch will never hit a cache block that contains kernel code. Similarly, a stack access will not hit a cacheline that contains heap data. Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Shih-Lien Lu, Ravi R. Iyer 0001, Tong Li 0003 |
ISLPED | 3 |
| 2012 | Active memory controller
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Ali Ibrahim, Michael A. Parker, Xiaowei Jiang |
J. Supercomput. | 7 |
| 2011 | ACCESS: Smart scheduling for asymmetric cache CMPsabstractIn current Chip-multiprocessors (CMPs), a significant portion of the die is consumed by the last-level cache. Until recently, the balance of cache and core space has been primarily guided by the needs of single applications. However, as multiple applications or virtual machines (VMs) are consolidated on such a platform, researchers have observed that not all VMs or applications require significant amount of cache space. In order to take advantage of this phenomenon, we explore the use of asymmetric last-level caches in a CMP platform. While asymmetric cache CMPs provide the benefit of reduced power and area, it is important to build in hardware/software support to appropriately schedule applications on to cores with suitable cache capacity. In this paper, we address this problem with our ACCESS architecture comprising of: (a) asymmetric caches across a group of cores, (b) hardware support that enables prediction of cache performance on the different sized caches and (c) OS scheduler support to make use of the prediction capability and appropriately schedule applications on to core with suitable cache capacity. Measurements on a working prototype using SPEC2006 benchmarks show that our ACCESS architecture can effectively schedule jobs in an asymmetric cache CMP and provide 23% performance improvement compared to a naive scheduler, and is 97% close to an oracle scheduler in making schedules. Xiaowei Jiang, Asit K. Mishra, Li Zhao 0002, Ravi R. Iyer 0001, Zhen Fang 0002, Sadagopan Srinivasan, Srihari Makineni, Paul Brett, Chita R. Das |
HPCA | 1 |
| 2011 | Architectural framework for supporting operating system survivabilityabstractThe ever increasing size and complexity of Operating System (OS) kernel code bring an inevitable increase in the number of security vulnerabilities that can be exploited by attackers. A successful security attack on the kernel has a profound impact that may affect all processes running on it. In this paper we propose an architectural framework that provides survivability to the OS kernel, i.e. able to keep normal system operation despite security faults. It consists of three components that work together: (1) security attack detection, (2) security fault isolation, and (3) a recovery mechanism that resumes normal system operation. Through simple but carefully-designed architecture support, we provide OS kernel survivability with low performance overheads (<; 5% for kernel intensive benchmarks). When tested with real world security attacks, our survivability mechanism automatically prevents the security faults from corrupting the kernel state or affecting other processes, recovers the kernel state and resumes execution. Xiaowei Jiang, Yan Solihin |
HPCA | 1 |
| 2011 | Cost-effectively offering private buffers in SoCs and CMPsabstractHigh performance SoCs and CMPs integrate multiple cores and hardware accelerators such as network interface devices and speech recognition engines. Cores make use of SRAM organized as a cache. Accelerators make use of SRAM as special-purpose storage such as FIFOs, scratchpad memory, or other forms of private buffers. Dedicated private buffers provide benefits such as deterministic access, but are highly area inefficient due to the lower average utilization of the total available storage. Zhen Fang 0002, Li Zhao 0002, Ravi R. Iyer 0001, Carlos Flores Fajardo, German Fabila Garcia, Bin Li 0018, Steve R. King, Xiaowei Jiang, Srihari Makineni |
ICS | 9 |
| 2010 | CHOP: Adaptive filter-based DRAM caching for CMP server platformsabstractAs manycore architectures enable a large number of cores on the die, a key challenge that emerges is the availability of memory bandwidth with conventional DRAM solutions. To address this challenge, integration of large DRAM caches that provide as much as 5× higher bandwidth and as low as 1/3rd of the latency (as compared to conventional DRAM) is very promising. However, organizing and implementing a large DRAM cache is challenging because of two primary tradeoffs: (a) DRAM caches at cache line granularity require too large an on-chip tag area that makes it undesirable and (b) DRAM caches with larger page granularity require too much bandwidth because the miss rate does not reduce enough to overcome the bandwidth increase. In this paper, we propose CHOP (Caching HOt Pages) in DRAM caches to address these challenges. We study several filter-based DRAM caching techniques: (a) a filter cache (CHOP-FC) that profiles pages and determines the hot subset of pages to allocate into the DRAM cache, (b) a memory-based filter cache (CHOP-MFC) that spills and fills filter state to improve the accuracy and reduce the size of the filter cache and (c) an adaptive DRAM caching technique (CHOP-AFC) to determine when the filter cache should be enabled and disabled for DRAM caching. We conduct detailed simulations with server workloads to show that our filter-based DRAM caching techniques achieve the following: (a) on average over 30% performance improvement over previous solutions, (b) several magnitudes lower area overhead in tag space required for cache-line based DRAM caches, (c) significantly lower memory bandwidth consumption as compared to page-granular DRAM caches. Xiaowei Jiang, Niti Madan, Li Zhao 0002, Mike Upton, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell, Yan Solihin, Rajeev Balasubramonian |
HPCA | 1 |
| 2010 | Understanding how off-chip memory bandwidth partitioning in Chip Multiprocessors affects system performanceabstractChip Multi-Processor (CMP) architectures have recently become a mainstream computing platform. Recent CMPs allow cores to share expensive resources, such as the last level cache and off-chip pin bandwidth. To improve system performance and reduce the performance volatility of individual threads, last level cache and off-chip bandwidth partitioning schemes have been proposed. While how cache partitioning affects system performance is well understood, little is understood regarding how bandwidth partitioning affects system performance, and how bandwidth and cache partitioning interact with one another. In this paper, we propose a simple yet powerful analytical model that gives us an ability to answer several important questions: (1) How does off-chip bandwidth partitioning improve system performance? (2) In what situations the performance improvement is high or low, and what factors determine that? (3) In what way cache and bandwidth partitioning interact, and is the interaction negative or positive? (4) Can a theoretically optimum bandwidth partition be derived, and if so, what factors affect it? We believe understanding the answers to these questions is very valuable to CMP system designers in coming up with strategies to deal with the scarcity of off-chip bandwidth in future CMPs with many cores on a chip. Xiaowei Jiang, Yan Solihin |
HPCA | 2 |
| 2009 | Architecture Support for Improving Bulk Memory Copying and Initialization PerformanceabstractBulk memory copying and initialization is one of the most ubiquitous operations performed in current computer systems by both user applications and Operating Systems. While many current systems rely on a loop of loads and stores, there are proposals to introduce a single instruction to perform bulk memory copying. While such an instruction can improve performance due to generating fewer TLB and cache accesses, and requiring fewer pipeline resources, in this paper we show that the key to significantly improving the performance is removing pipeline and cache bottlenecks of the code that follows the instructions. We show that the bottlenecks arise due to (1) the pipeline clogged by the copying instruction, (2) lengthened critical path due to dependent instructions stalling while waiting for the copying to complete, and (3) the inability to specify (separately) the cacheability of the source and destination regions. We propose FastBCI, an architecture support that achieves the granularity efficiency of a bulk copying/ initialization instruction, but without its pipeline and cache bottlenecks. When applied to OS kernel buffer management, we show that on average FastBCI achieves anywhere between 23% to 32% speedup ratios, which is roughly 3x-4x of an alternative scheme, and 1.5x-2x of a highly optimistic DMA with zero setup and interrupt overheads. Xiaowei Jiang, Yan Solihin, Li Zhao 0002, Ravi R. Iyer 0001 |
PACT | 1 |
| 2009 | Scaling the bandwidth wall: challenges in and avenues for CMP scalingabstractAs transistor density continues to grow at an exponential rate in accordance to Moore's law, the goal for many Chip Multi-Processor (CMP) systems is to scale the number of on-chip cores proportionally. Unfortunately, off-chip memory bandwidth capacity is projected to grow slowly compared to the desired growth in the number of cores. This creates a situation in which each core will have a decreasing amount of off-chip bandwidth that it can use to load its data from off-chip memory. The situation in which off-chip bandwidth is becoming a performance and throughput bottleneck is referred to as the bandwidth wall problem. Brian Rogers, Anil Krishna, Gordon B. Bell, Ken V. Vu, Xiaowei Jiang, Yan Solihin |
ISCA | 5 |
| 2009 | Phylogenomic inference of functional divergence
Tom A. Williams, Brian E. Caffrey, Xiaowei Jiang, Christina Toft, Mario A. Fares |
BMC Bioinform. | 3 |
| 2006 | Comprehensively and efficiently protecting the heapabstractThe goal of this paper is to propose a scheme that provides comprehensive security protection for the heap. Heap vulnerabilities are increasingly being exploited for attacks on computer programs. In most implementations, the heap management library keeps the heap meta-data (heap structure information) and the application's heap data in an interleaved fashion and does not protect them against each other. Such implementations are inherently unsafe: vulnerabilities in the application can cause the heap library to perform unintended actions to achieve control-flow and non-control attacks.Unfortunately, current heap protection techniques are limited in that they use too many assumptions on how the attacks will be performed, require new hardware support, or require too many changes to the software developers' toolchain. We propose Heap Server, a new solution that does not have such drawbacks. Through existing virtual memory and inter-process protection mechanisms, Heap Server prevents the heap meta-data from being illegally overwritten, and heap data from being meaningfully overwritten. We show that through aggressive optimizations and parallelism, Heap Server protects the heap with nearly-negligible performance overheads even on heap-intensive applications. We also verify the protection against several real-world exploits and attack kernels. Mazen Kharbutli, Xiaowei Jiang, Yan Solihin, Guru Venkataramani, Milos Prvulovic |
ASPLOS | 2 |
| 2004 | Hosting the .NET Runtime in Microsoft SQL ServerabstractThe integration of the .NET Common Language Runtime (CLR) inside the SQL Server DBMS enables database programmers to write business logic in the form of functions, stored procedures, triggers, data types, and aggregates using modern programming languages such as C#, Visual Basic, C++, COBOL, and J++. This paper presents three main aspects of this work. First, it describes the architecture of the integration of the CLR inside the SQL Server database process to provide a safe, scalable, secure, and efficient environment to run user code. Second, it describes our approach to defining and enforcing extensibility contracts to allow a tight integration of types, aggregates, functions, triggers, and procedures written in modern languages with the DBMS. Finally, it presents initial performance results showing the efficiency of user-defined types and functions relative to equivalent native DBMS features. Alazel Acheson, Mason Bendixen, José A. Blakeley, Peter Carlin, Ebru Ersan, Xiaowei Jiang, Christian Kleinerman, Balaji Rathakrishnan, Gideon Schaller, Beysim Sezgin, Ramachandran Venkatesh |
SIGMOD Conference | 7 |
| 2000 | A Constraint-Based Framework for Prototyping Distributed Virtual Applications
Vineet Gupta 0001, Lalita Jategaonkar Jagadeesan, Radha Jagadeesan, Xiaowei Jiang, Konstantin Läufer |
CP | 4 |