Dawei Huang

dblp:33/2974 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 1 first-author · 2 since 2021Theory of computation · 6 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Time-Varying Target Predictive Entrapment Based on Gene Regulatory Network and Sliding Mode Control
abstract
To address slow convergence and formation maintenance challenges in swarm robotic entrapment tasks, a predictive entrapment control algorithm that combines gene regulatory network and sliding mode control (GRN-SMC) is proposed. First, to stabilize the target position information generated by the hierarchical GRN, a novel sorting rule is designed. Then, an artificial neural network (ANN) is employed to perform on-line prediction of the swarm robots’ kinematic states. These predicted values are then fed into a specifically designed sliding mode controller, which ultimately outputs the optimal control velocities for the swarm robots. Comparative simulation experiments with three state-of-the-art algorithms demonstrate that our method achieves significant improvements in tracking accuracy(error reduced by 82%), single-iteration execution time(reduced by 29%), and formation maintenance (formation integrity increased by 34%). Furthermore, physical robot experiments demonstrate that even in the presence of unknown external disturbances (such as ground slippage) and robot positioning errors (±0.1 m, ±5°), the proposed algorithm still exhibits excellent robustness.
Ziling Wen, Zhaojun Wang, Dawei Huang, Binghao Yang, Wenji Li, Zhun Fan, An-Min Zou
IEEE Internet Things J.4
2025 miMamba: EEG-Based Emotion Recognition With Multi-Scale Inverted Mamba Models
abstract
EEG-based emotion recognition holds significant potential in the field of brain-computer interfaces. A key challenge is extracting discriminative spatiotemporal features from electroencephalogram (EEG) signals. Existing studies often rely on domain-specific time-frequency features and analyze temporal dependencies and spatial characteristics separately, neglecting the local-global relationships and the interaction in spatiotemporal dynamics. To address this, we propose a novel network called Multi-scale Inverted Mamba (miMamba), which consists of Multi-Scale Temporal Blocks (MSTB) and Temporal-Spatial Fusion Blocks (TSFB). Specifically, MSTBs are designed to capture both local details and global temporal dependencies across different scale subsequences. The TSFBs, implemented with an inverted Mamba structure, focus on the interaction between dynamic temporal dependencies and spatial characteristics. The primary advantage of miMamba lies in its ability to leverage transformed multi-scale EEG sequences, exploiting the interaction between temporal and spatial features without the need for domain-specific time-frequency feature extraction. Experiments show that using only four EEG channels, miMamba achieves remarkable average recognition accuracies for Valence and Arousal classification: 94.86% on the DEAP dataset, 94.94% on the DREAMER dataset, and 91.36% on the SEED dataset. These results underscore the model's superior performance in multidimensional emotion recognition tasks and its potential for practical applications in resource-constrained affective computing scenarios.
Dawei Huang, Xiaojiang Peng
IEEE Trans. Affect. Comput.2
2024 SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
abstract
Monolithic large language models (LLMs) like GPT-4 have paved the way for modern generative AI applications. Training, serving, and maintaining monolithic LLMs at scale, however, remains prohibitively expensive and challenging. The disproportionate increase in compute-to-memory ratio of modern AI accelerators have created a memory wall, necessitating new methods to deploy AI. Recent research has shown that a composition of many smaller expert models, each with several orders of magnitude fewer parameters, can match or exceed the capabilities of monolithic LLMs. Composition of Experts (CoE) is a modular approach that lowers the cost and complexity of training and serving. However, this approach presents two key challenges when using conventional hardware: (1) without fused operations, smaller models have lower operational intensity, which makes high utilization more challenging to achieve; and (2) hosting a large number of models can be either prohibitively expensive or slow when dynamically switching between them. In this paper, we describe how combining CoE, streaming dataflow, and a three-tier memory system scales the AI memory wall. We describe Samba-CoE, a CoE system with 150 experts and a trillion total parameters. We deploy Samba-CoE on the SambaNova SN40L Reconfigurable Dataflow Unit (RDU) -a commercial dataflow accelerator architecture that has been codesigned for enterprise inference and training applications. The chip introduces a new three-tier memory system with on-chip distributed SRAM, on-package HBM, and off-package DDR DRAM. A dedicated inter-RDU network enables scaling up and out over multiple sockets. We demonstrate speedups ranging from 2× to 13× on various benchmarks running on eight RDU sockets compared with an unfused baseline. We show that for CoE inference deployments, the 8-socket RDU Node reduces machine footprint by up to 19 ×, speeds up model switching time by 15× to 31×, and achieves an overall speedup of 3.7× over a DGX H100 and 6.6× over a DGX A100.
Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi, Mingran Wang, Kejie Zhang, Tianren Gao, Angela Wang, Yongning Sheng, Joshua Brot, Denis Sokolov, Apurv Vivek, Calvin Leung, Arjun Sabnis, Jiayu Bai, Tuowen Zhao, Mark Gottscho, Mark Luttrell, Manish K. Shah, Zhengyu Chen 0002, Kaizhao Liang, Swayambhoo Jain, Urmish Thakker, Dawei Huang, Sumti Jairath, Kevin J. Brown, Kunle Olukotun
MICRO27
2022 Approximate Generalized Matching: f-Matchings and f-Edge Covers
Dawei Huang, Seth Pettie
Algorithmica1
2022 Three-dimensional quota matching-based latency-sensitive task offloading for multi-mode green IoT in smart buildings
abstract
Abstract The green internet of things with heterogeneous communication technologies can provide data transmission and computing services for low‐carbon operation of smart buildings. However, latency‐sensitive task offloading in smart buildings for multi‐mode green internet of things still faces several challenges such as coupling between multi‐mode channel and multiple gateway selection, diversified quality of service requirement guarantee, and contradiction of long‐term performance guarantee and short‐term optimisation objectives. To address these challenges, a three‐dimensional quota matching‐based latency‐sensitive task offloading algorithm is proposed to minimise the weighted difference between energy consumption and throughput under the long‐term queuing delay constraints. Specifically, the minimisation problem is decoupled by Lyapunov optimisation. The three‐dimensional quota matching among devices, gateways, and channels is employed to solve the conflicts between gateway selection and channel selection. Finally, the three‐dimensional quota matching is converted to a two‐side quota matching to further reduce complexity and solved iteratively. Numerical results demonstrate that compared with H3CG and MMCS, the proposed algorithm improves the weighted difference between energy consumption and throughput by 21.85% and 27.91%, respectively, and reduces the sensor‐side average queuing delay by 30.82% and 16.83%, and gateway‐side average queuing delay by 16.57% and 26.71%, respectively.
Sunxuan Zhang, Ruiqiuyu Wang, Zhenyu Zhou 0001, Zhong Gan, Xianjiong Yao, Zhaoyang You, Dawei Huang, Guoxiang Hua, Shahid Mumtaz
IET Commun.9
2022 A 5-mm Untethered Crawling Robot via Self-Excited Electrostatic Vibration
abstract
The grand challenge toward the miniaturization of untethered robots is the lack of a proper actuation system that can effectively transform the limited onboard energy into an untethered movement of the whole body. Here, in this article, we present an insect scale, 5-mm-long, and 9.4 mg in body mass untethered crawling robot powered by a 37 mg onboard ceramic capacitor. With high electromechanical efficiency by utilizing self-excited electrostatic vibration, the prototype device can be directly powered by an onboard capacitor without any boost or frequency conversion circuitry, and forward movements have been achieved at an average moving speed of 5.9 mm/s for 1.68 s. The maneuverability of locomotion from solid ground to water and different moving directional controls via the leg designs have been demonstrated for the advancements in the field of millimeter-scale robotics research.
Yangsheng Zhu, Mingjing Qi, Jianmei Huang, Dawei Huang, Xiaojun Yan
IEEE Trans. Robotics5
2021 The Communication Complexity of Set Intersection and Multiple Equality Testing
abstract
In this paper we explore fundamental problems in randomized communication complexity such as computing SetIntersection on sets of size $k$ and EqualityTesting between vectors of length $k$. Sağlam and Tardos [ Proceedings of the 54 th Annual IEEE Symposium on Foundations of Computer Science, 2013, pp. 678--687] and Brody et al. [ Algorithmica, 76 (2016), pp. 796--845] showed that for these types of problems, one can achieve optimal communication volume of $O(k)$ bits, with a randomized protocol that takes $O(\log^* k)$ rounds. They also proved that this is one point along the optimal round-communication trade-off curve. Aside from rounds and communication volume, there is a third parameter of interest, namely the error probability $p_{{err}}$, which we write $2^{-E}$. It is straightforward to show that protocols for SetIntersection or EqualityTesting need to send at least $\Omega(k + E)$ bits, regardless of the number of rounds. Is it possible to simultaneously achieve optimality in all three parameters, namely $O(k + E)$ communication and $O(\log^* k)$ rounds? In this paper we prove that there is no universally optimal algorithm, and we complement the existing round-communication trade-offs [M. Sağlam and G. Tardos, Proceedings of the 54 th Annual IEEE Symposium on Foundations of Computer Science, 2013, pp. 678--687; J. Brody et al., Algorithmica, 76 (2016), pp. 796--845] with a new trade-off between rounds, communication, and probability of error. In particular, any protocol for solving multiple EqualityTesting in $r$ rounds with failure probability $p_{{err}} = 2^{-E}$ has communication volume $\Omega(Ek^{1/r})$. We present several algorithms for multiple EqualityTesting (and its variants) that match or nearly match our lower bound and the lower bound of [M. Sağlam and G. Tardos, Proceedings of the 54 th Annual IEEE Symposium on Foundations of Computer Science, 2013, pp. 678--687; J. Brody et al., Algorithmica, 76 (2016), pp. 796--845]. Lower bounds on EqualityTesting extend to SetIntersection for every $r, k,$ and $p_{{err}}$ (which is trivial); in the reverse direction, we prove that upper bounds on EqualityTesting for $r, k, p_{{err}}$ imply similar upper bounds on SetIntersection with parameters $r+1, k,$ and $p_{{err}}$. Our original motivation for considering $p_{{err}}$ as an independent parameter came from the problem of enumerating triangles in distributed (${CONGEST}$) networks having maximum degree $\Delta$. We prove that this problem can be solved in $O(\Delta/\log n + \log\log \Delta)$ time with high probability $1-1/{poly}(n)$. This beats the trivial (deterministic) $O(\Delta)$-time algorithm and is superior to the $\tilde{O}(n^{1/3})$ algorithm of [Y. Chang, S. Pettie, and H. Zhang, Proceedings of the 30 th Annual ACM-SIAM Symposium on Discrete Algorithms, 2019, pp. 821--840; Y. Chang and T. Saranurak, Proceedings of the ACM Symposium on Principles of Distributed Computing, 2019, pp. 66--73] when $\Delta=\tilde{O}(n^{1/3})$.
Dawei Huang, Seth Pettie, Zhijun Zhang 0007
SIAM J. Comput.1
2020 The Communication Complexity of Set Intersection and Multiple Equality Testing
abstract
In this paper we explore fundamental problems in randomized communication complexity such as computing Set Intersection on sets of size k and Equality Testing between vectors of length k. Brody et al. [BCK+ 16] and Sağlam and Tardos [ST13] showed that for these types of problems, one can achieve optimal communication volume of O(k) bits, with a randomized protocol that takes O(log* k) rounds. They also proved [BCK+ 16, ST13] that this is one point along the optimal round-communication tradeoff curve. Aside from rounds and communication volume, there is a third parameter of interest, namely the error probability perr. It is straightforward to show that protocols for Set Intersection or Equality Testing need to send bits. Is it possible to simultaneously achieve optimality in all three parameters, namely communication and O(log* k) rounds? In this paper we prove that there is no universally optimal algorithm, and complement the existing round-communication trade-offs [BCK+ 16, ST13] with a new tradeoff between rounds, communication, and probability of error. In particular: Any protocol for solving Multiple Equality Testing in r rounds with failure probability perr = 2−E has communication volume Ω(Ek1/r). There exists a protocol for solving Multiple Equality Testing in r + log* (k/E) rounds with O(k + rEk1/r) communication, thereby essentially matching our lower bound and that of [BCK+ 16, ST13]. Lower bounds on Equality Testing extend to Set Intersection, for every r, k, and perr (which is trivial); in the reverse direction, upper bounds on Equality Testing for r, k, perr imply similar upper bounds on Set Intersection with parameters r + 1, k, and perr. Our original motivation for considering perr as an independent parameter came from the problem of enumerating triangles in distributed (CONGEST) networks having maximum degree Δ. We prove that this problem can be solved in O(Δ/log n + log log Δ) time with high probability 1 – 1/poly(n). This beats the trivial (deterministic) O(Δ)-time algorithm and is superior to the Õ(n1/3) algorithm of [CPZ19, CS19] when Δ = Õ(n1/3).
Dawei Huang, Seth Pettie, Zhijun Zhang 0007
SODA1
2019 Join on Samples: A Theoretical Guide for Practitioners
abstract
Despite decades of research on AQP (approximate query processing), our understanding of sample-based joins has remained limited and, to some extent, even superficial. The common belief in the community is that joining random samples is futile. This belief is largely based on an early result showing that the join of two uniform samples is not an independent sample of the original join, and that it leads to quadratically fewer output tuples. Unfortunately, this early result has little applicability to the key questions practitioners face. For example, the success metric is often the final approximation's accuracy, rather than output cardinality. Moreover, there are many non-uniform sampling strategies that one can employ. Is sampling for joins still futile in all of these settings? If not, what is the best sampling strategy in each case? To the best of our knowledge, there is no formal study answering these questions. This paper aims to improve our understanding of sample-based joins and offer a guideline for practitioners building and using real-world AQP systems. We study limitations of offline samples in approximating join queries: given an offline sampling budget, how well can one approximate the join of two tables? We answer this question for two success metrics: output size and estimator variance. We show that maximizing output size is easy, while there is an information-theoretical lower bound on the lowest variance achievable by any sampling strategy. We then define a hybrid sampling scheme that captures all combinations of stratified, universe, and Bernoulli sampling, and show that this scheme with our optimal parameters achieves the theoretical lower bound within a constant factor. Since computing these optimal parameters requires shuffling statistics across the network, we also propose a decentralized variant in which each node acts autonomously using minimal statistics. We also empirically validate our findings on popular SQL and AQP engines.
Dawei Huang, Dong Young Yoon, Seth Pettie, Barzan Mozafari
Proc. VLDB Endow.1
2017 Meta-activity recognition: A wearable approach for logic cognition-based activity sensing
abstract
Activity sensing has become a key technology for many ubiquitous applications, such as exercise monitoring and elder care. Most traditional approaches track the human motions and perform activity recognition based on the waveform matching schemes in the raw data representation level. In regard to the complex activities with relatively large moving range, they usually fail to accurately recognize these activities, due to the inherent variations in human activities. In this paper, we propose a wearable approach for logic cognition-based activity sensing scheme in the logical representation level, by leveraging the meta-activity recognition. Our solution extracts the angle profiles from the raw inertial measurements, to depict the angle variation of limb movement in regard to the consistent body coordinate system. It further extracts the meta-activity profiles to depict the sequence of small-range activity units in the complex activity. By leveraging the least edit distance-based matching scheme, our solution is able to accurately perform the activity sensing. Based on the logic cognition-based activity sensing, our solution achieves lightweight-training recognition, which requires a small quantity of training samples to build the templates, and user-independent recognition, which requires no training from the specific user. The experiment results in real settings shows that our meta-activity recognition achieves an average accuracy of 92% for user-independent activity sensing.
Lei Xie 0004, Wei Wang 0002, Dawei Huang
INFOCOM4
2017 Fully Dynamic Connectivity in O(log n(log log n)2) Amortized Expected Time
abstract
Computing the strongly connected Components (SCCs) in a graph $G=(V,E)$ is known to take only $O(m + n)$ time using an algorithm by Tarjan [SIAM J. Comput., 1 (1972), pp. 146--160] where $m = |E|$, $n=|V|$. For fully dynamic graphs, conditional lower bounds provide evidence that the update time cannot be improved by polynomial factors over recomputing the SCCs from scratch after every update. Nevertheless, substantial progress has been made to find algorithms with fast update time for decremental graphs, i.e., graphs that undergo edge deletions. In this paper, we present the first algorithm for general decremental graphs that maintains the SCCs in total update time $\tilde{O}(m)$, thus only a polylogarithmic factor from the optimal running time. (We use $\tilde{O}(f(n))$ notation to suppress logarithmic factors, i.e., $g(n) = \tilde{O}(f(n))$ if $g(n) = O(f(n) {polylog}(n)).$) Our result also yields the fastest algorithm for the decremental single-source reachability (SSR) problem which can be reduced to decrementally maintaining SCCs. Using a well-known reduction, we use our decremental result to achieve new update/query-time trade-offs in the fully dynamic setting. We can maintain the reachability of pairs $S \times V$, $S \subseteq V$ in fully dynamic graphs with update time $\tilde{O}(\frac{|S|m}{t})$ and query time $O(t)$ for all $t \in [1,|S|]$; this matches to polylogarithmic factors the best all-pairs reachability algorithm for $S = V$.
Shang-En Huang, Dawei Huang, Tsvi Kopelowitz, Seth Pettie
SODA2
2016 Disturbance observer-based robust control for trajectory tracking of wheeled mobile robots
Dawei Huang, Junyong Zhai, Wei-qing Ai, Shumin Fei
Neurocomputing1
2014 Preliminary study on diurnal variation of pulse signals in TCM
abstract
Objective: To compare the signals of pulse — diagnosis of healthy volunteer in different times in one day. Methods: After collecting the pulse waves of 42 healthy volunteers in 4 specific periods in one day, do pretreatment, parameter extracting basing on harmonic fitting, modeling, and identification by unsupervised learning and supervised learning with cross-validation step by step for analysis of the 4 groups. Finally, paired T-test and ANOVA were used in feature mining. Results: There are significant differences among the pulse-diagnosis signals of healthy volunteers in different times, and the accuracy rate is about 63%∼ 84%. Pulse rate, F1zuocun and F2zuocun, etc. are key features in classification. Conclusion: Signals of pulse-diagnosis in TCM of healthy human have circadian rhythmicity.
Nanyue Wang, Youhua Yu, Dawei Huang, Yanping Chen 0001, Xueli Yuan, Tongda Li
BIBM3
2013 Comparative study of pulse-diagnosis signals between 2 kinds of liver disease patients based on the combination of unsupervised learning and supervised learning
abstract
Objective: To compare the signals of pulse-diagnosis of 2 kinds of liver disease patients: Fatty Liver Disease (FLD) and Cirrhosis. Methods: After collecting the pulse waves of patients with Fatty Liver, Cirrhosis, we do pretreatment, parameter extracting basing on harmonic fitting, modeling, and identification by unsupervised learning and supervised learning with cross-validation step by step for analysis. Results: There is significant difference between the pulse-diagnosis signals of patients with FLD and Cirrhosis, and the result was confirmed by 3 analysis methods. The identification accurate is 72%–91%. Conclusion: Pulse waves collected from radial artery basing on the theory of TCM are specific for different pathological conditions. And the analysis methods we built in this study might offer some confidence for the realization of computer-aided diagnosis by pulse-diagnosis in TCM.
Nanyue Wang, Youhua Yu, Dawei Huang, Tongda Li, Zengyu Shan, Yanping Chen 0001, Liyuan Xue
BIBM3
2011 Live Migration of Multiple Virtual Machines with Resource Reservation in Cloud Computing Environments
abstract
Virtualization technology is currently becoming increasingly popular and valuable in cloud computing environments due to the benefits of server consolidation, live migration, and resource isolation. Live migration of virtual machines can be used to implement energy saving and load balancing in cloud data center. However, to our knowledge, most of the previous work concentrated on the implementation of migration technology itself while didn't consider the impact of resource reservation strategy on migration efficiency. This paper focuses on the live migration strategy of multiple virtual machines with different resource reservation methods. We first describe the live migration framework of multiple virtual machines with resource reservation technology. Then we perform a series of experiments to investigate the impacts of different resource reservation methods on the performance of live migration in both source machine and target machine. Additionally, we analyze the efficiency of parallel migration strategy and workload-aware migration strategy. The metrics such as downtime, total migration time, and workload performance overheads are measured. Experiments reveal some new discovery of live migration of multiple virtual machines. Based on the observed results, we present corresponding optimization methods to improve the migration efficiency.
Kejiang Ye, Xiaohong Jiang 0002, Dawei Huang, Jianhai Chen
IEEE CLOUD3
2011 An Optimal Hysteretic Control Policy for Energy Saving in Cloud Computing
abstract
The information and communication technology (ICT) industry has emerged as one of the major sources of world energy consumption due to its explosive growth. As a result, energy saving in ICT industry has attracted more attention. Meanwhile, cloud computing is becoming a disruptive technology with profound implications for ICT industry. Its emergence promises the on-demand provisioning of resources as a service. In this paper, we study the energy saving issue in cloud computing. In a scenario where a data center has multiple data servers to deal with jobs, the servers are switched into sleeping mode in periods of low traffic load to reduce energy consumption while guaranteeing the quality of service in terms of job blocking probability. The problem is formulated as a Markov decision process. It is proved that the optimal policy has a double threshold structure. Numerical and simulation results show that our proposed policy can significantly reduce the energy consumption.
Zexi Yang, Meng-Hsi Chen, Zhisheng Niu, Dawei Huang
GLOBECOM4
2011 Virt-LM: a benchmark for live migration of virtual machine
abstract
Virtualization technology has been widely applied in data centers and IT infrastructures, with advantages of server consolidation and live migration. Through live migration, data centers could flexibly move virtual machines among different physical machines to balance workloads, reduce energy consumption and enhance service availability.
Dawei Huang, Deshi Ye, Qinming He, Jianhai Chen, Kejiang Ye
ICPE1
2011 Iterative Soft Joint Detection Algorithm for OFDM Systems
abstract
In this paper, a novel iterative soft joint detection algorithm for orthogonal frequency division multiplexing (OFDM) systems is proposed. Without the need of channel estimation, it directly generates the a posteriori probabilities (APPs) of data symbols by very efficient computation. With the increase of iterations, the proposed algorithm gradually achieves near-optimal performance, while its complexity is only linear in the number of subcarriers.
Zhendong Luo, Dawei Huang
IEEE Trans. Inf. Theory2
2010 Analyzing and Modeling the Performance in Xen-Based Virtual Cluster Environment
abstract
Virtualization technology is currently widely used due to its benefits on high resource utilization, flexible manageability and powerful system security. However, its use for high performance computing (HPC) is still not popular due to the unclearness of the virtualization overheads. It's worthy to evaluate the virtualization cost and to find the performance bottleneck when running HPC applications in virtual cluster. We first evaluate the basic performance overheads due to virtualization. Then we create a 16-node virtual cluster and perform a performance evaluation for both para-virtualization and full virtualization. After that, we evaluate the MPI (Message Passing Interface) scalability to investigate the impact of MPI and network communication between virtual machines. In addition to the macro assessment, we use the Oprofile/Xenoprof to investigate the architecture characterization like CPU cycle, L2 cache misses, DTLB misses and ITLB misses which is an auxiliary explanation to the performance bottleneck. Experimental results indicate that performance overheads of virtualization are acceptable for HPC, para-virtualization is very suitable for HPC due to the high virtualization efficiency and efficient inter-domain communication. Finally, we use the non-linear regression modeling technology to present a performance model for network latency and bandwidth to predict the performance in virtual cluster environment.
Kejiang Ye, Xiaohong Jiang 0002, Siding Chen, Dawei Huang
HPCC4
2010 Two Optimization Mechanisms to Improve the Isolation Property of Server Consolidation in Virtualized Multi-core Server
abstract
Virtualization brings many benefits such as improving system utilization and reducing cost through server consolidation. However, it also introduces isolation problem when running multiple virtual machine workloads in one physical platform. Additionally, with the advent of multi-core technology, more and more cores are built into one die in today's data center that will share and compete for the resource like cache. It's worthy to study the isolation of server consolidation in modern multi-core platform. However, to our knowledge there are few work done on the isolation property especially the fault isolation property when one of the virtual machine workloads is attacked in server consolidation. In this paper, we study the isolation property from performance perspective and provide two optimization methods to improve the isolation property. We first define the isolation property and quantify the performance isolation in consolidation and propose a VM-level optimization method. Then we study the fault isolation by introducing a misbehavior virtual machine in server consolidation scenario and propose a core-level cache-aware optimization method to improve the fault isolation. Experimental results show that our two optimization methods can effectively improve the performance isolation and fault isolation with 29.39% and 19.52% respectively. What's more, Oprofile/Xenoprof toolkits are used to find out the factors affecting isolation property from the hardware events level.
Kejiang Ye, Xiaohong Jiang 0002, Deshi Ye, Dawei Huang
HPCC4
2010 Monotonic Directed Designs
abstract
The notion of a monotonic directed design was introduced to construct difference triangle sets by Chu, Colbourn, and Golomb [SIAM J. Discrete Math., 18 (2005), pp. 741–748]. In this paper, we describe various constructions for monotonic directed designs and establish the necessary and sufficient conditions for the existence of a monotonic directed design with block size 3, and with block size 4 leaving two definite exceptions and six possible exceptions.
Gennian Ge, Dawei Huang, Ying Miao 0001
SIAM J. Discret. Math.2
2009 Integrated MAP detection for V-BLAST systems without channel estimation
abstract
Generally, channel estimation is indispensable for signal detection in vertical Bell Laboratories layered space-time (V-BLAST) systems. However, channel estimation always cannot perfectly obtain the channel state information (CSI), and the inaccurate CSI may result in large performance degradation. In this paper, we propose a novel soft V-BLAST detector, called integrated maximum a posteriori (IMAP) detector. Unlike conventional detectors, the IMAP detector can avoid estimating the CSI and directly generate a posteriori probability of data symbols by efficient iterative computations. Moreover, we design a list-based fast searching algorithm to further reduce the complexity of the proposed detector in the cases of large numbers of transmit antennas or high order modulations. Computer simulation shows that the IMAP detector can approximately achieve the performance of the optimal detector using the perfect CSI.
Zhendong Luo, Dawei Huang
IEEE Trans. Wirel. Commun.2
2008 Joint MAP Detection for MIMO-OFDM Systems
abstract
In this paper, a joint maximum a posteriori (MAP) detection algorithm is proposed for multiple-input multiple-output orthogonal frequency division multiplexing (MIMO- OFDM) systems. Without the need of estimating the channel state information (CSI), it can directly compute the a posteriori probability (APP) of the transmitted data and generate the detection results under the MAP criterion by efficient iterative processing. Based on the idea of list sphere decoding, we also present a fast searching algorithm to further reduce its complexity in the cases of large numbers of transmit antennas and high data rates. Computer simulation shows that even without the help of channel estimation, the proposed algorithm with a few iterations can approximately attain the performance of the optimal detector using the perfect CSI.
Zhendong Luo, Fan Yang 0097, Dawei Huang
GLOBECOM3
2008 Integrated MAP detector for V-BLAST systems
abstract
In this paper, we propose a novel soft-decision detector, called integrated maximum a posteriori (IMAP) detector, for vertical Bell Laboratories layered space-time (V-BLAST) systems. The proposed detector can avoid estimating channel state information (CSI) and directly generate a posteriori probability of data symbols via efficient iterative computations. Moreover, a list-based fast searching algorithm is developed to further reduce the complexity of this detector in the cases of large numbers of transmit antennas or high data rates. Computer simulation shows that the IMAP detector can approximately achieve the performance of the optimal detector using the perfect CSI.
Zhendong Luo, Dawei Huang
ISIT2
2008 Optimal and robust MMSE channel estimation for MIMO-OFDM systems
abstract
In this paper, the minimum mean square error (MMSE) channel estimation for multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems is investigated. We first propose an optimal MMSE channel estimation algorithm, which can fully exploit the channel correlations over space, time, and frequency domains to estimate the channel frequency response. By the principle of maximum entropy, we then design an efficient robust MMSE channel estimation algorithm, which does not need to know spatial and time correlations and has a complexity only linear in the number of subcarriers. Moreover, computer simulation shows that the robust algorithm has no significant performance degradation in comparison with the optimal algorithm.
Zhendong Luo, Dawei Huang
PIMRC2
2008 General MMSE Channel Estimation for MIMO-OFDM Systems
abstract
In this paper, a general minimum mean square error (MMSE) channel estimation algorithm is proposed for multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems in spatially correlated multipath fading channels. It can make full use of the channel correlations in space, time, and frequency to estimate the channel state information for various systems, including pilot-symbol- assisted systems, pilot-embedded systems, and blind systems. To improve the computational efficiency of the proposed algorithm, we then derive a fast implementation algorithm with the complexity only linear in the number of subcarriers. Finally, we analyze the performance of the proposed algorithm for different systems in spatially independent and correlated channels.
Zhendong Luo, Dawei Huang
VTC Fall2
2008 Iterative Soft Multiuser Detection for MIMO MC-CDMA Systems
abstract
In this paper, an iterative soft multiuser detection (MUD) algorithm is proposed for multiple-input multiple-output multi-carrier code-division multiple access (MIMO MC-CDMA) systems. Based on an equivalent system model, an efficient soft MUD algorithm is first derived by the central limit theorem, the property of the complex Gaussian distribution, and fast matrix computation formulas. By iteratively using the soft MUD algorithm with some updated probability information, we then obtain the proposed iterative soft MUD algorithm. Finally, it is shown that this algorithm achieves near-optimal performance with very efficient iterative computations.
Zhendong Luo, Dawei Huang
VTC Fall2
2007 Lazy flooding: a new technique for information dissemination in distributed network systems
Caixia Chi, Dawei Huang, XiaoRong Sun
IEEE/ACM Trans. Netw.2
2005 Testing throughput computing interconnect topologies with Tbits/sec bandwidth in manufacturing and in field
abstract
The next generation of throughput computing systems designed by Sun Microsystems require interconnect with bandwidth of the order of Tbits/sec. Interconnect topologies for such high bandwidth are based on a few hundred SerDes I/Os on chips operating at multi-Gbps. Testing of these I/Os at only the component level is inadequate. In this paper, we describe the design-for-testability features for system manufacturing and on-line test of such I/Os and interconnect
Ishwar Parulkar, Dawei Huang, Leandro Chua Jr., Drew Doblar
ITC2
2002 Time-Series Investigation of Anomalous CRC Error Patterns in Fiber Channel Arbitrated Loops
Kenny C. Gross, Wendy Lu, Dawei Huang
ICMLA3
2002 Perceptual audio coding using adaptive pre- and post-filters and lossless compression
abstract
This paper proposes a versatile perceptual audio coding method that achieves high compression ratios and is capable of low encoding/decoding delay. It accommodates a variety of source signals (including both music and speech) with different sampling rates. It is based on separating irrelevance and redundancy reductions into independent functional units. This contrasts traditional audio coding where both are integrated within the same subband decomposition. The separation allows for the independent optimization of the irrelevance and redundancy reduction units. For both reductions, we rely on adaptive filtering and predictive coding as much as possible to minimize the delay. A psycho-acoustically controlled adaptive linear filter is used for the irrelevance reduction, and the redundancy reduction is carried out by a predictive lossless coding scheme, which is termed weighted cascaded least mean squared (WCLMS) method. Experiments are carried out on a database of moderate size which contains mono-signals of different sampling rates and varying nature (music, speech, or mixed). They show that the proposed WCLMS lossless coder outperforms other competing lossless coders in terms of compression ratios and delay, as applied to the pre-filtered signal. Moreover, a subjective listening test of the combined pre-filter/lossless coder and a state-of-the-art perceptual audio coder (PAC) shows that the new method achieves a comparable compression ratio and audio quality with a lower delay.
Gerald Schuller, Bin Yu 0001, Dawei Huang, Bernd Edler
IEEE Trans. Speech Audio Process.3
2001 Low Delay Perpetually Lossless Coding of Audio Signals
abstract
A novel predictive lossless coding scheme is proposed. The prediction is based on a new weighted cascaded least mean squared (WCLMS) method. To obtain both a high compression ratio and a very low encoding and decoding delay, the residuals from the prediction are encoded using either a variant of adaptive Huffman coding or a version of adaptive arithmetic coding. WCLMS is especially designed for music/speech signals. It can be used either in combination with psycho-acoustically pre-filtered signals to obtain perceptually lossless coding, or as a stand-alone lossless coder. Experiments on a database of moderate size and a variety of pre-filtered mono-signals show that the proposed lossless coder (which needs about 2 bit/sample for pre-filtered signals) outperforms competing lossless coders, such as ppmz, bzip2, Shorten, and LPAC, in terms of compression ratios. The combination of WCLMS with either of the adaptive coding schemes is also shown to achieve better compression ratios and lower delay than an earlier scheme combining WCLMS with Huffman coding over blocks of 4096 samples.
Sean Dorward, Dawei Huang, Serap A. Savari, Gerald Schuller, Bin Yu 0001
Data Compression Conference2
2001 Lossless coding of audio signals using cascaded prediction
abstract
A novel predictive lossless coding scheme is proposed. The prediction is based on a new weighted cascaded least mean squared (WCLMS) method. WCLMS is especially designed for music/speech signals. It can be used either in combination with psycho-acoustically pre-filtered signals to obtain perceptually lossless coding, or as a stand-alone lossless coder. Experiments on a database of moderate size and a variety of pre-filtered mono-signals show that the proposed lossless coder (which needs about 2 bit/sample for pre-filtered signals) outperforms competing lossless coders, WaveZip, Shorten, LTAC and LPAC, in terms of compression ratios.
Gerald Schuller, Bin Yu 0001, Dawei Huang
ICASSP3
2000 The stationary phase error distribution of a digital phase-locked loop
abstract
The stationary properties of a first-order digital phase-locked loop based on the extended Kalman filter (EKF-PLL) are investigated. A discrete Markov chain approximation of the phase error process is used to derive the asymptotic distribution of the phase error as well as the distribution of the time at which the EKF-PLL is first "out of lock", given that it is in a steady state to begin with. These approximations are compared with large computer simulations.
Glen Skiller, Dawei Huang
IEEE Trans. Commun.2
1999 Least squares estimation of polynomial phase signals via stochastic tree-search
abstract
Estimating the parameters for a constant amplitude, polynomial phase signal with additive Gaussian noise is considered. The difficulty in this problem is that there are many unobserved integers when a linear regression model is used for wrapped phases. Analysing the least squares target function based on the regression model, we use the differencing approach to simplify it. Thus a tree-search algorithm can be used to find the solution of the least squares problem. To reduce the computational complexity, statistical inference methods are applied. Then an attractive recursive algorithm is derived. Simulation results show that this algorithm works at a lower SNR than that for existing methods.
Dawei Huang, Simon Sando, Lian Wen
ICASSP1
1999 Sufficient output conditions for identifiability in blind equalization
abstract
The problem of input identifiability in blind deconvolution is considered where the input belongs to a known discrete alphabet. Input identifiability is an algorithm independent property, which does not necessarily imply channel identifiability. Sufficient conditions for input identifiability are derived in terms of algebraic relations on the observed output. It is shown how these new results relate to and unify other known sufficient conditions.
Dawei Huang, Fredrik Gustafsson
IEEE Trans. Commun.1
1998 Computing Joint Distributions of 2D Moving Median Filters With Applications to Detection of Edges
abstract
This paper derives the joint distribution of medians over moving windows in a two dimensional noisy image. The general formulation presented allows derivation of the probability distribution, needed to evaluate the probability of failing to detect an edge when present ("edge miss probability") and the probability of falsely detecting a nonexistent edge.
Dawei Huang, William T. M. Dunsmuir
IEEE Trans. Pattern Anal. Mach. Intell.1