VLDB 2026 Research / reviewers in the wild / expert
Weiguo Wu
dblp:78/1632
· DBLP profile ↗
66ranked-venue papers
6as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 45 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Computer networks · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV CacheabstractThe high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the memory pressure caused by KV Cache. However, existing methods either rely on static one-size-fits-all precision allocation or fail to dynamically prioritize critical KV in long-context tasks, forcing memory-accuracy-throughput tradeoffs. In this work, we propose a novel mixed-precision quantization method for KV Cache named KVmix. KVmix leverages gradient-based importance analysis to evaluate how individual Key and Value projection matrices affect the model loss, enabling layer-specific bit-width allocation for mix-precision quantization. It dynamically prioritizes higher precision for important layers while aggressively quantizing less influential ones, achieving a tunable balance between accuracy and efficiency. KVmix introduces a dynamic long-context optimization strategy that adaptively keeps full-precision KV pairs for recent pivotal tokens and compresses older ones, achieving high-quality sequence generation with low memory usage. Additionally, KVmix provides efficient low-bit quantization and CUDA kernels to optimize computational overhead. On LLMs such as Llama and Mistral, KVmix achieves near-lossless inference performance with extremely low quantization configuration (Key 2.19bit Value 2.38bit), while delivering a remarkable 4.9× memory compression and a 5.3× speedup in inference throughput. Fei Li 0042, Song Liu 0007, Weiguo Wu, Shiqiang Nie, Jinyu Wang 0002 |
AAAI | 3 |
| 2026 | EADA: Efficient adaptive data augmentation
Song Liu 0007, Weiguo Wu, Jinyu Wang 0002, Shiqiang Nie |
Comput. Vis. Image Underst. | 4 |
| 2026 | BLSA: A cache-aware balanced load scheduling approach on task graphs
Song Liu 0007, Fei Li 0042, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu |
Future Gener. Comput. Syst. | 6 |
| 2026 | FTL optimization for 3D NAND flash considering layer RBER variation and data precision
Shiqiang Nie, Ping She, Weiguo Wu |
Future Gener. Comput. Syst. | 3 |
| 2026 | FDSR: Efficient Model Training via Adaptive Tensor Quantization Based on Frequency Domain Division and Similarity Data ReuseabstractAs deep neural networks (DNNs) continue to grow in scale and complexity, GPU memory limitations have become a significant challenge for DNN model training, especially on resource-constrained commercial GPUs. While model quantization facilitates memory-efficient training, it often necessitates a tradeoff between quantization granularity and model accuracy. And quantization imposes additional computational overhead, which adversely affects the training throughput and apportions out the performance gains it brings. In this article, we propose FDSR, an adaptive tensor quantization method that leverages frequency domain division and similarity-based data reuse to break the memory bottleneck in visual model training. FDSR leverages the frequency-domain characteristics of tensors in terms of memory consumption and model accuracy, and proposes a fine-grained tensor quantization with different quantization bit-widths. It adaptively optimizes the quantization parameters according to model accuracy during training while employing sparsification according to data frequency-domain features, minimizing memory consumption and accuracy loss. To counteract the computational cost, FDSR incorporates a novel similarity-based reuse strategy that avoids redundant quantization/dequantization computations, further enhanced by a tailored Locality-Sensitive Hashing (LSH) mechanism and optimized kernels. Experimental results demonstrate that FDSR achieves an average of 10.20× activation memory compression with only 1.10% average accuracy loss across various models on the commercial GPU. Compared to the state-of-the-art quantization methods, FDSR improves memory optimization by up to 68.6% and increases throughput by up to 25.55%, with consistent performance improvements on different GPU architectures. Song Liu 0007, Fei Li 0042, Qin Xia, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu |
ACM Trans. Archit. Code Optim. | 7 |
| 2026 | IFFS: An Interlaced Magnetic Recording Friendly File SystemabstractRecently, the emerging Interlaced Magnetic Recording (IMR) technology has substantially improved the areal density of disks by implementing interlaced track layout. While this track layout enhances disk storage capacity, it impairs the flexibility of write positioning. To maintain stable write performance of IMR disks, operations such as Read-Modify-Write (RMW) or Garbage Collection (GC) must be introduced, which inevitably incur extra I/Os. Especially in write-intensive workloads, the excessive additional I/Os exacerbate the write amplification effect, thereby severely degrading the overall performance of IMR disks. Although existing data management strategies strive to reduce extra I/Os via device drivers or system middleware, the semantic disparity between the disk and file system inherently limits these strategies to achieve optimal performance. To address the aforementioned challenge,this paper proposes IMR-Friendly File System (IFFS), an innovative file system tailored for IMR disks. First, we propose a semantic-aware hotness identification algorithm based on Online K-means, which redefines the data hotness metric by exploiting file system semantics to reduce data migration induced by inaccurate data classification. Second, we introduce an I/O-overhead-minimized data placement strategy that adaptively selects between in-place and out-of-place writing modes based on data hotness metrics. Furthermore, this strategy employs a log-transition mechanism to dynamically adjust write positions, effectively mitigating I/O overhead caused by RMWs and scattered read requests. Finally, we implement a multi-factor garbage collection mechanism that incorporates intra-zone data hotness, data layout, fragmentation levels, and other contextual attributes to optimize file system data management efficiency, thereby enhances the read performance on IMR disk-based storage system performance. Experimental results show that IFFS achieves an average bandwidth improvement of 12.96× over EXT4, XFS, F2FS, and the state-of-the-art work in Fio evaluation. Under YCSB workloads, IFFS improves bandwidth by an average of 56.89% while reducing latency by 34.46%. Fangxing Yu, Chi Zhang 0095, Shiqiang Nie, Zhike Li, Weiguo Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | ZeroCopy: file system assisted container buffer migration in cloud computing systemabstractAbstract In cloud computing data centers, containerized tasks are regularly scheduled from one physical host to another due to resource management requirements such as handling machine failures, rebalancing server resources, and upgrading/scaling applications. After the container running in the source host is scheduled to the target host, it suffers from I/O performance degradation until the DRAM buffer is fully rebuilt. However, migrating the DRAM buffer from the source host to the target host could also introduce intolerable downtime of containerized tasks. Especially, as the DRAM buffer capacity of the application already increases to about dozens or hundreds of GB, the cost of downtime due to container migration becomes unacceptable. Many researchers have devoted themselves to developing an effective DRAM buffer warm-up scheme to avoid the cold bootstrap issue after container migration, such as pre-copy and post-copy schemes. However, the cold bootstrap and large-capacity buffer migration issues of container scheduling are still an open research problem. In this paper, motivated by the observation that the DRAM buffer is always flushed to the storage backend before starting the container in the target host, we proposed a scheme named ZeroCopy to utilize the file system to assist the DRAM buffer migration. ZeroCopy traverses the files in the DRAM buffer and flags these files when these files are flushed into the file system, and reloads these files into DRAM after starting the container in the target host. By this scheme, the container migration procedure does not require migrating data buffers and can start within an acceptable time. We conduct a series of experiments with public cloud traces to measure several key metrics on container migration. The results show that ZeroCopy outperforms these existing schemes. The average data transmission volume is reduced by about 6.25 times compared with state-of-the-art, and the downtime of container migration is also reduced by 31.8%. Shiqiang Nie, Tingshen Ruan, Ruijia Chen, Song Liu 0007, Weiguo Wu |
CCF Trans. High Perform. Comput. | 6 |
| 2025 | Olsync: Object-level tiering and coordination in tiered storage systems based on software-defined network
Zhike Li, Shiqiang Nie, Jinyu Wang 0002, Chi Zhang 0095, Fangxing Yu, Zhankun Zhang, Song Liu 0007, Weiguo Wu |
Future Gener. Comput. Syst. | 9 |
| 2025 | Time-constrained persistent deletion for key-value store engine on ZNS SSD
Shiqiang Nie, Jie Niu, Qihan Hu, Song Liu 0007, Weiguo Wu |
Future Gener. Comput. Syst. | 6 |
| 2025 | ZoomDB: Building cost-effective key-value store engine on ZNS SSD and SMR HDD
Shiqiang Nie, Chi Zhang 0095, Fangxing Yu, Yaming Li, Weiguo Wu |
J. Syst. Archit. | 6 |
| 2025 | DTB+: An enhanced data management strategy for efficient RMW reduction in IMR drivesabstractThe emerging Interlaced Magnetic Recording (IMR) technology not only achieves higher storage density than SMR, but also significantly reduces rewrite overhead by dividing tracks into bottom and top tracks and organizing them in an interlaced fashion. However, frequent updates to the bottom track can trigger a large number of Read-Modify-Write (RMW) operations during high disk space utilization, which can severely degrade the I/O performance. Addressing this issue, this paper proposes an interlaced translation layer named DTB+ to improve the write performance of IMR disks. Firstly, a workload-sensitive track heat analysis mechanism is introduced to intelligently place data to reduce track rewrite probability. Simultaneously, the zero-incremental cost region is selectively used to construct a twin-buffer architecture to reduce RMW operations. In addition, an adaptive space allocation engine based on reinforcement learning was developed to flexibly allocate and reclaim space within the twin-buffer, improving disk resource utilization . Finally, establish a flexible evicted-data transfer zone to delay the writeback operations of interference data, further reducing the additional overhead. Experimental results indicate that compared with the state-of-the-art studies, DTB+ can reduce RMWs by 63.00% and additional I/O operations by 57.41%, decrease the average write latency by 37.77%, and lower the tail latency by 53.95%. Fangxing Yu, Chi Zhang 0095, Zhike Li, Shiqiang Nie, Weiguo Wu |
J. Syst. Archit. | 6 |
| 2025 | Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache HierarchyabstractFor scientific computing applications that consist of many loosely coupled tasks, efficient scheduling is critical to achieve high performance and good quality of service (QoS). One of the challenges for co-running tasks is the frequent contention for shared cache hierarchy of multi-core processors. Such contention significantly increases cache miss rate and therefore, results in performance deterioration for computational tasks. This paper presents Scalpel, a contention-aware task grouping and co-scheduling approach for efficient task scheduling on shared cache hierarchy. Scalpel utilizes the shared cache access features of tasks to group them in a heuristic way, which reduces the contention within groups by achieving equal shared cache locality, while maintaining load balancing between groups. Based thereon, it proposes a two-level scheduling strategy to schedule groups to processors and assign tasks to available cores in a timely manner, while considering the impact of task scheduling on shared cache locality to minimize task execution time. Experiments show that Scalpel reduces the shared cache miss rate by up to 2.14× and optimizes the execution time by up to 1.53× for scientific computing benchmarks, compared to several baseline approaches. Song Liu 0007, Zengyuan Zhang, Xinhe Wan, Bo Zhao 0019, Weiguo Wu |
IEEE Trans. Computers | 6 |
| 2025 | DPUSwap: building an infinite swap with DPU for cloud computing system
Shiqiang Nie, Jianqiang Ma, Jiaxin Shi, Weiguo Wu |
J. Supercomput. | 6 |
| 2025 | Constructing a scalable key-value store engine on multidisk system
Shiqiang Nie, Jie Niu, Fangxing Yu, Jianqiang Ma, Xingxing Zhu, Weiguo Wu |
J. Supercomput. | 6 |
| 2025 | Adaptive Read Level Recording for Read Performance Improvement in 3-D NAND FlashabstractWhile low-density parity-check code has been adopted in 3-D flash for improving chip reliability, it suffers from severe read latency due to the increasing number of read retries. Recent studies propose read-level recording to mitigate the performance loss from failed read retries. However, existing schemes induce large updating overhead and achieve suboptimal results, making it critical to develop better tradeoffs among storage overhead, process variation, and performance improvement. In this article, we propose AR$^{2}$, an adaptive read-level recording scheme to improve read performance for 3-D NOT AND (NAND) flash. It consists of two designs: AR$^{2}$-win and AR$^{2}$-pre. AR$^{2}$-win records the number of read levels that fit the majority of the last$N$reads, which prevents the worst page from dominating the read level recording. AR$^{2}$-pre predicts the number of read levels for the next read based on the recorded one and a simple machine learning model, which prevents using stale recorded levels in large-capacity solid-state drives (SSDs). Our experimental results show that AR$^{2}$significantly improves the read performance for 3-D NAND flash and achieves on average 15% or more read latency reduction over the state-of-the-art. Shiqiang Nie, Zhike Li, Fangxing Yu, Song Liu 0007, Weiguo Wu |
IEEE Trans. Reliab. | 5 |
| 2024 | DTB: A Novel Reinforcement Learning-Assisted Data Management Strategy in Interlaced Magnetic RecordingabstractShingled Magnetic Recording (SMR) technology, employing a shingled track layout, has significantly enhanced areal density capability. However, this layout imposes severe write penalties when dealing with non-sequential writes. The emerging Interlaced Magnetic Recording (IMR) technology not only achieves higher storage density than SMR, but also significantly reduces rewrite overhead by dividing all tracks into bottom and top tracks and organizing them in an interlaced fashion. However, frequent updates to the bottom track can trigger a large number of Read-Modify-Write (RMW) operations during high disk space utilization, which can severely affect the I/O performance of the disk. Addressing this issue, this paper proposes an interlaced translation layer named DTB to improve the write performance of IMR disks. Firstly, a workload-sensitive track heat analysis mechanism is introduced to intelligently place data to reduce track rewrite probability. Simultaneously, the zero-incremental cost region is selectively used to construct a Twin-Buffer architecture to reduce RMW operations triggered by frequent writeback of hot data, thereby effectively curtailing the rewriting overhead. In addition, an adaptive space allocation engine was developed by analyzing the data characteristics in the buffer, and we designed a dynamic configuration model based on reinforcement learning to flexibly allocate and reclaim space within the Twin-Buffer, improving disk resource utilization and I/O performance. Experimental results indicate that DTB can reduce the number of RMWs by 63.45%, decrease the average write latency by 44.37%, and lower the tail latency by 56.84% compared with state-of-the-art studies. Fangxing Yu, Chi Zhang 0095, Zhike Li, Shiqiang Nie, Weiguo Wu |
HPCC | 6 |
| 2024 | DIR: Dynamic Request Interleaving for Improving the Read Performance of Aged Solid-State Drives
Shiqiang Nie, Weiguo Wu |
J. Comput. Sci. Technol. | 3 |
| 2023 | An efficient computation offloading and resource allocation algorithm in RIS empowered MEC
Xiangjun Zhang, Weiguo Wu, Song Liu 0007, Jinyu Wang 0002 |
Comput. Commun. | 2 |
| 2023 | TurboStencil: You only compute once for stencil computation
Song Liu 0007, Xinhe Wan, Zengyuan Zhang, Bo Zhao 0019, Weiguo Wu |
Future Gener. Comput. Syst. | 5 |
| 2023 | DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUsabstractStencil computation is an important class of computational modes in scientific computing applications. Loop tiling techniques have been widely studied to accelerate stencil computations on different architectures by exploiting parallelism and data locality. Recent advanced tiling methods enable the tile-wise concurrent start-up to improve the execution performance. However, such methods statically partition all dimensions of iteration space into tiles with predetermined complex shapes and sizes, and thus lead to low thread utilization and memory access efficiency on GPUs. In this paper, we present DHTS, a novel dynamic hybrid tiling strategy for stencil computations. DHTS employs static tiling on the outer dimensions to achieve concurrent start-up parallelism, while proposes a dynamic rectangular tiling method on the inner dimensions to improve thread utilization and memory access efficiency. By deriving tile size constraints, DHTS adaptively achieves equal-size workload of tiles, and therefore reducing idle threads and increasing coalesced memory accesses within tiles. We implement the proposed strategy with different complex tile shapes. Experimental results on Titan V and Tesla V100 GPUs show that DHTS effectively improves the execution performance of 2D/3D stencils compared to state-of-the-art tiling methods, and achieves the best improvement of 28×. Song Liu 0007, Zengyuan Zhang, Weiguo Wu |
IEEE Trans. Computers | 3 |
| 2023 | MCB: a multidevice cooperative buffer management strategy for boosting the write performance of the SSD-SMR hybrid storage
Chi Zhang 0095, Shiqiang Nie, Jinyu Wang 0002, Song Liu 0007, Weiguo Wu |
J. Supercomput. | 5 |
| 2023 | BiLSTM-based Federated Learning Computation Offloading and Resource Allocation Algorithm in MECabstractMobile edge computing (MEC) driven by 5G cellular systems has recently emerged as a promising paradigm, enabling mobile devices (MDs) with limited computing resources to offload various computation-intensive tasks (such as autopilot, online game) to edge servers to enhance the data processing capabilities of MDs. However, the uncertainty of wireless channel state and data volume of offloading tasks, as well as the data security privacy of offloading tasks, bring serious challenges to computation offloading in MEC. In this article, we consider a time-varying MEC scenario and formalize the delay and energy consumption during the computation offloading process as a joint optimization problem. Then the optimization problem is decomposed into two sub-problems: intelligent task prediction and resource allocation. Different from traditional methods, we improve the federated learning (FL) algorithm and propose a thoughtful cloud-edge-client FL task prediction mechanism based on Bidirectional Long Short-Term Memory. Each participating MD trains the model locally without uploading data to the server, and periodically aggregates the model in the edge and in the cloud. The algorithm both eliminates the need to solve complex optimization problems and ensures user privacy security. Finally, experimental results show that our proposed algorithm significantly outperforms other benchmark algorithms in energy efficiency. Xiangjun Zhang, Weiguo Wu, Jinyu Wang 0002, Song Liu 0007 |
ACM Trans. Sens. Networks | 2 |
| 2022 | Status, challenges and trends of data-intensive supercomputing
Jia Wei 0002, Pei Ren, Yujia Lei, Yuqi Qu, Qiyu Jiang, Xiaoshe Dong, Weiguo Wu, Qiang Wang 0062, Xingjun Zhang |
CCF Trans. High Perform. Comput. | 9 |
| 2021 | SLA: A Cache Algorithm for SSD-SMR Storage System with Minimum RMWs
Xuda Zheng, Chi Zhang 0095, Keqiang Duan, Weiguo Wu |
ICA3PP (3) | 4 |
| 2021 | DUPRFloor: Dynamic Modeling and Floorplanning for Partially Reconfigurable FPGAsabstractNowadays, field-programmable gate array (FPGA) devices have been widely used in various fields. However, modules of circuits to be executed on FPGAs are placed within rectangular reconfigurable regions (RRs) with current floorplanners, leading to internal fragments, and lower utilization of resources. To address this, a dynamic description model of RRs and the corresponding floorplanner named dynamic union partial reconfiguration floorplan (DUPRFloor) are proposed in this article. The RR dynamic description modeling adds an anchor within a rectangular RR to reduce internal fragments. In this way, the modeling can represent both rectangular and nonrectangular shapes. Then, to find the optimal anchor, a clipping method is devised by constraining width and height of the candidate region. Finally, the mixed-integer linear programming (MILP) is used to optimize an objective function which considers the resources utilization and communication costs to obtain a desirable floorplanning result. The proposed method has been validated by simulation on three kinds of devices. And experimental results show that reconfigurable resources can be saved as much as 19.16% compared to rectangular modeling method. The DUPRFloor is also validated on the Microelectronics Center of North Carolina standard benchmark data sets. Results show that DUPRFloor can reduce 18.65% global wire length at most with almost the same execution time compared to state-of-the-art algorithms. Our approach is tested on a FPGA implemented software-defined radio (SDR) and reduced 29.41% wasted configurable frames, and to the overall design, 2% configurable frames are saved at most. Jinyu Wang 0002, Yifei Kang, Weiguo Wu, Guoliang Xing, Linlin Tu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Data Pattern Aware Reliability Enhancement Scheme for 3D Solid-State Drivesabstract3D charge-trap (CT) NAND flash-based SSD has been used widely for its large capacity, low cost per bit, and high endurance. One-shot program (OSP) scheme, as a variation of incremental step pulse programming (ISPP) scheme, has been employed to program data for CT flash, whose program unit is the Word-Line (WL) instead of the page. The existing program optimization schemes either make trade-offs among program latency and reliability by adjusting the program step voltage on demand; or remap the most error-prone cell states to others by re-encoding programmed data. However, the data pattern, which represents the ratio of 1s in data values, has not been thoroughly studied. In this paper, we observe that most small files do not contain uniform 1s and 0s among these common file types (i.e., image, audio, text, executable file), leading to programming WL cells in different states unevenly. Some cell states dominate over the WL, while others are not. Based on this observation, we propose a flexible reliability enhancement scheme based on the OSP scheme. This scheme programs the cells into different states with varied , i.e., these cells in one state, whose number is the largest in one WL, are programmed with a fine-grained (namely slow write). In contrast, the minority are programmed with a coarse-grained (namely fast write). So the reliability is improved due to averaging the major enhanced cells with the minor degraded cells without program latency overhead. A series of experiments have been conducted, and the results indicate that the proposed scheme achieves 34% read performance improvement and 16% lifetime elongation on average. Shiqiang Nie, Weiguo Wu, Chi Zhang 0095 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2020 | Layer RBER Variation Aware Read Performance Optimization for 3D Flash Memoriesabstract3D NAND flash enables the construction of large capacity Solid-State Drives (SSDs) for modern computer systems. While effectively reducing per bit cost, 3D NAND flash exhibits non-negligible process variations and thus RBER (raw bit error rate) difference across layers, which leads to sub-optimal read performance for applications with either small or large I/O requests. In this paper, we propose LRR, Layer RBER variation aware Read optimization schemes, to address the challenge. LRR consists of two schemes - LRR subpage read scheduling (SRS) and LRR fullpage allocation (FPA). SRS groups small read requests from the layers with similar RBERs to reduce the average read latency of subpage sized read requests. FPA distributes the data of a large write to multiple layers, which improves the read latency when reading from layers with large RBERs. Our experimental results show that our proposed scheme LRR reduces 46% read latency on average over the state-of-the-art. Shiqiang Nie, Youtao Zhang, Weiguo Wu, Jun Yang 0002 |
DAC | 3 |
| 2020 | Relevance assignation feature selection method based on mutual information for machine learning
Liyang Gao, Weiguo Wu |
Knowl. Based Syst. | 2 |
| 2020 | Structured mesh-oriented framework design and optimization for a coarse-grained parallel CFD solver based on hybrid MPI/OpenMP programming
Xiaoshe Dong, Nianjun Zou, Weiguo Wu, Xingjun Zhang |
J. Supercomput. | 4 |
| 2020 | Harnessing Hardware Defects for Improving Wireless Link PerformanceabstractThe design trade-offs of transceiver hardware are crucial to the performance of wireless systems. In this paper, we present an in-depth study to characterize the surprisingly notable systemic impacts of low-pass filter (LPF) design, which is a small yet indispensable component used for shaping spectrum and rejecting interference. Using a bottom-up approach, we examine how signal-level distortions caused by the trade-off of LPF design propagate to the upper-layers of wireless communication, reshaping bit error patterns and degrading link performance of today's 802.11 systems. Moreover, we propose a novel algorithm that harnesses LPF defects for improving video streaming, which substantially enhances video quality in mobile environments. Alireza Ameli Renani, Jun Huang 0001, Guoliang Xing, Abdol-Hossein Esfahanian, Weiguo Wu |
IEEE/ACM Trans. Netw. | 5 |
| 2019 | Accelerating Lattice Boltzmann Method by Fully Exposing Vectorizable Loops
Bin Qu, Song Liu 0007, Jiajun Yuan, Weiguo Wu |
ICA3PP (1) | 6 |
| 2019 | Revisiting the Parallel Strategy for DOACROSS Loops
Song Liu 0007, Yuanzhen Cui, Nianjun Zou, Weiguo Wu |
J. Comput. Sci. Technol. | 6 |
| 2018 | ShadowGC: Cooperative garbage collection with multi-level buffer for performance improvement in NAND flash-based SSDsabstractGarbage collection, an essential background activity in NAND flash based SSDs, often introduces large runtime overhead. Recent studies showed that it is beneficial to separate the flash pages that have dirty copies in the write buffers from those that do not. However, the existing schemes exploring this observation have limitations, which prevent them from maximizing the performance improvement. In this paper, we address the above challenge through ShadowGC, a novel GC design that exploits the pages in both host-side and device-side write buffers and adopts different read and write strategies to minimize the GC overhead. When garbage collecting flash pages that have dirty copies in the device-side write buffer, ShadowGC reads data from the write buffer. When garbage collecting flash pages that have dirty copies in the host-side write buffer, ShadowGC moves them to dedicated blocks and speeds up the movement with fast-write operations. Our experimental results show that, on average, ShadowGC reduces the write amplification by 16.2% and the GC latency by 20.5% over the state-of-the-art. Jinhua Cui 0001, Youtao Zhang, Jianhang Huang, Weiguo Wu, Jun Yang 0002 |
DATE | 4 |
| 2018 | A Dynamic Parallel Strategy for DOACROSS LoopsabstractMany parallelization methods work on exposing the pipeline/wave-front parallelism of DOACROSS loops through loop transformations. However, these methods statically assign iterations to available threads for parallel execution, and thus causing the waste of computing resources in synchronization among threads, especially in a multithreading environment. This paper proposes a brand-new parallel strategy that achieves wave-front parallelism with reduced dependences and provides dynamic tile assignment for DOACROSS loops, which has better ability to avoid threads from waiting in synchronization and utilize computing resources. The experimental results demonstrate that the proposed strategy outperforms two advanced strategies which are based on implicit barriers and POST/WAIT operations over six benchmarks on a multi-core server. The strategy also has better scalability for the increasing number of threads. Yuanzhen Cui, Song Liu 0007, Nianjun Zou, Weiguo Wu |
HPC Asia | 4 |
| 2018 | Kinetic Energy Attenuation Method for Posture Balance Control of Humanoid Biped Robot under Impact DisturbanceabstractFor the situation of the robot's foot soles rotating around their edges, the current posture balance control methods only used linear controllers and ignored the system's nonlinearity. For this problem, we propose a posture balance control method, which actively attenuate kinetic energy of the robot. Based on simplified dynamics model established for the foot rotating situation, our control law is deduced to adjust the amplitude of the ground contact force (GCF) and the angular acceleration of the robot foot rotation simultaneously. Then the posture balance controller against the impact force disturbance is designed. Stability of the foot rotation phase space is analyzed for our controller and other two existing controllers. Also, the situation of no balance controller is analyzed as a presentation of the system's passive characteristic. Then area of the stable region is compared between these three controllers. Simulations under frontal impact disturbances are also conducted. The results demonstrate that, our controller is able to resist stronger impact disturbance and make the robot recover its balance in a shorter time. This indicates that, compared with other posture control methods, the proposed control method makes better use of the robot joint motion to attenuate the robot's kinetic energy. Liyang Gao, Weiguo Wu |
IECON | 2 |
| 2018 | An efficient tile size selection model based on machine learning
Song Liu 0007, Yuanzhen Cui, Weiguo Wu |
J. Parallel Distributed Comput. | 5 |
| 2018 | ApproxFTL: On the Performance and Lifetime Improvement of 3-D NAND Flash-Based SSDsabstract3-D NAND flash is one of the most prospective advances in flash memory industry. While 3-D flash improves cell density and reduces lithography cost through die stacking, it suffers from severe program disturbance, which leads to significant performance and lifetime degradation for 3-D flash-based SSDs. To address the above challenge, we propose ApproxFTL, an approximate-write aware flash translation layer design, that uses approximate-write operations to store error-resilient data of modern applications. By reducing the maximal threshold voltage and tightening the guard bands between multilevel cell states, approximate write operations not only finish early but also exhibit large disturbance reduction, which can be exploited to alleviate disturbance in physical blocks that save both precise and approximate data. ApproxFTL maximizes the disturbance mitigation through approximate-write aware data placement, wear leveling, and garbage collection enhancements. Our experimental results show that ApproxFTL, while preserving high data quality, improves the read and write response time of flash accesses by 41.38% and 45.64% on average, respectively, and extends the lifetime of 3-D flash-based SSDs by 5.75% when comparing to the state-of-the-art. Jinhua Cui 0001, Youtao Zhang, Liang Shi 0001, Chun Jason Xue, Weiguo Wu, Jun Yang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | DLV: Exploiting Device Level Latency Variations for Performance Improvement on Flash Memory Storage SystemsabstractNAND flash has been widely adopted in storage systems due to its better read and write performance and lower power consumption over traditional mechanical hard drives. To meet the increasing performance demand of modern applications, recent studies speed up flash accesses by exploiting access latency variations at the device level. Unfortunately, existing flash access schedulers are still oblivious to such variations, leading to suboptimal I/O performance improvements. In this paper, we propose DLV, a novel flash access scheduler for exploring scheduling opportunities due to device level access latency variations. DLV improves flash access speeds based on process variations and data retention time difference across flash blocks. More importantly, DLV integrates access speed optimization with access scheduling such that the average access response time can be effectively reduced on flash memory storage systems. Our experimental results show that DLV achieves an average of 41.5% performance improvement over the state-of-the-art. Jinhua Cui 0001, Youtao Zhang, Weiguo Wu, Jun Yang 0002, Yinfeng Wang, Jianhang Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | An Efficient Locality-Aware Task Assignment Algorithm for Minimizing Shared Cache ContentionabstractTask scheduling can improve the performance of parallel execution through optimizing the utilization of on-chip computing resources, and thus it has been widely studied. Most of the previous work uses data access locality to predict cache behaviors for task scheduling, but usually suffering accuracy and computational time complexity issues. This paper proposes an efficient task assignment algorithm to minimize the contention for shared caches on multi-core processors among parallel independent process level tasks. The proposed algorithm leverages the property of footprint to approximately estimate the locality parameter of parallel tasks, choosing the best grouping of tasks with minimum locality value in a quick way for task assignment. The calculation time is therefore significantly reduced and the algorithm complexity is O(nlog2n). Meanwhile, the algorithm accuracy is very high. On an Intel 8 cores dual-processor system, the experimental results show that the task assignment algorithm achieves over 99% of the actual optimal performance on average and outperforms the default Linux task scheduling method by an average of over 5% for two sets of different parallel tasks. Song Liu 0007, Xiao Xie, Yuanzhen Cui, Weiguo Wu |
PDCAT | 4 |
| 2016 | Exploiting latency variation for access conflict reduction of NAND flash memoryabstractNAND flash memory has been widely used in storage systems by offering greater read/write performance and lower power consumption than mechanical hard drives. Recently, the tradeoff between endurance, write speed, and read speed has been exploited from many ways for I/O performance improvement, which also induce the read/write latency variation. In this paper, the latency variation is exploited in I/O scheduling for access characteristic guided read and write latency minimization. First, with the understanding of the relationship among read latency, write latency and raw bit error rates (RBER), different ways to exploit the relationship for read and write latency reduction is discussed. Then, an I/O scheduling scheme is proposed by using hotness and retention age of accessed data to determine the speed of writes or reads, giving scheduling priority to fast writes and fast reads for conflict reduction. Experiments with various traces reveal that the proposed technique achieves significant read and write performance improvements. Jinhua Cui 0001, Weiguo Wu, Xingjun Zhang, Jianhang Huang, Yinfeng Wang |
MSST | 2 |
| 2016 | VIOS: A Variation-Aware I/O Scheduler for Flash-Based Storage Systems
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang |
NPC | 2 |
| 2016 | Exploiting Cross-Layer Hotness Identification to Improve Flash Memory System Performance
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang |
NPC | 2 |
| 2016 | A Statistics Based Prediction Method for Rendering Application
Qian Li 0013, Weiguo Wu, Jianhang Huang, Mingxia Feng |
NPC | 2 |
| 2015 | Economy-Oriented Deadline Scheduling Policy for Render System Using IaaS Cloud
Qian Li 0013, Weiguo Wu, Zeyu Sun 0002, Jianhang Huang |
ICA3PP (3) | 2 |
| 2015 | Thread Count Prediction Model: Dynamically Adjusting Threads for Heterogeneous Many-Core SystemsabstractDetermining an appropriate thread count for a multithread application running on a heterogeneous many-core system is crucial for improving computing performance and reducing energy consumption. This paper investigates the interrelation between thread count and computing performance of applications, and designs a prediction model of the optimum thread count on the basis of Amdahl's law combined with regression analysis theory to improve computing performance and reduce energy consumption. The prediction model can estimate the optimum tread count relying on the program running behaviors and the architecture characteristics of heterogeneous many-core system. Using the estimated optimum thread count, the number of the active hardware threads and processing cores on the many-core processor is dynamically adjusted in the process of thread mapping to improve the energy efficiency of entire heterogeneous many-core system. The experimental results show that, using this paper proposed thread count prediction model, on an average, the computing performance is improved by 48.6%, energy consumption is reduced by 59%, and additional overhead introduced is 2.03% compared with that of the traditional thread mapping for the PARSEC benchmark programs run on an Intel MIC heterogeneous many-core system. Tao Ju 0002, Weiguo Wu, Heng Chen 0002, Zhengdong Zhu, Xiaoshe Dong |
ICPADS | 2 |
| 2014 | A Utility-Maximizing Tasks Assignment Method for Rendering Cluster SystemabstractA utility-maximizing tasks assignment method for rendering cluster system based on feedback is proposed to solve the problem that traditional task-centered assignment method and naïve load balancing strategy cannot make full use of resources. The method uses feedback on resource usage to choose an appropriate number of threads for the renderer and then divides computing nodes of the rendering cluster system into fine grain computing units. After that, the method takes advantage of frame-to-frame coherence to assign tasks to computing units with a new static load balancing strategy. Experiments on two scene models with different complexity and comparisons with the naïve method and the fixed-threads method show that the utility-maximizing method not only reduces the rendering time of a rendering job, but also balances the load between computing units. The scalability of the proposed method is also verified on different number of computing nodes. Jianhang Huang, Weiguo Wu, Qian Li 0013 |
ISPA | 2 |
| 2013 | Coverage Algorithm in Wireless Sensor Network Base on Point Set OptimizationabstractDuring the process that wireless sensor network try to cover its target area, a lot of redundant nodes are generated, resulting in excessive network energy consumption and incomplete node coverage. Against this problem, an efficient covering algorithm based on point set optimization is proposed. In the algorithm, the Gaussian normal density function and coverage area probability function are utilized during optimization of the point set, so that an optimal point set which meets certain coverage requirements is given through the quantitative relationship between the node sensing radius and number of nodes in the network, and furthermore, the network resource is optimized, the lifetime and Qos of the network are improved, and the network overhead is also reduced. Simulation result shows that: during test results against different coverage area and comparison to SCCP algorithm, the effectiveness and stability of the algorithm are verified. Zeyu Sun 0002, Weiguo Wu |
MSN | 2 |
| 2011 | Stable Adaptive Work-Stealing for Concurrent Multi-core Runtime SystemsabstractThe proliferation of multi-core architectures has led to explosive development of parallel applications using programming models, such as OpenMP, TBB, and Cilk, etc. With increasing number of cores, however, it becomes harder to efficiently schedule parallel applications on these resources since current multi-core runtime systems still lack efficient mechanisms to support collaborative scheduling of these applications. In this paper, we study feedback-driven adaptive scheduling based on work stealing, which provides an efficient solution for concurrently executing a set of applications on multi-core systems. To dynamically estimate the number of cores desired by each application, a stable feedback algorithm, called A-Deque, is proposed using the length of active deques, which more precisely captures the parallelism variation of the applications. Furthermore, a prototype system is built by extending the Cilk runtime system, and the experimental results show that feedback-driven scheduling algorithms have more advantages for scheduling parallel applications with dynamic changing parallelism, and better overall performances are achieved with more accurate and stable feedback mechanism. Compared with existing algorithms, A-Deque improves the performances by up to 19.13\% and 28.96\% with respect to average response time and processor utilization respectively. Yangjie Cao, Hongyang Sun 0001, Depei Qian 0001, Weiguo Wu |
HPCC | 4 |
| 2011 | Video scene segmentation using a novel boundary evaluation criterion and dynamic programmingabstractVideo scene segmentation is a fundamental step for video summarization and browsing, which is a very promising application of multimedia analysis. There are two key elements, namely, boundary evaluation and boundary searching, in a scene segmentation algorithm. In this paper, we propose a novel boundary evaluation criterion, including the multiple normalized min-max cut scores, which consider not only neighboring but non-neighboring scene similarities with a memory-fading model, and the maximal cross-boundary strict shot similarity, which considers both color and structure similarities. Dynamic programming with a heuristic search scheme is adopted to quickly find the global optimal scene boundary sequence. Moreover, a Monte Carlo method is adopted to improve the stability of the searching process. Experimental results on a dataset of 40 diversified videos have proven the algorithm efficient, robust, and superior to the existent methods. Weiguo Wu |
ICME | 2 |
| 2010 | How context helps: A discriminative codeword selection method for object detectionabstractWe first propose in this paper to localize objects in images based on the models learned from the weakly labeled images. This task is termed as region of interest (ROI) detection. Local features such as SIFT or HOG are extracted and the discriminative words from clustered codewords based on SIFT and HOG are selected to model the objects. Then how to find the discriminative words to model the object is important. Existing ROI detection methods consider the information from the foreground objects by selecting the words appearing more in the images belonging to one specific image class. Considering the information from background/context is also helpful for object detection and classification, we propose to select the discriminative words which appear more in the foreground/object and less in the background/context. Second, another task is to give the class label (object in this setting) for a given image and also give the position of the object appearing in the image. This task is termed as objection detection. A normal way for this task after ROI is to extract features from the detected regions and not from the whole image. Since the discriminative words extracted during ROI detection has good discriminative ability, we propose to use these words for object detection. Experimental results on PASCAL VOC 2006 dataset and a larger dataset containing 29 classes demonstrate the effectiveness of the proposed method. Renzhong Wei, Hong Lu 0001, Yingbin Zheng, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Weiguo Wu |
ICIP | 7 |
| 2010 | Scalable Hierarchical Scheduling for Multiprocessor Systems Using Adaptive Feedback-Driven PoliciesabstractThis work addresses the problem of allocating resource-intensive parallel jobs on multicore- and multiprocessor-based systems, where the performance gains largely depend on effectively exploiting application parallelization across the available parallel computing resources. The objective is to find efficient allocation approaches that minimize the parallel jobs' completion time, i.e. makespan. Integrating feedback-driven adaptive strategies, we present a general hierarchical scheduling framework and show that two hierarchical scheduling algorithms: ABG-DS and AG-DS achieve scalable performance in term of makespan regardless of the number of hierarchical levels. Specifically, we prove that both ABG-DS and AG-DS have O(1)-competitive ratio for batched parallel jobs. Extending an existing tool, called Malleable-Lab, we evaluate the performance and scalability of our proposed algorithms and compare with that of well-known EQUI-based strategies. The simulation results demonstrate that both ABG-DS and AG-DS generally outperforms EQUI-EQUI for a wide range of parallel workloads. Moreover, feedback-driven adaptive scheduling algorithms show better scalability when the number of levels increases in the scheduling hierarchy. Yangjie Cao, Hongyang Sun 0001, Depei Qian 0001, Weiguo Wu |
ISPA | 4 |
| 2009 | A general framework for automatic on-line replay detection in sports videoabstractReplay detection is a pivotal step for sports video highlight extraction, which is a very promising application of multimedia analysis. In this paper, a general framework, which is based on a Bayesian network, is proposed to make full use of the multiple clues, including shot structure, gradual transition pattern, slow-motion, and sports scene. A novel algorithm based on motion vector reliability classification is proposed to analyze the gradual transition patterns, so that the replay detector can meet the requirements of automatic on-line applications. This is the first integrated general replay detection framework proposed in the literature. Extensive experiments on diversified sports games have proven the scheme efficient, accurate and robust. Zhenghua Chen, Chang Liu 0087, Weiguo Wu |
ACM Multimedia | 5 |
| 2008 | A Distributed Trust Management Based on Authorizing Negotiation in Open and Dynamic EnvironmentsabstractTrust has been recognized as an important factor for information security in open and dynamic environments, such as Internet applications, p2p systems etc. On the basis of analyzing existing trust management systems, this paper proposes a Distributed Trust Management based on Authorizing Negotiation (DTMAN). DTMAN presents a number of innovative features. First, it can authorize strangers by using authorizing negotiation, so it is very suitable for open and dynamic environments. Second, a high efficient algorithm for compliance checking is developed to support DTMAN, whose time complexity and space complexity are both O(n) (where n is the cardinality of the set of authorization credentials). The experimental result shows that the algorithm of DTMAN is more efficient than others. Shangyuan Guan, Xiaoshe Dong, Weiguo Wu, Yiduo Mei, Guofu Feng |
AINA | 3 |
| 2008 | A Minimum Coverage Dynamic Adaptive Method for Concurrent Web Service CompositionabstractTo address the problems of concurrent Web services composition, this paper applies minimum coverage in logic theory to the semantic matching during the period of service composition, and presents a service composition satisfaction degree model. The least and most appropriate Web services match is obtained dynamically and adaptively by using the model to set the composition satisfaction degree, and the semantic match relation graph of Web services is structured. Based on the above graph and the userpsilas requirement, composition Web services can been gotten directly. This paper describes its the prototype system, and experimental results show that the method can be used to composite concurrent Web service successfully, and the success rate of service composition will increase by setting the appropriate composition satisfaction degree. At the same time, it can improve both the query nicety rate and the composition efficiency. Zhengdong Zhu, Xuehan Dong, Yahong Hu, Weiguo Wu, Zengzhi Li |
APSCC | 4 |
| 2008 | Directional entropy feature for human detectionabstractIn this paper we propose a novel feature, called directional entropy feature (DEF), to improve the performance of human detection under complicated background in images. DEF describe the regularity of region by computing the entropy value of edge pointspsila spatial distribution in specific direction, so DEF has the discriminating power for regular and random pattern. We combine histogram of oriented gradient (HOG) feature with DEF to construct a human detection classifier to test DEFpsilas performance. Experimental results show that DEF can help HOG to decreases false alarms caused by random complicated and rigid shaped background. Long Meng, Shuqi Mei, Weiguo Wu |
ICPR | 4 |
| 2007 | SDRD: A Novel Approach to Resource Discovery in Grid Environments
Yiduo Mei, Xiaoshe Dong, Weiguo Wu, Shangyuan Guan, Junyang Li 0004 |
APPT | 3 |
| 2007 | Spatial Map Data Share and Parallel Dissemination System Based on Distributed Network Services and Digital Watermark
Depei Qian 0001, Weiguo Wu, Ailong Liu, Xuewei Yang, Pen Han |
NPC | 3 |
| 2006 | AOCMS: An Adaptive and Scalable Monitoring System for Large-Scale ClustersabstractIn this paper, we present the design and implementation of AOCMS, an adaptive, scalable and efficient monitoring system for a large-scale cluster. We describe an adaptive architecture of AOCMS in detail, and focus on the discussion about some techniques as to enhancing the adaptation, scalability and efficiency of AOCMS. These techniques include: a solution to monitor a heterogeneous cluster; a universal applet-servlet communicating controller responsible for communication between the clients and the Web server; adaptive pools providing threads or connections to the database for the monitoring tasks on demand; and an AOP-based alarm decoupling the alarming logic from the monitoring logic. Moreover, we measured the performance of AOCMS. The results show that AOCMS runs with low overheads and responds to clients quickly Zhenghua Xue, Xiaoshe Dong, Weiguo Wu |
APSCC | 3 |
| 2006 | Research on the Walking Modes Shifting Based on the Variable ZMP and 3-D.O.F Inverted Pendulum Model for a Humanoid and Gorilla RobotabstractThe walking modes shifting of a gorilla robot is a kind of movements between biped standing state and quadruped landing state. In this paper, the robot mechanism is reduced to a 3-D.O.F inverted pendulum model with variable pendulum length, and the variable ZMP is defined reasonably as a function related to the inverted pendulum angle. Base on dynamic balance theory, the trajectory equation of robot's mass centre during the walking mode shifting is deduced. Furthermore, through inverse kinematics analysis for robot's mass center, the trajectories of joint are obtained. Thus method of trajectory generation about walking mode shifting for humanoid and gorilla robot is proposed. In order to verify the correctness of the method, a calculation example of trajectory generation is provided, and the continuing action simulation including biped walking and quadruped landing action, quadruped walking and standing up action is successfully realized. On the basis of above work, the continuing action experiment including biped walking and walking modes transitions for a humanoid and gorilla robot "GoRoBoT" developed by us is also finished Weiguo Wu, Yunzhong Pan |
IROS | 1 |
| 2005 | Eye-contact visual communication with virtual view synthesisabstractIn this paper, we propose a new visual communication system where eye contact is possible by using a virtual image. The virtual image is obtained by view synthesis with stereo matching from two real camera views. We developed a region based dynamic programming (DP) approach with improved matching cost, occlusion cost and vertical smoothness constraint. We also proposed a fast view interpolation method. To achieve real time performance, we developed a hardware system. Furthermore, to avoid the reordering problem in the foreground region, a view change approach with disorder detection is adopted. Experimental results demonstrate the validity of our improved DP matching algorithm and eye-contact visual communication system. Yuyu Liu, Keisuke Yamaoka, Akira Nakamura, Yoshiaki Iwai, Ken-ichiro Ooi, Weiguo Wu, Takayuki Yoshigahara |
CCNC | 7 |
| 2005 | Design, simulation and walking experiments for a humanoid and gorilla robot with multiple locomotion modesabstractAim at studying the whole humanoid and gorilla robot, a small integrated combinational gorilla robot with multiple locomotion modes and a humanoid head with facial expressions is presented in the paper. In order to meet the requirement on quadruped walking, a hand with multiple fingers that can also be used as the front foot is designed. A method of generating steady dynamic biped walking patterns based on variable ZMP is proposed. The motion simulation on the biped walking and quadruped walking of the robot by shifting the robot hands as feet is finished by means of a software for dynamics analysis, so that the basic walking ability of the robot can be ensured. On the basis of theoretical research and simulation, the experiments for biped walking and quadruped walking of the gorilla robot are finished successfully. Weiguo Wu, Yuedong Lang, Fuhai Zhang, Bingyin Ren |
IROS | 1 |
| 2005 | Omni-directional quadruped walking gaits and simulation for a gorilla robotabstractGorilla robot is a new-type robot with multiple locomotion modes including biped walking, quadruped walking and brachiate locomotion. In comparison with conventional four-legged walking machines, the mechanism of the robot has particularity and complexity. In order to realize quadruped walking of the gorilla robot, we propose the gaits of forward walking, sideways walking, translational walking and turning walking, respectively. The results of simulation show that these gaits are feasible and efficient. Fuhai Zhang, Weiguo Wu, Yuedong Lang, Bingyin Ren |
IROS | 2 |
| 2000 | Stereo matching method with deformable window and its application to 3D measurement of the human face
Weiguo Wu, Atsushi Yokoyama, Teruyuki Ushiro, Takayuki Yoshigahara, Yoko Miwa |
VCIP | 1 |
| 1999 | Optimal Motion Planning for a Wheeled Mobile RobotabstractConcerns the time optimal motion planning problem under kinematic and dynamic constraints for a 2-DOF wheeled mobile robot (WMR). The dynamic model of a WMR is derived using a Newton-Euler method and its constraints are analyzed. Kinematic constraints are imposed by its nonholonomy and structural limits while dynamic constraints are due to motor saturation. The motion planning problem is formulated as two stage planning. First, path planning under kinematic constraints is transformed into a pure geometric problem. The shortest path composed of circular arcs and straight lines is obtained. Then, combined with dynamic characteristics of the WMR, a time optimal velocity profile is generated under dynamic constraints. Since constraints of a WMR are fully exploited, the proposed method is simple and effective for motion planning. Simulation results illustrate the capability of the planning scheme. Weiguo Wu, Ping Jiang 0001, Huitang Chen |
ICRA | 1 |
| 1999 | A novel global tracking control method for mobile robotsabstractConcerns trajectory tracking control of mobile robots. In order to overcome the local stability resulting from design methods adapted from linearization, a global asymptotic stable controller, which both achieves global stability margin and avoids winding phenomenon, is designed using a backstepping method. This method breaks down nonlinear systems into low dimensional systems and simplifies the controller design using virtual control inputs and partial Lyapunov functions. The stability of the system is easily proven via the Lyapunov function. Abundant simulation results validate the theoretical analysis. Weiguo Wu, Huitang Chen, Yuejuan Wang |
IROS | 1 |
| 1999 | Backstepping design for path tracking of mobile robotsabstractFrom the practical engineering point of view, path tracking control for mobile robots has been investigated. It comes from planning trajectory with the planned geometric path and enables direct tracking of the geometric path. The kinematic model of mobile robots and the desired path are described in polar coordinates. The influence on control performance resulting from the factitious choice of desired reference points is eliminated by considering the polar angle as a parameter. The controller is designed based on a backstepping method, which is systematic and flexible. The convergence of the system is throughout the design procedure. Simulation results are given to verify the proposed control laws. Weiguo Wu, Huitang Chen, Yuejuan Wang |
IROS | 1 |