VLDB 2026 Research / reviewers in the wild / expert
Hengzhu Liu
dblp:00/5381
· DBLP profile ↗
51ranked-venue papers
2as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 10 since 2021Artificial intelligence and machine learning · 15 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LCA-Med: A lightweight cross-modal adaptive feature processing module for detecting imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Neural Networks | 9 |
| 2026 | Leveraging VLMs for MUDA: Category-specific prompt with multi-modal interactive LoRA
Xihuai He, Xueqiong Li, Wanrong Huang, Hengzhu Liu, Huibin Tan |
Neural Networks | 5 |
| 2025 | A Framework for Modeling Cognitive Processes in Intelligent Agents Using Behavior Trees
Kejia Wan, Yuntao Liu 0004, Hengzhu Liu, Xinhai Xu, Jinlong Tian |
CogSci | 3 |
| 2025 | Mental Model Alignment: Building Cognitive Interfaces for Explainable Reinforcement Learning
Kejia Wan, Yuntao Liu 0004, Hengzhu Liu, Xinhai Xu, Jinlong Tian |
CogSci | 3 |
| 2025 | Anchor-Prompt-based Segmentation and Embedding ModelabstractTackling multi-object tracking and segmentation (MOTS) can be attributed to a multi-task learning task, i.e., performing Segmentation and Identity Embedding jointly (SIEJ). Unfortunately, achieving optimal SIEJ is non-trivial, as it relies on different spatiotemporal features of objects. Besides, the lack of labeled data raises the difficulty of balancing the two objectives in SIEJ. Empowered by the Segment Anything Model (SAM), we propose an Anchor-Prompt-based Segmentation and Embedding Model (APSEM) towards optimal SIEJ by introducing redundant anchors and designing an embedding decoder. On the one hand, our APSEM allows redundant anchors for the same object to identify occluded objects and refine the pixel-wise edge mask. On the other hand, our proposed embedding decoder is designed to address identity learning for redundant anchors by minimizing differences between the same identities in redundant prompts. We also create a parameter-efficient fine-tuning strategy, which helps combine the embedding module into the foundation model through a bit of data. Experimental results on MOTSChallenge datasets validate the effectiveness of the proposed APSEM method for MOTS tasks. Such results also demonstrate that each module improves segmentation and identity embedding performance through joint training. Shuman Li, Haotian Wang 0001, Wenjing Yang 0002, Hengzhu Liu |
ICASSP | 5 |
| 2025 | Multi-layer Network Disintegration via Deep Reinforcement LearningabstractMulti-layer networks (MLN) effectively model interactions across layers, and the network disintegration (ND) problem yields significant importance in the analysis of MLN. Unfortunately, previous advances in ND for single-layer networks exhibits inefficiency and lack of scalability when extended to MLN, as MLN involve complex inter-layer dependencies and interactions that are absent in single-layer networks. To bridge this gap, we propose a pioneer framework named Multi-layer Network Disintegration via Deep Reinforcement Learning (MNDRL), by re-formulating the disintegration process into a preprocessing-encoding-decoding pipeline. To be specific, MNDRL consists of two Graph Neural Network (GNN) models to effectively capture intra-layer and inter-layer representations separately. Empowered with DRL, our MNDRL achieves near-optimal results in the ND task with the NP-hard complexity by autonomously learning efficient disintegration strategies. Extensive experiments demonstrate that our MNDRL outperforms the results of baseline disintegration methods by over 30% on real-world datasets and by over 10% on synthetic datasets with varying node counts. Zhenhua Liang, Xueqiong Li, Shaowu Yang, Hengzhu Liu |
ICASSP | 6 |
| 2025 | A visual state space Model-Based Cross-Domain adaptive detection method for imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Appl. Intell. | 9 |
| 2025 | Dragon Boat Optimization: A Meta-Heuristic for Intelligent SystemsabstractABSTRACT Dragon boat racing, a popular aquatic folklore team sport, is traditionally held during the Dragon Boat Festival. Inspired by this event, we propose a novel human‐based meta‐heuristic algorithm called dragon boat optimization (DBO) in this paper. It models the unique behaviours of each crew member on the dragon boat during the race by introducing social psychology mechanisms (social loafing, social incentive). Throughout this process, the focus is on the interaction and collaboration among the crew members, as well as their decision‐making in various situations. During each iteration, DBO implements different state updating strategies. By accurately modelling the crew's behaviour and employing adaptive state update strategies, DBO consistently achieves high optimization performance, as validated by comprehensive testing on 29 benchmark functions and 2 structural design problems. Experimental results indicate that DBO outperforms 7 and 16 state‐of‐the‐art meta‐heuristic algorithms across these test functions and problems, respectively. Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Expert Syst. J. Knowl. Eng. | 7 |
| 2025 | A survey on machine unlearning: Techniques and new emerged privacy risks
Hengzhu Liu, Ping Xiong 0001, Tianqing Zhu, Philip S. Yu |
J. Inf. Secur. Appl. | 1 |
| 2025 | Self-supervised re-identification for online joint multi-object trackingabstractRecently, the bottleneck of multi-object tracking is shifting from detection performance to association performance. However, research on association algorithms requires a large number of identity labels, which are more expensive than detection labels. To circumvent the need for identity labels, we propose a Self-supervised Re-identification module for online joint Multi-Object Tracking (SR-MOT). Specifically, we design an appearance discriminator to judge identities based solely on detection hypotheses and then associate the same identity with the final trajectory. To train the discriminator without using identity labels, we construct negative pairs by the detections that appear in the same video frame, as they definitely belong to different identities. Positive pairs are naturally constructed through several useful data augmentation strategies at the box level. In addition, our proposed method balances conflicting detection and re-ID tasks by using different output features and dynamically adjusts detection and re-ID loss weights based on the information content of the loss distribution to promote balance between the two tasks from the feature level and optimization methods. In our evaluation on the MOT Challenge benchmark, we show that our SR-MOT performs comparably to supervised methods and is significantly superior to other unsupervised methods. Our proposed method provides a practical solution for multi-object tracking without the need for identity labels, making it more accessible for real-world applications. Shuman Li, Longqi Yang 0002, Huibin Tan, Binglin Wang, Wanrong Huang, Hengzhu Liu, Wenjing Yang 0002, Long Lan |
Knowl. Inf. Syst. | 6 |
| 2025 | Tradeoff Performance and Energy Efficiency by Optimizing the Data Flow for PIM ArchitecturesabstractThe processing-in-memory (PIM) architecture becomes a promising candidate for deep learning accelerators by integrating computation and memory. Most PIM-based studies improve the performance and energy efficiency by using the weight stationary (WS) data flow due to its high parallelism. However, the WS data flow has some fundamental limitations. First, the WS data flow has huge activation movements between on-chip memory and off-chip memory due to the limited memory space of the resistive random-access memory (ReRAM) array. Second, the WS data flow needs to read the input activation repeatedly according to the convolution window. These data movements decrease the energy efficiency and performance of the PIM architecture. To address these issues, the input stationary (IS) data flow stores activations instead of weights to reduce data movements. But the IS data flow faces some challenges. First, the data dependency between adjacent layers limits the performance. Second, there are huge across-array computations due to the special mapping method. Third, the previous IS data flow cannot realize the high parallelism. Fourth, the IS data flow depends on the 3-D ReRAM structure. To address these issues, we propose a novel data flow for PIM architectures. We optimize the IS data flow to decrease the activation movement and propose a parallel computing method to realize high parallelism and reduce the across-array computations. We identify and analyze the fundamental limitations and impact of different interlayer data flows, including the WS-WS, IS-IS, WS-IS, and IS-WS. We also propose a method to build a hybrid data flow by combining these interlayer data flows to tradeoff performance and energy consumption. Our experimental results and analysis demonstrate the potential of our design. The performance and energy efficiency of our design reach 0.13–1.77 TFLOPS and 61–85 TOPS/J, respectively. Compared to the state-of-the-art design, the NEBULA, our design can improve performance by$1.4\times $,$2.3\times $, and$3.5\times $for deploying the MobileNet-V1, ResNet-18, and VGG-16, and also can improve energy efficiency by$3.3\times $,$2\times $, and$2\times $, respectively. Yunping Zhao, Sheng Ma, Yuhua Tang, Hengzhu Liu, Dongsheng Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Game-Theoretic Machine Unlearning: Mitigating Extra Privacy LeakageabstractWith the extensive use of machine learning technologies, data providers encounter increasing privacy risks. Recent legislation, such as GDPR, obligates organizations to remove requested data and its influence from a trained model. Machine unlearning is an emerging technique designed to enable machine learning models to erase users’ private information. Although several efficient machine unlearning schemes have been proposed, these methods still have limitations. First, removing the contributions of partial data may lead to model performance degradation. Second, discrepancies between the original and generated unlearned models can be exploited by attackers to obtain target sample’s information, resulting in additional privacy leakage risks. To address above challenges, we proposed a game-theoretic machine unlearning algorithm that simulates the competitive relationship between unlearning performance and privacy protection. This algorithm comprises unlearning and privacy modules. The unlearning module possesses a loss function composed of model distance and classification error, which is used to derive the optimal strategy. The privacy module aims to make it difficult for an attacker to infer membership information from the unlearned data, thereby reducing the privacy leakage risk during the unlearning process. Additionally, the experimental results on real-world datasets demonstrate that this game-theoretic unlearning algorithm’s effectiveness and its ability to generate an unlearned model with a performance similar to that of the retrained one while mitigating extra privacy leakage risks. Hengzhu Liu, Tianqing Zhu, Lefeng Zhang, Ping Xiong 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Bayesian Procedures for Modeling Truck Route ChoicesabstractThis study examines logit models applied to the truck route choice problem using GPS trucking data from the Dallas metropolitan area. Instead of assuming a constant coefficient for each variable in the conventional multinomial logit model, the proposed mixed c-logit model assumes a certain probability distribution for each coefficient, in an attempt to better reflect the drivers’ preference heterogeneity. A commonality factor is introduced in the model to address roadway travel time correlations due to route overlaps. Three Bayesian models with different hierarchy levels are introduced and are solved using mean-field variational inference with the block coordinate algorithm. In the reduced subnetwork of the examined area, the drivers are assumed to make their route choices in three groups of routes, referred to as choice groups. Attributes that would affect the truck driver’s route choice decisions are different among choice groups. With this setting, the proposed Bayesian models are then tested with the three truck route choice groups respectively. Generally, the study finds that the factors considered in truckers’ route choice vary with context. Xiubin Wang, Huibin Tan, Hengzhu Liu, Weixia Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | FAMS: A FrAmework of Memory-Centric Mapping for DNNs on Systolic Array AcceleratorsabstractIn recent years, deep neural networks (DNNs) have experienced rapid development. These DNNs demonstrate significant variations in architecture and scale, creating a substantial demand for domain-specific accelerators that are optimized for both high performance and low energy consumption. Systolic array accelerators, due to their efficient dataflow and parallel processing capabilities, offer significant advantages when performing computations for DNNs. Existing studies frequently overlook various hardware constraints in systolic array accelerators when representing mapping strategies. This oversight includes ignoring the differences in delays between communication and computation operations, as well as overlooking the capacities of multilevel memory hierarchies. Such omissions can lead to inaccuracies in predicting accelerator performance and inefficiencies in system design. We propose the FAMS framework, which introduces a memory-centric notation capable of fully representing the mapping of DNN operations on systolic array accelerators. Memory-centric notation moves away from the idealized assumptions of previous notations and considers various hardware constraints, thereby expanding the effective design and mapping spaces. The FAMS framework also includes a cycle-accurate simulator, which takes the hardware configurations, task descriptions, and mapping strategy represented by memory-centric notation as inputs, providing various metrics such as latency and energy consumption. The experimental results demonstrate that our proposed FAMS framework reduces latency by up to 29.7% and increases throughput by 42.4% compared to the state-of-the-art TENET framework. Additionally, under hardware configurations with a MAC delay of 2 and 3 clock cycles, the FAMS framework enhances performance by 12.0% and 25.4%, respectively. Hao Sun 0023, Junzhong Shen, Zhongyi Tang, Changwu Zhang, Yang Shi 0008, Hengzhu Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2024 | Enhancing Privacy in Machine Unlearning: Posterior Perturbation Against Membership Inference Attack
Hengzhu Liu, Huanhuan Chi, Ping Xiong 0001 |
ICA3PP (6) | 2 |
| 2024 | Radar Recognition in the Wild: Enhancing Radar Emitter Recognition through Auto-Correlation Model-Agnostic Meta LearningabstractIn Electronic Support Measure (ESM) systems, the recognition of radar emitters stands as a pivotal yet intricate task. The complex electromagnetic environments, however, often hinders the collection of clean radar signal data, and results in data with different noise levels. Consequently, formulating a robust recognition model with limited data becomes a big challenge, further compounded by the demand for generalizability across scenarios with different noise levels. While Model-Agnostic Meta Learning (MAML) has proven its effectiveness in solving few-shot learning problems in computer vision, its application in radar signal processing has remained unexplored deeply. This paper pioneers the incorporation of MAML and autocorrelation into radar signal processing. To fit MAML to radar signals, we introduce a novel loss function, termed AC-Loss, designed to facilitate learning effective signal representation by retaining the periodicity of the radar pulses which is the key feature for recognizing different Pulse Repetition Intervals (PRIs). This proposed Autocorrelation Model-Agnostic Meta Learning (AC-MAML) enhances its recognition capabilities while using only a sparse number of signal samples in both source and target domains. Empirical results show the superiority of AC-MAML, achieving an impressive average recognition accuracy of 90.4% across seven diverse target domain scenarios. Yixian Luo, Shaowu Yang, Huibin Tan, Ruochun Jin, Hengzhu Liu, Xueqiong Li |
ICASSP | 6 |
| 2024 | Unraveling Explainable Reinforcement Learning Using Behavior Tree StructuresabstractThe black-box characteristic of deep reinforcement learning restricts the safe and scalable application of decision models in practical deployment. Existing interpretability methods for deep reinforcement learning models are often inadequate in providing comprehensive insights and generating logical sequential decisions. In this study, we propose an innovative framework called XRLBT, which introduces the behavior tree structure to explainable reinforcement learning. XRLBT clusters state space by aggregating temporally related states. Based on these clustered states, XRLBT constructs behavior tree structures that align with the target deep reinforcement learning model. In this way, XRLBT incorporates an exploration technique specifically designed for temporal policies in DRL. Through extensive experimentation across six benchmark environments, we showcase the superiority of the discovered behavior tree structures over state-of-the-art algorithms. This work represents a significant stride towards addressing the challenges of explainability and performance in DRL applications. Kejia Wan, Yuntao Liu 0004, Hengzhu Liu, Xinhai Xu |
ICASSP | 3 |
| 2024 | OLSATM: Online Learning Based State-Aware Task Migration on S-NUCA Many-CoresabstractTask migration maximizes performance while maintaining thermal safety in many-cores systems. Existing techniques exploit offline learning which requires tremendous training data and fixed-cycle migration which causes threads to miss the optimal migration timing. This paper presents Online Learning based State-Aware Task Migration (OLSATM). It pretrains a neural network (NN) with a small set of data and updates the model online to substitute the laborious data collection and model training of offline learning. It is state-aware and detects the timing when migration is needed, overcoming the shortcomings of periodical migration. Experimental results show that OLSATM enhances the performance by 3.7% and reduces the number of migration judgments by 26 % on average compared to the state-of-the-art task migration. Yandong He, Guangda Zhang, Hengzhu Liu, Renzhi Chen |
ICCD | 4 |
| 2024 | LLM-Enhanced Theorem Proving with Term Explanation and Tactic Parameter Repair✱abstractThere has been emerging researches on leveraging large language models (LLMs) to improve the automation of theorem proving. However, they are still suffering from low accuracy and efficiency. In this paper, we propose to strengthen the existing approach by enhancing a language agent, which provides automatic explanation of terms and repair of tactics parameters. Term explanation explains terms specific to the proof obligations formally and tactic parameter repair complements the potentially correct proof tactics as much as possible. Similar to the existing approach, the agent uses GPT-4 as query objects in a search policy. During the search, we add term explanation to the prompt, and then the policy selects a proof tactic and repairs it. The repaired tactics interact with the theorem prover (Coq), and the execution result is fed back to build the prompt for the next policy invocation. We evaluate our approach on subsets of the CompCert project implemented using Coq. Our approach proves 8.11% more theorems than the existing language agent COPRA, and demonstrates faster search and proof speed. Besides, when term explanation and tactic parameter repair are applied, the performance of the SOTA method PROVERBOT9001 can be also improved. Xingpeng Liu, Hengzhu Liu, Xiaodong Yi 0002, Ji Wang 0001 |
Internetware | 2 |
| 2024 | Improving the Ability of Thermal Radiation Based Hardware Trojan Detection
Ting Su 0009, Lusi Zhang, Simin Feng, Jialong Song, Yongkang Tang, Shaoqing Li, Yang Guo 0003, Hengzhu Liu |
USENIX Security Symposium | 12 |
| 2024 | EAFP-Med: An efficient adaptive feature processing module based on prompts for medical image detectionabstractThe rapid proliferation of medical imaging technologies presents a significant challenge for cross-domain adaptive image detection, as lesion representations can vary dramatically across technologies. To address this issue, we draw inspiration from large language models to propose EAFP-Med, an efficient adaptive feature processing module based on prompts for medical image detection. EAFP-Med incorporates a prompt-driven dynamic parameter update mechanism, empowering it to extract cross-domain multi-scale lesion features from medical images of diverse modalities adaptively. This exceptional flexibility liberates it from the constraints of any particular imaging technique, fostering great adaptability. Furthermore, EAFP-Med can also serve as a feature preprocessing module connected to any model front-end to enhance the lesion features in input images. Moreover, we propose a novel adaptive disease detection model named EAFP-Med ST, which utilizes the Swin Transformer V2 – Tiny (SwinV2-T) as its backbone and connects it to EAFP-Med. We have compared our method to nine state-of-the-art methods. Experimental results show that the overall accuracy of EAFP Med ST on chest X-ray, brain magnetic resonance imaging, and skin image datasets is 98.47%, 97.60%, and 99.06%, respectively, superior to all the compared state-of-the-art methods. Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Expert Syst. Appl. | 7 |
| 2024 | SAC: An Ultra-Efficient Spin-based Architecture for Compressed DNNsabstractDeep Neural Networks (DNNs) have achieved great progress in academia and industry. But they have become computational and memory intensive with the increase of network depth. Previous designs seek breakthroughs in software and hardware levels to mitigate these challenges. At the software level, neural network compression techniques have effectively reduced network scale and energy consumption. However, the conventional compression algorithm is complex and energy intensive. At the hardware level, the improvements in the semiconductor process have effectively reduced power and energy consumption. However, it is difficult for the traditional Von-Neumann architecture to further reduce the power consumption, due to the memory wall and the end of Moore’s law. To overcome these challenges, the spintronic device based DNN machines have emerged for their non-volatility, ultra low power, and high energy efficiency. However, there is no spin-based design that has achieved innovation at both the software and hardware level. Specifically, there is no systematic study of spin-based DNN architecture to deploy compressed networks. In our study, we present an ultra-efficient Spin-based Architecture for Compressed DNNs (SAC), to substantially reduce power consumption and energy consumption. Specifically, we propose a One-Step Compression algorithm (OSC) to reduce the computational complexity with minimum accuracy loss. We also propose a spin-based architecture to realize better performance for the compressed network. Furthermore, we introduce a novel computation flow that enables the reuse of activations and weights. Experimental results show that our study can reduce the computational complexity of compression algorithm from 𝒪( Tk 3 to 𝒪( k 2 log k ), and achieve 14× ∼ 40× compression ratio. Furthermore, our design can attain a 2× enhancement in power efficiency and a 5× improvement in computational efficiency compared to the Eyeriss. Our models are available at an anonymous link https://bit.ly/39cdtTa . Yunping Zhao, Sheng Ma, Hengzhu Liu, Libo Huang 0002 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | SAL: Optimizing the Dataflow of Spin-based Architectures for Lightweight Neural NetworksabstractAs the Convolutional Neural Network (CNN) goes deeper and more complex, the network becomes memory-intensive and computation-intensive. To address this issue, the lightweight neural network reduces parameters and Multiplication-and-Accumulation (MAC) operations by using the Depthwise Separable Convolution (DSC) to improve speed and efficiency. Nonetheless, the energy efficiency of classical Von Neumann architectures for CNNs is limited due to the memory wall challenge. Spin-based architectures have the potential to address this challenge thanks to the integration of memory and computing with ultra-high energy efficiency. However, deploying the DSC on spin-based architectures with the traditional dataflow leads to huge activation movements and low hardware utilization. Moreover, the inter-layer data dependency of neural networks increases latency. These factors become the bottleneck of improving energy efficiency and performance. Inspired by these challenges, we propose a novel dataflow on Spin-based Architectures for Lightweight neural networks (SAL). The novel dataflow replaces convolution unrolling by selecting activations in the crossbar according to the convolution window and also realizes the inter-layer data reuse. Moreover, the novel dataflow also reduces the latency due to the data dependency between layers, realizing higher performance. To the best of our knowledge, this is the first design to use hybrid dataflow for the PIM architecture. We also optimize the structure of the spin-based crossbar and the pipeline based on the dataflow to achieve better data reuse and computational parallelism. For deploying the MobileNet V1, the novel dataflow improves the hardware utilization by 23×∼ 105× and reduces the data traffic by 1.09×∼ 18.6×. Compared with the NEBULA, a spin-based non-Von Neumann architecture, the SAL reduces the energy consumption by 4× and improves the performance by 7.3×, which are 0.32 mJ and 10.43 GOPs -1 , respectively. Moreover, the SAL improves power efficiency over 29 times more than the NEBULA. Compared with the Eyeriss, the SAL improves the energy efficiency by four orders of magnitude. Yunping Zhao, Sheng Ma, Hengzhu Liu, Dongsheng Li 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | EPHA: An Energy-efficient Parallel Hybrid Architecture for ANNs and SNNsabstractArtificial neural networks (ANNs) and spiking neural networks (SNNs) are two general approaches to achieve artificial intelligence (AI). The former have been widely used in academia and industry fields; the latter, SNNs, are more similar to biological neural networks and can realize ultra-low power consumption, thus have received widespread research attention. However, due to their fundamental differences in computation formula and information coding, the two methods often require different and incompatible platforms. Alongside the development of AI, a general platform that can support both ANNs and SNNs is necessary. Moreover, there are some similarities between ANNs and SNNs, which leaves room to deploy different networks on the same architecture. However, there is little related research on this topic. Accordingly, this article presents an energy-efficient, scalable, and non-Von Neumann architecture (EPHA) for ANNs and SNNs. Our study combines device-, circuit-, architecture-, and algorithm-level innovations to achieve a parallel architecture with ultra-low power consumption. We use the compensated ferrimagnet to act as both synapses and neurons to store weights and perform dot-product operations, respectively. Moreover, we propose a novel computing flow to reduce the operations across multiple crossbar arrays, which enables our design to conduct large and complex tasks. On a suite of ANN and SNN workloads, the EPHA is 1.6× more power-efficient than a state-of-the-art design, NEBULA, in the ANN mode. In the SNN mode, our design is 4 orders of magnitude more than the Loihi in power efficiency. Yunping Zhao, Sheng Ma, Hengzhu Liu, Libo Huang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Heterogeneous Ensemble Federated Learning With GAN-Based Privacy PreservationabstractMulti-party collaborative learning has become a paradigm for large-scale knowledge discovery in the era of big data. As a typical form of collaborative learning, federated learning (FL) has received widespread research attention in recent years. In practice, however, FL faces a range of challenges such as objective inconsistency, communication and synchronization issues, due to the heterogeneity in the clients' local datasets and devices. In this paper, we propose EnsembleFed, a novel ensemble framework for heterogeneous FL. The proposed framework first allows each client to train a local model with full autonomy and without having to consider the heterogeneity of local datasets. The confidence scores of training samples output by each local model are then perturbed to defend against membership inference attacks, after which they are submitted to the server for use in constructing the global model. We apply a GAN-based method to generate calibrated noise for confidence perturbation. Benefiting from the ensemble framework, EnsembleFed disengages from the restriction of real-time synchronization and achieves collaborative learning with lower communication costs than traditional FL. Experiments on real-world datasets demonstrate that the proposed EnsembleFed can significantly improve the performance of the global model while also effectively defending against membership inference attacks. Hengzhu Liu, Huanhuan Chi, Ping Xiong 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2023 | Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained RobotabstractDubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. Therefore, this paper presents a collision-free CPP approach (CFC) for the obstacle-constrained environment, which enhances time efficiency by constructing the variable-speed Dubins paths and ensures robot safety by building a risk potential surface for representing the possibility of collision. Furthermore, CFC models the CPP problem as an asymmetric traveling salesman problem (ATSP) and utilizes a graph pruning strategy to reduce the computational cost. Comparison tests with other Dubins coverage methods demonstrate that CFC provides shorter coverage times and better runtimes than the other Dubins coverage methods while preventing collision risk between the robot and obstacles. Physical experiments in a laboratory setting demonstrate the applicability of CFC to the physical robot. Lin Li 0075, Dian-xi Shi, Songchang Jin, Yixuan Sun, Xing Zhou 0004, Shaowu Yang, Hengzhu Liu |
ICRA | 7 |
| 2023 | Decentralized privacy-preserving truth discovery for crowd sensing
Ping Xiong 0001, Guirong Li, Hengzhu Liu, Yiyi Hu |
Inf. Sci. | 3 |
| 2023 | A Survey of Memory-Centric Energy Efficient Computer ArchitectureabstractEnergy efficient architecture is essential to improve both the performance and power consumption of a computer system. However, modern computers suffer from the severe “memory wall” problem due to the significant performance gap between the processor technology and the memory technology. Thus, the computer architecture community is evolving from compute-centric to memory-centric designs to reduce the data movement overhead. This paper presents a comprehensive survey of the main challenges and recent advances in memory-centric energy efficient computer architecture. We summarize two research directions: improving the memory technology and processing closer to memory. The former focuses on optimizing the conventional memory technology and exploiting emerging non-volatile memory (NVM) technology. The latter talks about currently popular processing in memory (PIM) technology, including near-memory processing (NMP) and in-memory processing (IMP). Moreover, some other topics like hardware for machine learning (ML), ML for hardware, security, privacy, and reliability are gaining increasing attention and should be considered seriously in the design phase of a computer system. The community is facing various challenges and opportunities simultaneously, requiring researchers to have a more comprehensive understanding of this field which is also the goal of this paper. Changwu Zhang, Hao Sun 0023, Shuman Li, Hengzhu Liu |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature ConstancyabstractFor CNNs based stereo matching methods, cost volumes play an important role in achieving good matching accuracy. In this paper, we present an end-to-end trainable convolution neural network to fully use cost volumes for stereo matching. Our network consists of three sub-modules, i.e., shared feature extraction, initial disparity estimation, and disparity refinement. Cost volumes are calculated at multiple levels using the shared features, and are used in both initial disparity estimation and disparity refinement sub-modules. To improve the efficiency of disparity refinement, multi-scale feature constancy is introduced to measure the correctness of the initial disparity in feature space. These sub-modules of our network are tightly-coupled, making it compact and easy to train. Moreover, we investigate the problem of developing a robust model to perform well across multiple datasets with different characteristics. We achieve this by introducing a two-stage finetuning scheme to gently transfer the model to target datasets. Specifically, in the first stage, the model is finetuned using both a large synthetic dataset and the target datasets with a relatively large learning rate, while in the second stage the model is trained using only the target datasets with a small learning rate. The proposed method is tested on several benchmarks including the Middlebury 2014, KITTI 2015, ETH3D 2017, and SceneFlow datasets. Experimental results show that our method achieves the state-of-the-art performance on all the datasets. The proposed method also won the 1st prize on the Stereo task of Robust Vision Challenge 2018. Zhengfa Liang, Yulan Guo, Yiliu Feng, Wei Chen 0009, Linbo Qiao, Li Zhou 0009, Hengzhu Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | Fine-Grained Multi-human Parsing
Jian Zhao 0006, Jianshu Li, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
Int. J. Comput. Vis. | 3 |
| 2019 | Look across Elapse: Disentangled Representation Learning and Photorealistic Cross-Age Face Synthesis for Age-Invariant Face RecognitionabstractDespite the remarkable progress in face recognition related technologies, reliably recognizing faces across ages still remains a big challenge. The appearance of a human face changes substantially over time, resulting in significant intraclass variations. As opposed to current techniques for ageinvariant face recognition, which either directly extract ageinvariant features for recognition, or first synthesize a face that matches target age before feature extraction, we argue that it is more desirable to perform both tasks jointly so that they can leverage each other. To this end, we propose a deep Age-Invariant Model (AIM) for face recognition in the wild with three distinct novelties. First, AIM presents a novel unified deep architecture jointly performing cross-age face synthesis and recognition in a mutual boosting way. Second, AIM achieves continuous face rejuvenation/aging with remarkable photorealistic and identity-preserving properties, avoiding the requirement of paired data and the true age of testing samples. Third, we develop effective and novel training strategies for end-to-end learning the whole deep architecture, which generates powerful age-invariant face representations explicitly disentangled from the age variation. Extensive experiments on several cross-age datasets (MORPH, CACD and FG-NET) demonstrate the superiority of the proposed AIM model over the state-of-the-arts. Benchmarking our model on one of the most popular unconstrained face recognition datasets IJB-C additionally verifies the promising generalizability of AIM in recognizing faces in the wild. Jian Zhao 0006, Yu Cheng 0009, Yang Yang 0002, Fang Zhao 0006, Jianshu Li, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
AAAI | 7 |
| 2019 | Multi-Prototype Networks for Unconstrained Set-based Face RecognitionabstractIn this paper, we address the challenging unconstrained set-based face recognition problem where each subject face is instantiated by a set of media (images and videos) instead of a single image. Naively aggregating information from all the media within a set would suffer from the large intra-set variance caused by heterogeneous factors (e.g., varying media modalities, poses and illumination) and fail to learn discriminative face representations. A novel Multi-Prototype Network (MP- Net) model is thus proposed to learn multiple prototype face representations adaptively from the media sets. Each learned prototype is representative for the subject face under certain condition in terms of pose, illumination and media modality. Instead of handcrafting the set partition for prototype learn- ing, MPNet introduces a Dense SubGraph (DSG) learning sub-net that implicitly untangles inconsistent media and learns a number of representative prototypes. Qualitative and quantitative experiments clearly demonstrate the superiority of the proposed model over state-of-the-arts. Jian Zhao 0006, Jianshu Li, Xiaoguang Tu, Fang Zhao 0006, Yuan Xin, Junliang Xing, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
IJCAI | 7 |
| 2019 | Closer to Optimal Angle-Constrained Path PlanningabstractPlanning on grids and planning via sampling are the two classical mainstreams of path planning for intelligent agents, whose respective representatives are A* and RRT, including their variants, Theta* and RRT*. However, in the nonholonomic path planning, such us being under angle constraints, Theta* and Lazy Theta* may fail to generate a feasible path because the line-of-sight check (LoS-Check) will modify the original orientation of a state, which makes the planning process incomplete (cannot visit all possible states). Then, we propose a more delayed evaluation algorithm called Late LoS-Check A* (LLA*) to relax the angle constraints. Due to the nature of random sampling, RRT* is asymptotically optimal but still not optimal, then we propose LoS-Check RRT* (LoS-RRT*). In order to solve the problems caused by improper settings of the planning resolution, we propose the LoS-Slider (LoSS) smoothing method. Through experimental comparison, it can be found that angle-constrained versions of LLA* and LoS-RRT* can both generate the near-optimal paths. Meanwhile, the experiment result shows that LLA* performs better than Theta* and Lazy Theta* under angle constraints. The planned path will be even closer to the optimal (shortest) solution after the smoothing of LoSS algorithm. Changwu Zhang, Hengzhu Liu |
IJCNN | 2 |
| 2018 | Learning for Disparity Estimation Through Feature ConstancyabstractStereo matching algorithms usually consist of four steps, including matching cost calculation, matching cost aggregation, disparity calculation, and disparity refinement. Existing CNN-based methods only adopt CNN to solve parts of the four steps, or use different networks to deal with different steps, making them difficult to obtain the overall optimal solution. In this paper, we propose a network architecture to incorporate all steps of stereo matching. The network consists of three parts. The first part calculates the multi-scale shared features. The second part performs matching cost calculation, matching cost aggregation and disparity calculation to estimate the initial disparity using shared features. The initial disparity and the shared features are used to calculate the feature constancy that measures correctness of the correspondence between two input images. The initial disparity and the feature constancy are then fed into a sub-network to refine the initial disparity. The proposed method has been evaluated on the Scene Flow and KITTI datasets. It achieves the state-of-the-art performance on the KITTI 2012 and KITTI 2015 benchmarks while maintaining a very fast running time. Source code is available at http://github.com/leonzfa/iResNet. Zhengfa Liang, Yiliu Feng, Yulan Guo, Hengzhu Liu, Wei Chen 0009, Linbo Qiao, Li Zhou 0009 |
CVPR | 4 |
| 2017 | Learning Generic Features from DEH Channels for Object Proposals GenerationabstractDepth information plays an important role in the human visual system, however it is not yet well- explored in existing proposal generation models. In this paper we propose a new geocentric embedding for depth images that encodes depth to the camera, the structural edges and height above ground-plane for each pixel named DEH channels. We demonstrate that this geocentric embedding works can be use to generate high quality object proposals with convolutional neural networks. Our experiments show significant performance improvements over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. We also exploit Fast R- CNN [8] on top of these proposals to perform object detection with RGB images, our approach obtains state-of-the-art results on all three KITTI object classes. Yiliu Feng, Zhengfa Liang, Hengzhu Liu |
SMARTCOMP | 3 |
| 2016 | A refinement framework for background subtraction based on color and depth dataabstractWe present a refinement framework for background subtraction based on color and depth data. The foreground objects are segmented based on color and depth data independently, in which all of the existed background subtraction (BGS) methods can be applied. The two detected foregrounds will be very inaccurate in some situations such as shadowing and color camouflage. We focus our works on refining the inaccurate results by a supervised learning way. We propose to re-extract features from the source color and depth data. The features together with the initial detection results are fed to classifiers to obtain a better foreground detection. Experiments show that our method can take full advantage of the both information to detect foreground in color camouflage and shadowing situations, giving a promising result which is robust to the inaccurate initial detections, and outperforming the state-of-art algorithms that based on color and depth data. Zhengfa Liang, Hengzhu Liu |
ICIP | 3 |
| 2016 | Large Page Address Mapping in Massive Parallel Processor SystemsabstractLarge and sparse are the prominent characteristics of small-world graph. In many scientific domains, such as biomedical science and scientific computing, as small-word graph grows in scale, processing small-world graph poses severe challenges to address mapping in Massive Parallel Processing. Data driven computation, unstructured data organization, which are poor in spatial and temporal locality, which are high frequency of memory access, leads to large RAM footprint on address mapping management. This paper proposes a novel approach with special Implementation for massive parallel processors. In our technique, The block level address mapping table is stored in large pages in DDR3 memory. Considering the highly frequency in accessing memory, we maintain a big cache in RAM to store address mapping entries of data array recently searching. The search algorithm in searching the cache is binary search. The goal is to reduce address mapping overhead without excessively compromising system response time. This scheme is designed for our massive parallel coprocessor system. For reducing power consumption, we have an attempt to implement address mapping of each massive parallel coprocessor in Field-Programmable Gate Array(FPGA). The experiment have been conducted on a real System on chip(Soc). The result shows that when the number of processor node is 4096 and its frequency is 233MHz, The RAM cost is 2.4 MB in each processor, when there is missing, the largest response time is 160us, which is less than the mainstream software implementation in address translation. In the case of making full use of available storage resources, The hit ratio in graph problem could be achieve 100%. Yichun Sun, Xiaodong Yi 0002, Hengzhu Liu |
ICPADS | 3 |
| 2016 | CORDIC-Based Enhanced Systolic Array Architecture for QR DecompositionabstractMultiple input multiple output (MIMO) with orthogonal frequency division multiplexing (OFDM) systems typically use orthogonal-triangular (QR) decomposition. In this article, we present an enhanced systolic array architecture to realize QR decomposition based on the Givens rotation (GR) method for a 4 × 4 real matrix. The coordinate rotation digital computer (CORDIC) algorithm is adopted and modified to speed up and simplify the process of GR. To verify the function and evaluate the performance, the proposed architectures are validated on a Virtex 5 FPGA development platform. Compared to a commercial implementation of vectoring CORDIC, the enhanced vectoring CORDIC is presented that uses 37.7% less hardware resources, dissipates 71.6% less power, and provides a 1.8 times speedup while maintaining the same computation accuracy. The enhanced QR systolic array architecture based on the enhanced vectoring CORDIC saves 24.5% in power dissipation, provides a factor of 1.5-fold improvement in throughput, and the hardware efficiency is improved 1.45-fold with no accuracy penalty when compared to our previously proposed QR systolic array architecture. Paul Chow, Hengzhu Liu |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2015 | FPGA implementation of low-power and high-PSNR DCT/IDCT architecture based on adaptive recoding CORDICabstractThe discrete cosine transform (DCT) and its inverse (IDCT) are widely used in image and video compression standards. In this paper, we propose a novel unified architecture for DCT and IDCT based on adaptive recoding coordinate rotation digital computer (ARC). The proposed architecture requires two types of ARC rotators. In addition, an efficient adder and shifter-based scale factor approximation is used in the proposed architecture. To verify the function and evaluate the performance, the proposed architecture is validated on a Virtex 5 FPGA development platform. Under DCT-only mode, compared with the proposed architecture, a state-of-the-art DCT architecture uses 12% more hardware resources, increases the critical path delay by 7.12%, consumes 10.1% more power and decreases 4.8 dB in PSNR. Under DCT/IDCT mode, the latest unified DCT/IDCT architecture has a factor of 2.17-fold in latency, needs 74.9% more hardware resources and dissipates 52.5% more power when compared to the proposed architecture. In addition, PSNR of the proposed architecture is better by 2 dB. Paul Chow, Hengzhu Liu |
FPT | 3 |
| 2015 | An Enhanced Adaptive Recoding Rotation CORDICabstractThe Conventional Coordinate Rotation Digital Computer (CORDIC) algorithm has been widely used in many applications, particularly in Direct Digital Frequency Synthesizers (DDS) and Fast Fourier Transforms (FFT). However, CORDIC is constrained by the excessive number of iterations, angle data path, and scaling factor compensation. In this article, an enhanced adaptive recoding CORDIC (EARC) is proposed. It uses the enhanced adaptive recoding method to reduce the required iterations and adopts the trigonometric transformation scheme to scale up the rotation angles. Computing sine and cosine is used first to compare the core functionality of EARC with basic CORDIC; then a 16-bit DDS and a 1,024-point FFT based on EARC are evaluated to demonstrate the benefits of EARC in larger applications. All the proposed architectures are validated on a Virtex 5 FPGA development platform. Compared with a commercial implementation of CORDIC, EARC requires 33.3% less hardware resources, provides a twofold speedup, dissipates 70.4% less power, and improves accuracy in terms of the Bit Error Position (BEP). Compared to the state-of-the-art Hybrid CORDIC, EARC reduces latency by 11.1% and consumes 17% less power. Compared with a commercial implementation of DDS, the dissipated power of the proposed DDS is reduced by 27.2%. The proposed DDS improves Spurious-Free Dynamic Range (SFDR) by nearly 7 dBc and dissipates 21.8% less power when compared with a recently published DDS circuit. The FFT based on EARC dissipates a factor of 2.05 less power than the commercial FFT even when choosing the 100% toggle rate for the FFT based on EARC and the 12.5% toggle rate for the commercial FFT. Compared with a recently published FFT, the FFT based on EARC improves Signal-to-Noise Ratio (SNR) by 8.9 dB and consumes 7.78% less power. Paul Chow, Hengzhu Liu |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2014 | Flexible Virtual Channel Power-Gating for High-Throughput and Low-Power Network-on-ChipabstractPower-gating is a representative circuit level technique to mitigate leakage power. While in low-power Network-on-Chip (NoC) design, the former fine-grained power-gating methods will decrease network performance due to serial wake-up latency and head-of-line blocking. Therefore, we propose a flexible Virtual Channel (VC) management scheme for fine-grained power-gating to achieve high throughput and low-power. The proposed power-gating method with the early wake-up is evaluated by using some synthetic workloads. When compared with an optimized early wake-up power-gating technique, it can improve performance effectively in medium and high network loads, and increases the network throughput by 15.7%~44.1% for different synthetic loads, while keeps network power consumption as low as the optimized method. For the PARSEC application traces of token based protocol, it can significantly decrease packet latency by 20.3% on average, however only increases less than 3.6% peak power when compared with the optimized method. Xiantuo Tang, Zuocheng Xing, Hengzhu Liu |
DSD | 5 |
| 2014 | An efficient FPGA implementation of QR decomposition using a novel systolic array architecture based on enhanced vectoring CORDICabstractMultiple input multiple output (MIMO) - Orthogonal frequency division multiplexing (OFDM) systems typically use Orthogonal-triangular (QR) decomposition. In this paper, we present a novel systolic array architecture to realize QR decomposition based on the Givens rotation method for a 4 × 4 real matrix. The coordinate rotation digital computer (CORDIC) algorithm is adopted and modified to speed up and simplify the Givens rotation. To verify the function and evaluate the performance, the proposed architectures are validated on a Virtex 5 FPGA development platform. Compared to a commercial implementation of vectoring CORDIC, an enhanced vectoring CORDIC is presented that uses 37.7% less hardware resources, dissipates 76.8% less power and provides a 1.8 times speed-up while maintaining the same computation accuracy. The novel QR systolic array architecture based on the enhanced vectoring CORDIC saves 5% in hardware and the throughput is improved by a factor of two with no accuracy penalty when compared with the best previous version of the QR systolic array. Paul Chow, Hengzhu Liu |
FPT | 3 |
| 2011 | Power-efficient tree-based multicast support for Networks-on-ChipabstractIn this paper, a novel hardware support for multicast on mesh Networks-on-Chip (NoC) is proposed. It supports multicast routing on any shape of tree-based paths. Two power-efficient tree-based multicast routing algorithms, Optimized tree (OPT) and Left-XY-Right-Optimized tree (LXYROPT) are also proposed. XY tree-based (XYT) algorithm and multiple unicast copies (MUC) are also implemented on the router as baselines. Along with the increase of the destination size, compared with MUC, OPT and LXYROPT achieve a remarkable improvement in both latency and throughput while the average power consumption is reduced by 50% and 45%, respectively. Compared with XYT, OPT is 10% higher in latency but gains 17% saving in power consumption. LXYROPT is 3% lower in latency and 8% lower in power consumption. In some cases, OPT and LXYROPT give power saving up to 70% less than the XYT. Wenmin Hu, Zhonghai Lu, Axel Jantsch, Hengzhu Liu |
ASP-DAC | 4 |
| 2011 | Network-on-Chip multicasting with low latency path setupabstractA low-latency path setup approach with multiple setup packets for parallel set is presented. It reduces the header overhead compared to multiaddress encoding. Further, we propose four variants of deadlock-free multicast routing algorithms using different subpath generation methods, different destination partitioning, and channel sharing strategies. Experimental results show that the quatuor partitions path-like tree outperforms other algorithms. Wenmin Hu, Zhonghai Lu, Axel Jantsch, Hengzhu Liu, Botao Zhang 0002, Dongpei Liu |
VLSI-SoC | 4 |
| 2010 | Domain specific architecture for next generation wireless communicationabstractIn order to solve the challenges in processor design for the next generation wireless communication systems, this paper first proposes a system level design flow for communication domain specific processor, and then proposes a novel processor architecture for the next generation wireless communication named GAEA using this design flow. GAEA is a shared memory multi-core SoC based on Software Controlled Time Division Multiplexing Bus, with which programmers can easily explore memory-level parallelism of applications by proper instructions and scheduling algorithms. MPE, which is the kernel component of GAEA, adopts hybrid parallel processing scheme to explore instruction-level and data-level parallelism. The pipeline and instruction set of GAEA are also optimized for the next generation wireless communication systems. The evaluation and implementation results show that GAEA architecture is suitable for the next generation wireless communication systems. Botao Zhang 0002, Hengzhu Liu, Fangzheng Mo |
DATE | 2 |
| 2010 | YHFT-QDSP: High-Performance Heterogeneous Multi-Core DSP
Shuming Chen, Jianghua Wan, Jianzhuang Lu, Hai-Yan Sun, Yong-Jie Sun, Hengzhu Liu, Xiang-Yuan Liu, Zhentao Li |
J. Comput. Sci. Technol. | 7 |
| 2009 | Low Complexity DVB-S2 LDPC DecoderabstractThe decoding complexity is the main issue in designing DVB-S2 Low Density Parity Check (LDPC) decoder. This paper proposes a low complexity decoder, which is based on a new row message passing Offset Min-Sum algorithm. The proposed algorithm can reduce the complexity of the decoder with no performance loss. The simplified node update units based on the proposed algorithm and the memory organization optimized across different code rates play the key role in reducing the complexity. The synthesized area of the decoder is 9.6 mm2in Chartered 90 nm COMS technology. When the code rate is 9/10 and the work frequency is 320 MHz, the net throughput of the decoder is 998 Mbps. Botao Zhang 0002, Hengzhu Liu, Xucan Chen, Dongpei Liu, Xiaofei Yi |
VTC Spring | 2 |
| 2006 | A Parallel Mutual Information Based Image Registration Algorithm for Applications in Remote Sensing
Haifang Zhou, Panfeng Wang, Xuejun Yang, Hengzhu Liu |
ISPA | 5 |
| 2006 | GPGC: a Grid-enabled parallel algorithm of geometric correction for remote-sensing applicationsabstractAbstract ChinaGrid is an important project sponsored by the China Ministry of Education, aiming to provide high‐performance services in a Grid computing environment. In this paper, one of the applications offered by ChinaGrid, parallel remote‐sensing image processing, is described. Geometric correction is a basic step during the processing of remote‐sensing imagery, which is traditionally a computation‐intensive and communication‐intensive application if in parallel mode. In order to move this application into a Grid, a new Grid‐enabled parallel algorithm of geometric correction is proposed, called GPGC. GPGC changes the frequent and fine‐grain communication mode of the existing parallel method into a delayed but concentrated exchanging mode by computing an irregular local output area. This change means no communication or synchronization happens during resampling that occupies most of the execution time. To prove its efficiency, the complexity of GPGC is analyzed in theory. Finally, performance testing of GPGC and its application in ChinaGrid are given. Experimental results show that our algorithm is more suitable for a Grid platform, excelling the old method in both performance and salability. Copyright © 2006 John Wiley & Sons, Ltd. Haifang Zhou, Xuejun Yang, Hengzhu Liu, Yu Tang 0014 |
Concurr. Comput. Pract. Exp. | 3 |
| 2005 | A Proposal of Parallel Strategy for Global Wavelet-Based Registration of Remote-Sensing Images
Haifang Zhou, Yu Tang 0014, Xuejun Yang, Hengzhu Liu |
ICA3PP | 4 |
| 2005 | First Evaluation of Parallel Methods of Automatic Global Image Registration Based on WaveletsabstractWith the increasing importance of multiple multiplatform remote sensing missions, fast and automatic integration of digital data from disparate sources has become critical to the success of these endeavors. Firstly, an overview of development of automatic and parallel global image registration is given. And then, based on the analyses of existing three parallel methods of wavelet-based global registration, a new parallel strategy is proposed. Moreover, towards the quantitative evaluation, first results of the intercomparision of four parallel global registration algorithms are presented in theory and in experiments. Haifang Zhou, Xuejun Yang, Hengzhu Liu, Yu Tang 0014 |
ICPP | 3 |