VLDB 2026 Research / reviewers in the wild / expert
Zhiwu Huang
dblp:47/7711
· DBLP profile ↗
105ranked-venue papers
26as first author
55since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 15 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 12 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 18 · 4 since 2021Systems, architecture and hardware · 16 · 7 first-author · 12 since 2021Computer networks · 9 · 5 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Penny-Wise and Pound-Foolish in AI-Generated Image DetectionabstractThe rise of AI-generated images has sparked serious concerns about their potential misuse across various domains, prompting the urgent need for robust detection methods. Despite advancements, many current approaches prioritize short-term gains at the expense of long-term effectiveness. This paper critiques the overly specialized approach of fine-tuning pre-trained models for short-term gains on a single AI image dataset, while disregarding the long-term imperative of achieving generalization and knowledge retention. To address this trade-off issue, we propose a novel learning framework (PoundNet) for the generalization of AI-generated image detection on a pre-trained vision-language model. PoundNet incorporates a learnable prompt design and a balanced objective to preserve broad knowledge from upstream tasks (object classification) while enhancing generalization for downstream tasks (AI-generated image detection). We train PoundNet on a single standard AI image dataset, following common practice in the literature. We then evaluate its performance across 10 large-scale public AI-generated image detection datasets with 5 main evaluation metrics, forming the largest benchmark test set for assessing the generalization ability of AI-generated image detection models, to our knowledge. The comprehensive benchmark evaluation demonstrates that PoundNet successfully balances generalization with knowledge retention, achieving a remarkable relative improvement of 19% in AI-generated image detection performance compared to state-of-the-art methods, while maintaining a strong performance of 63% on object classification tasks. Yabin Wang 0001, Zhiwu Huang, Zhou Su 0001, Adam Prügel-Bennett, Xiaopeng Hong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Linguistic profiling of deepfakes: An open database for next-Generation deepfake detection
Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhiwu Huang |
Pattern Recognit. | 5 |
| 2026 | Physics-Aware Spatial-Temporal Vehicle Trajectory Prediction With Discriminative LearningabstractAccurate prediction of vehicle trajectories in complex traffic environments is essential for path planning and safety decisions in autonomous driving systems. However, purely data-driven models lack physical constraints, making it challenging to ensure reliability and consistency in dynamic traffic scenarios, while physics models face challenges in maintaining long-term prediction reliability under complex traffic conditions. To address these issues, a trajectory prediction method is proposed by combining a data-driven model based on graph neural networks and Informer with a physics model, fused through discriminative learning. Firstly, graph neural networks are utilized to extract spatial information, and the Informer is used for long-term trajectory encoding and decoding to capture the temporal dynamics of trajectories. Then, to improve the physical plausibility of trajectory predictions, a kinematic model with Cubature Kalman Filtering is employed to estimate the trajectory distribution. Furthermore, discriminative learning is designed to fuse a physics model into the decoder of the data-driven model using a generative adversarial approach. This paper conducts ablation experiments and comparative tests on real-world highway trajectory data from the NGSIM dataset. The evaluation confirms that the proposed method achieves consistent and high-quality prediction results across multiple scenarios. Zhiwu Huang, Yicong He, Guoyu Gu, Yongjie Liu, Zhuozhuo Zhang, Xiaoyong Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2026 | Energy-Efficient Multi-UAV Navigation for Cooperative Data Sensing and TransmissionabstractUnmanned aerial vehicles (UAVs) hold significant potential for sensing services in a large scope of area, thanks to their wide coverage and adaptable deployment. Considering the complex environment dynamics and limited sensing range, navigating multiple UAVs in a distributed way becomes challenging to implement cooperative data sensing and transmission tasks. In this paper, we optimize the trajectory design of UAVs by jointly considering the collected data volume, geographical fairness and limited energy reserve during their service period. To achieve the long-term serving objective, a memory augmented multi-agent deep reinforcement learning approach is presented to ensure energy-efficient distributed trajectory design with partial observations. Specifically, the intrinsic criterion is developed to enhance UAV spatial exploration when reaching the boundary of explored regions. Then, to address the information loss caused by incomplete observations, the spatial-temporal memory augmented actor-critic architecture is designed to extract historical contextual features for multi-UAV cooperative navigation. Furthermore, the prioritized experience replay mechanism is incorporated to enhance important experience exploitation for UAV collaboration. Extensive simulations using two real-world datasets in Shenzhen and Beijing demonstrate that the proposed method outperforms the state-of-the-art methods in terms of data collection ratio, geographical fairness, and energy consumption ratio. Hu He 0003, Jun Peng 0001, Lin Cai 0001, Weirong Liu 0001, Xin Gu 0002, Zhiwu Huang |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | Continual Conceptual Entity Learning for Text-to-Image Generative ModelsabstractCurrent Text-to-Image generative models struggle to continuously learn multiple distinct entities or concepts, limiting their scalability and hindering practical deployment in dynamic environments. We formulate this task as Continual Conceptual Entity Learning (CEL) and propose a novel framework called Continual Entity Adapter Learning (CEAL). CEAL leverages a compact set of tunable parameters, termed SuperLoRA, to efficient and scalable learning of new entities. We propose a dynamic rank-increasing strategy to train the SuperLoRA, balancing computational efficiency with performance. To evaluate our method, we create three benchmarks encompassing generic objects, human faces, and artistic styles. Experimental results demonstrate that CEAL effectively learns new entities while preserving prior knowledge, outperforming existing methods in both entity fidelity and parameter efficiency. Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhou Su 0001, Zhiwu Huang |
IEEE Trans. Multim. | 6 |
| 2025 | OpenSDI: Spotting Diffusion-Generated Images in the Open WorldabstractThis paper identifies OpenSDI, a challenge for spotting diffusion-generated images in open-world settings. In response to this challenge, we define a new benchmark, the OpenSDI dataset (OpenSDID), which stands out from existing datasets due to its diverse use of large vision-language models that simulate open-world diffusion-based manipulations. Another outstanding feature of OpenSDID is its inclusion of both detection and localization tasks for images manipulated globally and locally by diffusion models. To address the OpenSDI challenge, we propose a Synergizing Pretrained Models (SPM) scheme to build up a mixture of foundation models. This approach exploits a collaboration mechanism with multiple pretrained foundation models to enhance generalization in the OpenSDI context, moving beyond traditional training by synergizing multiple pretrained models through prompting and attending strategies. Building on this scheme, we introduce MaskCLIP, an SPM-based model that aligns Contrastive Language-Image Pre-Training (CLIP) with Masked Autoencoder (MAE). Extensive evaluations on OpenSDID show that MaskCLIP significantly outperforms current state-of-the-art methods for the OpenSDI challenge, achieving remarkable relative improvements of 14.23% in IoU (14.11% in F1) and 2.05% in accuracy (2.38% in F1) compared to the second-best model in localization and detection tasks, respectively. Our dataset and code are available at https://github.com/iamwangyabin/OpenSDI. Yabin Wang 0001, Zhiwu Huang, Xiaopeng Hong |
CVPR | 2 |
| 2025 | Vehicle Trajectory Prediction with Driving Style-Aware Spatial-Temporal Fusion NetworkabstractVehicle trajectory prediction is a critical and complex task in autonomous driving systems, where accurate prediction is essential to ensure both safety and comfort. Given that driving style impacts future trajectory prediction, integrating driving style information is crucial. In this paper, a trajectory prediction framework is proposed, in which a spatial-temporal information fusion network incorporating driving style is leveraged. Driving style is captured through a denoising Transformer autoencoder for dimensionality reduction and refined using fuzzy k-means++ clustering. Temporal information and spatial interactions of the vehicle are dynamically extracted using Transformer and graph neural network, and the future trajectory distribution is generated using a Transformer decoder enhanced by the KAN network with a sine function. Ablation and comparison experiments are carried out on the public NGSIM dataset. The results demonstrate that our model outperforms others in prediction accuracy, with improvements of up to 29.7% across evaluation metrics. Zhiwu Huang, Yicong He, Guoyu Gu, Heng Li 0005, Yongjie Liu |
IECON | 1 |
| 2025 | Sequential Intention-driven Vehicle Trajectory Prediction Integrated with Spatial-Temporal FeaturesabstractOur framework employs a hybrid model of Bidirectional Temporal Convolutional Network and Bidirectional Gated Recurrent Unit for sequential intention prediction. Additionally, a TransformerConv-based Graph Attention Network captures spatial interactions from historical frames and incorporates temporal features to generate spatial-temporal features, which are further processed by an intention-inspired context extraction attention mechanism to generate inputs for final trajectory prediction. On the NGSIM US-101 and I-80 datasets, our model achieves a 93.20% accuracy in predicting sequential driving intentions at a prediction horizon of 3 seconds. By incorporating these predicted intentions into the trajectory prediction network, it reduces the RMSE by 11.5% over a prediction horizon of 5 seconds compared to state-of-the-art methods, demonstrating its effectiveness in highway scenarios. Zhiwu Huang, Zhuozhuo Zhang, Zini Wang, Heng Li 0005, Yongjie Liu |
IECON | 1 |
| 2025 | Battery Health and Shifting-Aware Gear Ratio Optimization for Distributed Drive Electric TrucksabstractGear ratio optimization is essential for improving transmission efficiency and dynamic performance of four-wheel distributed drive electric heavy trucks. This study proposes a gear ratio optimization method that integrates battery health and shifting-induced energy losses and considers shift frequency. The method employs particle swarm optimization to optimize front and rear axle gear ratios under realistic truck operating conditions, followed by dynamic programming to determine the optimal gear-shifting sequence and real-time torque allocation. Through iterative refinement, the proposed method achieves optimal gear ratios of [32.17, 18.16] for the front axle and [35.53, 14.61] for the rear axle. Simulation results demonstrate that the optimized configuration reduces annual operational costs by 0.5%-1.9% compared to other optimization methods, yielding savings of 103,328 RMB per year while mitigating battery degradation and kinetic energy loss during gear shifts. Shaokun Li, Zhiwu Huang, Yue Wu 0024, Xiaoyong Zhang 0001 |
IECON | 2 |
| 2025 | Cooperative Reinforcement Learning for Car-Following and Energy Management Optimization of Dual-Motor Electric VehiclesabstractFor distributed drive electric vehicles, energy consumption is affected by the power demand and energy management strategy. In this paper, an adaptive cruise control and energy management strategy cooperative framework for dual-motor electric vehicles is proposed based on the deep deterministic policy gradient algorithm. Firstly, the energy management problem in the car-following scenario is decomposed into two subproblems: the adaptive cruise control governs vehicle acceleration, while the energy management strategy allocates driving torque. Then, based on the cooperative architecture, the speed trajectory and torque allocation strategy are co-optimized to realize cooperation between the adaptive cruise control and the energy management strategy. Finally, the proposed cooperative strategy is compared with the traditional hierarchical strategy under the worldwide harmonized light vehicles test cycle. Results show that the proposed cooperative strategy can reduce the maximum acceleration and maximum deceleration by 10.8% and 10.4%, respectively, and improve energy consumption by 4.3% compared with the traditional hierarchical strategy. Zhiwu Huang, Yue Wu 0024, Shaokun Li, Xiaoyong Zhang 0001 |
IECON | 2 |
| 2025 | Air-Ground Collaborative Mobile Crowdsensing by Predictive Multi-Agent Deep Reinforcement LearningabstractMobile crowdsensing (MCS) by human participants and unmanned aerial vehicles (UAVs) is an emerging air-ground collaborative data collection paradigm by navigating a group of UAVs to collaborate with human participants to provide large-scale and fine-grained sensing services. In this paper, we aim to optimize the trajectory design of UAVs by jointly considering the collected data volume, geographical sensing fairness, and limited energy reserve during the serving period. To achieve the long-term serving objective, we propose a human participants distribution prediction based multi-agent deep reinforcement learning method for efficient UAV navigation to collaborate with human participants in performing MCS tasks. Specifically, we first introduce a region division based human participant spatial distribution prediction method to help UAVs to collaborate with human participants by the predictive mobility flows. Then, we present the multi-agent proximal policy optimization (MAPPO) based method for efficient UAV navigation decision-making. Extensive simulations and trajectory visualization using the real-world mobility dataset in KAIST show that the proposed method consistently outperforms the state-of-the-art in terms of the energy efficiency when varying the number of UAVs and human participants. Hu He 0003, Jun Peng 0001, Lin Cai 0001, Weirong Liu 0001, Zhiwu Huang |
VTC2025-Fall | 5 |
| 2025 | Mobility and Context-Aware Precaching Strategy Using Spatial-Temporal Informer for Vehicular ServiceabstractWith the rapid development of vehicle-to-everything technology, vehicular edge caching has emerged as a crucial component for managing frequently accessed content at the network’s edge. However, due to vehicles’ high mobility, it is challenging to determine where and which content needs to be cached. To address this issue, a mobility and context-aware precache strategy is proposed to proactively prefetch and replace content in two steps. First, by integrating the traffic features from vehicles and roads, a spatial-temporal informer-based model is designed to predict long-term vehicle trajectories. Subsequently, a proactive context-aware precache strategy is proposed. By analyzing the context of different cache types, the required content can be further accurately estimated according to the cache type and workload. Extensive simulations based on real-world mobility scenarios are conducted to validate the performance of the proposed method. The results show that the proposed method can improve prediction accuracy and cache hit rate by 34.56% and 18.89%, and reduce mean response time and total energy cost by 6.1% and 2.65% compared to the existing precaching methods. Jun Peng 0001, Lin Cai 0001, Weirong Liu 0001, Hu He 0003, Zhiwu Huang |
IEEE Internet Things J. | 7 |
| 2024 | A Rapid Charging Strategy Based on Joint Optimization of Charging Time and Aging DegradationabstractLithium-ion batteries are widely used in portable devices and mobile medical equipment due to their high energy density and long cycle life. However, long charging times for lithium-ion batteries can limit their usability. This paper proposes a multi-stage constant current charging protocol. Kaifu Guan, Zhiwu Huang, Yongjie Liu, Yue Wu 0024, Yunsheng Fan, Heng Li 0005 |
HealthCom | 2 |
| 2024 | Lateral Control of Autonomous Vehicles Using Barrier Lyapunov FunctionabstractGuaranteed safety and performance under different cases have significant influence on the development of lateral control of autonomous vehicles. This paper proposes a robust nonlinear controller using barrier Lyapunov function within the constraints. In the case of unknown bounded uncertainty, the barrier Lyapunov function is appropriately arranged into the controller to restrict the state variables of the designed safe region in the process. Thus, the proposed method is composed of the nonlinear controller using barrier Lyapunov function and kalman filter. The kalman filter is designed as an observer to estimate the state variables. At the meantime, the method we proposed meets the constraints of output and the external disturbance. Moreover, the validity of the proposed method is validated in the co-simulation of the MATLAB/Simulink and CarSim. Zhiwu Huang, Liuye Shao, Bin Chen 0017, Yue Wu 0024, Boyu Shu, Heng Li 0005 |
HPCC | 1 |
| 2024 | Cooperative Cell Balancing For Supercapacitors With Reinforcement LearningabstractWith the rapid advancement of technologies such as electric vehicles, the demand for energy storage devices has surged, leading to the widespread adoption of supercapacitors due to their numerous advantages. In practical applications, supercapacitors are often arranged in series or parallel configurations to form capacitor banks, which cater to higher voltage or capacity requirements. However, inconsistencies in the manufacturing processes and materials can lead to variations in the electrical performance of individual supercapacitors, necessitating effective balance management. Existing balancing methods are generally categorized into passive and active approaches. While passive balancing circuits are simple and cost-effective, they tend to be inefficient. On the other hand, active balancing methods, although offering high control precision, are typically more complex and expensive. This paper introduces a collaborative balancing strategy based on Deep Deterministic Policy Gradient (DDPG) using a switch resistor circuit, which serves as an intermediate approach between passive and active methods by combining their respective advantages. By integrating deep reinforcement learning with the switch resistor circuit for supercapacitor balancing, the proposed method addresses the slow balancing speed of traditional circuits under significant voltage disparities, enhancing the robustness of the balancing process and achieving superior performance. A simulation environment is established in Simulink to evaluate the effectiveness of the proposed method under various initial voltage conditions. The results demonstrate that the supercapacitor bank achieves balance within a short time frame. Moreover, comparative experiments indicate that the collaborative strategy significantly reduces overshoot and improves the robustness of supercapacitor balancing compared to noncollaborative approaches. Zhiwu Huang, Yundong Song, Yunsheng Fan, Shilong Zhuo, Taozhen Chang, Heng Li 0005 |
HPCC | 1 |
| 2024 | Core Temperature-Aware Optimal Preheating Strategy for Lithium-ion BatteryabstractLithium-ion batteries are the crucial energy source for electric vehicles. However, they experience capacity degeneration when used in low-temperature environments. It is necessary to preheat them before using. In this paper, a core temperature-aware optimal preheating strategy, featuring a multi-stage constant-current discharge heating method, is proposed to heat lithium-ion batteries in low-temperature environments. Firstly, this paper builds an internal battery temperature distribution model based on Fourier’s law of heat conduction. Secondly, the temperature distribution model is coupled within the battery model to display the comprehensive performance of the battery. Thirdly, decreasing heating time and reducing capacity loss jointly formulate a multi-objective optimization problem solved by dynamic programming(DP) algorithm. Judging by simulation results, heating time is downsized and the capacity loss is reduced at the same time, proving the progressiveness of the proposed strategy. Zhiwu Huang, Yongjie Liu, Kaifu Guan, Lisen Yan |
HPCC | 1 |
| 2024 | Reinforcement Learning-Driven Relay Selection for Enhanced V2V Communication in Vehicle PlatoonsabstractIn truck platoons with a bidirectional-leader topology, variations in channel conditions result in unreliability and high latency in vehicle-to-vehicle (V2V) communications. This paper proposes an adaptive relay selection strategy based on Q-learning (QL). The strategy ensures that all vehicles in the platoon receive safety messages from the lead vehicle quickly and reliably. Firstly, relay selection is modeled as a Markov decision process (MDP). The lead vehicle and the relays act as intelligent agents. Agents make decisions adaptively based on real-time state observations in a dynamic communication environment. Secondly, a reward function is designed based on platoon topology and channel state information statistics (CSI). The purpose is to drive the proposed strategy to learn the optimal strategy for message transmission under different environments. Lastly, the simulation results demonstrate the effectiveness and robustness of the proposed algorithm. In various channel attenuation environments, the strategy has been demonstrated to enhance the packet delivery ratio (PDR) for the platoon tail and significantly increase the platoon’s throughput. Xiaoyong Zhang 0001, Xin Gu 0002, Jun Peng 0001, Heng Li 0005, Zhiwu Huang, Weirong Liu 0001 |
HPCC | 6 |
| 2024 | A Counterfactual Reasoning-based Trajectory Prediction Model for Multiple AgentsabstractAccurate trajectory prediction is crucial in autonomous driving to ensure safe and efficient navigation, yet effectively modeling complex interactions among multiple agents remains a significant challenge. Many existing methods still suffer from over-reliance on HD maps, high computational cost, and a lack of interpretability in interaction reasoning. In response, the proposed model innovatively incorporates counterfactual reasoning into social interaction modeling to tackle the challenges of interaction-aware multi-agent trajectory prediction, prioritizing both accuracy and efficiency. In light of the spatiotemporal interaction mechanism and the inherent human cognition governing agents’ motion, the approach simultaneously generates multi-modal trajectories for all agents in a scenario, providing a novel perspective for modeling social interactions through causal reasoning. The results demonstrate that our map-free, lightweight trajectory prediction model rivals the performance of state-of-the-art methods and shows notable improvements over various baselines on publicly available real-world datasets. Zhiwu Huang, Xinshu Yang, Heng Li 0005, Hongjiang He, Jing Wang 0005 |
IECON | 1 |
| 2024 | AI Robust Anomaly Localization for DC Microgrid Using Adversarial Autoencoder
Jieqi Rong, Weirong Liu 0001, Heng Li 0005, Lisen Yan, Jun Peng 0001, Zhiwu Huang |
MobiQuitous | 7 |
| 2024 | Optimal Operator-based Modeling for Open Circuit Voltage Hysteresis of LiFePO4 BatteriesabstractAccurate modeling of open circuit voltage hysteresis for LiFePO4batteries is crucial for establishing an advanced battery model. However, existing hysteresis modeling methods often yield suboptimal results due to inadequate parameterization. This paper proposes an optimal modeling method for open circuit voltage hysteresis based on the Prandtl-Ishlinskii model and an associated parameterization method. First, an asymmetric operator with cubic envelope functions is designed to enhance the classical Prandtl-Ishlinskii model, which originally features a symmetric and linear operator. This modification enables the proposed model to accurately capture intricate hysteresis. Second, a hierarchical parameterization method is proposed to identify optimal parameters. Specifically, an improved grey wolf optimizer is employed to determine the operator-related parameters. Then, the remaining parameters are calculated using the least squares algorithm, enhancing computational efficiency. Finally, the proposed model is validated on the experimental hysteresis data from three distinct scenarios. The modeling error of the proposed model decreased by 66.57 % and 32.51 % compared with two other benchmark models. Lisen Yan, Jun Peng 0001, Yue Wu 0024, Heng Li 0005, Zhiwu Huang |
SMC | 6 |
| 2024 | A novel multi-scale competitive network for fault diagnosis in rotating machinery
Zhiwu Huang, Xinlong Zhao |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Adaptive Log-Euclidean Metrics for SPD Matrix LearningabstractSymmetric Positive Definite (SPD) matrices have received wide attention in machine learning due to their intrinsic capacity to encode underlying structural correlation in data. Many successful Riemannian metrics have been proposed to reflect the non-Euclidean geometry of SPD manifolds. However, most existing metric tensors are fixed, which might lead to sub-optimal performance for SPD matrix learning, especially for deep SPD neural networks. To remedy this limitation, we leverage the commonly encountered pullback techniques and propose Adaptive Log-Euclidean Metrics (ALEMs), which extend the widely used Log-Euclidean Metric (LEM). Compared with the previous Riemannian metrics, our metrics contain learnable parameters, which can better adapt to the complex dynamics of Riemannian neural networks with minor extra computations. We also present a complete theoretical analysis to support our ALEMs, including algebraic and Riemannian properties. The experimental and theoretical results demonstrate the merit of the proposed metrics in improving the performance of SPD neural networks. The efficacy of our metrics is further showcased on a set of recently developed Riemannian building blocks, including Riemannian batch normalization, Riemannian Residual blocks, and Riemannian classifiers. Ziheng Chen 0001, Yue Song 0002, Tianyang Xu 0001, Zhiwu Huang, Xiaojun Wu 0001, Nicu Sebe |
IEEE Trans. Image Process. | 4 |
| 2024 | An Optimized Prediction Horizon Energy Management Method for Hybrid Energy Storage Systems of Electric VehiclesabstractModel predictive control is a real-time energy management method for hybrid energy storage systems, whose performance is closely related to the prediction horizon. However, a longer prediction horizon also means a higher computation burden and more predictive uncertainties. This paper proposed a predictive energy management strategy with an optimized prediction horizon for the hybrid energy storage system of electric vehicles. Firstly, the receding horizon optimization problem is formulated to minimize the battery degradation cost and traction electricity cost for the electric vehicle operation. Then, the optimal control sequence is solved to obtain the power allocation between the battery and the supercapacitor. Furthermore, the effect of different horizons on the optimization results is analyzed under diverse operating conditions, determining the optimal horizon to balance the system costs and computation burden. Compared with the short horizon, the optimal horizon can achieve 5.2%$\sim$8.5% performance improvement with the acceptable computation time approaching 1 s. Zini Wang, Zhiwu Huang, Yue Wu 0024, Weirong Liu 0001, Heng Li 0005, Jun Peng 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | AI-Enabled Spatial-Temporal Mobility Awareness Service Migration for Connected VehiclesabstractIn the future 6G intelligent transportation system, the edge server will bring great convenience to the timely computing service for connected vehicles. To guarantee the quality of service, the time-critical services need to be migrated according to the future location of the vehicle. However, predicting vehicle mobility is challenging due to the time-varying of road traffic and the complex mobility patterns of vehicles. To address this issue, a spatial-temporal awareness proactive service migration strategy is proposed in this paper. First, a spatial-temporal neural network is designed to obtain accurate mobility by using gated recurrent units and graph convolutional layers extracting features from spatial road traffic and multi-time scales driving data. Then a proactive migration method is proposed to guarantee the reliability of services and reduce energy consumption. Considering the reliability of services and the real-time workload of servers, the migration problem is modeled as a multi-objective optimization problem, and the Lyapunov optimization method is utilized to obtain utility-optimal migration decisions. Extensive simulations based on real-world datasets are performed to validate the performance of the proposed method. The results show that the proposed method achieved 6% higher prediction accuracy, 10% lower dropping rate, and 10% lower energy consumption compared to state-of-the-art methods. Jun Peng 0001, Lin Cai 0001, Weirong Liu 0001, Xin Gu 0002, Zhiwu Huang |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | Resource Reservation Coordination for Vehicle Platooning in C-V2X NetworksabstractHigh-reliability and low-latency communication is essential for timely information exchange in vehicle platooning. As a key enabler of this, the cellular vehicular-to-everything (C-V2X) network uses a sensing-based semi-persistent scheduling (SPS) protocol, where radio resources are reserved for a number of transmissions with reduced resource re-allocation and control overhead. However, consecutive access collisions may be caused by reservation conflict, which leads to long delay and threatens platoon’s stability and safety. In this paper, a coordinating resource reservation (CRR) protocol is proposed for vehicle platooning. By implementing error detection with coordination among platoon vehicles, the resource reservation is improved for reduced collisions and delay. Specifically, packet reception/loss information is sent out by platoon vehicles through their own packets. Such information is shared with transmitters and guides them to reserve new resources when access collision occurs. As a result, long delay is avoided while no extra feedback packet is introduced. Furthermore, Markov analysis is presented to evaluate the performance of SPS and the proposed CRR for vehicle platooning, providing the quantified performance gains. Finally, simulation results demonstrate the superiority of the proposed CRR in reducing packet loss and latency, compared with the legacy SPS and other state-of-the-art solutions. Xin Gu 0002, Jun Peng 0001, Lin Cai 0001, Weirong Liu 0001, Xiaoyong Zhang 0001, Zhiwu Huang |
IEEE Trans. Wirel. Commun. | 6 |
| 2024 | Proactive Bandwidth Allocation for V2X Networks With Multi-Attentional Deep Graph LearningabstractThe increasing number of connected vehicles exacerbates the scarcity of spectrum resources in vehicle-to-everything (V2X) communication. To optimize the utilization of wireless resources, it is crucial to allocate the limited spectrum blocks to each roadside unit (RSU) based on the real-time bandwidth demand of vehicles within their coverage. However, the complex mobility patterns of vehicles and dynamic traffic conditions make it challenging to accurately and promptly estimate the bandwidth demand. To address this issue, a spatial-temporal multi-attentional network (STMA-net) is designed to predict the future bandwidth demand of RSUs. Based on the predicted bandwidth demand, a prediction error-compensable proactive bandwidth allocation algorithm is proposed to adaptively allocate spectrum resources and narrow the discrepancy between predicted and actual demand. Experimental results with realistic traffic in Bologna demonstrate that the proposed STMA-net achieves 11.25% higher prediction accuracy compared to state-of-the-art methods. Furthermore, the proposed proactive bandwidth allocation method outperforms existing methods, providing the highest throughput and serving 5% more vehicles while reducing the service drop rate by an order of magnitude. Jun Peng 0001, Lin Cai 0001, Weirong Liu 0001, Shuo Li 0006, Hu He 0003, Zhiwu Huang |
IEEE Trans. Wirel. Commun. | 7 |
| 2023 | Riemannian Local Mechanism for SPD Neural NetworksabstractThe Symmetric Positive Definite (SPD) matrices have received wide attention for data representation in many scientific areas. Although there are many different attempts to develop effective deep architectures for data processing on the Riemannian manifold of SPD matrices, very few solutions explicitly mine the local geometrical information in deep SPD feature representations. Given the great success of local mechanisms in Euclidean methods, we argue that it is of utmost importance to ensure the preservation of local geometric information in the SPD networks. We first analyse the convolution operator commonly used for capturing local information in Euclidean deep networks from the perspective of a higher level of abstraction afforded by category theory. Based on this analysis, we define the local information in the SPD manifold and design a multi-scale submanifold block for mining local geometry. Experiments involving multiple visual tasks validate the effectiveness of our approach. Ziheng Chen 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Zhiwu Huang, Josef Kittler |
AAAI | 5 |
| 2023 | Isolation and Impartial Aggregation: A Paradigm of Incremental Learning without InterferenceabstractThis paper focuses on the prevalent stage interference and stage performance imbalance of incremental learning. To avoid obvious stage learning bottlenecks, we propose a new incremental learning framework, which leverages a series of stage-isolated classifiers to perform the learning task at each stage, without interference from others. To be concrete, to aggregate multiple stage classifiers as a uniform one impartially, we first introduce a temperature-controlled energy metric for indicating the confidence score levels of the stage classifiers. We then propose an anchor-based energy self-normalization strategy to ensure the stage classifiers work at the same energy level. Finally, we design a voting-based inference augmentation strategy for robust inference. The proposed method is rehearsal-free and can work for almost all incremental learning scenarios. We evaluate the proposed method on four large datasets. Extensive results demonstrate the superiority of the proposed method in setting up new state-of-the-art overall performance. Code is available at https://github.com/iamwangyabin/ESN. Yabin Wang 0001, Zhiheng Ma, Zhiwu Huang, Yaowei Wang 0001, Zhou Su 0001, Xiaopeng Hong |
AAAI | 3 |
| 2023 | Freestyle Layout-to-Image SynthesisabstractTypical layout-to-image synthesis (LIS) models generate images for a closed set of semantic classes, e.g., 182 common objects in COCO-Stuff. In this work, we explore the freestyle capability of the model, i.e., how far can it generate unseen semantics (e.g., classes, attributes, and styles) onto a given layout, and call the task Freestyle LIS (FLIS). Thanks to the development of large-scale pre-trained language-image models, a number of discriminative models (e.g., image classification and object detection) trained on limited base classes are empowered with the ability of unseen class prediction. Inspired by this, we opt to leverage large-scale pre-trained text-to-image diffusion models to achieve the generation of unseen semantics. The key challenge of FLIS is how to enable the diffusion model to synthesize images from a specific layout which very likely violates its pre-learned knowledge, e.g., the model never sees “a unicorn sitting on a bench” during its pre-training. To this end, we introduce a new module called Rectified Cross-Attention (RCA) that can be conveniently plugged in the diffusion model to integrate semantic masks. This “plug-in” is applied in each cross-attention layer of the model to rectify the attention maps between image and text tokens. The key idea of RCA is to enforce each text token to act on the pixels in a specified region, allowing us to freely put a wide variety of semantics from pre-trained knowledge (which is general) onto the given layout (which is specific). Extensive experiments show that the proposed diffusion network produces realistic and freestyle layout-to-image generation results with diverse text inputs, which has a high potential to spawn a bunch of interesting applications. Code is available at https://github.com/essunny310/FreestyleNet. Zhiwu Huang, Qianru Sun, Li Song 0001, Wenjun Zhang 0001 |
CVPR | 2 |
| 2023 | Spatial-Temporal Data-Driven Speed Prediction for Energy Management of Battery/Supercapacitor Electric VehiclesabstractAccurate speed prediction plays a critical role in the predictive energy management of electric vehicles. This paper proposes a spatial-temporal data-driven speed prediction method for the predictive energy management of battery/supercapacitor electric vehicles. The proposed speed prediction method is performed using a long short-term memory network and validated on a real-world commuting data set in China. Different from existing prediction methods based only on speed and acceleration, we take spatial information as an additional input to improve speed prediction accuracy. The predicted future speed is then leveraged by a model predictive control-based energy management strategy to minimize the battery degradation cost. Quantitative comparisons illustrate that the proposed speed prediction method can reduce the root mean square error and mean absolute error by 10.01-19.15% compared with no spatial information prediction method. The more accurate prediction can further improve the optimality of the predictive energy management strategy, i.e., reduce the battery capacity loss and yield closer results to model predictive control with completely accurate prediction. Yue Wu 0024, Zhiwu Huang, Yunhong Che, Zini Wang, Jun Peng 0001 |
IECON | 2 |
| 2023 | Exploring the Hysteresis Effect of Li-ion Batteries: A Machine Learning based ApproachabstractWith the rapid development of electric vehicle industry, the battery management system of electric vehicle is the focus of research. Battery management is not only related to the safe driving of electric vehicles, but also the basis of intelligent driving of electric vehicles. The state-of-charge (SoC) estimation of battery is very important in battery management system. The battery is in a state of power consumption when the electric vehicle is running, but when the electric vehicle is braked, the kinetic energy will also be converted into electric energy to charge the battery. The acceleration and braking of electric vehicles are frequently switched. Therefore, the working conditions of electric vehicle batteries are complex, and the influence of battery hysteresis on the accuracy of SoC estimation cannot be ignored. In this paper, a lithium ion battery model considering hysteresis effect based on machine learning is proposed. The experiment was designed to collect the data of small cycle charge and discharge of the battery. The data were used to train the long short-term memory (LSTM) neural network model, and a battery model with hysteresis effect was obtained. It is verified that the model performs well in the test set, and the error of hysteresis voltage can be reduced to 0.002V. This model can be used for SoC estimation considering hysteresis effect. Sijie Zhang, Heng Li 0005, Yaoxin Xia, Lisen Yan, Zhiwu Huang |
IJCNN | 6 |
| 2023 | A Continual Deepfake Detection Benchmark: Dataset, Methods, and EssentialsabstractThere have been emerging a number of benchmarks and techniques for the detection of deepfakes. However, very few works study the detection of incrementally appearing deepfakes in the real-world scenarios. To simulate the wild scenes, this paper suggests a continual deepfake detection benchmark (CDDB) over a new collection of deepfakes from both known and unknown generative models. The suggested CDDB designs multiple evaluations on the detection over easy, hard, and long sequence of deepfake tasks, with a set of appropriate measures. In addition, we exploit multiple approaches to adapt multiclass incremental learning methods, commonly used in the continual visual recognition, to the continual deepfake detection problem. We evaluate existing methods, including their adapted ones, on the proposed CDDB. Within the proposed benchmark, we explore some commonly known essentials of standard continual learning. Our study provides new insights on these essentials in the context of continual deepfake detection. The suggested CDDB is clearly more challenging than the existing benchmarks, which thus offers a suitable evaluation avenue to the future research. Both data and code are available at https://github.com/Coral79/CDDB. Chuqiao Li, Zhiwu Huang, Danda Pani Paudel, Yabin Wang 0001, Mohamad Shahbazi, Xiaopeng Hong, Luc Van Gool |
WACV | 2 |
| 2023 | An Efficient Recurrent Adversarial Framework for Unsupervised Real-Time Video EnhancementabstractAbstract Video enhancement is a challenging problem, more than that of stills, mainly due to high computational cost, larger data volumes and the difficulty of achieving consistency in the spatio-temporal domain. In practice, these challenges are often coupled with the lack of example pairs, which inhibits the application of supervised learning strategies. To address these challenges, we propose an efficient adversarial video enhancement framework that learns directly from unpaired video examples. In particular, our framework introduces new recurrent cells that consist of interleaved local and global modules for implicit integration of spatial and temporal information. The proposed design allows our recurrent cells to efficiently propagate spatio-temporal information across frames and reduces the need for high complexity networks. Our setting enables learning from unpaired videos in a cyclic adversarial manner, where the proposed recurrent units are employed in all architectures. Efficient training is accomplished by introducing one single discriminator that learns the joint distribution of source and target domain simultaneously. The enhancement results demonstrate clear superiority of the proposed video enhancer over the state-of-the-art methods, in all terms of visual quality, quantitative metrics, and inference speed. Notably, our video enhancer is capable of enhancing over 35 frames per second of FullHD video (1080x1920). Dario Fuoli, Zhiwu Huang, Danda Pani Paudel, Luc Van Gool, Radu Timofte |
Int. J. Comput. Vis. | 2 |
| 2023 | Multi-agent actor-critic with time dynamical opponent modelabstractIn multi-agent reinforcement learning, multiple agents learn simultaneously while interacting with a common environment and each other. Since the agents adapt their policies during learning, not only the behavior of a single agent becomes non-stationary, but also the environment as perceived by the agent. This renders it particularly challenging to perform policy improvement. In this paper, we propose to exploit the fact that the agents seek to improve their expected cumulative reward and introduce a novel Time Dynamical Opponent Model (TDOM) to encode the knowledge that the opponent policies tend to improve over time. We motivate TDOM theoretically by deriving a lower bound of the log objective of an individual agent and further propose Multi-Agent Actor-Critic with Time Dynamical Opponent Model (TDOM-AC). We evaluate the proposed TDOM-AC on a differential game and the Multi-agent Particle Environment. We show empirically that TDOM achieves superior opponent behavior prediction during test time. The proposed TDOM-AC methodology outperforms state-of-the-art Actor-Critic methods on the performed tasks in cooperative and especially in mixed cooperative-competitive environments. TDOM-AC results in a more stable training and a faster convergence. Our code is available at https://github.com/Yuantian013/TDOM-AC. Yuan Tian 0014, Klaus-Rudolf Kladny, Qin Wang 0013, Zhiwu Huang, Olga Fink |
Neurocomputing | 4 |
| 2023 | Adaptive and Scalable Caching With Erasure Codes in Distributed Cloud-Edge Storage SystemsabstractErasure codes have been widely used to enhance data resiliency with low storage overheads. However, in geo-distributed cloud storage systems, erasure codes may incur high service latency as they require end users to access remote storage nodes to retrieve data. An elegant solution to achieving low latency is to deploy caching services at the edge servers close to end users. In this paper, we propose adaptive and scalable caching schemes to achieve low latency in the cloud-edge storage system. Based on the measured data popularity and network latencies in real time, an adaptive content replacement scheme is proposed to update caching decisions upon the arrival of requests. Theoretical analysis shows that the reduced data access latency of the replacement scheme is at least 50% of the maximum reducible latency. With the low computation complexity of our design, nearly no extra overheads will be introduced when handling intensive data flows. For further performance improvements without sacrificing its efficiency, an adaptive content adjustment scheme is presented to replace the subset of cached contents that incur the aforementioned performance loss. Driven by real-world data traces, extensive experiments based on Amazon Simple Storage Service demonstrate the effectiveness and efficiency of our design. Kaiyang Liu, Jun Peng 0001, Jingrong Wang, Zhiwu Huang, Jianping Pan 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Generative Flows with Invertible AttentionsabstractFlow-based generative models have shown an excellent ability to explicitly learn the probability density function of data via a sequence of invertible transformations. Yet, learning attentions in generative flows remains understudied, while it has made breakthroughs in other domains. To fill the gap, this paper introduces two types of invertible attention mechanisms, i.e., map-based and transformer-based attentions, for both unconditional and conditional generative flows. The key idea is to exploit a masked scheme of these two attentions to learn long-range data dependencies in the context of generative flows. The masked scheme allows for invertible attention modules with tractable Jacobian determinants, enabling its seamless integration at any positions of the flow-based models. The proposed attention mechanisms lead to more efficient generative flows, due to their capability of modeling the long-term data dependencies. Evaluation on multiple image synthesis tasks shows that the proposed attention flows result in efficient models and compare favorably against the state-of-the-art unconditional and conditional generative flows. Rhea Sanjay Sukthanker, Zhiwu Huang, Suryansh Kumar 0001, Radu Timofte, Luc Van Gool |
CVPR | 2 |
| 2022 | Energy Management Strategy for Hybrid Energy Storage System using Optimized Velocity Predictor and Model Predictive ControlabstractReasonable power distribution between battery and supercapacitor in electric vehicles is a crucial problem to improve energy consumption and economy. An online energy management strategy based on model predictive control (MPC) is proposed in this paper. Firstly, a radial basis function neural network optimized by particle swarm algorithm is presented to generate the short-term future velocity, i.e., the reference trajectory of the MPC. Then, a cost function considering the battery degradation cost and the electricity cost is constructed and optimized within each prediction horizon while maintaining the state of charge of the supercapacitor. Simulation results on the UDDS driving cycle show that the total cost of the proposed strategy is reduced by 6.3% and 3.9% compared with the near-optimal rule-based strategy and the none optimized velocity predictor-MPC, respectively, indicating that the velocity prediction accuracy has a significant impact on the performance of real-time energy management. Zhiwu Huang, Pei Huang 0020, Yue Wu 0024, Heng Li 0005, Jun Peng 0001 |
IV | 1 |
| 2022 | S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningabstractState-of-the-art deep neural networks are still struggling to address the catastrophic forgetting problem in continual learning. In this paper, we propose one simple paradigm (named as S-Prompting) and two concrete approaches to highly reduce the forgetting degree in one of the most typical continual learning scenarios, i.e., domain increment learning (DIL). The key idea of the paradigm is to learn prompts independently across domains with pre-trained transformers, avoiding the use of exemplars that commonly appear in conventional methods. This results in a win-win game where the prompting can achieve the best for each domain. The independent prompting across domains only requests one single cross-entropy loss for training and one simple K-NN operation as a domain identifier for inference. The learning paradigm derives an image prompt learning approach and a novel language-image prompt learning approach. Owning an excellent scalability (0.03% parameter increase per domain), the best of our approaches achieves a remarkable relative improvement (an average of about 30%) over the best of the state-of-the-art exemplar-free methods for three standard DIL tasks, and even surpasses the best of them relatively by about 6% in average when they use exemplars. Source code is available at https://github.com/iamwangyabin/S-Prompts. Yabin Wang 0001, Zhiwu Huang, Xiaopeng Hong |
NeurIPS | 2 |
| 2022 | Battery Aging-Robust Driving Range Prediction of Electric BusabstractThe prediction of driving range is very important for electric bus, but there is usually a difficulty: battery aging affects the accuracy of driving range prediction. In order to solve this problem, this paper proposes a driving range prediction method for electric bus, which is robust to the battery aging effect. Firstly, we extract the features that affect the driving range from the real-world dataset, quantify the correlation between them and the driving range by grey correlation analysis. Then through the feature enhancement technology, the time window processing is used to mitigate the influence of battery aging, and the time information hidden in the historical period sequence is deeply excavated. On this basis, we establish the driving range prediction model based on k-nearest neighbors regression, where the key parameters are optimized with the particle swarm optimization algorithm. Numerous experimental results show that compared with the classical methods, the method proposed in this paper has higher prediction accuracy especially when the batteries undergo significant aging effects. Heng Li 0005, Yongting Liu, Rui Zhang 0041, Jun Peng 0001, Zhiwu Huang |
TrustCom | 7 |
| 2022 | A Thermal-Aware Digital Twin Model of Permanent Magnet Synchronous Motors (PMSM) Based on BP Neural NetworksabstractEstimating accurate torque and speed is critical to control the operation of permanent magnet synchronous motors (PMSM). But the temperature factors are usually neglected in existing studies, which degrades estimation accuracy. In this paper, a thermal-aware digital twin model is proposed for PMSM to estimate motor torque and speed with the motor temperature and d-q axis current and voltage. Firstly, the motor parameters related to torque and speed are extracted by the Spearman correlation coefficients. Moreover, the stator winding temperature is selected as the input feature. Secondly, a digital model based on BP neural networks (BPNN) is established to estimate torque and speed. Thirdly, the parameters of the BPNN model are optimized by the whale optimization algorithm to accelerate the convergence speed and avoid local optima. Finally, experimental results show that the mean square error (MSE) of the BPNN model considering the temperature factors is reduced by 8.3%, which verifies that there is an effect of temperature on the torque and speed estimation. The MSE of the proposed method is reduced by 11.7% on average, which confirmed the higher accuracy of the proposed method compared with the classical BPNN model. Heng Li 0005, Peinan He, Yingze Yang, Bin Chen 0017, Jun Peng 0001, Zhiwu Huang |
TrustCom | 7 |
| 2022 | Markov Analysis of C-V2X Resource Reservation for Vehicle PlatooningabstractVehicle platooning utilizes automated driving and communication to let a group of vehicles travel closely, which improves road safety, traffic efficiency and fuel economy. In a platoon system, a critical task is to guarantee reliable communication among vehicles with efficient medium access control (MAC). This paper focuses on the feasibility of the distributed resource reservation MAC for communications among platoon vehicles using the cellular vehicle-to-everything (C-V2X) technology. For this purpose, a Markov chain-based model is proposed, which precisely estimates the network performance with different information flow topologies and system configurations. The state transition matrix is deduced and the stable state distribution is obtained. Given the information flow topology, we derive the probability that a platoon vehicle successfully delivers packets to all of the designated receivers. Finally, simulation results validate the analysis. To better implement the MAC protocol in practice, we also discuss the success probability for various information flow topologies in platoon communication. Xin Gu 0002, Jun Peng 0001, Lin Cai 0001, Xiaoyong Zhang 0001, Zhiwu Huang |
VTC Spring | 5 |
| 2022 | Neural Architecture Search for Efficient Uncalibrated Deep Photometric StereoabstractWe present an automated machine learning approach for uncalibrated photometric stereo (PS). Our work aims at discovering lightweight and computationally efficient PS neural networks with excellent surface normal accuracy. Unlike previous uncalibrated deep PS networks, which are handcrafted and carefully tuned, we leverage differentiable neural architecture search (NAS) strategy to find uncalibrated PS architecture automatically. We begin by defining a discrete search space for a light calibration network and a normal estimation network, respectively. We then perform a continuous relaxation of this search space and present a gradient-based optimization strategy to find an efficient light calibration and normal estimation network. Directly applying the NAS methodology to uncalibrated PS is not straightforward as certain task-specific constraints must be satisfied, which we impose explicitly. Moreover, we search for and train the two networks separately to account for the Generalized Bas-Relief (GBR) ambiguity. Extensive experiments on the DiLiGenT dataset show that the automatically searched neural architectures performance compares favorably with the state-of-the-art uncalibrated PS methods while having a lower memory footprint. Francesco Sarno, Suryansh Kumar 0001, Berk Kaya, Zhiwu Huang, Vittorio Ferrari, Luc Van Gool |
WACV | 4 |
| 2022 | A Learning-Based Data Placement Framework for Low Latency in Data Center NetworksabstractLow-latency data service is an increasingly critical challenge for data center applications. In modern distributed storage systems, proper data placement helps reduce the data movement delay, which can contribute to the service latency reduction tremendously. Existing data placement solutions have often assumed the prior distribution of data requests or discovered it via trace analysis. However, data placement is a difficult online decision-making problem faced with dynamic network conditions and time-varying user request patterns. The conventional static model-based solutions are less effective to handle the dynamic system. With an overall consideration of data movement and analytical latency, we develop a reinforcement learning-based framework DataBot+, automatically learning the optimal placement policies. DataBot+ adopts neural networks, trained with a variant of$Q$-learning, whose input is the real-time data flow measurements and whose output is a value function estimating the near-future latency. For instantaneous decision making, DataBot+ is decoupled into two asynchronous production and training components, ensuring that the training delay will not introduce extra overheads to handle the data flows. Evaluation results driven by real-world traces demonstrate the effectiveness of our design. Kaiyang Liu, Jun Peng 0001, Jingrong Wang, Boyang Yu 0001, Zhuofan Liao, Zhiwu Huang, Jianping Pan 0001 |
IEEE Trans. Cloud Comput. | 6 |
| 2022 | Distributed Group Coordination of Multiagent Systems in Cloud Computing Systems Using a Model-Free Adaptive Predictive Control StrategyabstractThis article studies the group coordinated control problem for distributed nonlinear multiagent systems (MASs) with unknown dynamics. Cloud computing systems are employed to divide agents into groups and establish networked distributed multigroup-agent systems (ND-MGASs). To achieve the coordination of all agents and actively compensate for communication network delays, a novel networked model-free adaptive predictive control (NMFAPC) strategy combining networked predictive control theory with model-free adaptive control method is proposed. In the NMFAPC strategy, each nonlinear agent is described as a time-varying data model, which only relies on the system measurement data for adaptive learning. To analyze the system performance, a simultaneous analysis method for stability and consensus of ND-MGASs is presented. Finally, the effectiveness and practicability of the proposed NMFAPC strategy are verified by numerical simulations and experimental examples. The achievement also provides a solution for the coordination of large-scale nonlinear MASs. Haoran Tan, Yaonan Wang 0001, Min Wu 0002, Zhiwu Huang, Zhiqiang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Neural Architecture Search as Sparse SupernetabstractThis paper aims at enlarging the problem of Neural Architecture Search (NAS) from Single-Path and Multi-Path Search to automated Mixed-Path Search. In particular, we model the NAS problem as a sparse supernet using a new continuous architecture representation with a mixture of sparsity constraints. The sparse supernet enables us to automatically achieve sparsely-mixed paths upon a compact set of nodes. To optimize the proposed sparse supernet, we exploit a hierarchical accelerated proximal gradient algorithm within a bi-level optimization framework. Extensive experiments on Convolutional Neural Network and Recurrent Neural Network search demonstrate that the proposed method is capable of searching for compact, general and powerful neural architectures. Yan Wu 0019, Aoming Liu, Zhiwu Huang, Luc Van Gool |
AAAI | 3 |
| 2021 | Spectral Tensor Train Parameterization of Deep Learning LayersabstractWe study low-rank parameterizations of weight matrices with embedded spectral properties in the Deep Learning context. The low-rank property leads to parameter efficiency and permits taking computational shortcuts when computing mappings. Spectral properties are often subject to constraints in optimization problems, leading to better models and stability of optimization. We start by looking at the compact SVD parameterization of weight matrices and identifying redundancy sources in the parameterization. We further apply the Tensor Train (TT) decomposition to the compact SVD components, and propose a non-redundant differentiable parameterization of fixed TT-rank tensor manifolds, termed the Spectral Tensor Train Parameterization (STTP). We demonstrate the effects of neural network compression in the image classification setting, and both compression and improved training stability in the generative adversarial training setting. Project website: www.obukhov.ai/sttp Anton Obukhov, Maksim Rakhuba, Alexander Liniger, Zhiwu Huang, Stamatios Georgoulis, Dengxin Dai, Luc Van Gool |
AISTATS | 4 |
| 2021 | Efficient Conditional GAN Transfer With Knowledge Propagation Across ClassesabstractGenerative adversarial networks (GANs) have shown impressive results in both unconditional and conditional image generation. In recent literature, it is shown that pre-trained GANs, on a different dataset, can be transferred to improve the image generation from a small target data. The same, however, has not been well-studied in the case of conditional GANs (cGANs), which provides new opportunities for knowledge transfer compared to unconditional setup. In particular, the new classes may borrow knowledge from the related old classes, or share knowledge among themselves to improve the training. This motivates us to study the problem of efficient conditional GAN transfer with knowledge propagation across classes. To address this problem, we introduce a new GAN transfer method to explicitly propagate the knowledge from the old classes to the new classes. The key idea is to enforce the popularly used conditional batch normalization (BN) to learn the class-specific information of the new classes from that of the old classes, with implicit knowledge sharing among the new ones. This allows for an efficient knowledge propagation from the old classes to the new ones, with the BN parameters increasing linearly with the number of new classes. The extensive evaluation demonstrates the clear superiority of the proposed method over state-of-the-art competitors for efficient conditional GAN transfer tasks. The code is available at: https://github.com/mshahbazi72/cGANTransfer Mohamad Shahbazi, Zhiwu Huang, Danda Pani Paudel, Ajad Chhatkuli, Luc Van Gool |
CVPR | 2 |
| 2021 | GANmut: Learning Interpretable Conditional Space for Gamut of EmotionsabstractHumans can communicate emotions through a plethora of facial expressions, each with its own intensity, nuances and ambiguities. The generation of such variety by means of conditional GANs is limited to the expressions encoded in the used label system. These limitations are caused either due to burdensome labelling demand or the confounded label space. On the other hand, learning from inexpensive and intuitive basic categorical emotion labels leads to limited emotion variability. In this paper, we propose a novel GAN-based framework that learns an expressive and interpretable conditional space (usable as a label space) of emotions, instead of conditioning on handcrafted labels. Our framework only uses the categorical labels of basic emotions to learn jointly the conditional space as well as emotion manipulation. Such learning can benefit from the image variability within discrete labels, especially when the intrinsic labels reside beyond the discrete space of the defined. Our experiments demonstrate the effectiveness of the proposed framework, by allowing us to control and generate a gamut of complex and compound emotions while using only the basic categorical emotion labels during training. Our source code is available at https://github.com/stefanodapolito/GANmut. Stefano d'Apolito, Danda Pani Paudel, Zhiwu Huang, Andrés Romero, Luc Van Gool |
CVPR | 3 |
| 2021 | Direct Differentiable Augmentation SearchabstractData augmentation has been an indispensable tool to improve the performance of deep neural networks, however the augmentation can hardly transfer among different tasks and datasets. Consequently, a recent trend is to adopt AutoML technique to learn proper augmentation policy without extensive hand-crafted tuning. In this paper, we propose an efficient differentiable search algorithm called Direct Differentiable Augmentation Search (DDAS). It exploits meta-learning with one-step gradient update and continuous relaxation to the expected training loss for efficient search. Our DDAS can achieve efficient augmentation search without relying on approximations such as Gumbel-Softmax or second order gradient approximation. To further reduce the adverse effect of improper augmentations, we organize the search space into a two level hierarchy, in which we first decide whether to apply augmentation, and then determine the specific augmentation policy. On standard image classification benchmarks, our DDAS achieves state-of-the-art performance and efficiency tradeoff while reducing the search cost dramatically, e.g. 0.15 GPU hours for CIFAR-10. In addition, we also use DDAS to search augmentation for object detection task and achieve comparable performance with AutoAugment [8], while being 1000× faster. Code will be released in https://github.com/zxcvfd13502/DDAS_code Aoming Liu, Zehao Huang, Zhiwu Huang, Naiyan Wang |
ICCV | 3 |
| 2021 | Neural Architecture Search of SPD Manifold NetworksabstractIn this paper, we propose a new neural architecture search (NAS) problem of Symmetric Positive Definite (SPD) manifold networks, aiming to automate the design of SPD neural architectures. To address this problem, we first introduce a geometrically rich and diverse SPD neural architecture search space for an efficient SPD cell design. Further, we model our new NAS problem with a one-shot training process of a single supernet. Based on the supernet modeling, we exploit a differentiable NAS algorithm on our relaxed continuous search space for SPD neural architecture search. Statistical evaluation of our method on drone, action, and emotion recognition tasks mostly provides better results than the state-of-the-art SPD networks and traditional NAS algorithms. Empirical results show that our algorithm excels in discovering better performing SPD network design and provides models that are more than three times lighter than searched by the state-of-the-art NAS algorithms. Rhea Sanjay Sukthanker, Zhiwu Huang, Suryansh Kumar 0001, Erik Goron Endsjo, Yan Wu 0019, Luc Van Gool |
IJCAI | 2 |
| 2021 | An Optimal Pulse Heating Strategy for Lithium-ion Battery Considering both Capacity Fade and Heating TimeabstractThe driving performance of electric vehicles seriously degrades due to the deterioration of lithium-ion batteries at low temperatures. Preheating lithium-ion batteries can effectively improve the driving range of electric vehicles at subzero temperatures. In this paper, an optimal pulse heating strategy is proposed for low-temperature heating of lithiumion battery. Firstly, this paper establishes a coupling model to describe the electro-thermal-aging behavior of battery. Secondly, the heating time and capacity loss jointly form a multi-objective optimization problem with the current constraint. The optimization problem is solved by using the particle swarm optimization(PSO) algorithm and the effect of weighting coefficient on heating performance is discussed to obtain the optimal pulse current. The results show that the proposed strategy can effectively reduce heating time without causing serious capacity reduction. Honglang Jiang, Zhiwu Huang, Yongjie Liu, Dianzhu Gao, Heng Li 0005, Weirong Liu 0001, Jun Peng 0001 |
SMC | 2 |
| 2021 | Optimal Charging of Supercapacitors with Limited Charging TimeabstractSupercapacitors have recieved increasing attentions in emerging portable power applications. The charging process of supercapacitors significantly affects the performance of both supercapacitors and chargers. Considering the charging time of supercapacitors is typically limited in practical applications, in this paper, we propose an optimal charging method for supercapacits with the limited charging time. Firstly, we analyze existing cell balancing and charging circuits, and adopt the switched resistor circuit. Then, we design a user-interactive optimal charging method for supercapacitors where the charging time can be specified by the users. The energy efficiency maximization of the proposed charging method is proved rigorously. A simulation charging platform has been established to verify the effectiveness of the proposed charging method. The simulation results show that the proposed charging method can effectively improve the energy efficiency under charging time constraints when compared with existing methods. Heng Li 0005, Dianzhu Gao, Jun Peng 0001, Zhiwu Huang |
SMC | 5 |
| 2021 | Performance Analysis on Access Collision in Semi-Persistent Scheduling of C-V2X Mode 4abstractFor autonomous vehicles and smart transportation services, information exchange and fusion with low latency and high reliability is critical. The 3rd Generation Partnership Project has released the cellular vehicle-to-everything (C-V2X) Mode 4 to enable direct vehicle-to-vehicle communications regardless of the cellular coverage. Mode 4 uses the sensing-based semi-persistent resource scheduling (SPS) to support autonomous resource selection by vehicles. However, channel access collisions lead to packet losses, especially in crowded scenarios. Thus, an accurate analytical model is essential to quantify the system performance, reveal how to mitigate collision and ensure system reliability and scalability. This paper focuses on the analytical modeling of the SPS and derives the access collision ratio considering both the sensed and hidden terminals in V2X. Extended simulations are conducted to verify the correctness of the analytical framework. In addition, we investigate the impact of system parameters on performance, which provides important guidelines for improving the system configuration. Xin Gu 0002, Jun Peng 0001, Yijun Cheng, Xiaoyong Zhang 0001, Weirong Liu 0001, Zhiwu Huang, Lin Cai 0001 |
VTC Fall | 6 |
| 2021 | Facial Emotion Recognition with Noisy Multi-task AnnotationsabstractHuman emotions can be inferred from facial expressions. However, the annotations of facial expressions are often highly noisy in common emotion coding models, including categorical and dimensional ones. To reduce human labelling effort on multi-task labels, we introduce a new problem of facial emotion recognition with noisy multi-task annotations. For this new problem, we suggest a formulation from the point of joint distribution match view, which aims at learning more reliable correlations among raw facial images and multi-task labels, resulting in the reduction of noise influence. In our formulation, we exploit a new method to enable the emotion prediction and the joint distribution learning in a unified adversarial learning game. Evaluation throughout extensive experiments studies the real setups of the suggested new problem, as well as the clear superiority of the proposed method over the state-of-the-art competing methods on either the synthetic noisy labeled CIFAR-10 or practical noisy multi-task labeled RAF and AffectNet. The code is available at https://github.com/sanweiliti/noisyFER. Zhiwu Huang, Danda Pani Paudel, Luc Van Gool |
WACV | 2 |
| 2021 | An Instance Reservation Framework for Cost Effective Services in Geo-Distributed Data CentersabstractInfrastructure-as-a-Service clouds in geo-distributed data centers offer various pricing options, including on-demand and reserved instances, which provide an elastic and cost-effective infrastructure to support High Performance Computing (HPC) applications. In this paper, we propose an instance reservation based cloud service framework, modeling the cost-minimizing reservation decision issue as an NP-hard integer programming problem for distributed data centers. To ease its computation complexity, two algorithms are proposed to minimize the HPC service cost with the worst-case performance guarantees: an offline heuristic-greedy algorithm, and a rolling-horizon based online algorithm when only short-term demand prediction is available. Facing fluctuating demands, instance reservation in a single data center may incur the highly underutilized capacity. To address this issue for further cost reduction, we extend the scheme with a novel cloud broker federation based resource sharing mechanism, reallocating already reserved but unused instances to computation-intensive and short-lived tasks for continuous execution without interruption. Extensive evaluations driven by large-scale trace-based datasets demonstrate that the proposed mechanism can effectively handle large volumes of service requests, saving considerable service costs with higher reservation resource utilization. Kaiyang Liu, Jun Peng 0001, Boyang Yu 0001, Weirong Liu 0001, Zhiwu Huang, Jianping Pan 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2020 | Weakly Paired Multi-Domain Image Translation
Marc Yanlong Zhang, Zhiwu Huang, Danda Pani Paudel, Janine Thoma, Luc Van Gool |
BMVC | 2 |
| 2020 | Off-Policy Reinforcement Learning for Efficient and Effective GAN Architecture Search
Yuan Tian 0014, Qin Wang 0013, Zhiwu Huang, Wen Li 0001, Dengxin Dai, Jun Wang 0012, Olga Fink |
ECCV (7) | 3 |
| 2020 | Foreign Objects Intrusion Detection Using Millimeter Wave Radar on Railway CrossingsabstractThe safety of railway crossings are of great important for rail and road transportation, because serious accidents occur in this area. Therefore, it is necessary to carry out foreign objects detection on railway crossings in order to improve the safety. Traditionally, video surveillance is one such solution, but it suffer from weather and illumination conditions. Under the hard environment conditions, the image of railway crossings is failed to capture by the camera. We propose a foreign objects detection system based on millimeter wave radar which has a higher detection accuracy, without the limitation of weather and light. Unlike vision-based approach, it can operate in darkness, high or low light intensity environment. With a millimeter wave radar, we first obtain the reflected signal from objects or ground and perform signal processing algorithm to extract the targets and suppress the clutter from received signal. We evaluate the detection capabilities of the millimeter wave radar in level crossings of railway. Huiling Cai, Dianzhu Gao, Yingze Yang, Shuo Li 0006, Kai Gao 0010, Aina Qin, Zhiwu Huang |
SMC | 9 |
| 2020 | Logistics Distribution Path Planning Based on Fireworks Differential AlgorithmabstractLogistics distribution is an important link in logistics. Whether the logistics distribution path can be effectively optimized will directly affect the efficiency of the logistics distribution system. To plan the logistics distribution path reasonably, to reduce the cost of logistics management, for the multi-object path planning problem in logistics distribution, the fireworks differential evolution algorithm is used to design an optimization scheme. To achieve the overall goal of saving logistics and distribution costs, real number coding is used for each distribution point, and actual road information is obtained through the Gaode API. Aiming at the defects of the standard fireworks algorithm, the differential evolution algorithm is introduced based on the fireworks algorithm to plan the distribution route. The simulation results show that the firework differential evolution algorithm can effectively plan the optimal distribution path, and compared with the original firework algorithm, the ant colony algorithm and particle swarm optimization algorithm have a better improvement in the optimization accuracy. Xiaoyong Zhang 0001, Dianzhu Gao, Kai Gao 0010, Mengfei Wen, Zhiwu Huang |
SMC | 6 |
| 2020 | A Traffic Flow Adaptive Energy Saving Scheme for Smart Lighting SystemsabstractTraditional lighting systems suffer from the problem of low energy efficiency and low illumination quality due to its disappointing management. To address this issue, in this paper, a novel traffic-flow adaptive scheme of smart lighting systems is proposed on the basis of the cyber-physical cloud system. The cyber-physical cloud system consists of the digital twin and cyber-physical system. The operation of the lighting system is simulated in the counterpart twin system with the digital twin technology. The cyber-physical system realizes data collection, information interaction, analysis, and processing, as well as complex computation and remote control. The traffic adaptive scheme works according to the brightness sequence to improves the energy efficiency of the lighting system and provide higher illumination quality for drivers. Extensive simulation results verify the proposed control scheme could improve the energy efficiency of lighting systems. Yunsheng Fan, Zhiwu Huang, Yue Wu 0024, Yongjie Liu, Yingze Yang, Weirong Liu 0001, Jun Peng 0001 |
SMC | 2 |
| 2020 | Optimal Filter-Based Energy Management for Hybrid Energy Storage Systems with Energy Consumption MinimizationabstractThe filter-based real-time energy management method has been proved practical and widely utilized in hybrid energy storage systems. However, the determination for the cutoff frequency of the energy-split filter is challenging. In this paper, an optimal filter-based energy management strategy is proposed for a battery/ultracapacitor electric vehicle to minimize the total energy consumption. A cost function of energy consumption for the cutoff frequency is established first. Considering the working condition of ultracapacitors, dynamic programming is adopted to obtain the optimal cutoff frequency series, i.e., the optimal energy distribution between batteries and ultracapacitors. Such an off-line optimization process is carried out under different driving cycles, e.g., urban and highway road conditions. Optimization results are used to determine the optimal cutoff frequency of a real-time filter-based energy management strategy. Simulation results indicate that the proposed strategy can minimize the total energy consumption of the hybrid energy storage system with ultracapacitors state of charge limitations being guaranteed. Compared with the existing real-time energy management strategies, the energy consumption is reduced 23.85% under aggressive acceleration conditions and 7.08% under urban conditions by the proposed strategy. Zhiwu Huang, Yue Wu 0024, Hongtao Liao, Yongjie Liu, Heng Li 0005, Mengfei Wen, Jun Peng 0001 |
SMC | 2 |
| 2020 | Car-Following Safe Headway Strategy with Battery-Health Conscious: A Reinforcement Learning ApproachabstractThis paper proposes an optimal car-following strategy for pure electric vehicles (EVs) with the aim of keeping an expected headway of the leader and reducing vehicle battery loss. In particular, a car-following system model is established. The primary task of the automatic vehicle is to follow the trajectory of the preceding car and maintain an expected headway. Then, the paper analyzes the powertrain of the electric vehicle. The loss of battery life over a period of time is proportional to the acceleration, so it takes the battery life into consideration. The Q-learning algorithm is conducted for the optimal car-following strategy using system data instead of system dynamics information. It utilizes reward function and greedy strategy to select actions to train the following vehicle to achieve car-following safety. When there is no collision in these two cars, acceleration is considered into reward function to reduce battery loss. Finally, it is verified by simulation that the proposed car-following strategy can keep good tracking, maintain the expected headway from the preceding vehicle, and reduce battery loss. Xi Jia, Jun Peng 0001, Yongjie Liu, Mengfei Wen, Zhiwu Huang |
SMC | 8 |
| 2020 | A Hierarchical State of Charge Estimation Method for Lithium-ion Batteries via XGBoost and Kalman FilterabstractDifferent from previous data-driven methods for lithium-ion battery State-of-Charge (SoC) estimation, this paper aims to develop a hierarchical SoC estimation method to address the data dependency issue and measurement noise interferences. In the off-line training layer, aging-aware features are extracted to improve SoC estimation accuracy throughout the entire battery life cycle. Extreme gradient boosting (XGBoost) is introduced to map the relationship between the extracted features and SoC for its strong nonlinear fitting ability. In the on-line estimation layer, Ampere-hour integral method is utilized to provide SoC reference to guarantee the stability of the proposed method. Meanwhile, to suppress the measurement noise, we adopt Kalman filter to correct the SoC value estimated by XGBoost. The superiority of the proposed method is proved under the random walk discharging experiment by comparing with the results of XGBoost, i.e., without Kalman filter. The proposed method improved the accuracy of lithium-ion battery SoC by 4% to 10%. Shiyu Song, Xiaoyong Zhang 0001, Dianzhu Gao, Yue Wu 0024, Yadong Gong, Zhiwu Huang |
SMC | 9 |
| 2020 | An Adaptive Deep Q-learning Service Migration Decision Framework for Connected VehiclesabstractThe vehicular service support with adaptability, real-time, and low delay is crucial for connected vehicles. However, due to limited coverage of mobile edge computing servers and data processing capability of connected vehicles, vehicular services need to be offloaded to the edge server and adaptively migrate as the connected vehicle moves. Aiming at the adaptive migration service, a deep Q-learning service migration decision algorithm is proposed in this paper. The proposed algorithm can dynamically adjust the vehicular service migration decision according to traffic information. Furthermore, a service migration framework consisting of neural networks is proposed in this paper to improve the adaptability and real-time performance of the algorithm. By using this framework, training and decision-making can be carried out simultaneously in different places. Finally, compared with the two existing algorithms, extensive simulations are conducted to verify the effectiveness of the proposed algorithm. Jun Peng 0001, Xiaoyong Zhang 0001, Weirong Liu 0001, Xin Gu 0002, Zhiwu Huang |
SMC | 7 |
| 2020 | Observer-Driven Charging of SupercapacitorsabstractCell balancing is crucial for charging supercapacitor cells to prevent cells from over-charging. Most existing cell-balancing charging methods typically adopt an output feedback control, i.e., the terminal voltages of cells are directly utilized in the controller design. One limitation of these methods is the voltage drop effect when the charging is terminated, which degrades the system capacity and results in cell imbalance. To address this challenge, in this article, we propose an observer-driven charging method for supercapacitors. The switched resistor circuit is applied and is further modeled using the switched systems theory, where the RC model of cells is considered. The communication interactions among cells is modeled using the graph theory. A switching Luenberger observer is designed to estimate the voltage of the equivalent capacitor of each cell, and a consensus-based switching control law is designed to charge and balance supercapacitors. The closed-loop system model is derived using the block diagram. A laboratory testbed has been built to verify the effectiveness of the proposed charging method. Experimental results show that the proposed method can effectively alleviate the voltage drop effect when compared with existing charging methods. Heng Li 0005, Jun Peng 0001, Jianping He 0001, Zhiwu Huang, Jing Wang 0005 |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | Scalable and Adaptive Data Replica Placement for Geo-Distributed Cloud StoragesabstractIn geo-distributed cloud storage systems, data replication has been widely used to serve the ever more users around the world for high data reliability and availability. How to optimize the data replica placement has become one of the fundamental problems to reduce the inter-node traffic and the system overhead of accessing associated data items. In the big data era, traditional solutions may face the challenges of long running time and large overheads to handle the increasing scale of data items with time-varying user requests. Therefore, novel offline community discovery and online community adjustment schemes are proposed to solve the replica placement problem in a scalable and adaptive way. The offline scheme can find a replica placement solution based on the average read/write rates for a certain period of time. The scalability can be achieved as 1) the computation complexity is linear to the amount of data items and 2) the data-node communities can evolve in parallel for a distributed replica placement. Furthermore, the online scheme is adaptive to handle the bursty data requests, without the need to completely override the existing replica placement. Driven by real-world data traces, extensive performance evaluations demonstrate the effectiveness of our design to handle large-scale datasets. Kaiyang Liu, Jun Peng 0001, Jingrong Wang, Weirong Liu 0001, Zhiwu Huang, Jianping Pan 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | Manifold-Valued Image Generation with Wasserstein Generative Adversarial NetsabstractGenerative modeling over natural images is one of the most fundamental machine learning problems. However, few modern generative models, including Wasserstein Generative Adversarial Nets (WGANs), are studied on manifold-valued images that are frequently encountered in real-world applications. To fill the gap, this paper first formulates the problem of generating manifold-valued images and exploits three typical instances: hue-saturation-value (HSV) color image generation, chromaticity-brightness (CB) color image generation, and diffusion-tensor (DT) image generation. For the proposed generative modeling problem, we then introduce a theorem of optimal transport to derive a new Wasserstein distance of data distributions on complete manifolds, enabling us to achieve a tractable objective under the WGAN framework. In addition, we recommend three benchmark datasets that are CIFAR-10 HSV/CB color images, ImageNet HSV/CB color images, UCL DT image datasets. On the three datasets, we experimentally demonstrate the proposed manifold-aware WGAN model can generate more plausible manifold-valued images than its competitors. Zhiwu Huang, Jiqing Wu, Luc Van Gool |
AAAI | 1 |
| 2019 | Sliced Wasserstein Generative ModelsabstractIn generative modeling, the Wasserstein distance (WD) has emerged as a useful metric to measure the discrepancy between generated and real data distributions. Unfortunately, it is challenging to approximate the WD of high-dimensional distributions. In contrast, the sliced Wasserstein distance (SWD) factorizes high-dimensional distributions into their multiple one-dimensional marginal distributions and is thus easier to approximate. In this paper, we introduce novel approximations of the primal and dual SWD. Instead of using a large number of random projections, as it is done by conventional SWD approximation methods, we propose to approximate SWDs with a small number of parameterized orthogonal projections in an end-to-end deep learning fashion. As concrete applications of our SWD approximations, we design two types of differentiable SWD blocks to equip modern generative frameworks---Auto-Encoders (AE) and Generative Adversarial Networks (GAN). In the experiments, we not only show the superiority of the proposed generative models on standard image synthesis benchmarks, but also demonstrate the state-of-the-art performance on challenging high resolution image and video generation in an unsupervised manner. Jiqing Wu, Zhiwu Huang, Dinesh Acharya 0001, Wen Li 0001, Janine Thoma, Danda Pani Paudel, Luc Van Gool |
CVPR | 2 |
| 2019 | A Novel Adhesion Force Estimation for Railway Vehicles Using an Extended State ObserverabstractThe accurate estimation of adhesion force between wheels and rails is an important task as it helps to the wheel-slip prevention (WSP) system preventing the wheels from locking and reducing the stopping distance. Influenced by a changing external environment, the adhesion force estimation process is complex. Thus, an extended state observer (ESO) is proposed to accurately estimate the adhesion force of railway vehicles with the modeling deviation and measurement noise. With the estimated modeling error information, an auxiliary compensation part is designed to eliminate the steady-state estimation error causing by the modeling deviation. Further, a Fal function filter is added to the ESO to deal with the effect of measurement noise. The convergence of the proposed estimation method is analyzed theoretically. The effectiveness of the designed algorithm is corroborated by simulation comparisons to other standard approaches. Bin Chen 0017, Zhiwu Huang, Weirong Liu 0001, Rui Zhang 0041, Feng Zhou 0002, Jun Peng 0001 |
IECON | 2 |
| 2019 | Adaptive Precision Automatic Train Stop Control based on Pneumatic Brake SystemsabstractPrecision stopping of trains requires special attention to brake control because a pneumatic brake system of a train is highly nonlinear, hybrid and uncertain. Existing solutions to automatic train stop control ignore the pneumatic brake system or simply treat it as a system delay, which is quite far apart from the real train stopping dynamics. Moreover, the service life of pneumatic brake systems decreases fast due to the frequent changes in output of existing controllers. Thus, an adaptive nonlinear sliding mode control method is developed in this paper, which has strong applicability to nonlinear and hybrid system control synthesis due to its natural variable structure characteristic. A nonlinear integral sliding surface with adaptive updating parameters is proposed to improve the control precision and the robustness. Extensive simulations are performed to validate the effectiveness of the proposed method. The results show that the proposed algorithm outperforms a PID control algorithm in terms of stopping error and expected lifetime of pneumatic brake systems. Rui Zhang 0041, Jun Peng 0001, Feng Zhou 0002, Bin Chen 0017, Weirong Liu 0001, Zhiwu Huang |
IECON | 6 |
| 2018 | Building Deep Networks on Grassmann ManifoldsabstractLearning representations on Grassmann manifolds is popular in quite a few visual recognition tasks. In order to enable deep learning on Grassmann manifolds, this paper proposes a deep network architecture by generalizing the Euclidean network paradigm to Grassmann manifolds. In particular, we design full rank mapping layers to transform input Grassmannian data to more desirable ones, exploit re-orthonormalization layers to normalize the resulting matrices, study projection pooling layers to reduce the model complexity in the Grassmannian context, and devise projection mapping layers to respect Grassmannian geometry and meanwhile achieve Euclidean forms for regular output layers. To train the Grassmann networks, we exploit a stochastic gradient descent setting on manifolds of the connection weights, and study a matrix generalization of backpropagation to update the structured data. The evaluations on three visual recognition tasks show that our Grassmann networks have clear advantages over existing Grassmann learning methods, and achieve results comparable with state-of-the-art approaches. Zhiwu Huang, Jiqing Wu, Luc Van Gool |
AAAI | 1 |
| 2018 | Wasserstein Divergence for GANs
Jiqing Wu, Zhiwu Huang, Janine Thoma, Dinesh Acharya 0001, Luc Van Gool |
ECCV (5) | 2 |
| 2018 | Cross Euclidean-to-Riemannian Metric Learning with Application to Face Recognition from VideoabstractRiemannian manifolds have been widely employed for video representations in visual classification tasks including video-based face recognition. The success mainly derives from learning a discriminant Riemannian metric which encodes the non-linear geometry of the underlying Riemannian manifolds. In this paper, we propose a novel metric learning framework to learn a distance metric across a Euclidean space and a Riemannian manifold to fuse average appearance and pattern variation of faces within one video. The proposed metric learning framework can handle three typical tasks of video-based face recognition: Video-to-Still, Still-to-Video and Video-to-Video settings. To accomplish this new framework, by exploiting typical Riemannian geometries for kernel embedding, we map the source Euclidean space and Riemannian manifold into a common Euclidean subspace, each through a corresponding high-dimensional Reproducing Kernel Hilbert Space (RKHS). With this mapping, the problem of learning a cross-view metric between the two source heterogeneous spaces can be converted to learning a single-view Euclidean distance metric in the target common Euclidean space. By learning information on heterogeneous data with the shared label, the discriminant metric in the common space improves face recognition from videos. Extensive experiments on four challenging video face databases demonstrate that the proposed framework has a clear advantage over the state-of-the-art methods in the three classical video-based face recognition scenarios. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Luc Van Gool, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Geometry-Aware Similarity Learning on SPD Manifolds for Visual RecognitionabstractSymmetric positive definite (SPD) matrices have been employed for data representation in many visual recognition tasks. The success is mainly attributed to learning discriminative SPD matrices encoding the Riemannian geometry of the underlying SPD manifolds. In this paper, we propose a geometry-aware SPD similarity learning (SPDSL) framework to learn discriminative SPD features by directly pursuing a manifold-manifold transformation matrix of full column rank. Specifically, by exploiting the Riemannian geometry of the manifolds of fixed-rank positive semidefinite (PSD) matrices, we present a new solution to reduce optimization over the space of column full-rank transformation matrices to optimization on the PSD manifold, which has a well-established Riemannian structure. Under this solution, we exploit a new supervised SPDSL technique to learn the manifold-manifold transformation by regressing the similarities of selected SPD data pairs to their ground-truth similarities on the target SPD manifold. To optimize the proposed objective function, we further derive an optimization algorithm on the PSD manifold. Evaluations on three visual classification tasks show the advantages of the proposed approach over the existing SPD-based discriminant learning methods. Zhiwu Huang, Ruiping Wang 0001, Xianqiu Li, Wenxian Liu, Shiguang Shan, Luc Van Gool, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Discriminant Analysis on Riemannian Manifold of Gaussian Distributions for Face Recognition With Image SetsabstractTo address the problem of face recognition with image sets, we aim to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as the Gaussian mixture model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. Since in the light of information geometry, the Gaussians lie on a specific Riemannian manifold, this paper presents a method named discriminant analysis on Riemannian manifold of Gaussian distributions (DARG). We investigate several distance metrics between Gaussians and accordingly two discriminative learning frameworks are presented to meet the geometric and statistical characteristics of the specific manifold. The first framework derives a series of provably positive definite probabilistic kernels to embed the manifold to a high-dimensional Hilbert space, where conventional discriminant analysis methods developed in Euclidean space can be applied, and a weighted Kernel discriminant analysis is devised which learns discriminative representation of the Gaussian components in GMMs with their prior probabilities as sample weights. Alternatively, the other framework extends the classical graph embedding method to the manifold by utilizing the distance metrics between Gaussians to construct the adjacency graph, and hence the original manifold is embedded to a lower-dimensional and discriminative target manifold with the geometric structure preserved and the interclass separability maximized. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB, and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art.To address the problem of face recognition with image sets, we aim to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as the Gaussian mixture model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. Since in the light of information geometry, the Gaussians lie on a specific Riemannian manifold, this paper presents a method named discriminant analysis on Riemannian manifold of Gaussian distributions (DARG). We investigate several distance metrics between Gaussians and accordingly two discriminative learning frameworks are presented to meet the geometric and statistical characteristics of the specific manifold. The first framework derives a series of provably positive definite probabilistic kernels to embed the manifold to a high-dimensional Hilbert space, where conventional discriminant analysis methods developed in Euclidean space can be applied, and a weighted Kernel discriminant analysis is devised which learns discriminative representation of the Gaussian components in GMMs with their prior probabilities as sample weights. Alternatively, the other framework extends the classical graph embedding method to the manifold by utilizing the distance metrics between Gaussians to construct the adjacency graph, and hence the original manifold is embedded to a lower-dimensional and discriminative target manifold with the geometric structure preserved and the interclass separability maximized. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB, and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art. Wen Wang 0019, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Distributed Economic Dispatch in Microgrids Based on Cooperative Reinforcement LearningabstractMicrogrids incorporated with distributed generation (DG) units and energy storage (ES) devices are expected to play more and more important roles in the future power systems. Yet, achieving efficient distributed economic dispatch in microgrids is a challenging issue due to the randomness and nonlinear characteristics of DG units and loads. This paper proposes a cooperative reinforcement learning algorithm for distributed economic dispatch in microgrids. Utilizing the learning algorithm can avoid the difficulty of stochastic modeling and high computational complexity. In the cooperative reinforcement learning algorithm, the function approximation is leveraged to deal with the large and continuous state spaces. And a diffusion strategy is incorporated to coordinate the actions of DG units and ES devices. Based on the proposed algorithm, each node in microgrids only needs to communicate with its local neighbors, without relying on any centralized controllers. Algorithm convergence is analyzed, and simulations based on real-world meteorological and load data are conducted to validate the performance of the proposed algorithm. Weirong Liu 0001, Peng Zhuang, Hao Liang 0002, Jun Peng 0001, Zhiwu Huang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | A Riemannian Network for SPD Matrix LearningabstractSymmetric Positive Definite (SPD) matrix learning methods have become popular in many image and video processing tasks, thanks to their ability to learn appropriate statistical representations while respecting Riemannian geometry of underlying SPD manifolds. In this paper we build a Riemannian network architecture to open up a new direction of SPD matrix non-linear learning in a deep model. In particular, we devise bilinear mapping layers to transform input SPD matrices to more desirable SPD matrices, exploit eigenvalue rectification layers to apply a non-linear activation function to the new SPD matrices, and design an eigenvalue logarithm layer to perform Riemannian computing on the resulting SPD matrices for regular output layers. For training the proposed deep network, we exploit a new backpropagation with a variant of stochastic gradient descent on Stiefel manifolds to update the structured connection weights and the involved SPD matrix data. We show through experiments that the proposed SPD matrix network can be simply trained and outperform existing SPD matrix learning and state-of-the-art methods in three typical visual classification tasks. Zhiwu Huang, Luc Van Gool |
AAAI | 1 |
| 2017 | Deep Learning on Lie Groups for Skeleton-Based Action RecognitionabstractIn recent years, skeleton-based action recognition has become a popular 3D classification problem. State-of-the-art methods typically first represent each motion sequence as a high-dimensional trajectory on a Lie group with an additional dynamic time warping, and then shallowly learn favorable Lie group features. In this paper we incorporate the Lie group structure into a deep network architecture to learn more appropriate Lie group features for 3D action recognition. Within the network structure, we design rotation mapping layers to transform the input Lie group features into desirable ones, which are aligned better in the temporal domain. To reduce the high feature dimensionality, the architecture is equipped with rotation pooling layers for the elements on the Lie group. Furthermore, we propose a logarithm mapping layer to map the resulting manifold data into a tangent space that facilitates the application of regular output layers for the final classification. Evaluations of the proposed network for standard 3D human action recognition datasets clearly demonstrate its superiority over existing shallow Lie group feature learning methods as well as most conventional deep learning methods. Zhiwu Huang, Chengde Wan, Thomas Probst, Luc Van Gool |
CVPR | 1 |
| 2017 | An optimal task decision method for a warehouse robot with multiple tasks based on linear temporal logicabstractCurrently, the robot is playing an increasingly significant role in managing a warehouse. This paper proposes an optimal method to help a warehouse robot make task decisions, which aims at minimizing the whole cost of completing multiple tasks. Firstly, Abstract Transition System (ATS) is used to model the warehouse environment, and Linear Temporal Logic (LTL) formula is used to formulate the tasks of warehouse robot. Then based on the ATS and the Büchi automaton translated from the LTL formula, a Min-cost Task Decision Algorithm is proposed to obtain the task decision for the warehouse robot. The decision points out the optimal order and path for the robot to do its tasks. The effectiveness of the proposed method is validated through case studies with two kinds of tasks. Zhiwu Huang, Lulu Wang 0012, Rui Zhang 0041, Xiaoyong Zhang 0001, Jun Peng 0001 |
SMC | 2 |
| 2017 | Consensus control for state-of-energy balancing between the supercapacitor modules in cyber-physical energy systemabstractRecent advancement in the field of electrical technology and cyber-physical energy system (CPES) has brought the key towards challenging issues regarding transparency of information management and efficient allocation of energy. This paper is dedicated to a CPES that deals with an electric light rail that involves large number of distributed and globally interconnected supercapacitor energy storage modules, with the aim of efficient fusion of information, control protocol and energy. The autonomous modules estimate local state of energy and share the information via the communication network. A consensus control strategy is proposed to reach the balanced state of energy between these distributed supercapacitor modules with the advantages in terms of increasing efficiency and reducing time consumption of energy transfer. The proposed method enforces the CPES constraints specific to the particular supercapacitor modules in the electric light rails. Experimental results are provided to verify the effectiveness of the proposed state of energy balancing method. Chengzhang Lyu, Zhiwu Huang, Heng Li 0005, Jun Peng 0001, Yingze Yang |
SMC | 2 |
| 2017 | Robust and accurate state-of-charge estimation for lithium-ion batteries using generalized extended state observerabstractWith the wide application of Lithium-ion (Li-ion) batteries in electric vehicles and unmanned aerial vehicles (UAVs), it is becoming more important and urgent to estimate the battery state to extend the operation range of electric vehicles or UAVs. Existing state of charge (SOC) estimation methods are highly model-based, which are difficult to be implemented in different scenarios. In this paper, we propose a generalized extended state observer (GESO) based SOC estimation method, where the accurate model is unnecessary, which can be effectively tracked by the observer. Thus, a first order RC model is utilized in GESO to capture the characteristics of Li-ion batteries. By appropriately designing a disturbance compensation gain, the GESO is applied for the nonintegral-chain system that is subject to uncertainties and nonlinear parameters of Li-ion batteries. Experiment results show that the proposed method has a good performance and robustness on SOC estimation of the battery. The SOC estimation results are found to be consistent with the reference SOC with less chattering than sliding mode observer, where the error is within 2% under both the known and unknown initial SOC value cases. Weirong Liu 0001, Heng Li 0005, Zhiwu Huang |
SMC | 5 |
| 2017 | Cooperative Neural Fitted Learning for Distributed Energy Management in Microgrids via Wireless NetworksabstractWith the proliferation of renewable energy sources and the elevation of environmental concerns, it is expected that microgrids will become one of the major means for residential energy supply. However, the distributed nature of microgrid operation brings new technical challenges to energy management. Endowing wireless communication capability to the distributed generation (DG) units and energy storage (ES) devices in a microgrid is beneficial for their cooperation without a centralized controller. Yet, how to establish distributed energy management without \emph{a priori} statistical information for all the DG units and loads still requires extensive research. In this paper, a reinforcement learning algorithm with cooperative neural fitting iteration is proposed for distributed energy management in microgrids via wireless networks. The reinforcement learning algorithm leverages a distributed actor- critic structure to adopt the continuous states and action spaces of a microgrid. A diffusion strategy is incorporated in the reinforcement learning algorithm to coordinate the actions of DG units and ES devices by exchanging their evaluations and decisions via a wireless network. Simulation results based on realistic renewable power generation and load data are presented to evaluate the performance of the proposed algorithm. Weirong Liu 0001, Peng Zhuang, Yuan Liu 0006, Hao Liang 0002, Zhiwu Huang, Jun Peng 0001 |
VTC Fall | 5 |
| 2017 | Decentralized event-triggered cooperative control for multi-agent systems with uncertain dynamics using local estimators
Feng Zhou 0002, Zhiwu Huang, Yingze Yang, Jing Wang 0005, Liran Li, Jun Peng 0001 |
Neurocomputing | 2 |
| 2016 | Genetic Based Data Placement for Geo-Distributed Data-Intensive Applications in Cloud Computing
Weifeng Fan, Jun Peng 0001, Xiaoyong Zhang 0001, Zhiwu Huang |
APSCC | 4 |
| 2016 | A Combinatorial Optimization for Energy-Efficient Mobile Cloud Offloading over Cellular NetworksabstractRecently, mobile cloud offloading is a promising technique to deal with the increasingly complex applications on mobile devices, meeting the ever- increasing energy requirements. However, cloud offloading with multiple mobile devices may cause considerable mutual interference, which may result in intolerable time delay and more energy consumption. In this paper, a novel offloading decision method is investigated to minimize the total energy consumption of mobile devices over cellular networks. Generally, mobile devices can execute a sequence of tasks in parallel with different characteristics, i.e., communication- intensive and computation-intensive. And recent advances show that only computation-intensive tasks are applicable to be offloaded for energy saving. The offloading decision issue is formulated as a NP- hard combinatorial optimization problem with the time deadline and communication quality constraints. Combining the problem linearization method and decision variables mapping from integer to the real domain, a rapid and efficient iterative approximation method is proposed, helping the cloud controller to select the best tasks for offloading aiming at minimizing the total energy consumption. Numerical simulation demonstrates that considerable energy can be saved with the proposed task offloading method in mobile cloud scenarios. Kaiyang Liu, Jun Peng 0001, Xiaoyong Zhang 0001, Zhiwu Huang |
GLOBECOM | 4 |
| 2016 | A Hybrid Particle Swarm Ant Colony Based Resource Reservation for Geo-Distributed Cloud ServiceabstractIn cloud market, cloud providers offer diverse service options, including on-demand instances and reserved instances. Generally reserved instance price is cheaper than on-demand instance, but excessive reservation may result in high capacity underutilization. To improve resource utilization rate and reduce providers' service cost, a steady broker federation is necessary to propose. Broker federation coordinate cloud providers service demands in geo-distributed data centers through deciding when and how many instances to reserve. Firstly, the service cost optimization problem is formulated as a nonlinear integer programming model. Then a hybrid algorithm combining ant colony with particle swarm optimization is presented to reduce computational complexity and providers service cost. Extensive simulations driven by large-scale Parallel Workloads Archive demonstrate the effectiveness and efficiency of the hybrid algorithm. Yazhen Song, Jun Peng 0001, Kaiyang Liu, Weirong Liu 0001, Zhiwu Huang |
GLOBECOM | 6 |
| 2016 | A combinatorial double auction based resource allocation mechanism with multiple rounds for geo-distributed data centersabstractWith the explosion of application of big data, it becomes inefficient and infeasible to process big data stream by using conventional data service infrastructure and management system. Cloud computing platform having multiple geo-distributed data centers is expected to be the most efficient platform to process the big data stream. In this paper, a multi-round combinational double auction based mechanism is proposed to allocate the resources of geo-distributed data centers to multiple users with large data stream processing tasks. This mechanism combines the advantages of combinatorial auction and double auction. In the proposed mechanism, the different types of VMs can be integrated into a bundle to be bid. The auction is double and conducted from both users and data centers. Different from existed double auctions, the QoS level is taken into consideration. In addition, the multiple rounds mode is adopted, so the failed users and data centers have the chance to adjust bids and asks to participate next auction round, increasing the ratio of successful transactions. Simulation results validate the effectiveness of the proposed mechanism. Yeru Zhao, Zhiwu Huang, Weirong Liu 0001, Jun Peng 0001 |
ICC | 2 |
| 2015 | Projection Metric Learning on Grassmann Manifold with Application to Video based Face RecognitionabstractIn video based face recognition, great success has been made by representing videos as linear subspaces, which typically lie in a special type of non-Euclidean space known as Grassmann manifold. To leverage the kernel-based methods developed for Euclidean space, several recent methods have been proposed to embed the Grassmann manifold into a high dimensional Hilbert space by exploiting the well established Project Metric, which can approximate the Riemannian geometry of Grassmann manifold. Nevertheless, they inevitably introduce the drawbacks from traditional kernel-based methods such as implicit map and high computational cost to the Grassmann manifold. To overcome such limitations, we propose a novel method to learn the Projection Metric directly on Grassmann manifold rather than in Hilbert space. From the perspective of manifold learning, our method can be regarded as performing a geometry-aware dimensionality reduction from the original Grassmann manifold to a lower-dimensional, more discriminative Grassmann manifold where more favorable classification can be achieved. Experiments on several real-world video face datasets demonstrate that the proposed method yields competitive performance compared with the state-of-the-art algorithms. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2015 | Face video retrieval with image query via hashing across Euclidean space and Riemannian manifoldabstractRetrieving videos of a specific person given his/her face image as query becomes more and more appealing for applications like smart movie fast-forwards and suspect searching. It also forms an interesting but challenging computer vision task, as the visual data to match, i.e., still image and video clip are usually represented quite differently. Typically, face image is represented as point (i.e., vector) in Euclidean space, while video clip is seemingly modeled as a point (e.g., covariance matrix) on some particular Riemannian manifold in the light of its recent promising success. It thus incurs a new hashing-based retrieval problem of matching two heterogeneous representations, respectively in Euclidean space and Riemannian manifold. This work makes the first attempt to embed the two heterogeneous spaces into a common discriminant Hamming space. Specifically, we propose Hashing across Euclidean space and Riemannian manifold (HER) by deriving a unified framework to firstly embed the two spaces into corresponding reproducing kernel Hilbert spaces, and then iteratively optimize the intra- and inter-space Hamming distances in a max-margin framework to learn the hash functions for the two spaces. Extensive experiments demonstrate the impressive superiority of our method over the state-of-the-art competitive hash learning methods. Yan Li 0014, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
CVPR | 3 |
| 2015 | Discriminant analysis on Riemannian manifold of Gaussian distributions for face recognition with image setsabstractThis paper presents a method named Discriminant Analysis on Riemannian manifold of Gaussian distributions (DARG) to solve the problem of face recognition with image sets. Our goal is to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as Gaussian Mixture Model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. In the light of information geometry, the Gaussians lie on a specific Riemannian manifold. To encode such Riemannian geometry properly, we investigate several distances between Gaussians and further derive a series of provably positive definite probabilistic kernels. Through these kernels, a weighted Kernel Discriminant Analysis is finally devised which treats the Gaussians in GMMs as samples and their prior probabilities as sample weights. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art. Wen Wang 0019, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
CVPR | 3 |
| 2015 | Log-Euclidean Metric Learning on Symmetric Positive Definite Manifold with Application to Image Set ClassificationabstractThe manifold of Symmetric Positive Definite (SPD) matrices has been successfully used for data representation in image set classification. By endowing the SPD manifold with Log-Euclidean Metric, existing methods typically work on vector-forms of SPD matrix logarithms. This however not only inevitably distorts the geometrical structure of the space of SPD matrix logarithms but also brings low efficiency especially when the dimensionality of SPD matrix is high. To overcome this limitation, we propose a novel metric learning approach to work directly on logarithms of SPD matrices. Specifically, our method aims to learn a tangent map that can directly transform the matrix logarithms from the original tangent space to a new tangent space of more discriminability. Under the tangent map framework, the novel metric learning can then be formulated as an optimization problem of seeking a Mahalanobis-like matrix, which can take the advantage of traditional metric learning techniques. Extensive evaluations on several image set classification tasks demonstrate the effectiveness of our proposed metric learning method. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xianqiu Li, Xilin Chen 0001 |
ICML | 1 |
| 2015 | Face recognition on large-scale video in the wild with hybrid Euclidean-and-Riemannian metric learning
Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 1 |
| 2015 | A Benchmark and Comparative Study of Video-Based Face Recognition on COX Face DatabaseabstractFace recognition with still face images has been widely studied, while the research on video-based face recognition is inadequate relatively, especially in terms of benchmark datasets and comparisons. Real-world video-based face recognition applications require techniques for three distinct scenarios: 1) Videoto-Still (V2S); 2) Still-to-Video (S2V); and 3) Video-to-Video (V2V), respectively, taking video or still image as query or target. To the best of our knowledge, few datasets and evaluation protocols have benchmarked for all the three scenarios. In order to facilitate the study of this specific topic, this paper contributes a benchmarking and comparative study based on a newly collected still/video face database, named COX(1) Face DB. Specifically, we make three contributions. First, we collect and release a largescale still/video face database to simulate video surveillance with three different video-based face recognition scenarios (i.e., V2S, S2V, and V2V). Second, for benchmarking the three scenarios designed on our database, we review and experimentally compare a number of existing set-based methods. Third, we further propose a novel Point-to-Set Correlation Learning (PSCL) method, and experimentally show that it can be used as a promising baseline method for V2S/S2V face recognition on COX Face DB. Extensive experimental results clearly demonstrate that video-based face recognition needs more efforts, and our COX Face DB is a good benchmark database for evaluation. Zhiwu Huang, Shiguang Shan, Ruiping Wang 0001, Haihong Zhang, Shihong Lao, Alifu Kuerban, Xilin Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Hybrid Euclidean-and-Riemannian Metric Learning for Image Set Classification
Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 1 |
| 2014 | Learning Euclidean-to-Riemannian Metric for Point-to-Set ClassificationabstractIn this paper, we focus on the problem of point-to-set classification, where single points are matched against sets of correlated points. Since the points commonly lie in Euclidean space while the sets are typically modeled as elements on Riemannian manifold, they can be treated as Euclidean points and Riemannian points respectively. To learn a metric between the heterogeneous points, we propose a novel Euclidean-to-Riemannian metric learning framework. Specifically, by exploiting typical Riemannian metrics, the Riemannian manifold is first embedded into a high dimensional Hilbert space to reduce the gaps between the heterogeneous spaces and meanwhile respect the Riemannian geometry of the manifold. The final distance metric is then learned by pursuing multiple transformations from the Hilbert space and the original Euclidean space (or its corresponding Hilbert space) to a common Euclidean subspace, where classical Euclidean distances of transformed heterogeneous points can be measured. Extensive experiments clearly demonstrate the superiority of our proposed approach over the state-of-the-art methods. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2014 | Dynamic resource reservation via broker federation in cloud service: A fine-grained heuristic-based approachabstractIn cloud computing, Infrastructure-as-a-Service (IaaS) cloud providers can offer two types of purchasing plans for cloud users, including on-demand plan and reservation plan. Generally reservation price is cheaper than on-demand price, while reservation plan may cause highly underutilized capacity problem. How to joint optimize the service cost and the resource utilization for clouds is a critical issue. To address this issue, a novel steady broker federation is developed to coordinate service demands in this paper. And the optimal reservation problem can be formulated as a nonlinear integer programming model. Then a fine-grained heuristic algorithm is proposed to reduce its computational complexity and obtain quasi-optimal solutions. Numerical simulations driven by large-scale Parallel Workloads Archive demonstrate that the proposed approach can save considerable costs for cloud users and improves the resource utilization for IaaS cloud providers. Kaiyang Liu, Jun Peng 0001, Weirong Liu 0001, Pingping Yao, Zhiwu Huang |
GLOBECOM | 5 |
| 2014 | Combining Multiple Kernel Methods on Riemannian Manifold for Emotion Recognition in the WildabstractIn this paper, we present the method for our submission to the Emotion Recognition in the Wild Challenge (EmotiW 2014). The challenge is to automatically classify the emotions acted by human subjects in video clips under real-world environment. In our method, each video clip can be represented by three types of image set models (i.e. linear subspace, covariance matrix, and Gaussian distribution) respectively, which can all be viewed as points residing on some Riemannian manifolds. Then different Riemannian kernels are employed on these set models correspondingly for similarity/distance measurement. For classification, three types of classifiers, i.e. kernel SVM, logistic regression, and partial least squares, are investigated for comparisons. Finally, an optimal fusion of classifiers learned from different kernels and different modalities (video and audio) is conducted at the decision level for further boosting the performance. We perform an extensive evaluation on the challenge data (including validation set and blind test set), and evaluate the effects of different strategies in our pipeline. The final recognition accuracy achieved 50.4% on test set, with a significant gain of 16.7% above the challenge baseline 33.7%. Ruiping Wang 0001, Shaoxin Li 0001, Shiguang Shan, Zhiwu Huang, Xilin Chen 0001 |
ICMI | 5 |
| 2014 | A high efficient and reliable DC-DC converter for Electronically Controlled Pneumatic brake system applicationsabstractThe Electronically Controlled Pneumatic (ECP) brake system is being applied to the heavy-haul trains for its control accracy and real time performance. A high power supply with high efficiency and reliability is essential to the safe operation of the ECP brake system. In this paper, a push-pull forward converter with a secondary full bridge rectifier is proposed for the ECP power supply. The voltage stress of the rectifier diodes can be reduced by a modified nodissipative snubber. Furthermore, hybrid control methods and criteria are utilized to improve the converter performance, including forced flux balancing circuit design, power supply output impedance and noise requirements. Finally, A 2.5kW prototype of the converter with efficiency up to 93% verifies the proposed circuit and theoretical analysis. Zhiwu Huang, Xiaohui Qu, Weirong Liu 0001, Kai Gao 0010 |
IECON | 1 |
| 2013 | Coupling Alignments with Recognition for Still-to-Video Face RecognitionabstractThe Still-to-Video (S2V) face recognition systems typically need to match faces in low-quality videos captured under unconstrained conditions against high quality still face images, which is very challenging because of noise, image blur, low face resolutions, varying head pose, complex lighting, and alignment difficulty. To address the problem, one solution is to select the frames of `best quality' from videos (hereinafter called quality alignment in this paper). Meanwhile, the faces in the selected frames should also be geometrically aligned to the still faces offline well-aligned in the gallery. In this paper, we discover that the interactions among the three tasks-quality alignment, geometric alignment and face recognition-can benefit from each other, thus should be performed jointly. With this in mind, we propose a Coupling Alignments with Recognition (CAR) method to tightly couple these tasks via low-rank regularized sparse representation in a unified framework. Our method makes the three tasks promote mutually by a joint optimization in an Augmented Lagrange Multiplier routine. Extensive experiments on two challenging S2V datasets demonstrate that our method outperforms the state-of-the-art methods impressively. Zhiwu Huang, Shiguang Shan, Ruiping Wang 0001, Xilin Chen 0001 |
ICCV | 1 |
| 2013 | Partial least squares regression on grassmannian manifold for emotion recognitionabstractIn this paper, we propose a method for video-based human emotion recognition. For each video clip, all frames are represented as an image set, which can be modeled as a linear subspace to be embedded in Grassmannian manifold. After feature extraction, Class-specific One-to-Rest Partial Least Squares (PLS) is learned on video and audio data respectively to distinguish each class from the other confusing ones. Finally, an optimal fusion of classifiers learned from both modalities (video and audio) is conducted at decision level. Our method is evaluated on the Emotion Recognition In The Wild Challenge (EmotiW 2013). The experimental results on both validation set and blind test set are presented for comparison. The final accuracy achieved on test set outperforms the baseline by 26%. Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
ICMI | 3 |
| 2012 | Cross-view Graph Embedding
Zhiwu Huang, Shiguang Shan, Haihong Zhang, Shihong Lao, Xilin Chen 0001 |
ACCV (2) | 1 |
| 2012 | Benchmarking Still-to-Video Face Recognition via Partial and Local Linear Discriminant Analysis on COX-S2V Dataset
Zhiwu Huang, Shiguang Shan, Haihong Zhang, Shihong Lao, Alifu Kuerban, Xilin Chen 0001 |
ACCV (2) | 1 |
| 2011 | A Novel Energy-Efficient Routing Algorithm in Multi-sink Wireless Sensor NetworksabstractIn wireless sensor networks, node energy resources are so limited that how to reduce energy consumption and prolong network lifetime become the primary factor that should be taken into account for the design of wireless sensor network routing protocols. In the single sink wireless sensor networks, the failure of sink will lead to entire network paralysis and reduce network reliability. In multi-sink wireless sensor networks, based on the mechanism of saving node energy and balancing flow of base stations, the zone of every sink in the network is divided. In each subnetwork, a location and energy based dynamical pre-clustering algorithm LEBDPC is proposed. LEBDPC not only balances energy consumption of nodes in the same cluster, but also energy consumption of nodes in different clusters. Simulation results show that compared with other clustering algorithms, LEBDPC algorithm can effectively reduce energy consumption and prolong network lifetime. Zhiwu Huang, Weirong Liu 0001 |
TrustCom | 1 |
| 2005 | On the stability and stabilization of linear neutral time-delay systemsabstractIn this paper, the problem of the stability and stabilization analysis for linear neutral time-delay systems is investigated. The time-delays considered here are assumed bounded but no information to be available. Using Lyapunov functional method, both delay-dependent and delay-independent conditions for the stability and stabilization of the systems are presented in terms of linear matrix inequality (LMI). Examples are given to illustrate the main result of this paper and compare with the results presented in literatures. Xiaohong Nian, Weihua Gui 0001, Zhiwu Huang |
SMC | 3 |
| 2005 | Robust H∞ control of linear uncertain neutral type systems with time-varying delayabstractThe problems of robust stability and robust H/sub /spl infin// control for a class of uncertain neutral systems with time-varying delay are investigated. The class describes linear state models with norm-bounded uncertain system parameters and time-varying delay. First, a sufficient condition for robust stability independent of time-varying delay is developed. Then, a sufficient condition for designing a memoryless state-feedback controller which stabilizes the uncertain neutral system under consideration and guarantees an H/sub /spl infin//-norm bound constraint on the disturbance attenuation for all admissible uncertainties is derived. In both problems, the results are expressed in the form of LMI. Xiaohong Nian, Zhiwu Huang, Weihua Gui 0001 |
SMC | 2 |