EDBT 2026 Demo / reviewers in the wild / expert
Zhixiang Huang
dblp:38/11054
· DBLP profile ↗
31ranked-venue papers
0as first author
30since 2021 · last 2025
0000-0002-8023-9075ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent-space diffusion models for stealthy and transferable adversarial attacks on object detection
Wenxuan Wang 0003, Huihui Qi, Zhixiang Huang, Bangjie Yin, Peng Wang 0015 |
Neurocomputing | 3 |
| 2025 | Visible-thermal multiple object tracking: Large-scale video dataset and progressive fusion approach
Yabin Zhu, Qianwu Wang, Chenglong Li 0002, Jin Tang 0001, Chengjie Gu, Zhixiang Huang |
Pattern Recognit. | 6 |
| 2025 | CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event CamerasabstractExisting datasets for RGB-DVS tracking are collected with DVS346 camera and their resolution ($346 \times 260$) is low for practical applications. Actually, only visible cameras are deployed in many practical systems, and the newly designed neuromorphic cameras may have different resolutions. The latest neuromorphic sensors can output high-definition event streams, but it is very difficult to achieve strict alignment between events and frames on both spatial and temporal views. Therefore, how to achieve accurate tracking with unaligned neuromorphic and visible sensors is a valuable but unresearched problem. In this work, we formally propose the task of object tracking using unaligned neuromorphic and visible cameras. We build the first unaligned frame-event dataset CRSOT collected with a specially built data acquisition system, which contains 1,030 high-definition RGB-Event video pairs, 304,974 video frames. In addition, we propose a novel unaligned object tracking framework that can realize robust tracking even using the loosely aligned RGB-Event data. This proposed method utilizes uncertainty perception techniques, which can effectively reduce the negative impact of noise (especially noise in event data) on tracking performance. Specifically, we extract the template and search regions of RGB and Event data and feed them into a unified ViT backbone for feature embedding. Next, we propose uncertainty perception modules to encode the RGB and Event features, respectively, then, we propose a modality uncertainty fusion module to aggregate the two modalities. These three branches are jointly optimized in the training phase. Extensive experiments demonstrate that our tracker can collaborate the dual modalities for high-performance tracking even without strictly temporal and spatial alignment. Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Lin Zhu 0012, Zhixiang Huang, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | The Causal Impact of Credit Lines on Spending DistributionsabstractConsumer credit services offered by electronic commerce platforms provide customers with convenient loan access during shopping and have the potential to stimulate sales. To understand the causal impact of credit lines on spending, previous studies have employed causal estimators, (e.g., direct regression (DR), inverse propensity weighting (IPW), and double machine learning (DML)) to estimate the treatment effect. However, these estimators do not treat the spending of each individual as a distribution that can capture the range and pattern of amounts spent across different orders. By disregarding the outcome as a distribution, valuable insights embedded within the outcome distribution might be overlooked. This paper thus develops distribution valued estimators which extend from existing real valued DR, IPW, and DML estimators within Rubin’s causal framework. We establish their consistency and apply them to a real dataset from a large electronic commerce platform. Our findings reveal that credit lines generally have a positive impact on spending across all quantiles, but consumers would allocate more to luxuries (higher quantiles) than necessities (lower quantiles) as credit lines increase. Yijun Li 0005, Cheuk Hang Leung, Xiangqian Sun, Chaoqun Wang 0007, Yiyan Huang, Qi Wu 0009, Dongdong Wang 0005, Zhixiang Huang |
AAAI | 9 |
| 2024 | National Image Resources Recommendation for Targeted International Communication via E-CARGO ModelabstractIt is important to select appropriate national image resources to build a nation’s images for targeted international communication of national images. Existing research only focusing on the methodologies, lacks the systematic modeling and solving of national image resources recommendation. A collaborative recommendation approach to national image resources is proposed. In the approach, an evaluation model of national image resources and an evaluation model of communication audiences are put forward, and an evaluation mechanism is proposed to measure the comprehensive compatibility between national image resources and communication audiences. By innovatively introducing the role-based collaboration (RBC) theory and the environment-classes, agents, roles, groups, and objects (E-CARGO) model, the national image resources recommendation is formalized as a collaborative optimization problem. The mathematical model is built and solved via an optimization package. Finally, the case study and experiments show that the approach is efficient, feasible, and conducive to enhancing the efficiency of national image resources recommendation. It offers a novel research paradigm for targeted international communication of national images. Hua Ma 0002, Xiangru Fu, Zhixiang Huang, Hong-Yu Zhang 0001 |
CSCWD | 4 |
| 2024 | ADR: An Adversarial Approach to Learn Decomposed Representations for Causal Inference
Guogang Tian, Zhixiang Huang |
ECML/PKDD (2) | 4 |
| 2024 | UTBoost: Gradient Boosted Decision Trees for Uplift Modeling
Dongdong Wang 0005, Zhixiang Huang, Bangqi Zheng |
PRICAI (1) | 4 |
| 2024 | Constrained squared sine derived adaptive algorithm: Performance and analysis
Liping Li 0001, Yingsong Li 0001, Zhixiang Huang |
Signal Process. | 4 |
| 2024 | RGBT Tracking via Progressive Fusion Transformer With Dynamically Guided LearningabstractExisting Transformer-based RGB-Thermal (RGBT) tracking methods either use cross-attention to fuse the two modalities, or use self-attention and cross-attention to model both modality-specific and modality-sharing information. However, the significant appearance gap between modalities limits the feature representation ability of certain modalities during the fusion process. To address this problem, we propose a novel Progressive Fusion Transformer called ProFormer, which progressively integrates single-modality information into the multimodal representation for robust RGBT tracking. In particular, ProFormer first uses a self-attention module to collaboratively extract the multimodal representation. Then, ProFormer introduces two cross-attention modules to interact it with the features of the dual modalities for enhancing modality-specific information in the multimodal representation. In addition, we propose a dynamically guided learning algorithm that adaptively employs the well-performing branches to guide the learning of other branches, to improve the representation ability of each branch. Extensive experiments demonstrate that our proposed ProFormer achieves a new state-of-the-art performance on RGBT210, RGBT234, LasHeR, and VTUAV datasets. Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Zhixiang Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | SARGap: A Full-Link General Decoupling Automatic Pruning Algorithm for Deep Learning-Based SAR Target DetectorsabstractSynthetic aperture radar (SAR) target detectors based on deep learning have difficulty finding a good balance between accuracy and speed. Current pruning methods are usually used for backbone consistent pruning and seldom directly for the whole structure of deep learning target detectors; therefore, for edge-end applications, this article proposes a new full-link general automatic pruning algorithm for SAR target detectors, referred to as SARGap. First, SARGap automatically analyzes the network structure by creating a dependency graph, divides the pair-coupled network structure into the same group, and prunes the same channel for the same group of network structures so that the algorithm can be applied to a variety of complex target detectors. Second, an automatic pruning rate search method (APRS) is designed to search for the optimal pruning rate of each group of network structures in the target detector. Finally, to find a good balance between precision and speed in the automatic search of the pruning rate, a multiobjective optimization loss function (MOOL) is constructed as the APRS objective function. A series of experiments based on SSDD and HRSID, two large-scale SAR target detection datasets, are carried out to prove the superiority of this method. Using Yolov5s as the baseline, SARGap can compress parameters by 84.29%/82.86% and flops by 80.50%/81.93% on two datasets with almost no loss of accuracy. In addition, SARGap can be applied to any deep learning target detector and match hardware computing resources to achieve optimal full-link pruning. Jingqian Yu, Jie Chen 0035, Huiyao Wan, Yice Cao, Zhixiang Huang, Yingsong Li 0001, Bocai Wu, Baidong Yao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Long-Term Motion-Assisted Remote Sensing Object TrackingabstractRemote sensing object tracking has gained significant attention due to its wide range of applications including surveillance and motion analysis. However, it faces various challenges such as low resolution, low contrast, blurring, and occlusion, which impede its development at a significantly slower pace compared to object tracking methods for general scenes. The challenges of low resolution, low contrast, and blurring result in weak target features, while the occlusion challenge poses a problem for target search range and tracker discrimination in subsequent frames. To address these issues, we propose a novel long-term motion-assisted framework, which can effectively mine long-term motion information and use an evaluation scheme for robust remote sensing object tracking. Specifically, we design a long-term motion feature mining module (LMFM), which efficiently calculates the long-term motion information by integrating previous motion features in a temporal-iterative manner to alleviate the problem of weak features caused by low resolution, low contrast, and blurring. Moreover, we design an evaluation scheme that combines the motion trajectory model, target classification scores, and predicted target positions to handle the issue of massive occlusion or target loss. Extensive experiments on the SatSOT, SV248S, and VISO datasets show that our approach outperforms state-of-the-art (SOTA) trackers. The source code, trained models, and raw results are released athttps://github.com/zhaoxingle/LMANet. Yabin Zhu, Xingle Zhao, Chenglong Li 0002, Jin Tang 0001, Zhixiang Huang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Deep Into the Domain Shift: Transfer Learning Through Dependence RegularizationabstractClassical domain adaptation methods acquire transferability by regularizing the overall distributional discrepancies between features in the source domain (labeled) and features in the target domain (unlabeled). They often do not differentiate whether the domain differences come from the marginals or the dependence structures. In many business and financial applications, the labeling function usually has different sensitivities to the changes in the marginals versus changes in the dependence structures. Measuring the overall distributional differences will not be discriminative enough in acquiring transferability. Without the needed structural resolution, the learned transfer is less optimal. This article proposes a new domain adaptation approach in which one can measure the differences in the internal dependence structure separately from those in the marginals. By optimizing the relative weights among them, the new regularization strategy greatly relaxes the rigidness of the existing approaches. It allows a learning machine to pay special attention to places where the differences matter the most. Experiments on three real-world datasets show that the improvements are quite notable and robust compared to various benchmark domain adaptation models. Shumin Ma, Zhiri Yuan, Qi Wu 0009, Yiyan Huang, Xixu Hu, Cheuk Hang Leung, Dongdong Wang 0005, Zhixiang Huang |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Tiny Object Tracking: A Large-Scale Dataset and a BaselineabstractTiny objects, frequently appearing in practical applications, have weak appearance and features, and receive increasing interests in many vision tasks, such as object detection and segmentation. To promote the research and development of tiny object tracking, we create a large-scale video dataset, which contains 434 sequences with a total of more than 217K frames. Each frame is carefully annotated with a high-quality bounding box. In data creation, we take 12 challenge attributes into account to cover a broad range of viewpoints and scene complexities, and annotate these attributes for facilitating the attribute-based performance analysis. To provide a strong baseline in tiny object tracking, we propose a novel multilevel knowledge distillation network (MKDNet), which pursues three-level knowledge distillations in a unified framework to effectively enhance the feature representation, discrimination, and localization abilities in tracking tiny objects. Extensive experiments are performed on the proposed dataset, and the results prove the superiority and effectiveness of MKDNet compared with state-of-the-art methods. The dataset, the algorithm code, and the evaluation code are available at https://github.com/mmic-lcl/Datasets-and-benchmark-code. Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Zhixiang Huang |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Towards Balanced Representation Learning for Credit Policy EvaluationabstractCredit policy evaluation presents profitable opportunities for E-commerce platforms through improved decision-making. The core of policy evaluation is estimating the causal effects of the policy on the target outcome. However, selection bias presents a key challenge in estimating causal effects from real-world data. Some recent causal inference methods attempt to mitigate selection bias by leveraging covariate balancing in the representation space to obtain the domain-invariant features. However, it is noticeable that balanced representation learning can be accompanied by a failure of domain discrimination, resulting in the loss of domain-related information. This is referred to as the over-balancing issue. In this paper, we introduce a novel objective for representation balancing methods to do policy evaluation. In particular, we construct a doubly robust loss based on the predictions of treatment and outcomes, serving as a prerequisite for covariate balancing to deal with the over-balancing issue. In addition, we investigate how to improve treatment effect estimations by exploiting the unconfoundedness assumption. The extensive experimental results on benchmark datasets and a newly introduced credit dataset show a general outperformance of our method compared with existing methods. Yiyan Huang, Cheuk Hang Leung, Shumin Ma, Zhiri Yuan, Qi Wu 0009, Dongdong Wang 0005, Zhixiang Huang |
AISTATS | 8 |
| 2023 | SARNas: A Hardware-Aware SAR Target Detection Algorithm via Multi-Objective Neural Architecture SearchabstractMost of the existing deep learning-based SAR target detection algorithms rely on manual experience to repeatedly adjust structure and parameters to design deep learning models that meet specific scenarios or tasks. The implementation of the above methods is complicated, the design efficiency is low, and most of the designed models are local optimal. In addition, edge-oriented applications have relatively limited computing resources, and it is difficult for existing deep learning compression methods to ensure a balance between accuracy and complexity. In this paper, we innovatively propose a hardware-aware SAR target detection algorithm via multi-objective NAS, referred to as SARNas. Our SARNas method takes object detection accuracy and model computational complexity as joint guidance objectives to automatically search for an optimal SAR object detection model end-to-end. Experimental results show that SARNas can automatically search for a better SAR target detector with balanced detection accuracy and computational complexity. Wentian Du, Jie Chen 0035, Zhixiang Huang |
IGARSS | 3 |
| 2023 | DeLELSTM: Decomposition-based Linear Explainable LSTM to Capture Instantaneous and Long-term Effects in Time SeriesabstractTime series forecasting is prevalent in various real-world applications. Despite the promising results of deep learning models in time series forecasting, especially the Recurrent Neural Networks (RNNs), the explanations of time series models, which are critical in high-stakes applications, have received little attention. In this paper, we propose a Decomposition-based Linear Explainable LSTM (DeLELSTM) to improve the interpretability of LSTM. Conventionally, the interpretability of RNNs only concentrates on the variable importance and time importance. We additionally distinguish between the instantaneous influence of new coming data and the long-term effects of historical data. Specifically, DeLELSTM consists of two components, i.e., standard LSTM and tensorized LSTM. The tensorized LSTM assigns each variable with a unique hidden state making up a matrix h(t), and the standard LSTM models all the variables with a shared hidden state H(t). By decomposing the H(t) into the linear combination of past information h(t-1) and the fresh information h(t)-h(t-1), we can get the instantaneous influence and the long-term effect of each feature. In addition, the advantage of linear regression also makes the explanation transparent and clear. We demonstrate the effectiveness and interpretability of DeLELSTM on three empirical datasets. Extensive experiments show that the proposed method achieves competitive performance against the baseline methods and provides a reliable explanation relative to domain knowledge. Chaoqun Wang 0007, Yijun Li 0005, Xiangqian Sun, Qi Wu 0009, Dongdong Wang 0005, Zhixiang Huang |
IJCAI | 6 |
| 2023 | A high-isolation coupled-fed building block for metal-rimmed 5G smartphonesabstractA compact coupled-fed dual-antenna building block has been constructed in this study. The building block is simple in structure and easy to process, and has a high degree of isolation. The dual-antenna building block is composed of a coupled-fed loop antenna and a coupled-fed slot antenna that completely overlap. Based on this dual-antenna module, an eight-element MIMO system is designed, and the fabricated eight-element MIMO array is measured. The measured isolation of the designed eight-element MIMO system is >18.5 dB without any decoupling element. In addition, the MIMO array has good measured efficiencies, with a measured efficiency variation range of 43%–54% in the entire working frequency band. The measured ECC of the MIMO system is <0.02. Therefore, the designed MIMO array has great potential in 5G metal-rimmed mobile phone applications. Aidi Ren, Chengwei Yu, Lixia Yang, Zhixiang Huang |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2023 | Squared Sine Adaptive Algorithm and Its Performance AnalysisabstractThe squared sine adaptive (SSA) algorithm is presented for identification scenarios, such as acoustic-echo cancellation (AEC) applications, in non-Gaussian environments. To devise the SSA algorithm, a novel cost function is constructed by exerting a sliding window-type squared sine function on the estimation error vector, which provides robustness in impulsive-noise environments and speeds up convergence when the input is colored. Theoretical results are presented for predicting the mean-weight, convergence, transient excess-mean-square-error (EMSE), and tracking behaviour. Moreover, the minimum EMSE and the optimum step size for tracking are presented. The computational complexity of the SSA algorithm has also been investigated. Numerical experiments demonstrate that results of the theoretical analysis match the simulated results very well and the proposed SSA algorithm outperforms known algorithms in AEC applications. Xinqi Huang, Yingsong Li 0001, Yuriy V. Zakharov, Yongchun Miao, Zhixiang Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | SARNas: A Hardware-Aware SAR Target Detection Algorithm via Multiobjective Neural Architecture SearchabstractMost of the existing deep learning-based SAR target detection algorithms rely on manual experience to repeatedly adjust structures and parameters to design models suitable for specific scenarios or tasks. The implementation of the above methods is complicated, the design efficiency is low, and it is difficult to ensure the balance between accuracy and complexity. We innovatively propose a hardware-aware SAR target detection algorithm via multiobjective neural architecture search (NAS), referred to as SARNas. First, we design a flexible and efficient search space, a supernet search strategy and a subnet contribution evaluation strategy. Furthermore, we construct a new NAS loss function, called SARMI-Loss, to guide the learning of a SAR object detector that balances accuracy and computational complexity. Our SAR-Nas method can address the resource limitations of edge devices and automatically search for the optimal SAR target detector in an end-to-end manner for any deep learning-based SAR baseline model. A series of comparative experiments on three SAR image object detection datasets (SSDD, HRSID and MSAR) demonstrate the superiority of our method. The experimental results with YOLOV5 as the benchmark model show that the detection accuracy of the target detection networks automatically found by using the SARNas method on the SSDD, HRSID, and MSAR datasets can reach 98.5%, 92.8%, and 91.8% in mean average precision (mAP) with only 2.31M, 1.99M, 2.21M parameters, respectively. The number of model parameters is reduced by 88.9%, 90.46%, and 68.5%, respectively, and the inference speed is increased by 51.6%, 46.1%, and 13.9% without losing accuracy. Wentian Du, Jie Chen 0035, Chaochen Zhang, Po Zhao, Huiyao Wan, Yice Cao, Zhixiang Huang, Yingsong Li 0001, Bocai Wu |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | Orientation Detector for Ship Targets in SAR Images Based on Semantic Flow Feature Alignment and Gaussian Label MatchingabstractTo address the challenges in synthetic aperture radar (SAR) ship target detection, this paper proposes a SAR ship small target orientation detector named FADet based on semantic flow feature alignment and Gaussian label matching. First, to solve the feature misalignment problem caused by feature extraction downsampling and residual connections, we introduce the FAM module into FPN, which automatically aligns deep and shallow fine-grained semantics information through semantic flow alignment. Second, due to the scattering characteristics of SAR imaging, the boundary information of SAR targets is not obvious, we combining attention mechanisms design an adaptive boundary enhancement module to enhance the target boundary information. Finally, to solve the problem that small targets have difficulty matching positive samples under IOU rules, we design a label matching strategy based on Gaussian distribution. This matching strategy can still learn regression information when two boxes do not intersect. Based on the SSDD+ and RSDD-SAR datasets, the effectiveness of each module in FADet is verified by ablation experiments. Additionally, through comparison experiments with the latest orientation detection methods, FADet achieves a good compromise between accuracy and inference speed. The AP50 and AP75 on the SSDD+ and RSDD-SAR is 91.03, 59.94 and 90.78, 59.91 respectively, and the FPS is 19.83. Huiyao Wan, Jie Chen 0035, Zhixiang Huang, Wentian Du, Feng Xu 0001, Feng Wang 0022, Bocai Wu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | HRLE-SARDet: A Lightweight SAR Target Detection Algorithm Based on Hybrid Representation Learning EnhancementabstractIn recent years, deep learning has been widely used in remote sensing, especially in the field of synthetic aperture radar (SAR) image target detection. However, all of these deep learning models continue increasing the network’s depth and width without maintaining a good balance between accuracy and speed. Therefore, in this article, we propose a hybrid representation learning-enhanced SAR target detection algorithm based on the unique features of SAR images from a lightweight perspective called HRLE-SARDet. First, we design a lightweight and scattering feature extraction backbone that is more suitable for SAR image data. Second, for the multiscale feature discrepancy, we design a new multiscale feature fusion neck. Next, to better extract the scattering information from small targets of SAR images and improve the detection accuracy, we design a lightweight hybrid representation learning enhancement module. Finally, to better fit target detection for SAR image datasets, we redesign a more flexible loss function, which allows for an easy adjustment of the importance of polynomial bases according to the target task and dataset. Extensive experimental results on three SAR image ship target datasets (SSDD, AIR-SARShip-2.0, and HRSID) and a newly released large multiclass target SAR dataset (MSAR-1.0) show that our HRLE-SARDet achieves 98.4%, 79.2%, 92.5%, and 88.4% mean average precision (mAP) with only 1.09 M parameters and 2.5 G floating-point operations (FLOPs) on the SSDD, AIR-SARShip-2.0, HRSID, and MSAR-1.0 datasets, respectively, which is an excellent performance. Jie Chen 0035, Zhixiang Huang, Jianming Lv, Honglin Luo, Bocai Wu, Yingsong Li 0001, Paulo S. R. Diniz |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | FPT: Fine-Grained Detection of Driver Distraction Based on the Feature Pyramid Vision TransformerabstractAccording to the surveys of the World Health Organization, distracted driving is one of main causes of road traffic accidents. To improve road traffic safety, real-time detection of drivers’ driving behavior is very important for the development of highly reliable Advanced Driver Assistance System (ADAS). At present, the deep learning architecture based on a Convolutional Neural Network (CNN) has disadvantages such as large number of parameters and weak global feature extraction ability. Therefore, this paper proposes an innovative driver distraction detection model based on the fusion of a transformer and a CNN, referred to as FPT, which is the first exploration in the field of driver distraction detection. First, we introduce the latest Twins transformer as a benchmark. Then, we design residual embedding to replace block embedding, which can further integrate the convolutional neural network with Transformer and improve the feature extraction ability. In addition, the Multilayer Perceptron (MLP) module with a large parameter occupancy rate in the original transformer structure is replaced with a lightweight group convolution module to reduce computational complexity. Finally, a cross-entropy loss function for label smoothing is designed to guide network learning with significantly differentiated features. Comparison results on two large-scale driver distraction detection datasets show that the proposed FPT offers a better compromise between computational cost and performance compared to the state-of-the-art CNN and Transformer architectures. Jie Chen 0035, Zhixiang Huang, Bing Li 0033, Jianming Lv, Jingmin Xi, Bocai Wu, Jun Zhang 0034, ZhongCheng Wu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Robust Causal Learning for the Estimation of Average Treatment EffectsabstractMany practical decision-making problems in economics and healthcare seek to estimate the average treatment effect (ATE) from observational data. The Double/Debiased Machine Learning (DML) is one of the prevalent methods to estimate ATE in the observational study. However, the DML estimators can suffer an error-compounding issue and even give an extreme estimate when the propensity scores are misspecified or very close to 0 or 1. Previous studies have overcome this issue through some empirical tricks such as propensity score trimming, yet none of the existing literature solves this problem from a theoretical standpoint. In this paper, we propose a Robust Causal Learning (RCL) method to offset the deficiencies of the DML estimators. Theoretically, the RCL estimators i) are as consistent and doubly robust as the DML estimators, and ii) can get rid of the error-compounding issue. Empirically, the comprehensive experiments show that i) the RCL estimators give more stable estimations of the causal parameters than the DML estimators, and ii) the RCL estimators outperform the traditional estimators and their variants when applying different machine learning models on both simulation and benchmark datasets. Yiyan Huang, Cheuk Hang Leung, Qi Wu 0009, Shumin Ma, Zhiri Yuan, Dongdong Wang 0005, Zhixiang Huang |
IJCNN | 8 |
| 2022 | Moderately-Balanced Representation Learning for Treatment Effects with Orthogonality Information
Yiyan Huang, Cheuk Hang Leung, Shumin Ma, Qi Wu 0009, Dongdong Wang 0005, Zhixiang Huang |
PRICAI (2) | 6 |
| 2022 | Ship Detection Method Based on Scattering Contribution for PolSAR ImageabstractDue to the differentiation of polarimetric scattering mechanisms between ships and sea surface, designing the ship detection method in polarimetric synthetic aperture radar (PolSAR) is a potential promising technique and has been paid extensive attention. The complexity of sea clutter and weak scattering of small ships result in a great challenge for high-precision ship detection. In this letter, we investigate the scattering mechanisms of ships to improve the detection performance and propose a novel ship detection method based on the principal contribution of scattering mechanisms. First, the seven-component model-based decomposition (SCMD) is used to analyze the scattering mechanisms of ships. Second, the primary scattering contribution and local contrast (SCLC) mechanism are used to enhance ships, especially small ships. Finally, the threshold segmentation is used to realize the extraction of ships. Experimental results by real PolSAR data not only verify the rationality and effectiveness of the constructed detection metric but also show the clear superiority of the proposed detection method, which can encourage further application of polarimetric scattering mechanisms in ship detection. Xueli Pan, Lixia Yang, Zhixiang Huang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | AFSar: An Anchor-Free SAR Target Detection Algorithm Based on Multiscale Enhancement Representation LearningabstractUnlike optical images, synthetic aperture radar (SAR) images have unique characteristics, such as few samples, strong scattering, sparseness, multiple scales, complex interference and background, and inconspicuous target edge contour information. Current SAR target detection algorithms have difficulty in balancing accuracy and speed, and the performance of these algorithms is relatively limited, thus making it difficult to deploy practical applications. To this end, this article proposes AFSar, an innovative anchor-free SAR target detection algorithm based on multiscale enhancement representation learning. First, we introduce the latest anchor-free architecture YOLOX as the basic framework. Second, to reduce the computational complexity of the model and to improve the ability of multiscale feature extraction, we redesigned the lightweight backbone, namely, MobileNetV2S. Furthermore, we propose an attention enhancement PAN module, called CSEMPAN, which highlights the unique strong scattering characteristics of SAR targets by integrating channel and spatial attention mechanisms. Finally, in view of the multiscale and strong sparse characteristics of SAR targets, we propose a new target detection head, namely, ESPHead. ESPHead extracts the features of targets with different scales by using dilated convolution with different dilated rates, so as to enhance the detection ability of the model for targets with different scales. The results of ablation experiments on the SSDD dataset show that the mAP of our algorithm reaches 0.977, while the Flops is only 9.86 G, achieving state of the art. Huiyao Wan, Jie Chen 0035, Zhixiang Huang, Runfan Xia, Bocai Wu, Baidong Yao, Mengdao Xing |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | FSODS: A Lightweight Metalearning Method for Few-Shot Object Detection on SAR ImagesabstractAt present, few-shot object detection research in the field of optical remote sensing images has been conducted, but few-shot object detection in the field of SAR images have rarely been explored. To this end, this paper proposes a lightweight meta-learning-based SAR image few-shot object detection method, which improves the accuracy and speed of SAR image few-shot object detection from a more balanced perspective. First, we introduce the latest FSODM method in optical remote sensing as a benchmark framework. Second, a lightweight meta-feature extractor named DarknetS is designed to enhance the feature representation of SAR images and improve detection timeliness. Furthermore, we build a new aggregation module called AggregationS, which encodes support features and query features into the same feature subspace via a novel transformer encoder. This module design can better extract the correlation and saliency between different classes in the support set, improve the detection accuracy of the query set, and enhance the detection generalization performance of new classes. Finally, we built several real-world SAR image few-shot object detection datasets to verify the effectiveness of the method. Experimental results show that FSODS can achieve a better object detection performance compared to the baseline model under the condition that only a small amount of labelled data is required for new classes of SAR image objects. Jie Chen 0035, Zhixiang Huang, Huiyao Wan, Pei Chang, Baidong Yao, Bocai Wu, Mengdao Xing |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A New Unsupervised Deep Learning Algorithm for Fine-Grained Detection of Driver DistractionabstractTraffic accidents caused by distracted drivers account for a large proportion of traffic accidents each year, and monitoring the driving state of drivers to avoid traffic accidents caused by distracted driving has become a very important research direction. At present, the field of driver distraction detection mainly adopts supervised learning methods, which have problems such as poor generalization ability, large labeling cost, and weak artificial intelligence. This paper is oriented toward driver distraction fine-grained detection and innovatively proposes a new unsupervised deep learning algorithm, which is referred to as UDL, to achieve a more human-like level of intelligence. First, we build a new unsupervised deep learning algorithm; furthermore, we integrate the multilayer perceptron (MLP) architecture to build a new backbone and projection head to strengthen feature extraction capabilities; and finally, a new loss function based on contrast learning and a stop-gradient strategy is designed to guide the model to learn more robust features. The comparison results on large-scale driver distraction detection datasets show that our UDL method can accurately detect driver distraction without labels and exhibits excellent generalization performance with a linear evaluation accuracy of 97.38%; In addition, after fine-tuning with fewer labels, our UDL method can achieve superior performance close to state-of-the-art supervised learning methods, achieving 99.07% accuracy after fine-tuning using only 50% of the labeled data, which greatly reduces the cost and limitations of manual annotation. Bing Li 0033, Jie Chen 0035, Zhixiang Huang, Jianming Lv, Jingmin Xi, Jun Zhang 0034, ZhongCheng Wu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | The Causal Learning of Retail DelinquencyabstractThis paper focuses on the expected difference in borrower's repayment when there is a change in the lender's credit decisions. Classical estimators overlook the confounding effects and hence the estimation error can be magnificent. As such, we propose another approach to construct the estimators such that the error can be greatly reduced. The proposed estimators are shown to be unbiased, consistent, and robust through a combination of theoretical analysis and numerical testing. Moreover, we compare the power of estimating the causal quantities between the classical estimators and the proposed estimators. The comparison is tested across a wide range of models, including linear regression models, tree-based models, and neural network-based models, under different simulated datasets that exhibit different levels of causality, different degrees of nonlinearity, and different distributional properties. Most importantly, we apply our approaches to a large observational dataset provided by a global technology firm that operates in both the e-commerce and the lending business. We find that the relative reduction of estimation error is strikingly substantial if the causal effects are accounted for correctly. Yiyan Huang, Cheuk Hang Leung, Qi Wu 0009, Nanbo Peng, Dongdong Wang 0005, Zhixiang Huang |
AAAI | 7 |
| 2021 | Fine-Grained Detection of Driver Distraction Based on Neural Architecture SearchabstractIn the future, vehicles will be equipped with increasingly advanced interactive intelligent electronic devices, which will induce drivers to conduct secondary tasks, thereby leading to distractions. Therefore, the detection and early warning of driver distraction are essential for improving driving safety and pose an important challenge in intelligent transportation systems. Previous studies used traditional machine learning and deep learning transfer models, which have the disadvantages of complicated and time-consuming manual feature engineering, strong subjectivity, and weak generalization performance. In this paper, we propose a fine-grained detection method for driver distraction based on neural architecture search. First, we design an automatic construction algorithm for deep convolutional neural networks based on neural architecture search, which automatically searches for the optimal deep convolutional neural network architecture without human involvement. In addition, we fuse driver-related multisource perception information, use an automatically constructed deep convolutional neural network to extract high-dimensional mapping features, and implement fine-grained detection of various types of driver distraction states. The results on a large-scale multimodal driver distraction dataset demonstrate that our method can efficiently search an optimal deep convolutional neural network, which can quickly converge, and can accurately detect the considered types of driver distraction states, the average detection accuracy reaches 99.7796%; moreover, it has satisfactory robustness. Jie Chen 0035, Zhixiang Huang, Xiaohui Guo, Bocai Wu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2014 | A novel image denoising algorithm using linear Bayesian MAP estimation based on sparse representation
Dong Sun 0003, Qingwei Gao, Yixiang Lu, Zhixiang Huang, Teng Li 0001 |
Signal Process. | 4 |