EDBT 2026 Demo / reviewers in the wild / expert
Xiaolin Hu 0001
dblp:60/6028-1
· DBLP profile ↗
148ranked-venue papers
28as first author
77since 2021 · last 2026
0000-0002-4907-7354ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 126 · 22 first-author · 67 since 2021Graphics, computer vision, multimedia, augmented reality and games · 56 · 1 first-author · 33 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FGNet: Leveraging Feature-Guided Attention to Refine SAM2 for 3D EM Neuron SegmentationabstractAccurate segmentation of neural structures in Electron Microscopy (EM) images is paramount for neuroscience. However, this task is challenged by intricate morphologies, low signal-to-noise ratios, and scarce annotations, limiting the accuracy and generalization of existing methods. To address these challenges, we seek to leverage the priors learned by visual foundation models on a vast amount of natural images to better tackle this task. Specifically, we propose a novel framework that can effectively transfer knowledge from Segment Anything 2 (SAM2), which is pre-trained on natural images, to the EM domain. We first use SAM2 to extract powerful, general-purpose features. To bridge the domain gap, we introduce a Feature-Guided Attention module that leverages semantic cues from SAM2 to guide a lightweight encoder, the Fine-Grained Encoder (FGE), in focusing on these challenging regions. Finally, a dual-affinity decoder generates both coarse and refined affinity maps. Experimental results demonstrate that our method achieves performance comparable to state-of-the-art (SOTA) approaches with the SAM2 weights frozen. Upon further fine-tuning on EM data, our method significantly outperforms existing SOTA methods. This study validates that transferring representations pre-trained on natural images, when combined with targeted domain-adaptive guidance, can effectively address the specific challenges in neuron segmentation. Zhenghua Li, Hang Chen 0004, Kai Li 0047, Xiaolin Hu 0001 |
AAAI | 5 |
| 2026 | VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG SamplesabstractRetrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating particular strength in real-time queries and Visual Question Answering tasks. However, the effectiveness of RAG is frequently hindered by the precision of the retriever: many retrieved samples fed into the generation phase are irrelevant or misleading, posing a critical bottleneck to LLMs’ performance. To address this challenge, we introduce \textbf{VaccineRAG}, a novel Chain-of-Thought-based retrieval-augmented generation dataset. On one hand, VaccineRAG employs a benchmark to evaluate models using data with varying positive/negative sample ratios, systematically exposing inherent weaknesses in current LLMs. On the other hand, it enhances models’ sample-discrimination capabilities by prompting LLMs to generate explicit Chain-of-Thought (CoT) analysis for each sample before producing final answers. Furthermore, to enhance the model’s ability to learn long-sequence complex CoT content, we propose \textbf{Partial-GRPO}. By modeling the outputs of LLMs as multiple components rather than a single whole, our model can make more informed preference selections for complex sequences, thereby enhancing its capacity to learn complex CoT. Comprehensive evaluations and ablation studies on VaccineRAG validate the effectiveness of the proposed scheme. Qixin Sun, Ziqin Wang, Hengyuan Zhao, Kaiyou Song, Si Liu 0001, Xiaolin Hu 0001, Qingpei Guo, Linjiang Huang |
AAAI | 7 |
| 2026 | Put the Space of LoRA Initialization to the Extreme to Preserve Pre-trained KnowledgeabstractLow-Rank Adaptation (LoRA) is the leading parameter-efficient fine-tuning method for Large Language Models (LLMs), but it still suffers from catastrophic forgetting. Recent work has shown that specialized LoRA initialization can alleviate catastrophic forgetting. There are currently two approaches to LoRA initialization aimed at preventing knowledge forgetting during fine-tuning: (1) making residual weights close to pre-trained weights, and (2) ensuring the space of LoRA initialization is orthogonal to pre-trained knowledge. The former is what current methods strive to achieve, while the importance of the latter is not sufficiently recognized. We find that the space of LoRA initialization is the key to preserving pre-trained knowledge rather than the residual weights. Existing methods like MiLoRA propose making the LoRA initialization space orthogonal to pre-trained weights. However, MiLoRA utilizes the null space of pre-trained weights. Compared to pre-trained weights, the input activations of pre-trained knowledge take into account the parameters of all previous layers as well as the input data, while pre-trained weights only contain information from the current layer. Moreover, we find that the effective ranks of input activations are much smaller than those of pre-trained weights. Thus, the null space of activations is more accurate and contains less pre-trained knowledge information compared to that of weights. Based on these, we introduce LoRA-Null, our proposed method that initializes LoRA in the null space of activations. Experimental results show that LoRA-Null effectively preserves the pre-trained world knowledge of LLMs while achieving good fine-tuning performance, as evidenced by extensive experiments. Pengwei Tang, Xiaolin Hu 0001, Yong Liu 0018, Lizhong Ding 0001, Debing Zhang |
AAAI | 2 |
| 2026 | A turbo-inference strategy for object detection and instance segmentation
Xiaolin Hu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2026 | Defending Against Patch-Based and Texture-Based Adversarial Attacks With Spectral DecompositionabstractAdversarial examples present significant challenges to the security of Deep Neural Network (DNN) applications. Specifically, there are patch-based and texture-based attacks that are usually used to craft physical-world adversarial examples, posing real threats to security-critical applications such as person detection in surveillance and autonomous systems, because those attacks are physically realizable. Existing defense mechanisms face challenges in the adaptive attack setting, i.e., the attacks are specifically designed against them. In this paper, we propose Adversarial Spectrum Defense (ASD), a defense mechanism that leverages spectral decomposition via Discrete Wavelet Transform (DWT) to analyze adversarial patterns across multiple frequency scales. The multi-resolution and localization capability of DWT enables ASD to capture both high-frequency (fine-grained) and low-frequency (spatially pervasive) perturbations. By integrating this spectral analysis with the off-the-shelf Adversarial Training (AT) model, ASD provides a comprehensive defense strategy against both patch-based and texture-based adversarial attacks. Extensive experiments demonstrate that ASD+AT achieved state-of-the-art (SOTA) performance against various attacks, outperforming the APs of previous defense methods by 21.73%, in the face of strong adaptive adversaries specifically designed against ASD. Wei Zhang 0370, Xinyu Chang, Xiao Li 0028, Xiaolin Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing TopologyabstractZeroth-order (ZO) optimization as the gradient-free method has become a powerful tool when the first-order gradient is unavailable or expensive to obtain, especially in decentralized learning scenarios where data and computational resources are distributed across multiple clients. There have been many efforts to analyze the optimization convergence rate of zeroth-order decentralized stochastic gradient descent (ZO-DSGD) algorithms. However, the generalization of these methods has not been well studied. In this paper, we provide a generalization analysis of ZO-DSGD with changing topology, where the clients run zeroth-order SGD with local data and communicate with each other according to time-varying topology. We systematically analyze the generalization error in convex, strongly convex, and non-convex cases. The obtained results in the convex and strongly convex cases with zeroth-order oracles recover the results of SGD. Moreover, the generalization bounds derived in non-convex cases align with that of DSGD. To capture the influence of communication topology on the generalization performance, we analyze local generalization bounds concerning local models held at different clients. The obtained results reflect the influence of the number of clients, local sample size, and topology on the generalization error. To the best of our knowledge, this is the first work that provides a generalization analysis of zeroth-order decentralized stochastic gradient descent methods and recovers the results of SGD. Xiaolin Hu 0001, Zixuan Gong, Gengze Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020 |
AAAI | 1 |
| 2025 | Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech SeparationabstractExisting causal speech separation models often under perform compared to non-causal models due to difficulties in retaining historical information. To address this, we propose the Time-Frequency Attention Cache Memory (TFACM) model, which effectively captures spatio-temporal relationships through an attention mechanism and cache memory (CM) for historical information storage. In TFACM, an LSTM layer captures frequency-relative positions, while causal modeling is applied to the time dimension using local and global representations. The CM module stores past information, and the causal attention refinement (CAR) module further enhances time-based feature representations for finer granularity. Experimental results showed that TFACM achieved comparable performance to the SOTA TF-GridNet-Causal model, with significantly lower complexity and fewer trainable parameters. For more details, visit the project page: https://anonymous.4open.science/w/TFACM-Page/. Kai Li 0047, Runxuan Yang, Xiaolin Hu 0001 |
ASRU | 4 |
| 2025 | PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuningabstractLow-rank adaptation (LoRA) and its variants have recently gained much interest due to their ability to avoid excessive inference costs. However, LoRA still encounters the following challenges: (1) Limitation of low-rank assumption; and (2) Its initialization method may be suboptimal. To this end, we propose PMSS(Pre-trained Matrices Skeleton Selection), which enables high-rank updates with low costs while leveraging semantic and linguistic information inherent in pre-trained weight. It achieves this by selecting skeletons from the pre-trained weight matrix and only learning a small matrix instead. Experiments demonstrate that PMSS outperforms LoRA and other fine-tuning methods across tasks with much less trainable parameters. We demonstrate its effectiveness, especially in handling complex tasks such as DROP benchmark(+3.4%/+5.9% on LLaMA2-7B/13B) and math reasoning (+12.89%/+5.61%/+3.11% on LLaMA2-7B, Mistral-7B and Gemma-7B of GSM8K).The code and model will be released soon. Qibin Wang, Xiaolin Hu 0001, Weikai Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004 |
COLING | 2 |
| 2025 | Improving Accuracy and Calibration via Differentiated Deep Mutual LearningabstractDeep Neural Networks (DNNs) have achieved remarkable success in a variety of tasks, particularly in terms of prediction accuracy. However, in real-world scenarios, especially in safety-critical applications, accuracy alone is insufficient; reliable uncertainty estimates are essential. Modern DNNs, often trained with cross-entropy loss, tend to exhibit overconfidence, especially on ambiguous samples. Many techniques aim to improve uncertainty calibration, yet they often come at the cost of reduced accuracy or increased computational demands. To address this challenge, we propose Differentiated Deep Mutual Learning (Diff-DML), an efficient ensemble approach that simultaneously enhances accuracy and uncertainty calibration. Diff-DML draws inspiration from Deep Mutual Learning (DML) while introducing two strategies to maintain prediction diversity: (1) Differentiated Training Strategy (DTS) and (2) Diversity-Preserving Learning Objective (DPLO). Our theoretical analysis shows that Diff-DML’s diversified learning framework not only leverages ensemble benefits but also avoids the loss of prediction diversity observed in traditional DML setups, which is crucial for improved calibration. Extensive evaluations on various benchmarks confirm the effectiveness of Diff-DML. For instance, on the CIFAR-100 dataset, Diff-DML on ResNet34/50 models achieved substantial improvements over the previous state-of-the-art method, MDCA, with absolute accuracy gains of 1.3%/3.1%, relative ECE reductions of 49.6%/43.8%, and relative classwise-ECE reductions of 7.7%/13.0%. Peng Cui 0007, Bingning Wang, Weipeng Chen, Jun Zhu 0001, Xiaolin Hu 0001 |
CVPR | 7 |
| 2025 | A Fast and Lightweight Model for Causal Audio-Visual Speech SeparationabstractAudio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on future context, operating offline, which renders them unsuitable for real-time applications. Inspired by the pipeline of RTFSNet, we propose a novel streaming AVSS model, named Swift-Net, which enhances the causal processing capabilities required for real-time applications. Swift-Net adopts a lightweight visual feature extraction module and an efficient fusion module for audio-visual integration. Additionally, Swift-Net employs Grouped SRUs to integrate historical information across different feature spaces, thereby improving the utilization efficiency of historical information. We further propose a causal transformation template to facilitate the conversion of non-causal AVSS models into causal counterparts. Experiments on three standard benchmark datasets (LRS2, LRS3, and VoxCeleb2) demonstrated that under causal conditions, our proposed Swift-Net exhibited outstanding performance, highlighting the potential of this method for processing speech in complex environments. Wendi Sang, Kai Li 0047, Runxuan Yang, Jianqiang Huang 0002, Xiaolin Hu 0001 |
ECAI | 5 |
| 2025 | PBCAT: Patch-Based Composite Adversarial Training Against Physically Realizable Attacks on Object DetectionabstractObject detection plays a crucial role in many security-sensitive applications. However, several recent studies have shown that object detectors can be easily fooled by physically realizable attacks, \eg, adversarial patches and recent adversarial textures, which pose realistic and urgent threats. Adversarial Training (AT) has been recognized as the most effective defense against adversarial attacks. While AT has been extensively studied in the $l_\infty$ attack settings on classification models, AT against physically realizable attacks on object detectors has received limited exploration. Early attempts are only performed to defend against adversarial patches, leaving AT against a wider range of physically realizable attacks under-explored. In this work, we consider defending against various physically realizable attacks with a unified AT method. We propose PBCAT, a novel Patch-Based Composite Adversarial Training strategy. PBCAT optimizes the model by incorporating the combination of small-area gradient-guided adversarial patches and imperceptible global adversarial perturbations covering the entire image. With these designs, PBCAT has the potential to defend against not only adversarial patches but also unseen physically realizable attacks such as adversarial textures. Extensive experiments in multiple settings demonstrated that PBCAT significantly improved robustness against various physically realizable attacks over state-of-the-art defense methods. Notably, it improved the detection accuracy by 29.7\% over previous defense methods under one recent adversarial texture attack. Xiao Li 0028, Wei Zhang 0370, Yingzhe He, Xiaolin Hu 0001 |
ICCV | 7 |
| 2025 | Efficient Neuron Segmentation in Electron Microscopy by Affinity-Guided QueriesabstractAccurate segmentation of neurons in electron microscopy (EM) images plays a crucial role in understanding the intricate wiring patterns of the brain. Existing automatic neuron segmentation methods rely on traditional clustering algorithms, where affinities are predicted first, and then watershed and post-processing algorithms are applied to yield segmentation results. Due to the nature of watershed algorithm, this paradigm has deficiency in both prediction quality and speed. Inspired by recent advances in natural image segmentation, we propose to use query-based methods to address the problem because they do not necessitate watershed algorithms. However, we find that directly applying existing query-based methods faces great challenges due to the large memory requirement of the 3D data and considerably different morphology of neurons. To tackle these challenges, we introduce affinity-guided queries and integrate them into a lightweight query-based framework. Specifically, we first predict affinities with a lightweight branch, which provides coarse neuron structure information. The affinities are then used to construct affinity-guided queries, facilitating segmentation with bottom-up cues. These queries, along with additional learnable queries, interact with the image features to directly predict the final segmentation results. Experiments on benchmark datasets demonstrated that our method achieved better results over state-of-the-art methods with a 2$\sim$3$\times$ speedup in inference. Code is available at https://github.com/chenhang98/AGQ. Hang Chen 0004, Chufeng Tang, Xiao Li 0028, Xiaolin Hu 0001 |
ICLR | 4 |
| 2025 | Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from GeneralizationabstractLarge language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: \textbf{(a) Limited \textit{i.i.d.} Setting.} Most studies focus on supervised function learning tasks where prompts are constructed with \textit{i.i.d.} input-label pairs. This \textit{i.i.d.} assumption diverges significantly from real language learning scenarios where prompt tokens are interdependent. \textbf{(b) Lack of Emergence Explanation.} Most literature answers \textbf{\textit{what}} ICL does from an implicit optimization perspective but falls short in elucidating \textbf{\textit{how}} ICL emerges and the impact of pre-training phase on ICL. In our paper, to extend (a), we adopt a more practical paradigm, \textbf{\textit{auto-regressive next-token prediction (AR-NTP)}}, which closely aligns with the actual training of language models. Specifically, within AR-NTP, we emphasize prompt token-dependency, which involves predicting each subsequent token based on the preceding sequence. To address (b), we formalize a systematic pre-training and ICL framework, highlighting the layer-wise structure of sequences and topics, alongside a two-level expectation. In conclusion, we present data-dependent, topic-dependent and optimization-dependent PAC-Bayesian generalization bounds for pre-trained LLMs, investigating that \textbf{\textit{ICL emerges from the generalization of sequences and topics}}. Our theory is supported by experiments on numerical linear dynamic systems, synthetic GINC and real-world language datasets. Zixuan Gong, Xiaolin Hu 0001, Huayi Tang, Yong Liu 0018 |
ICLR | 2 |
| 2025 | ADePT: Adaptive Decomposed Prompt Tuning for Parameter-Efficient Fine-tuningabstractPrompt Tuning (PT) enables the adaptation of Pre-trained Large Language Models (PLMs) to downstream tasks by optimizing a small amount of soft virtual tokens, which are prepended to the input token embeddings. Recently, Decomposed Prompt Tuning (DePT) has demonstrated superior adaptation capabilities by decomposing the soft prompt into a shorter soft prompt and a pair of low-rank matrices. The product of the pair of low-rank matrices is added to the input token embeddings to offset them. Additionally, DePT achieves faster inference compared to PT due to the shorter soft prompt. However, in this paper, we find that the position-based token embedding offsets of DePT restricts its ability to generalize across diverse model inputs, and that the shared embedding offsets across many token embeddings result in sub-optimization. To tackle these issues, we introduce \textbf{A}daptive \textbf{De}composed \textbf{P}rompt \textbf{T}uning (ADePT), which is composed of a short soft prompt and a shallow token-shared feed-forward neural network. ADePT utilizes the token-shared feed-forward neural network to learn the embedding offsets for each token, enabling adaptive embedding offsets that vary according to the model input and better optimization of token embedding offsets. This enables ADePT to achieve superior adaptation performance without requiring more inference time or additional trainable parameters compared to vanilla PT and its variants. In comprehensive experiments across 23 natural language processing tasks and 4 typical PLMs of different scales, ADePT consistently surpasses the leading parameter-efficient fine-tuning methods, and even outperforms the full fine-tuning in certain scenarios. We also provide a theoretical analysis towards ADePT. Code is available at https://github.com/HungerPWAY/ADePT. Pengwei Tang, Xiaolin Hu 0001, Yong Liu 0018 |
ICLR | 2 |
| 2025 | SPMamba: Leveraging Long-Sequence Modeling with State Space Models for Speech SeparationabstractExisting CNN-based speech separation models face local receptive field limitations and cannot effectively capture long time dependencies. Although LSTM and Transformer-based speech separation models can avoid this problem, their high complexity causes them to face the challenge of computational resources and inference efficiency when dealing with long audio. To address this challenge, we introduce an innovative speech separation method called SPMamba. This model builds upon the robust TF-GridNet architecture, replacing its traditional BLSTM modules with bidirectional Mamba modules. These modules effectively model the spatiotemporal relationships between the time and frequency dimensions, allowing SPMamba to capture long-range dependencies with linear computational complexity. Specifically, the bidirectional processing within the Mamba modules enables the model to utilize both past and future contextual information, thereby enhancing separation performance. Extensive experiments were conducted on public datasets, including the WSJ0-2Mix and WHAM! and Libri2Mix, as well as the newly constructed Echo2Mix dataset, demonstrated that SPMamba achieved superior results to previous state-of-the-art (SOTA) models with reduced computational complexity. These findings highlight the effectiveness of SPMamba in addressing the intricate challenges of speech separation in complex environments. The source code for SPMamba is publicly accessible at https://anonymous.4open.science/r/SPMamba-ICME/. Kai Li 0047, Runxuan Yang, Xiaolin Hu 0001 |
ICME | 4 |
| 2025 | Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and OptimizationabstractLarge Language Models (LLMs), built on Transformer architectures, exhibit remarkable generalization across a wide range of tasks. However, fine-tuning these models for specific tasks remains resource-intensive due to their extensive parameterization. In this paper, we explore two remarkable phenomena related to the attention mechanism during the fine-tuning of LLMs (where Wq, Wk, and Wv denote the weights of the query, key, and value layers, respectively). The first phenomenon, termed “Unequal Importance of Attention Matrices”, highlights the impact of fine-tuning different weight matrices. It shows that optimizing the Wv matrix yields significantly better performance than optimizing the Wk matrix. Fine-tuning only the Wq and Wv matrices is computationally efficient while delivering results comparable to, or even better than fine-tuning all three matrices (Wq, Wk, and Wv). The second phenomenon, “Attention Matrices with Customized Learning Rate Lead to Better Convergence”, emphasizes the importance of assigning distinct learning rates to these matrices. Specifically, a higher learning rate for the Wv matrix compared to Wq and Wk accelerates convergence and improves performance. Building on these insights, we propose a new strategy that improves fine-tuning efficiency in terms of both storage and time. Experimental results on benchmark datasets validate the effectiveness of this approach, supporting our theoretical findings. Our analysis lays the theoretical groundwork for configuring and improving algorithms in LLMs fine-tuning. Xinhao Yao, Hongjin Qian, Xiaolin Hu 0001, Gengze Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020 |
IJCAI | 3 |
| 2025 | DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
Xiaolin Hu 0001, Xiang Cheng 0007, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020 |
PAKDD (5) | 1 |
| 2025 | Physical Adversarial Examples for Person Detectors in Thermal Images Based on 3D ModelingabstractThermal Infrared detection is widely used in autonomous driving, medical AI, etc., but its security has only attracted attention recently. We propose infrared adversarial clothing designed to evade thermal person detectors in real-world scenarios. The design of the adversarial clothing is based on 3D modeling, which makes it easier to simulate multiangle scenes near the real world compared to 2D modeling. We optimized the patch layout pattern of 3D clothing based on the adversarial example technique and made physical adversarial clothing using the aerogel. The idea is to paste a set of square aerogel patches, which display black squares in thermal images, in the inner side of clothing at specific locations with specific orientations. To enhance realism, we propose a method to build infrared 3D models with real infrared photos and develop texture maps for 3D models to simulate varied infrared characteristics over time and location. In physical attacks, we achieved an attack success rate of 80.11% indoors and 76.85% outdoors against YOLOv9. In contrast, randomly placed patches yielded much lower success rates (26.53% indoors and 23.03% outdoors). The adversarial clothing also showed good transferability to unknown detectors with an ensemble attack method, demonstrating the effectiveness of our approach. Xiaopei Zhu, Zhanhao Hu, Jianmin Li 0001, Jun Zhu 0001, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | CEDNet: A cascade encoder-decoder network for dense prediction
Chufeng Tang, Jianmin Li 0001, Xiaolin Hu 0001 |
Pattern Recognit. | 5 |
| 2025 | On the Importance of Backbone to the Adversarial Robustness of Object DetectorsabstractObject detection is a critical component of various security-sensitive applications, such as autonomous driving and video surveillance. However, existing object detectors are vulnerable to adversarial attacks, which poses a significant challenge to their reliability and security. Through experiments, first, we found that existing works on improving the adversarial robustness of object detectors give a false sense of security. Second, we found that adversarially pre-trained backbone networks were essential for enhancing the adversarial robustness of object detectors. We then proposed a simple yet effective recipe for fast adversarial fine-tuning on object detectors with adversarially pre-trained backbones. Without any modifications to the structure of object detectors, our recipe achieved significantly better adversarial robustness than previous works. Finally, we explored the potential of different modern object detector designs for improving adversarial robustness with our recipe and demonstrated interesting findings, which inspired us to design state-of-the-art (SOTA) robust detectors. Our empirical results set a new milestone for adversarially robust object detection. Code and trained checkpoints are available athttps://github.com/thu-ml/oddefense. Xiao Li 0028, Hang Chen 0004, Xiaolin Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Language-Driven Anchors for Zero-Shot Adversarial RobustnessabstractDeep Neural Networks (DNNs) are known to be susceptible to adversarial attacks. Previous researches mainly fo-cus on improving adversarial robustness in the fully super-vised setting, leaving the challenging domain of zero-shot adversarial robustness an open question. In this work, we investigate this domain by leveraging the recent advances in large vision-language models, such as CLIP, to introduce zero-shot adversarial robustness to DNNs. We pro-pose LAAT, a Language-driven, Anchor-based Adversarial Training strategy. LAAT utilizes the features of a text en-coder for each category as fixed anchors (normalized feature embeddings) for each category, which are then employed for adversarial training. By leveraging the semantic consistency of the text encoders, LAAT aims to enhance the adversarial robustness of the image model on novel cate-gories. However, naively using text encoders leads to poor results. Through analysis, we identified the issue to be the high cosine similarity between text encoders. We then design an expansion algorithm and an alignment cross-entropy loss to alleviate the problem. Our experimental results demonstrated that LAAT significantly improves zero-shot adversarial robustness over state-of-the-art methods. LAAT has the potential to enhance adversarial robustness by large-scale multimodal models, especially when labeled data is unavailable during training. Code is available at https://github.com/LixiaoTHU/LAAT. Xiao Li 0028, Wei Zhang 0370, Zhanhao Hu, Bo Zhang 0010, Xiaolin Hu 0001 |
CVPR | 6 |
| 2024 | SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object DetectionabstractLiDAR-based 3D object detection plays an essential role in autonomous driving. Existing high-performing 3D object detectors usually build dense feature maps in the backbone network and prediction head. However, the computational costs introduced by the dense feature maps grow quadratically as the perception range increases, making these models hard to scale up to long-range detection. Some recent works have attempted to construct fully sparse detectors to solve this issue; nevertheless, the resulting models either rely on a complex multi-stage pipeline or exhibit inferior performance. In this work, we propose a fully sparse adaptive feature diffusion network (SAFDNet) for LiDAR-based 3D object detection. In SAFDNet, an adaptive feature diffusion strategy is designed to address the center feature missing problem. We conducted extensive experiments on Waymo Open, nuScenes, and Argoverse2 datasets. SAFDNet performed slightly better than the previous SOTA on the first two datasets but much better on the last dataset, which features long-range detection, verifying the efficacy of SAFDNet in scenarios where long-range detection is required. Notably, on Argoverse2, SAFDNet surpassed the previous best hybrid detector HEDNet by 2.6% mAP while being 2.1 × faster, and yielded 2.1% mAP gains over the previous best sparse detector FSDv2 while being 1.3 × faster. The code will be available at https://github.com/zhanggang001/HEDNet. Junnan Chen, Guohuan Gao, Jianmin Li 0001, Si Liu 0001, Xiaolin Hu 0001 |
CVPR | 6 |
| 2024 | Infrared Adversarial Car StickersabstractInfrared physical adversarial examples are of great significance for studying the security of infrared AI systems that are widely used in our lives such as autonomous driving. Previous infrared physical attacks mainly focused on 2D infrared pedestrian detection which may not fully manifest its destructiveness to AI systems. In this work, we propose a physical attack method against infrared detectors based on 3D modeling, which is applied to a real car. The goal is to design a set of infrared adversarial stickers to make cars invisible to infrared detectors at various viewing angles, distances, and scenes. We build a 3D infrared car model with real infrared characteristics and propose an infrared adversarial pattern generation method based on 3D mesh shadow. We propose a 3D control points-based mesh smoothing algorithm and use a set of smoothness loss functions to enhance the smoothness of adversarial meshes and facilitate the sticker implementation. Besides, We designed the aluminum stickers and conducted physical experiments on two real Mercedes-Benz A200L cars. Our adversarial stickers hid the cars from Faster RCNN, an object detector, at various viewing angles, distances, and scenes. The attack success rate (ASR) was 91.49% for real cars. In comparison, the ASRs of random stickers and no sticker were only 6.21% and 0.66%, respectively. In addition, the ASRs of the designed stickers against six unseen object detectors such as YOLOv3 and Deformable DETR were between 73.35%-95.80%, showing good transferability of the attack performance across detectors. Xiaopei Zhu, Yuqiu Liu, Zhanhao Hu, Jianmin Li 0001, Xiaolin Hu 0001 |
CVPR | 5 |
| 2024 | Controllable Navigation Instruction Generation with Chain of Thought Prompting
Xianghao Kong, Wenguan Wang, Hang Su 0006, Xiaolin Hu 0001, Yi Yang 0001, Si Liu 0001 |
ECCV (29) | 5 |
| 2024 | PartImageNet++ Dataset: Scaling Up Part-Based Models for Robust Recognition
Xiao Li 0028, Sitian Qin, Xiaolin Hu 0001 |
ECCV (71) | 5 |
| 2024 | RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationabstractAudio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA) models operate in the time domain. However, their overly simplistic approach to modeling acoustic features often necessitates larger and more computationally intensive models in order to achieve SOTA performance. In this paper, we present a novel time-frequency domain audio-visual speech separation method: Recurrent Time-Frequency Separation Network (RTFS-Net), which applies its algorithms on the complex time-frequency bins yielded by the Short-Time Fourier Transform. We model and capture the time and frequency dimensions of the audio independently using a multi-layered RNN along each dimension. Furthermore, we introduce a unique attention-based fusion technique for the efficient integration of audio and visual information, and a new mask separation approach that takes advantage of the intrinsic spectral nature of the acoustic features for a clearer separation. RTFS-Net outperforms the prior SOTA method in both inference speed and separation quality while reducing the number of parameters by 90% and MACs by 83%. This is the first time-frequency domain audio-visual speech separation method to outperform all contemporary time-domain counterparts. Samuel Pegg, Kai Li 0047, Xiaolin Hu 0001 |
ICLR | 3 |
| 2024 | IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationabstractRecent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without employing selective attention mechanisms, which is in sharp contrast with the brain. To address this, We propose a novel model called intra- and inter-attention network (IIANet), which leverages the attention mechanism for efficient audio-visual feature fusion. IIANet consists of two types of attention blocks: intra-attention (IntraA) and inter-attention (InterA) blocks, where the InterA blocks are distributed at the top, middle and bottom of IIANet. Heavily inspired by the way how human brain selectively focuses on relevant content at various temporal scales, these blocks maintain the ability to learn modality-specific features and enable the extraction of different semantics from audio-visual features. Comprehensive experiments on three standard audio-visual separation benchmarks (LRS2, LRS3, and VoxCeleb2) demonstrate the effectiveness of IIANet, outperforming previous state-of-the-art methods while maintaining comparable inference time. In particular, the fast version of IIANet (IIANet-fast) has only 7% of CTCNet’s MACs and is 40% faster than CTCNet on CPUs while achieving better separation quality, showing the great potential of attention mechanism for efficient and effective multimodal fusion. Kai Li 0047, Runxuan Yang, Fuchun Sun 0001, Xiaolin Hu 0001 |
ICML | 4 |
| 2024 | CSFuser: A Cascade Siamese Fusion Architecture for RGB-Infrared Object Detection
Zhigang Zeng, Xiaolin Hu 0001 |
ISNN | 4 |
| 2024 | Improve Adversarial Robustness of MNIST Classification via Topological Data Analysis
Xiao Li 0028, Sitian Qin, Xiaolin Hu 0001 |
ISNN | 4 |
| 2024 | nuScenesComplex: A More Rigorous Evaluation Framework for End-to-End Autonomous Driving Planning
Hujie Pan, Xiaolin Hu 0001 |
ISNN | 4 |
| 2024 | PhonHuBERT: A Phoneme Transcription Tool for Song Datasets
Amaury Prat, Runxuan Yang, Xiaolin Hu 0001 |
ISNN | 3 |
| 2024 | Neural Retrievers are Biased Towards LLM-Generated ContentabstractRecently, the emergence of large language models (LLMs) has revolutionized the paradigm of information retrieval (IR) applications, especially in web search, by generating vast amounts of human-like texts on the Internet. As a result, IR systems in the LLM era are facing a new challenge: the indexed documents are now not only written by human beings but also automatically generated by the LLMs. How these LLM-generated documents influence the IR systems is a pressing and still unexplored question. In this work, we conduct a quantitative evaluation of IR models in scenarios where both human-written and LLM-generated texts are involved. Surprisingly, our findings indicate that neural retrieval models tend to rank LLM-generated documents higher. We refer to this category of biases in neural retrievers towards the LLM-generated content as the source bias. Moreover, we discover that this bias is not confined to the first-stage neural retrievers, but extends to the second-stage neural re-rankers. Then, in-depth analyses from the perspective of text compression indicate that LLM-generated texts exhibit more focused semantics with less noise, making it easier for neural retrieval models to semantic match. To mitigate the source bias, we also propose a plug-and-play debiased constraint for the optimization objective, and experimental results show its effectiveness. Finally, we discuss the potential severe concerns stemming from the observed source bias and hope our findings can serve as a critical wake-up call to the IR community and beyond. To facilitate future explorations of IR in the LLM era, the constructed two new benchmarks are available at https://github.com/KID-22/Source-Bias. Sunhao Dai, Yuqi Zhou 0001, Liang Pang 0001, Weihao Liu 0001, Xiaolin Hu 0001, Yong Liu 0018, Xiao Zhang 0034, Gang Wang 0056, Jun Xu 0001 |
KDD | 5 |
| 2024 | Natural Language Induced Adversarial Images
Xiaopei Zhu, Peiyang Xu, Guanning Zeng, Yinpeng Dong, Xiaolin Hu 0001 |
ACM Multimedia | 5 |
| 2024 | CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object DynamicsabstractEnabling humanoid robots to clean rooms has long been a pursued dream within humanoid research communities. However, many tasks require multi-humanoid collaboration, such as carrying large and heavy furniture together. Given the scarcity of motion capture data on multi-humanoid collaboration and the efficiency challenges associated with multi-agent learning, these tasks cannot be straightforwardly addressed using training paradigms designed for single-agent scenarios. In this paper, we introduce **Coo**perative **H**uman-**O**bject **I**nteraction (**CooHOI**), a framework designed to tackle the challenge of multi-humanoid object transportation problem through a two-phase learning paradigm: individual skill learning and subsequent policy transfer. First, a single humanoid character learns to interact with objects through imitation learning from human motion priors. Then, the humanoid learns to collaborate with others by considering the shared dynamics of the manipulated object using centralized training and decentralized execution (CTDE) multi-agent RL algorithms. When one agent interacts with the object, resulting in specific object dynamics changes, the other agents learn to respond appropriately, thereby achieving implicit communication and coordination between teammates. Unlike previous approaches that relied on tracking-based methods for multi-humanoid HOI, CooHOI is inherently efficient, does not depend on motion capture data of multi-humanoid interactions, and can be seamlessly extended to include more participants and a wide range of object types. Jiawei Gao 0004, Ziqin Wang, Zeqi Xiao, Jingbo Wang 0003, Jinkun Cao, Xiaolin Hu 0001, Si Liu 0001, Jifeng Dai, Jiangmiao Pang |
NeurIPS | 7 |
| 2024 | Full-Distance Evasion of Pedestrian Detectors in the Physical WorldabstractMany studies have proposed attack methods to generate adversarial patterns for evading pedestrian detection, alarming the computer vision community about the need for more attention to the robustness of detectors. However, adversarial patterns optimized by these methods commonly have limited performance at medium to long distances in the physical world. To overcome this limitation, we identify two main challenges. First, in existing methods, there is commonly an appearance gap between simulated distant adversarial patterns and their physical world counterparts, leading to incorrect optimization. Second, there exists a conflict between adversarial losses at different distances, which causes difficulties in optimization. To overcome these challenges, we introduce a Full Distance Attack (FDA) method. Our physical world experiments demonstrate the effectiveness of our FDA patterns across various detection models like YOLOv5, Deformable-DETR, and Mask RCNN. Codes available at https://github.com/zhicheng2T0/Full-Distance-Attack.git Zhi Cheng, Zhanhao Hu, Yuqiu Liu, Jianmin Li 0001, Hang Su 0006, Xiaolin Hu 0001 |
NeurIPS | 6 |
| 2024 | Enhancing In-Context Learning Performance with just SVD-Based Weight Pruning: A Theoretical PerspectiveabstractPre-trained large language models (LLMs) based on Transformer have demonstrated striking in-context learning (ICL) abilities. With a few demonstration input-label pairs, they can predict the label for an unseen input without any parameter updates. In this paper, we show an exciting phenomenon that SVD-based weight pruning can enhance ICL performance, and more surprising, pruning weights in deep layers often results in more stable performance improvements than in shallow layers. However, the underlying mechanism of those findings still remains an open question. To reveal those findings, we conduct an in-depth theoretical analysis by presenting the implicit gradient descent (GD) trajectories of ICL and giving the mutual information based generalization bounds of ICL via full implicit GD trajectories. This helps us reasonably explain the surprising experimental findings. Besides, based on all our experimental and theoretical insights, we intuitively propose a simple, model-compression and derivative-free algorithm for downstream tasks in enhancing ICL inference. Experiments on benchmark datasets and open source LLMs display the method effectiveness. Xinhao Yao, Xiaolin Hu 0001, Shenzhi Yang, Yong Liu 0018 |
NeurIPS | 2 |
| 2024 | DHS-DETR: Efficient DETRs with dynamic head switching
Hang Chen 0004, Chufeng Tang, Xiaolin Hu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Hiding from thermal imaging pedestrian detectors in the physical world
Xiaopei Zhu, Xiao Li 0028, Jianmin Li 0001, Zheyao Wang, Xiaolin Hu 0001 |
Neurocomputing | 5 |
| 2024 | An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical CircuitsabstractAudio-visual approaches involving visual inputs have laid the foundation for recent progress in speech separation. However, the optimization of the concurrent usage of auditory and visual inputs is still an active research area. Inspired by the cortico-thalamo-cortical circuit, in which the sensory processing mechanisms of different modalities modulate one another via the non-lemniscal sensory thalamus, we propose a novel cortico-thalamo-cortical neural network (CTCNet) for audio-visual speech separation (AVSS). First, the CTCNet learns hierarchical auditory and visual representations in a bottom-up manner in separate auditory and visual subnetworks, mimicking the functions of the auditory and visual cortical areas. Then, inspired by the large number of connections between cortical regions and the thalamus, the model fuses the auditory and visual information in a thalamic subnetwork through top-down connections. Finally, the model transmits this fused information back to the auditory and visual subnetworks, and the above process is repeated several times. The results of experiments on three speech separation benchmark datasets show that CTCNet remarkably outperforms existing AVSS methods with considerably fewer parameters. These results suggest that mimicking the anatomical connectome of the mammalian brain has great potential for advancing the development of deep neural networks. Kai Li 0047, Fenghua Xie, Hang Chen 0004, Kexin Yuan, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | On the Privacy Effect of Data Enhancement via the Lens of MemorizationabstractMachine learning poses severe privacy concerns as it has been shown that the learned models can reveal sensitive information about their training data. Many works have investigated the effect of widely adopted data augmentation and adversarial training techniques, termed data enhancement in the paper, on the privacy leakage of machine learning models. Such privacy effects are often measured by membership inference attacks (MIAs), which aim to identify whether a particular example belongs to the training set or not. We propose to investigate privacy from a new perspective calledmemorization. Through the lens of memorization, we find that previously deployed MIAs produce misleading results as they are less likely to identify samples with higher privacy risks as members compared to samples with low privacy risks. To solve this problem, we deploy a recent attack that can capture individual samples’ memorization degrees for evaluation. Through extensive experiments, we unveil several findings about the connections between three essential properties of machine learning models, including privacy, generalization gap, and adversarial robustness. We demonstrate that the generalization gap and privacy leakage are less correlated than that of the previous results. Moreover, there is not necessarily a trade-off between adversarial robustness and privacy as stronger adversarial robustness does not make the model more susceptible to privacy attacks. Xiao Li 0028, Qiongxiu Li, Zhanhao Hu, Xiaolin Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Physically Realizable Natural-Looking Clothing Textures Evade Person Detectors via 3D ModelingabstractRecent works have proposed to craft adversarial clothes for evading person detectors, while they are either only effective at limited viewing angles or very conspicuous to humans. We aim to craft adversarial texture for clothes based on 3D modeling, an idea that has been used to craft rigid adversarial objects such as a 3D-printed turtle. Unlike rigid objects, humans and clothes are non-rigid, leading to difficulties in physical realization. In order to craft natural-looking adversarial clothes that can evade person detectors at multiple viewing angles, we propose adversarial camou-flage textures (AdvCaT) that resemble one kind of the typical textures of daily clothes, camouflage textures. We leverage the Voronoi diagram and Gumbel-softmax trick to parameterize the camouflage textures and optimize the parameters via 3D modeling. Moreover, we propose an efficient augmentation pipeline on 3D meshes combining topologically plausible projection (TopoProj) and Thin Plate Spline (TPS) to narrow the gap between digital and real-world objects. We printed the developed 3D texture pieces on fabric materials and tailored them into T-shirts and trousers. Experiments show high attack success rates of these clothes against multiple detectors. Zhanhao Hu, Wenda Chu, Xiaopei Zhu, Bo Zhang 0010, Xiaolin Hu 0001 |
CVPR | 6 |
| 2023 | Visual Recognition by RequestabstractHumans have the ability of recognizing visual semantics in an unlimited granularity, but existing visual recognition algorithms cannot achieve this goal. In this paper, we establish a new paradigm named visual recognition by request (ViRReq11We recommend the readers to pronounce ViRReqas/virik/.) to bridge the gap. The key lies in decomposing visual recognition into atomic tasks named requests and leveraging a knowledge base, a hierarchical and text-based dictionary, to assist task definition. ViRReq allows for (i) learning complicated whole-part hierarchies from highly incomplete annotations and (ii) inserting new concepts with minimal efforts. We also establish a solid baseline by integrating language-driven recognition into recent semantic and instance segmentation methods, and demonstrate its flexible recognition ability on CPP and ADE20K, two datasets with hierarchical whole-part annotations. Chufeng Tang, Lingxi Xie, Xiaopeng Zhang 0008, Xiaolin Hu 0001, Qi Tian 0001 |
CVPR | 4 |
| 2023 | Generalization Bounds for Federated Learning: Fast Rates, Unparticipating Clients and Unbounded Losses
Xiaolin Hu 0001, Yong Liu 0018 |
ICLR | 1 |
| 2023 | An efficient encoder-decoder architecture with top-down attention for speech separation
Kai Li 0047, Runxuan Yang, Xiaolin Hu 0001 |
ICLR | 3 |
| 2023 | NP-SemiSeg: When Neural Processes meet Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation involves assigning pixel-wise labels to unlabeled images at training time. This is useful in a wide range of real-world applications where collecting pixel-wise labels is not feasible in time or cost. Current approaches to semi-supervised semantic segmentation work by predicting pseudo-labels for each pixel from a class-wise probability distribution output by a model. If this predicted probability distribution is incorrect, however, it leads to poor segmentation results which can have knock-on consequences in safety critical systems, like medical images or self-driving cars. It is, therefore, important to understand what a model does not know, which is mainly achieved by uncertainty quantification. Recently, neural processes (NPs) have been explored in semi-supervised image classification, and they have been a computationally efficient and effective method for uncertainty quantification. In this work, we move one step forward by adapting NPs to semi-supervised semantic segmentation, resulting in a new model called NP-SemiSeg. We experimentally evaluated NP-SemiSeg on the public benchmarks PASCAL VOC 2012 and Cityscapes, with different training settings, and the results verify its effectiveness. Daniela Massiceti, Xiaolin Hu 0001, Vladimir Pavlovic 0001, Thomas Lukasiewicz |
ICML | 3 |
| 2023 | Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model
Héctor Martel, Julius Richter, Kai Li 0047, Xiaolin Hu 0001, Timo Gerkmann |
INTERSPEECH | 4 |
| 2023 | HEDNet: A Hierarchical Encoder-Decoder Network for 3D Object Detection in Point Cloudsabstract3D object detection in point clouds is important for autonomous driving systems. A primary challenge in 3D object detection stems from the sparse distribution of points within the 3D scene. Existing high-performance methods typically employ 3D sparse convolutional neural networks with small kernels to extract features. To reduce computational costs, these methods resort to submanifold sparse convolutions, which prevent the information exchange among spatially disconnected features. Some recent approaches have attempted to address this problem by introducing large-kernel convolutions or self-attention mechanisms, but they either achieve limited accuracy improvements or incur excessive computational costs. We propose HEDNet, a hierarchical encoder-decoder network for 3D object detection, which leverages encoder-decoder blocks to capture long-range dependencies among features in the spatial space, particularly for large and distant objects. We conducted extensive experiments on the Waymo Open and nuScenes datasets. HEDNet achieved superior detection accuracy on both datasets than previous state-of-the-art methods with competitive efficiency. The code is available at https://github.com/zhanggang001/HEDNet. Junnan Chen, Guohuan Gao, Jianmin Li 0001, Xiaolin Hu 0001 |
NeurIPS | 5 |
| 2023 | Hiding from infrared detectors in real world with adversarial clothes
Xiaopei Zhu, Zhanhao Hu, Jianmin Li 0001, Xiaolin Hu 0001, Zheyao Wang |
Appl. Intell. | 5 |
| 2023 | Recognizing Object by Components With Human Prior Knowledge Enhances Adversarial Robustness of Deep Neural NetworksabstractAdversarial attacks can easily fool object recognition systems based on deep neural networks (DNNs). Although many defense methods have been proposed in recent years, most of them can still be adaptively evaded. One reason for the weak adversarial robustness may be that DNNs are only supervised by category labels and do not have part-based inductive bias like the recognition process of humans. Inspired by a well-known theory in cognitive psychology - recognition-by-components, we propose a novel object recognition model ROCK (Recognizing Object by Components with human prior Knowledge). It first segments parts of objects from images, then scores part segmentation results with predefined human prior knowledge, and finally outputs prediction based on the scores. The first stage of ROCK corresponds to the process of decomposing objects into parts in human vision. The second stage corresponds to the decision process of the human brain. ROCK shows better robustness than classical recognition models across various attack settings. These results encourage researchers to rethink the rationality of currently widely-used DNN-based object recognition models and explore the potential of part-based models, once important but recently ignored, for improving robustness. Xiao Li 0028, Ziqi Wang 0003, Bo Zhang 0010, Fuchun Sun 0001, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Extracting Semantic Knowledge From GANs With Unsupervised LearningabstractRecently, unsupervised learning has made impressive progress on various tasks. Despite the dominance of discriminative models, increasing attention is drawn to representations learned by generative models and in particular, Generative Adversarial Networks (GANs). Previous works on the interpretation of GANs reveal that GANs encode semantics in feature maps in a linearly separable form. In this work, we further find that GAN's features can be well clustered with the linear separability assumption. We propose a novel clustering algorithm, named KLiSH, which leverages the linear separability to cluster GAN's features. KLiSH succeeds in extracting fine-grained semantics of GANs trained on datasets of various objects, e.g., car, portrait, animals, and so on. With KLiSH, we can sample images from GANs along with their segmentation masks and synthesize paired image-segmentation datasets. Using the synthesized datasets, we enable two downstream applications. First, we train semantic segmentation networks on these datasets and test them on real images, realizing unsupervised semantic segmentation. Second, we train image-to-image translation networks on the synthesized datasets, enabling semantic-conditional image synthesis without human annotations. Jianjin Xu, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Adjacency Constraint for Efficient Hierarchical Reinforcement LearningabstractGoal-conditioned Hierarchical Reinforcement Learning (HRL) is a promising approach for scaling up reinforcement learning (RL) techniques. However, it often suffers from training inefficiency as the action space of the high-level, i.e., the goal space, is large. Searching in a large goal space poses difficulty for both high-level subgoal generation and low-level policy learning. In this article, we show that this problem can be effectively alleviated by restricting the high-level action space from the whole goal space to a k-step adjacent region of the current state using an adjacency constraint. We theoretically prove that in a deterministic Markov Decision Process (MDP), the proposed adjacency constraint preserves the optimal hierarchical policy, while in a stochastic MDP the adjacency constraint induces a bounded state-value suboptimality determined by the MDP's transition structure. We further show that this constraint can be practically implemented by training an adjacency network that can discriminate between adjacent and non-adjacent subgoals. Experimental results on discrete and continuous control tasks including challenging simulated robot locomotion and manipulation tasks show that incorporating the adjacency constraint significantly boosts the performance of state-of-the-art goal-conditioned HRL approaches. Shangqi Guo, Tian Tan 0003, Xiaolin Hu 0001, Feng Chen 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A Fast High-Fidelity Source-Filter Vocoder With Lightweight Neural ModulesabstractThe quality of raw audio waveform generated by a vocoder could affect various audio generative tasks. In recent years, the dominance of source-filter vocoders was greatly challenged by neural vocoders as the latter presents far superior synthesized audio quality. Meanwhile, neural vocoders introduced unprecedented limitations including low runtime efficiency as well as unstable pitch especially in those without explicit periodic excitation input, while these have never been a problem in source-filter vocoders. We present in this paper a novel approach that takes the best from both parties. We start by an in-depth examination of every building block in WORLD – one of the best-performing source-filter vocoders based on plain signal processing algorithms, looking for ones that do not work well, and we replace them with small, lightweight and task-specific neural network models. We also rearranged the vocoding pipeline for a smoother collaboration between building blocks. Our objective and subjective evaluations demonstrate that our methods present competitive synthesized audio quality even when compared against neural vocoders at a much lower computational cost, while keeping spectral envelope acoustic feature, high pitch accuracy as in conventional source-filter vocoders. Runxuan Yang, Yuyang Peng, Xiaolin Hu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Focal Distillation From High-Resolution Data to Low-Resolution Data for 3D Object DetectionabstractLiDAR-based 3D object detection plays an essential role in autonomous driving. Although the detector trained on high-resolution data has much better performance than the same detector trained on low-resolution data, the high-resolution LiDAR cannot be widely used due to its high price. In this work, we propose a new distillation method called Focal Distillation to bridge the gap between high-resolution detector (teacher model) and low-resolution detector (student model). It consists of three essential components: focal classification distillation (FCD), focal regression distillation (FRD) and focal feature distillation (FFD). Taking the low-resolution data as input, the student model can learn discriminative features and produce more accurate results with the assistance of the teacher model trained on high-resolution data. We conducted extensive experiments to validate the effectiveness of Focal Distillation. Evaluated on the KITTI validation set, a typical SECOND model trained with Focal Distillation outperformed its non-distilled counterpart by 3.37%, 7.52%, 11.35% mAP on the category Car, Pedestrian, and Cyclist of moderate level, respectively. Moreover, the remarkable improvements observed on different models and different datasets further demonstrate the generalization ability of our proposed method. Jiawei Shan, Chufeng Tang, Hujie Pan, Qiankun Yu, Guanhao Wu, Xiaolin Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Dense Contrastive Loss for Instance Segmentation
Hang Chen 0004, Chufeng Tang, Xiaolin Hu 0001 |
BMVC | 3 |
| 2022 | Adversarial Texture for Fooling Person Detectors in the Physical WorldabstractNowadays, cameras equipped with AI systems can capture and analyze images to detect people automatically. However, the AI system can make mistakes when receiving deliberately designed patterns in the real world, i.e., physical adversarial examples. Prior works have shown that it is possible to print adversarial patches on clothes to evade DNN-based person detectors. However, these adversarial examples could have catastrophic drops in the attack success rate when the viewing angle (i.e., the camera's angle towards the object) changes. To perform a multi-angle attack, we propose Adversarial Texture (AdvTexture). AdvTexture can cover clothes with arbitrary shapes so that people wearing such clothes can hide from person detectors from different viewing angles. We propose a generative method, named Toroidal-Cropping-based Expandable Generative Attack (TC-EGA), to craft AdvTexture with repetitive structures. We printed several pieces of cloth with AdvTexure and then made T-shirts, skirts, and dresses in the physical world. Experiments showed that these clothes could fool person detectors in the physical world. Zhanhao Hu, Xiaopei Zhu, Fuchun Sun 0001, Bo Zhang 0010, Xiaolin Hu 0001 |
CVPR | 6 |
| 2022 | Infrared Invisible Clothing: Hiding from Infrared Detectors at Multiple Angles in Real WorldabstractThermal infrared imaging is widely used in body temperature measurement, security monitoring, and so on, but its safety research attracted attention only in recent years. We proposed the infrared adversarial clothing, which could fool infrared pedestrian detectors at different angles. We simulated the process from cloth to clothing in the digital world and then designed the adversarial “QR code” pattern. The core of our method is to design a basic pattern that can be expanded periodically, and make the pattern after random cropping and deformation still have an adversarial effect, then we can process the flat cloth with an adversarial pattern into any 3D clothes. The results showed that the optimized “QR code” pattern lowered the Average Precision (AP) of YOLOv3 by 87.7%, while the random “QR code” pattern and blank pattern lowered the AP of YOLOv3 by 57.9% and 30.1%, respectively, in the digital world. We then manufactured an adversarial shirt with a new material: aerogel. Physical-world experiments showed that the adversarial “QR code” pattern clothing lowered the AP of YOLOv3 by 64.6%, while the random “QR code” pattern clothing and fully heat-insulated clothing lowered the AP of YOLOv3 by 28.3% and 22.8%, respectively. We used the model ensemble technique to improve the attack transferability to unseen models. Xiaopei Zhu, Zhanhao Hu, Jianmin Li 0001, Xiaolin Hu 0001 |
CVPR | 5 |
| 2022 | Active Pointly-Supervised Instance Segmentation
Chufeng Tang, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001, Xiaolin Hu 0001 |
ECCV (28) | 6 |
| 2022 | NP-Match: When Neural Processes meet Semi-Supervised LearningabstractSemi-supervised learning (SSL) has been widely explored in recent years, and it is an effective way of leveraging unlabeled data to reduce the reliance on labeled data. In this work, we adjust neural processes (NPs) to the semi-supervised image classification task, resulting in a new method named NP-Match. NP-Match is suited to this task for two reasons. Firstly, NP-Match implicitly compares data points when making predictions, and as a result, the prediction of each unlabeled data point is affected by the labeled data points that are similar to it, which improves the quality of pseudolabels. Secondly, NP-Match is able to estimate uncertainty that can be used as a tool for selecting unlabeled samples with reliable pseudo-labels. Compared with uncertainty-based SSL methods implemented with Monte Carlo (MC) dropout, NP-Match estimates uncertainty with much less computational overhead, which can save time at both the training and the testing phases. We conducted extensive experiments on four public datasets, and NP-Match outperforms state-of-theart (SOTA) results or achieves competitive results on them, which shows the effectiveness of NPMatch and its potential for SSL. Thomas Lukasiewicz, Daniela Massiceti, Xiaolin Hu 0001, Vladimir Pavlovic 0001, Alexandros Neophytou |
ICML | 4 |
| 2022 | On the Use of Deep Mask Estimation Module for Neural Source Separation SystemsabstractMost of the recent neural source separation systems rely on a masking-based pipeline where a set of multiplicative masks are estimated from and applied to a signal representation of the input mixture.The estimation of such masks, in almost all network architectures, is done by a single layer followed by an optional nonlinear activation function.However, recent literatures have investigated the use of a deep mask estimation module and observed performance improvement compared to a shallow mask estimation module.In this paper, we analyze the role of such deeper mask estimation module by connecting it to a recently proposed unsupervised source separation method, and empirically show that the deep mask estimation module is an efficient approximation of the so-called overseparation-grouping paradigm with the conventional shallow mask estimation layers. Kai Li 0047, Xiaolin Hu 0001 |
INTERSPEECH | 2 |
| 2022 | The MSR-Video to Text dataset with clean annotations
Haoran Chen 0011, Jianmin Li 0001, Simone Frintrop, Xiaolin Hu 0001 |
Comput. Vis. Image Underst. | 4 |
| 2022 | Improving Image Segmentation with Boundary Patch Refinement
Xiaolin Hu 0001, Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2022 | Amplification trojan network: Attack deep neural networks by amplifying their inherent weakness
Zhanhao Hu, Jun Zhu 0001, Bo Zhang 0010, Xiaolin Hu 0001 |
Neurocomputing | 4 |
| 2022 | Bridging the Functional and Wiring Properties of V1 Neurons Through Sparse CodingabstractThe functional properties of neurons in the primary visual cortex (V1) are thought to be closely related to the structural properties of this network, but the specific relationships remain unclear. Previous theoretical studies have suggested that sparse coding, an energy-efficient coding method, might underlie the orientation selectivity of V1 neurons. We thus aimed to delineate how the neurons are wired to produce this feature. We constructed a model and endowed it with a simple Hebbian learning rule to encode images of natural scenes. The excitatory neurons fired sparsely in response to images and developed strong orientation selectivity. After learning, the connectivity between excitatory neuron pairs, inhibitory neuron pairs, and excitatory-inhibitory neuron pairs depended on firing pattern and receptive field similarity between the neurons. The receptive fields (RFs) of excitatory neurons and inhibitory neurons were well predicted by the RFs of presynaptic excitatory neurons and inhibitory neurons, respectively. The excitatory neurons formed a small-world network, in which certain local connection patterns were significantly overrepresented. Bidirectionally manipulating the firing rates of inhibitory neurons caused linear transformations of the firing rates of excitatory neurons, and vice versa. These wiring properties and modulatory effects were congruent with a wide variety of data measured in V1, suggesting that the sparse coding principle might underlie both the functional and wiring properties of V1 neurons. Xiaolin Hu 0001, Zhigang Zeng |
Neural Comput. | 1 |
| 2022 | Inferring Mechanisms of Auditory Attentional Modulation with Deep Neural NetworksabstractHumans have an exceptional ability to extract specific audio streams of interest in a noisy environment; this is known as the cocktail party effect. It is widely accepted that this ability is related to selective attention, a mental process that enables individuals to focus on a particular object. Evidence suggests that sensory neurons can be modulated by top-down signals transmitted from the prefrontal cortex. However, exactly how the projection of attention signals to the cortex and subcortex influences the cocktail effect is unclear. We constructed computational models to study whether attentional modulation is more effective at earlier or later stages for solving the cocktail party problem along the auditory pathway. We modeled the auditory pathway using deep neural networks (DNNs), which can generate representational neural patterns that resemble the human brain. We constructed a series of DNN models in which the main structures were autoencoders. We then trained these DNNs on a speech separation task derived from the dichotic listening paradigm, a common paradigm to investigate the cocktail party effect. We next analyzed the modulation effects of attention signals during all stages. Our results showed that the attentional modulation effect is more effective at the lower stages of the DNNs. This suggests that the projection of attention signals to lower stages within the auditory pathway plays a more significant role than the higher stages in solving the cocktail party problem. This prediction could be tested using neurophysiological experiments. Ting-Yu Kuo, Yuanda Liao, Kai Li 0047, Xiaolin Hu 0001 |
Neural Comput. | 5 |
| 2022 | State-Temporal Compression in Reinforcement Learning With the Reward-Restricted Geodesic MetricabstractIt is difficult to solve complex tasks that involve large state spaces and long-term decision processes by reinforcement learning (RL) algorithms. A common and promising method to address this challenge is to compress a large RL problem into a small one. Towards this goal, the compression should be state-temporal and optimality-preserving (i.e., the optimal policy of the compressed problem should correspond to that of the uncompressed problem). In this paper, we propose a reward-restricted geodesic (RRG) metric, which can be learned by a neural network, to perform state-temporal compression in RL. We prove that compression based on the RRG metric is approximately optimality-preserving for the raw RL problem endowed with temporally abstract actions. With this compression, we design an RRG metric-based reinforcement learning (RRG-RL) algorithm to solve complex tasks. Experiments in both discrete (2D Minecraft) and continuous (Doom) environments demonstrated the superiority of our method over existing RL approaches. Shangqi Guo, Qi Yan 0005, Xiaolin Hu 0001, Feng Chen 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Convolutional Neural Networks With Gated Recurrent ConnectionsabstractThe convolutional neural network (CNN) has become a basic model for solving many computer vision problems. In recent years, a new class of CNNs, recurrent convolution neural network (RCNN), inspired by abundant recurrent connections in the visual systems of animals, was proposed. The critical element of RCNN is the recurrent convolutional layer (RCL), which incorporates recurrent connections between neurons in the standard convolutional layer. With increasing number of recurrent computations, the receptive fields (RFs) of neurons in RCL expand unboundedly, which is inconsistent with biological facts. We propose to modulate the RFs of neurons by introducing gates to the recurrent connections. The gates control the amount of context information inputting to the neurons and the neurons' RFs therefore become adaptive. The resulting layer is called gated recurrent convolution layer (GRCL). Multiple GRCLs constitute a deep model called gated RCNN (GRCNN). The GRCNN was evaluated on several computer vision tasks including object recognition, scene text recognition and object detection, and obtained much better results than the RCNN. In addition, when combined with other adaptive RF techniques, the GRCNN demonstrated competitive performance to the state-of-the-art models on benchmark datasets for these tasks. Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Fooling Thermal Infrared Pedestrian Detectors in Real World Using Small BulbsabstractThermal infrared detection systems play an important role in many areas such as night security, autonomous driving, and body temperature detection. They have the unique advantages of passive imaging, temperature sensitivity and penetration. But the security of these systems themselves has not been fully explored, which poses risks in applying these systems. We propose a physical attack method with small bulbs on a board against the state of-the-art pedestrian detectors. Our goal is to make infrared pedestrian detectors unable to detect real-world pedestrians. Towards this goal, we first showed that it is possible to use two kinds of patches to attack the infrared pedestrian detector based on YOLOv3. The average precision (AP) dropped by 64.12% in the digital world, while a blank board with the same size caused the AP to drop by 29.69% only. After that, we designed and manufactured a physical board and successfully attacked YOLOv3 in the real world. In recorded videos, the physical board caused AP of the target detector to drop by 34.48%, while a blank board with the same size caused the AP to drop by 14.91% only. With the ensemble attack techniques, the designed physical board had good transferability to unseen detectors. Xiaopei Zhu, Xiao Li 0028, Jianmin Li 0001, Zheyao Wang, Xiaolin Hu 0001 |
AAAI | 5 |
| 2021 | Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object DetectionabstractLocalization Quality Estimation (LQE) is crucial and popular in the recent advancement of dense object detectors since it can provide accurate ranking scores that benefit the Non-Maximum Suppression processing and improve detection performance. As a common practice, most existing methods predict LQE scores through vanilla convolutional features shared with object classification or bounding box regression. In this paper, we explore a completely novel and different perspective to perform LQE – based on the learned distributions of the four parameters of the bounding box. The bounding box distributions are inspired and introduced as "General Distribution" in GFLV1, which describes the uncertainty of the predicted bounding boxes well. Such a property makes the distribution statistics of a bounding box highly correlated to its real localization quality. Specifically, a bounding box distribution with a sharp peak usually corresponds to high localization quality, and vice versa. By leveraging the close correlation between distribution statistics and the real localization quality, we develop a considerably lightweight Distribution-Guided Quality Predictor (DGQP) for reliable LQE based on GFLV1, thus producing GFLV2. To our best knowledge, it is the first attempt in object detection to use a highly relevant, statistical representation to facilitate LQE. Extensive experiments demonstrate the effectiveness of our method. Notably, GFLV2 (ResNet101) achieves 46.2 AP at 14.6 FPS, surpassing the previous state-of-the-art ATSS baseline (43.6 AP at 14.6 FPS) by absolute 2.6 AP on COCO test-dev, without sacrificing the efficiency both in training and inference. Xiang Li 0041, Wenhai Wang, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003 |
CVPR | 3 |
| 2021 | Look Closer To Segment Better: Boundary Patch Refinement for Instance SegmentationabstractTremendous efforts have been made on instance segmentation but the mask quality is still not satisfactory. The boundaries of predicted instance masks are usually imprecise due to the low spatial resolution of feature maps and the imbalance problem caused by the extremely low proportion of boundary pixels. To address these issues, we propose a conceptually simple yet effective post-processing refinement framework to improve the boundary quality based on the results of any instance segmentation model, termed BPR. Following the idea of looking closer to segment boundaries better, we extract and refine a series of small boundary patches along the predicted instance boundaries. The refinement is accomplished by a boundary patch refinement network at higher resolution. The proposed BPR framework yields significant improvements over the Mask R-CNN baseline on Cityscapes benchmark, especially on the boundary-aware metrics. Moreover, by applying the BPR framework to the "PolyTransform + SegFix" baseline, we reached 1stplace on the Cityscapes leaderboard. Code is available at https://github.com/tinyalpha/BPR. Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
CVPR | 6 |
| 2021 | RSG: A Simple but Effective Module for Learning Imbalanced DatasetsabstractImbalanced datasets widely exist in practice and are a great challenge for training deep neural models with a good generalization on infrequent classes. In this work, we propose a new rare-class sample generator (RSG) to solve this problem. RSG aims to generate some new samples for rare classes during training, and it has in particular the following advantages: (1) it is convenient to use and highly versatile, because it can be easily integrated into any kind of convolutional neural network, and it works well when combined with different loss functions, and (2) it is only used during the training phase, and therefore, no additional burden is imposed on deep neural networks during the testing phase. In extensive experimental evaluations, we verify the effectiveness of RSG. Furthermore, by leveraging RSG, we obtain competitive results on Imbalanced CIFAR and new state-of-the-art results on Places-LT, ImageNet-LT, and iNaturalist 2018. The source code is available at https://github.com/Jianf-Wang/RSG. Thomas Lukasiewicz, Xiaolin Hu 0001, Jianfei Cai 0001, Zhenghua Xu 0001 |
CVPR | 3 |
| 2021 | RefineMask: Towards High-Quality Instance Segmentation With Fine-Grained FeaturesabstractThe two-stage methods for instance segmentation, e.g. Mask R-CNN, have achieved excellent performance recently. However, the segmented masks are still very coarse due to the downsampling operations in both the feature pyramid and the instance-wise pooling process, especially for large objects. In this work, we propose a new method called RefineMask for high-quality instance segmentation of objects and scenes, which incorporates fine-grained features during the instance-wise segmenting process in a multi-stage manner. Through fusing more detailed information stage by stage, RefineMask is able to refine high-quality masks consistently. RefineMask succeeds in segmenting hard cases such as bent parts of objects that are oversmoothed by most previous methods and outputs accurate boundaries. Without bells and whistles, RefineMask yields significant gains of 2.6, 3.4, 3.8 AP over Mask R-CNN on COCO, LVIS, and Cityscapes benchmarks respectively at a small amount of additional computational cost. Furthermore, our single-model result outperforms the winner of the LVIS Challenge 2020 by 1.3 points on the LVIS test-dev set and establishes a new state-of-the-art. Code will be available at https://github.com/zhanggang001/RefineMask. Xin Lu 0002, Jingru Tan, Jianmin Li 0001, Zhaoxiang Zhang 0001, Quanquan Li, Xiaolin Hu 0001 |
CVPR | 7 |
| 2021 | Attack on Practical Speaker Verification System Using Universal Adversarial PerturbationsabstractIn authentication scenarios, applications of practical speaker verification systems usually require a person to read a dynamic authentication text. Previous studies played an audio adversarial example as a digital signal to perform physical attacks, which would be easily rejected by audio replay detection modules. This work shows that by playing our crafted adversarial perturbation as a separate source when the adversary is speaking, the practical speaker verification system will misjudge the adversary as a target speaker. A two-step algorithm is proposed to optimize the universal adversarial perturbation to be text-independent and has little effect on the authentication text recognition. We also estimated room impulse response (RIR) in the algorithm which allowed the perturbation to be effective after being played over the air. In the physical experiment, we achieved targeted attacks with success rate of 100%, while the word error rate (WER) on speech recognition was only increased by 3.55%. And recorded audios could pass replay detection for the live person speaking. Shuning Zhao, Jianmin Li 0001, Xingliang Cheng, Thomas Fang Zheng, Xiaolin Hu 0001 |
ICASSP | 7 |
| 2021 | DAM: Discrepancy Alignment Metric for Face RecognitionabstractThe field of face recognition (FR) has witnessed remarkable progress with the surge of deep learning. The effective loss functions play an important role for FR. In this paper, we observe that a majority of loss functions, including the widespread triplet loss and softmax-based cross-entropy loss, embed inter-class (negative) similarity snand intra-class (positive) similarity spinto similarity pairs and optimize to reduce (sn− sp) in the training process. However, in the verification process, existing metrics directly take the absolute similarity between two features as the confidence of belonging to the same identity, which inevitably causes a gap between the training and verification process. To bridge the gap, we propose a new metric called Discrepancy Alignment Metric (DAM) for verification, which introduces the Local Inter-class Discrepancy (LID) for each face image to normalize the absolute similarity score. To estimate the LID of each face image in the verification process, we propose two types of LID Estimation (LIDE) methods, which are reference-based and learning-based estimation methods, respectively. The proposed DAM is plug-and-play and can be easily applied to the most existing methods. Extensive experiments on multiple popular face recognition benchmark datasets demonstrate the effectiveness of our proposed method. Yudong Wu, Yichao Wu, Chuming Li, Xiaolin Hu 0001, Ding Liang |
ICCV | 5 |
| 2021 | CloudAAE: Learning 6D Object Pose Regression with On-line Data Synthesis on Point CloudsabstractIt is often desired to train 6D pose estimation systems on synthetic data because manual annotation is expensive. However, due to the large domain gap between the synthetic and real images, synthesizing color images is expensive. In contrast, this domain gap is considerably smaller and easier to fill for depth information. In this work, we present a system that regresses 6D object pose from depth information represented by point clouds, and a lightweight data synthesis pipeline that creates synthetic point cloud segments for training. We use an augmented autoencoder (AAE) for learning a latent code that encodes 6D object pose information for pose regression. The data synthesis pipeline only requires texture-less 3D object models and desired viewpoints, and it is cheap in terms of both time and hardware storage. Our data synthesis process is up to three orders of magnitude faster than commonly applied approaches that render RGB image data. We show the effectiveness of our system on the LineMOD, LineMOD Occlusion, and YCB Video datasets. The implementation of our system is available at: https://github.com/GeeeG/CloudAAE. Mikko Lauri, Xiaolin Hu 0001, Jianwei Zhang 0001, Simone Frintrop |
ICRA | 3 |
| 2021 | Robust Logo Detection in E-Commerce Images by Data AugmentationabstractLogo detection is an important task in the intellectual property protection in e-commerce. In the paper, we introduce our solution for the ACM MM2021 Robust Logo Detection Grand Challenge. The competition requires the detection of logos (515 categories) in e-commerce images. This competition is challenged by long-tail distribution, small objects, and different types of noises. To overcome these challenges, we built a highly optimized and robust detector. We first tested many effective techniques for general object detection and then focused on data augmentation. We found that data augmentation was effective in improving the performance and robustness of logo detection. Based on the combination of these techniques, we achieved APs of 64.6% and 61.3% on the clean and noisy datasets respectively, which were improved by 8.1% and 19.5% relative to the official baseline. We ranked 5th among 36489 teams in the competition. Hang Chen 0004, Xiao Li 0028, Zefan Wang, Xiaolin Hu 0001 |
ACM Multimedia | 4 |
| 2021 | Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkabstractRecent advances in the design of neural network architectures, in particular those specialized in modeling sequences, have provided significant improvements in speech separation performance. In this work, we propose to use a bio-inspired architecture called Fully Recurrent Convolutional Neural Network (FRCNN) to solve the separation task. This model contains bottom-up, top-down and lateral connections to fuse information processed at various time-scales represented by stages. In contrast to the traditional approach updating stages in parallel, we propose to first update the stages one by one in the bottom-up direction, then fuse information from adjacent stages simultaneously and finally fuse information from all stages to the bottom stage together. Experiments showed that this asynchronous updating scheme achieved significantly better results with much fewer parameters than the traditional synchronous updating scheme on speech separation. In addition, the proposed model achieved competitive or better results with high efficiency as compared to other state-of-the-art approaches on two benchmark datasets. Xiaolin Hu 0001, Kai Li 0047, Jean-Marie Lemercier, Timo Gerkmann |
NeurIPS | 1 |
| 2021 | Vocabulary-Wide Credit Assignment for Training Image Captioning ModelsabstractReinforcement learning (RL) algorithms have been shown to be efficient in training image captioning models. A critical step in RL algorithms is to assign credits to appropriate actions. There are mainly two classes of credit assignment methods in existing RL methods for image captioning, assigning a single credit for the whole sentence and assigning a credit to every word in the sentence. In this article, we propose a new credit assignment method which is orthogonal to the above two. It assigns every word in vocabulary an appropriate credit at each generation step. It is called vocabulary-wide credit assignment. Based on this we propose a Vocabulary-Critical Sequence Training (VCST). VCST can be incorporated into existing RL methods for training image captioning models to achieve better results. Extensive experiments with many popular models validated the effectiveness of VCST. Jianmin Li 0001, Xiaolin Hu 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Dynamic Network Pruning with Interpretable Layerwise Channel SelectionabstractDynamic network pruning achieves runtime acceleration by dynamically determining the inference paths based on different inputs. However, previous methods directly generate continuous decision values for each weight channel, which cannot reflect a clear and interpretable pruning process. In this paper, we propose to explicitly model the discrete weight channel selections, which encourages more diverse weights utilization, and achieves more sparse runtime inference paths. Meanwhile, with the help of interpretable layerwise channel selections in the dynamic network, we can visualize the network decision paths explicitly for model interpretability. We observe that there are clear differences in the layerwise decisions between normal and adversarial examples. Therefore, we propose a novel adversarial example detection algorithm by discriminating the runtime decision features. Experiments show that our dynamic network achieves higher prediction accuracy under the similar computing budgets on CIFAR10 and ImageNet datasets compared to traditional static pruning methods and other dynamic pruning approaches. The proposed adversarial detection algorithm can significantly improve the state-of-the-art detection rate across multiple attacks, which provides an opportunity to build an interpretable and robust model. Xiaolin Hu 0001, Bo Zhang 0010, Hang Su 0006 |
AAAI | 3 |
| 2020 | Pruning from ScratchabstractNetwork pruning is an important research field aiming at reducing computational costs of neural networks. Conventional approaches follow a fixed paradigm which first trains a large and redundant network, and then determines which units (e.g., channels) are less important and thus can be removed. In this work, we find that pre-training an over-parameterized model is not necessary for obtaining the target pruned structure. In fact, a fully-trained over-parameterized model will reduce the search space for the pruned structure. We empirically show that more diverse pruned structures can be directly pruned from randomly initialized weights, including potential models with better performance. Therefore, we propose a novel network pruning pipeline which allows pruning from scratch with little training overhead. In the experiments for compressing classification models on CIFAR10 and ImageNet datasets, our approach not only greatly reduces the pre-training burden of traditional pruning methods, but also achieves similar or even higher accuracy under the same computation budgets. Our results facilitate the community to rethink the effectiveness of existing techniques used for network pruning. Lingxi Xie, Jun Zhou 0011, Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001 |
AAAI | 7 |
| 2020 | Online Knowledge Distillation via Collaborative LearningabstractThis work presents an efficient yet effective online Knowledge Distillation method via Collaborative Learning, termed KDCL, which is able to consistently improve the generalization ability of deep neural networks (DNNs) that have different learning capacities. Unlike existing two-stage knowledge distillation approaches that pre-train a DNN with large capacity as the ''teacher'' and then transfer the teacher's knowledge to another ''student'' DNN unidirectionally (i.e. one-way), KDCL treats all DNNs as ''students'' and collaboratively trains them in a single stage (knowledge is transferred among arbitrary students during collaborative training), enabling parallel computing, fast computations, and appealing generalization ability. Specifically, we carefully design multiple methods to generate soft target as supervisions by effectively ensembling predictions of students and distorting the input images. Extensive experiments show that KDCL consistently improves all the ''students'' on different datasets, including CIFAR-100 and ImageNet. For example, when trained together by using KDCL, ResNet-50 and MobileNetV2 achieve 78.2% and 74.0% top-1 accuracy on ImageNet, outperforming the original results by 1.4% and 2.0% respectively. We also verify that models pre-trained with KDCL transfer well to object detection and semantic segmentation on MS COCO dataset. For instance, the FPN detector is improved by 0.9% mAP. Qiushan Guo, Xinjiang Wang, Yichao Wu, Ding Liang, Xiaolin Hu 0001, Ping Luo 0002 |
CVPR | 6 |
| 2020 | Rotation Consistent Margin Loss for Efficient Low-Bit Face RecognitionabstractIn this paper, we consider the low-bit quantization problem of face recognition (FR) under the open-set protocol. Different from well explored low-bit quantization on closed-set image classification task, the open-set task is more sensitive to quantization errors (QEs). We redefine the QEs in angular space and disentangle it into class error and individual error. These two parts correspond to inter-class separability and intra-class compactness, respectively. Instead of eliminating the entire QEs, we propose the rotation consistent margin (RCM) loss to minimize the individual error, which is more essential to feature discriminative power. Extensive experiments on popular benchmark datasets such as MegaFace Challenge, Youtube Faces (YTF), Labeled Face in the Wild (LFW) and IJB-C show the superiority of proposed loss in low-bit FR quantization tasks. Yudong Wu, Yichao Wu, Ruihao Gong, Yuanhao Lv, Ding Liang, Xiaolin Hu 0001, Xianglong Liu 0001 |
CVPR | 7 |
| 2020 | Delving Deeper into the Decoder for Video CaptioningabstractVideo captioning is an advanced multi-modal task which aims to describe a video clip using a natural language sentence. The encoder-decoder framework is the most popular paradigm for this task in recent years. However, there exist some problems in the decoder of a video captioning model. We make a thorough investigation into the decoder and adopt three techniques to improve the performance of the model. First of all, a combination of variational dropout and layer normalization is embedded into a recurrent unit to alleviate the problem of overfitting. Secondly, a new online method is proposed to evaluate the performance of a model on a validation set so as to select the best checkpoint for testing. Finally, a new training strategy called professional learning is proposed which uses the strengths of a captioning model and bypasses its weaknesses. It is demonstrated in the experiments on Microsoft Research Video Description Corpus (MSVD) and MSR-Video to Text (MSR-VTT) datasets that our model has achieved the best results evaluated by BLEU, CIDEr, METEOR and ROUGE-L metrics with significant gains of up to 18% on MSVD and 3.5% on MSR-VTT compared with the previous state-of-the-art models. Haoran Chen 0011, Jianmin Li 0001, Xiaolin Hu 0001 |
ECAI | 3 |
| 2020 | Boosting Decision-Based Black-Box Adversarial Attacks with Random Sign Flip
Weilun Chen, Zhaoxiang Zhang 0001, Xiaolin Hu 0001, Baoyuan Wu |
ECCV (15) | 3 |
| 2020 | Dynamic Multi-path Neural NetworkabstractAlthough deeper and larger neural networks have achieved better performance, due to overwhelming burden on computation, they cannot meet the demands of deployment on resource-limited devices. An effective strategy to address this problem is to make use of dynamic inference mechanism, which changes the inference path for different samples at runtime. Existing methods only reduce the depth by skipping an entire specific layer, which may lose important information in this layer. In this paper, we propose a novel method called Dynamic Multipath Neural Network (DMNN), which provides more topology choices in terms of both width and depth on the fly. For better modelling the inference path selection, we further introduce previous state and object category information to guide the training process. Compared to previous dynamic inference techniques, the proposed method is more flexible and easier to incorporate into most modern network architectures. Experimental results on ImageNet and CIFAR-100 demonstrate the superiority of our method on both efficiency and classification accuracy. Yingcheng Su, Yichao Wu, Ding Liang, Xiaolin Hu 0001 |
ICPR | 5 |
| 2020 | 6D Object Pose Regression via Supervised Learning on Point CloudsabstractThis paper addresses the task of estimating the 6 degrees of freedom pose of a known 3D object from depth information represented by a point cloud. Deep features learned by convolutional neural networks from color information have been the dominant features to be used for inferring object poses, while depth information receives much less attention. However, depth information contains rich geometric information of the object shape, which is important for inferring the object pose. We use depth information represented by point clouds as the input to both deep networks and geometry-based pose refinement and use separate networks for rotation and translation regression. We argue that the axis-angle representation is a suitable rotation representation for deep learning, and use a geodesic loss function for rotation regression. Ablation studies show that these design choices outperform alternatives such as the quaternion representation and L2 loss, or regressing translation and rotation with the same network. Our simple yet effective approach clearly outperforms state-of-the-art methods on the YCB-video dataset. Mikko Lauri, Xiaolin Hu 0001, Jianwei Zhang 0001, Simone Frintrop |
ICRA | 4 |
| 2020 | Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionabstractOne-stage detector basically formulates object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is to introduce an \emph{individual} prediction branch to estimate the quality of localization, where the predicted quality facilitates the classification to improve detection performance. This paper delves into the \emph{representations} of the above three fundamental elements: quality estimation, classification and localization. Two problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference, and (2) the inflexible Dirac delta distribution for localization. To address the problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation, and use a vector to represent arbitrary distribution of box locations. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain \emph{continuous} labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFL) that generalizes Focal Loss from its discrete form to the \emph{continuous} version for successful optimization. On COCO {\tt test-dev}, GFL achieves 45.0\% AP using ResNet-101 backbone, surpassing state-of-the-art SAPD (43.5\%) and ATSS (43.6\%) with higher or comparable inference speed. Xiang Li 0041, Wenhai Wang, Shuo Chen 0003, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003 |
NeurIPS | 5 |
| 2020 | Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement LearningabstractGoal-conditioned hierarchical reinforcement learning (HRL) is a promising approach for scaling up reinforcement learning (RL) techniques. However, it often suffers from training inefficiency as the action space of the high-level, i.e., the goal space, is often large. Searching in a large goal space poses difficulties for both high-level subgoal generation and low-level policy learning. In this paper, we show that this problem can be effectively alleviated by restricting the high-level action space from the whole goal space to a k-step adjacent region of the current state using an adjacency constraint. We theoretically prove that the proposed adjacency constraint preserves the optimal hierarchical policy in deterministic MDPs, and show that this constraint can be practically implemented by training an adjacency network that can discriminate between adjacent and non-adjacent subgoals. Experimental results on discrete and continuous control tasks show that incorporating the adjacency constraint improves the performance of state-of-the-art HRL approaches in both deterministic and stochastic environments. Shangqi Guo, Tian Tan 0003, Xiaolin Hu 0001, Feng Chen 0007 |
NeurIPS | 4 |
| 2020 | Companion Guided Soft Margin for Face Recognition
Yingcheng Su, Yichao Wu, Zhenmao Li, Qiushan Guo, Ding Liang, Xiaolin Hu 0001 |
ECML/PKDD (3) | 8 |
| 2020 | PopMNet: Generating structured pop music melodies using neural networks
Xiaolin Hu 0001, Jun Zhu 0001 |
Artif. Intell. | 3 |
| 2020 | A Hierarchical Recurrent Neural Network for Symbolic Melody GenerationabstractIn recent years, neural networks have been used to generate symbolic melodies. However, the long-term structure in the melody has posed great difficulty to design a good model. In this article, we present a hierarchical recurrent neural network (HRNN) for melody generation, which consists of three long-short-term-memory (LSTM) subnetworks working in a coarse-to-fine manner along time. Specifically, the three subnetworks generate bar profiles, beat profiles, and notes, in turn, and the output of the high-level subnetworks are fed into the low-level subnetworks, serving as guidance to generate the finer time-scale melody components in the low-level subnetworks. Two human behavior experiments demonstrate the advantage of this structure over the single-layer LSTM which attempts to learn all hidden structures in melodies. Compared with the recently proposed models MidiNet and MusicVAE, the HRNN produces better melodies evaluated by humans. Changran Hu, Xiaolin Hu 0001, Jun Zhu 0001 |
IEEE Trans. Cybern. | 4 |
| 2020 | Interpret Neural Networks by Extracting Critical SubnetworksabstractIn recent years, deep neural networks have achieved excellent performance in many fields of artificial intelligence. The requirements for the interpretability and robustness of neural networks are also increasing. In this paper, we propose to understand the functional mechanism of neural networks by extracting critical subnetworks. Specifically, we denote the critical subnetworks as a group of important channels across layers such that if they were suppressed to zeros, the final test performance would deteriorate severely. This novel perspective can not only reveal the layerwise semantic behavior within the model but also present more accurate visual explanations appearing in the data through attribution methods. Moreover, we propose two adversarial example detection methods based on the properties of sample-specific and class-specific subnetworks, which provides the possibility for increasing the model robustness. Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Line-CNN: End-to-End Traffic Line Detection With Line Proposal UnitabstractThe task of traffic line detection is a fundamental yet challenging problem. Previous approaches usually conduct traffic line detection via a two-stage way, namely the line segment detection followed by a segment clustering, which is very likely to ignore the global semantic information of an entire line. To address the problem, we propose an end-to-end system called Line-CNN (L-CNN), in which the key component is a novel line proposal unit (LPU). The LPU utilizes line proposals as references to locate accurate traffic curves, which forces the system to learn the global feature representation of the entire traffic lines. We benchmark the proposed L-CNN on two public datasets including MIKKI and TuSimple, and the results suggest that L-CNN outperforms the state-of-the-art methods. In addition, L-CNN can run at approximately 30 f/s on a Titan X GPU, which indicates the practicability and effectiveness of L-CNN for real-time intelligent self-driving systems. Xiang Li 0041, Jun Li 0027, Xiaolin Hu 0001, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Learning Reliable Visual Saliency For Model ExplanationsabstractBy highlighting important features that contribute to model prediction, visual saliency is used as a natural form to interpret the working mechanism of deep neural networks. Numerous methods have been proposed to achieve better saliency results. However, we find that previous visual saliency methods are not reliable enough to provide meaningful interpretation through a simple sanity check: saliency methods are required to explain the output of non-maximum prediction classes, which are usually not ground-truth classes. For example, let the methods interpret an image of “dog” given a wrong class label “fish” as the query. This procedure can test whether these methods reliably interpret model's predictions based on existing features that appear in the data. Our experiments show that previous methods failed to pass the test by generating similar saliency maps or scattered patterns. This false saliency response can be dangerous in certain scenarios, such as medical diagnosis. We find that these failure cases are mainly due to the attribution vanishing and adversarial noise within these methods. In order to learn reliable visual saliency, we propose a simple method that requires the output of the model to be close to the original output while learning an explanatory saliency mask. To enhance the smoothness of the optimized saliency masks, we then propose a simple Hierarchical Attribution Fusion (HAF) technique. In order to fully evaluate the reliability of visual saliency methods, we propose a new task Disturbed Weakly Supervised Object Localization (D-WSOL) to measure whether these methods can correctly attribute the model's output to existing features. Experiments show that previous methods fail to meet this standard, and our approach helps to improve the reliability by suppressing false saliency responses. After observing a significant layout difference in saliency masks between real and adversarial samples. we propose to train a simple CNN on these learned hierarchical attribution masks to distinguish adversarial samples. Experiments show that our method can improve detection performance over other approaches significantly. Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Turbo Learning Framework for Human-Object Interactions Recognition and Human Pose EstimationabstractHuman-object interactions (HOI) recognition and pose estimation are two closely related tasks. Human pose is an essential cue for recognizing actions and localizing the interacted objects. Meanwhile, human action and their interacted objects’ localizations provide guidance for pose estimation. In this paper, we propose a turbo learning framework to perform HOI recognition and pose estimation simultaneously. First, two modules are designed to enforce message passing between the tasks, i.e. pose aware HOI recognition module and HOI guided pose estimation module. Then, these two modules form a closed loop to utilize the complementary information iteratively, which can be trained in an end-to-end manner. The proposed method achieves the state-of-the-art performance on two public benchmarks including Verbs in COCO (V-COCO) and HICO-DET datasets. Wei Feng 0016, Wentao Liu 0002, Chen Qian 0006, Xiaolin Hu 0001 |
AAAI | 6 |
| 2019 | Understanding the Disharmony Between Dropout and Batch Normalization by Variance ShiftabstractThis paper first answers the question ``why do the two most powerful techniques Dropout and Batch Normalization (BN) often lead to a worse performance when they are combined together in many modern neural networks, but cooperate well sometimes as in Wide ResNet (WRN)?'' in both theoretical and empirical aspects. Theoretically, we find that Dropout shifts the variance of a specific neural unit when we transfer the state of that network from training to test. However, BN maintains its statistical variance, which is accumulated from the entire learning procedure, in the test phase. The inconsistency of variances in Dropout and BN (we name this scheme ``variance shift'') causes the unstable numerical behavior in inference that leads to erroneous predictions finally. Meanwhile, the large feature dimension in WRN further reduces the ``variance shift'' to bring benefits to the overall performance. Thorough experiments on representative modern convolutional networks like DenseNet, ResNet, ResNeXt and Wide ResNet confirm our findings. According to the uncovered mechanism, we get better understandings in the combination of these two techniques and summarize guidelines for better practices. Xiang Li 0041, Shuo Chen 0003, Xiaolin Hu 0001, Jian Yang 0003 |
CVPR | 3 |
| 2019 | Selective Kernel NetworksabstractIn standard Convolutional Neural Networks (CNNs), the receptive fields of artificial neurons in each layer are designed to share the same size. It is well-known in the neuroscience community that the receptive field size of visual cortical neurons are modulated by the stimulus, which has been rarely considered in constructing CNNs. We propose a dynamic selection mechanism in CNNs that allows each neuron to adaptively adjust its receptive field size based on multiple scales of input information. A building block called Selective Kernel (SK) unit is designed, in which multiple branches with different kernel sizes are fused using softmax attention that is guided by the information in these branches. Different attentions on these branches yield different sizes of the effective receptive fields of neurons in the fusion layer. Multiple SK units are stacked to a deep network termed Selective Kernel Networks (SKNets). On the ImageNet and CIFAR benchmarks, we empirically show that SKNet outperforms the existing state-of-the-art architectures with lower model complexity. Detailed analyses show that the neurons in SKNet can capture target objects with different scales, which verifies the capability of neurons for adaptively adjusting their receptive field sizes according to the input. The code and models are available at https://github.com/implus/SKNet. Xiang Li 0041, Wenhai Wang, Xiaolin Hu 0001, Jian Yang 0003 |
CVPR | 3 |
| 2019 | Learning Sparse Hidden States in Long Short-Term Memory
Niange Yu, Cornelius Weber, Xiaolin Hu 0001 |
ICANN (2) | 3 |
| 2019 | Knowledge Distillation via Route Constrained OptimizationabstractDistillation-based learning boosts the performance of the miniaturized neural network based on the hypothesis that the representation of a teacher model can be used as structured and relatively weak supervision, and thus would be easily learned by a miniaturized model. However, we find that the representation of a converged heavy model is still a strong constraint for training a small student model, which leads to a higher lower bound of congruence loss. In this work, we consider the knowledge distillation from the perspective of curriculum learning by teacher's routing. Instead of supervising the student model with a converged teacher model, we supervised it with some anchor points selected from the route in parameter space that the teacher model passed by, as we called route constrained optimization (RCO). We experimentally demonstrate this simple operation greatly reduces the lower bound of congruence loss for knowledge distillation, hint and mimicking learning. On close-set classification tasks like CIFAR and ImageNet, RCO improves knowledge distillation by 2.14% and 1.5% respectively. For the sake of evaluating the generalization, we also test RCO on the open-set face recognition task MegaFace. RCO achieves 84.3% accuracy on one-to-million task with only 0.8 M parameters, which push the SOTA by a large margin. Baoyun Peng, Yichao Wu, Yu Liu 0015, Ding Liang, Xiaolin Hu 0001 |
ICCV | 8 |
| 2019 | Improving Pedestrian Attribute Recognition With Weakly-Supervised Multi-Scale Attribute-Specific LocalizationabstractPedestrian attribute recognition has been an emerging research topic in the area of video surveillance. To predict the existence of a particular attribute, it is demanded to localize the regions related to the attribute. However, in this task, the region annotations are not available. How to carve out these attribute-related regions remains challenging. Existing methods applied attribute-agnostic visual attention or heuristic body-part localization mechanisms to enhance the local feature representations, while neglecting to employ attributes to define local feature areas. We propose a flexible Attribute Localization Module (ALM) to adaptively discover the most discriminative regions and learns the regional features for each attribute at multiple levels. Moreover, a feature pyramid architecture is also introduced to enhance the attribute-specific localization at low-levels with high-level semantic guidance. The proposed framework does not require additional region annotations and can be trained end-to-end with multi-level deep supervision. Extensive experiments show that the proposed method achieves state-of-the-art results on three pedestrian attribute datasets, including PETA, RAP, and PA-100K. Chufeng Tang, Lu Sheng, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
ICCV | 4 |
| 2019 | A hierarchical sparse coding model predicts acoustic feature encoding in both auditory midbrain and cortexabstractThe auditory pathway consists of multiple stages, from the cochlear nucleus to the auditory cortex. Neurons acting at different stages have different functions and exhibit different response properties. It is unclear whether these stages share a common encoding mechanism. We trained an unsupervised deep learning model consisting of alternating sparse coding and max pooling layers on cochleogram-filtered human speech. Evaluation of the response properties revealed that computing units in lower layers exhibited spectro-temporal receptive fields (STRFs) similar to those of inferior colliculus neurons measured in physiological experiments, including properties such as sound onset and termination, checkerboard pattern, and spectral motion. Units in upper layers tended to be tuned to phonetic features such as plosivity and nasality, resembling the results of field recording in human auditory cortex. Variation of the sparseness level of the units in each higher layer revealed a positive correlation between the sparseness level and the strength of phonetic feature encoding. The activities of the units in the top layer, but not other layers, correlated with the dynamics of the first two formants (F1, F2) of all phonemes, indicating the encoding of phoneme dynamics in these units. These results suggest that the principles of sparse coding and max pooling may be universal in the human auditory pathway. Qingtian Zhang, Xiaolin Hu 0001, Bo Zhang 0010 |
PLoS Comput. Biol. | 2 |
| 2019 | Hierarchical Bayesian Inference and Learning in Spiking Neural NetworksabstractNumerous experimental data from neuroscience and psychological science suggest that human brain utilizes Bayesian principles to deal the complex environment. Furthermore, hierarchical Bayesian inference has been proposed as an appropriate theoretical framework for modeling cortical processing. However, it remains unknown how such a computation is organized in the network of biologically plausible spiking neurons. In this paper, we propose a hierarchical network of winner-take-all circuits which can carry out hierarchical Bayesian inference and learning through a spike-based variational expectation maximization (EM) algorithm. Particularly, we show how the firing activities of spiking neurons in response to the input stimuli and the spike-timing-dependent plasticity rule can be understood, respectively, as variational E-step and M-step of variational EM. Finally, we demonstrate the utility of this spiking neural network on the MNIST benchmark for unsupervised classification of handwritten digits. Shangqi Guo, Zhaofei Yu, Fei Deng 0001, Xiaolin Hu 0001, Feng Chen 0007 |
IEEE Trans. Cybern. | 4 |
| 2019 | Estimation of the Volume of the Left Ventricle From MRI Images Using Deep Neural NetworksabstractSegmenting human left ventricle (LV) in magnetic resonance imaging images and calculating its volume are important for diagnosing cardiac diseases. The latter task became the topic of the Second Annual Data Science Bowl organized by Kaggle. The dataset consisted of a large number of cases with only systole and diastole volume labels. We designed a system based on neural networks to solve this problem. It began with a detector to detect the regions of interest (ROI) containing LV chambers. Then a deep neural network named hypercolumns fully convolutional network was used to segment LV in ROI. The 2-D segmentation results were integrated across different images to estimate the volume. With ground-truth volume labels, this model was trained end-to-end. To improve the result, an additional dataset with only segmentation labels was used. The model was trained alternately on these two tasks. We also proposed a variance estimation method for the final prediction. Our algorithm ranked the fourth on the test set in this competition. Fangzhou Liao, Xiaolin Hu 0001, Sen Song |
IEEE Trans. Cybern. | 3 |
| 2019 | Topic-Oriented Image Captioning Based on Order-EmbeddingabstractWe present an image captioning framework that generates captions under a given topic. The topic candidates are extracted from the caption corpus. A given image's topics are then selected from these candidates by a CNN-based multi-label classifier. The input to the caption generation model is an image-topic pair, and the output is a caption of the image. For this purpose, a cross-modal embedding method is learned for the images, topics, and captions. In the proposed framework, the topic, caption, and image are organized in a hierarchical structure, which is preserved in the embedding space by using the order-embedding method. The caption embedding is upper bounded by the corresponding image embedding and lower bounded by the topic embedding. The lower bound pushes the images and captions about the same topic closer together in the embedding space. A bidirectional caption-image retrieval task is conducted on the learned embedding space and achieves the state-of-the-art performance on the MS-COCO and Flickr30K datasets, demonstrating the effectiveness of the embedding method. To generate a caption for an image, an embedding vector is sampled from the region bounded by the embeddings of the image and the topic, then a language model decodes it to a sentence as the output. The lower bound set by the topic shrinks the output space of the language model, which may help the model to learn to match images and captions better. Experiments on the image captioning task on the MS-COCO and Flickr30K datasets validate the usefulness of this framework by showing that the different given topics can lead to different captions describing specific aspects of the given image and that the quality of generated captions is higher than the control model without a topic as input. In addition, the proposed method is competitive with many state-of-the-art methods in terms of standard evaluation metrics. Niange Yu, Xiaolin Hu 0001, Binheng Song, Jian Yang 0003, Jianwei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Evaluate the Malignancy of Pulmonary Nodules Using the 3-D Deep Leaky Noisy-OR NetworkabstractAutomatic diagnosing lung cancer from computed tomography scans involves two steps: detect all suspicious lesions (pulmonary nodules) and evaluate the whole-lung/pulmonary malignancy. Currently, there are many studies about the first step, but few about the second step. Since the existence of nodule does not definitely indicate cancer, and the morphology of nodule has a complicated relationship with cancer, the diagnosis of lung cancer demands careful investigations on every suspicious nodule and integration of information of all nodules. We propose a 3-D deep neural network to solve this problem. The model consists of two modules. The first one is a 3-D region proposal network for nodule detection, which outputs all suspicious nodules for a subject. The second one selects the top five nodules based on the detection confidence, evaluates their cancer probabilities, and combines them with a leaky noisy-OR gate to obtain the probability of lung cancer for the subject. The two modules share the same backbone network, a modified U-net. The overfitting caused by the shortage of the training data is alleviated by training the two modules alternately. The proposed model won the first place in the Data Science Bowl 2017 competition. Fangzhou Liao, Zhe Li 0002, Xiaolin Hu 0001, Sen Song |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | UnrealStereo: Controlling Hazardous Factors to Analyze Stereo VisionabstractA reliable stereo algorithm is critical for many robotics applications. But textureless and specular regions can easily cause failure by making feature matching difficult. Understanding whether an algorithm is robust to these hazardous regions is important. Although many stereo benchmarks have been developed to evaluate performance, it is hard to quantify the effect of hazardous regions in real images because the location and severity of these regions are unknown. In this paper, we develop a synthetic image generation tool enabling to control hazardous factors, such as making objects more specular or transparent, to produce hazardous regions at different degrees. The densely controlled sampling strategy in virtual worlds enables to effectively stress test stereo algorithms by varying the types and degrees of the hazard. We generate a large synthetic image dataset with automatically computed hazardous regions and analyze algorithms on these regions. The observations from synthetic images are further validated by annotating hazardous regions in real-world datasets Middlebury and KITTI (which gives a sparse sampling of the hazards). Our synthetic image generation tool is based on a game engine Unreal Engine 4 and will be open-source along with the virtual scenes in our experiments. Many publicly available realistic game contents can be used by our tool to provide an enormous resource for development and evaluation of algorithms. Yi Zhang 0099, Weichao Qiu, Qi Chen 0014, Xiaolin Hu 0001, Alan L. Yuille |
3DV | 4 |
| 2018 | A Cascaded Inception of Inception Network With Attention Modulated Feature Fusion for Human Pose EstimationabstractAccurate keypoint localization of human pose needs diversified features: the high level for contextual dependencies and the low level for detailed refinement of joints. However, the importance of the two factors varies from case to case, but how to efficiently use the features is still an open problem. Existing methods have limitations in preserving low level features, adaptively adjusting the importance of different levels of features, and modeling the human perception process. This paper presents three novel techniques step by step to efficiently utilize different levels of features for human pose estimation. Firstly, an inception of inception (IOI) block is designed to emphasize the low level features. Secondly, an attention mechanism is proposed to adjust the importance of individual levels according to the context. Thirdly, a cascaded network is proposed to sequentially localize the joints to enforce message passing from joints of stand-alone parts like head and torso to remote joints like wrist or ankle. Experimental results demonstrate that the proposed method achieves the state-of-the-art performance on both MPII and LSP benchmarks. Wentao Liu 0002, Cheng Li 0009, Chen Qian 0006, Xiao Chu, Xiaolin Hu 0001 |
AAAI | 6 |
| 2018 | Boosting Adversarial Attacks With MomentumabstractDeep neural networks are vulnerable to adversarial examples, which poses security concerns on these algorithms due to the potentially severe consequences. Adversarial attacks serve as an important surrogate to evaluate the robustness of deep learning models before they are deployed. However, most of existing adversarial attacks can only fool a black-box model with a low success rate. To address this issue, we propose a broad class of momentum-based iterative algorithms to boost adversarial attacks. By integrating the momentum term into the iterative process for attacks, our methods can stabilize update directions and escape from poor local maxima during the iterations, resulting in more transferable adversarial examples. To further improve the success rates for black-box attacks, we apply momentum iterative algorithms to an ensemble of models, and show that the adversarially trained models with a strong defense ability are also vulnerable to our black-box attacks. We hope that the proposed methods will serve as a benchmark for evaluating the robustness of various deep models and defense methods. With this method, we won the first places in NIPS 2017 Non-targeted Adversarial Attack and Targeted Adversarial Attack competitions. Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su 0006, Jun Zhu 0001, Xiaolin Hu 0001 |
CVPR | 6 |
| 2018 | Defense Against Adversarial Attacks Using High-Level Representation Guided DenoiserabstractNeural networks are vulnerable to adversarial examples, which poses a threat to their application in security sensitive systems. We propose high-level representation guided denoiser (HGD) as a defense for image classification. Standard denoiser suffers from the error amplification effect, in which small residual adversarial noise is progressively amplified and leads to wrong classifications. HGD overcomes this problem by using a loss function defined as the difference between the target model's outputs activated by the clean image and denoised image. Compared with ensemble adversarial training which is the state-of-the-art defending method on large images, HGD has three advantages. First, with HGD as a defense, the target model is more robust to either white-box or black-box adversarial attacks. Second, HGD can be trained on a small subset of the images and generalizes well to other images and unseen classes. Third, HGD can be transferred to defend models other than the one guiding it. In NIPS competition on defense against adversarial attacks, our HGD solution won the first place and outperformed other models by a large margin. Fangzhou Liao, Yinpeng Dong, Tianyu Pang, Xiaolin Hu 0001, Jun Zhu 0001 |
CVPR | 5 |
| 2018 | Interpret Neural Networks by Identifying Critical Data Routing PathsabstractInterpretability of a deep neural network aims to explain the rationale behind its decisions and enable the users to understand the intelligent agents, which has become an important issue due to its importance in practical applications. To address this issue, we develop a Distillation Guided Routing method, which is a flexible framework to interpret a deep neural network by identifying critical data routing paths and analyzing the functional processing behavior of the corresponding layers. Specifically, we propose to discover the critical nodes on the data routing paths during network inferring prediction for individual input samples by learning associated control gates for each layer's output channel. The routing paths can, therefore, be represented based on the responses of concatenated control gates from all the layers, which reflect the network's semantic selectivity regarding to the input patterns and more detailed functional process across different layer levels. Based on the discoveries, we propose an adversarial sample detection algorithm by learning a classifier to discriminate whether the critical data routing paths are from real or adversarial samples. Experiments demonstrate that our algorithm can effectively achieve high defense rate with minor training overhead. Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001 |
CVPR | 4 |
| 2017 | Scale-Aware Face DetectionabstractConvolutional neural network (CNN) based face detectors are inefficient in handling faces of diverse scales. They rely on either fitting a large single model to faces across a large scale range or multi-scale testing. Both are computationally expensive. We propose Scale-aware Face Detection (SAFD) to handle scale explicitly using CNN, and achieve better performance with less computation cost. Prior to detection, an efficient CNN predicts the scale distribution histogram of the faces. Then the scale histogram guides the zoom-in and zoom-out of the image. Since the faces will be approximately in uniform scale after zoom, they can be detected accurately even with much smaller CNN. Actually, more than 99% of the faces in AFW can be covered with less than two zooms per image. Extensive experiments on FDDB, MALF and AFW show advantages of SAFD. Zekun Hao, Yu Liu 0015, Hongwei Qin, Xiu Li 0001, Xiaolin Hu 0001 |
CVPR | 6 |
| 2017 | Training the Hopfield Neural Network for Classification Using a STDP-Like Rule
Xiaolin Hu 0001, Tao Wang 0003 |
ICONIP (3) | 1 |
| 2017 | An STDP-Based Supervised Learning Algorithm for Spiking Neural Networks
Zhanhao Hu, Tao Wang 0003, Xiaolin Hu 0001 |
ICONIP (2) | 3 |
| 2017 | Delving deeper into convolutional neural networks for camera relocalizationabstractConvolutional Neural Networks (CNNs) have been applied to camera relocalization, which is to infer the pose of the camera given a single monocular image. However, there are still many open problems for camera relocalization with CNNs. We delve into the CNNs for camera relocalization. First, a variant of Euler angles named Euler6 is proposed to represent orientation. Then a data augmentation method named pose synthesis is designed to reduce spsarsity of poses in the whole pose space to cope with overfitting in training. Third, a multi-task CNN named BranchNet is proposed to deal with the complex coupling of orientation and translation. The network consists of several shared convolutional layers and splits into two branches which predict orientation and translation, respectively. Experiments on the 7Scenes dataset show that incorporating these techniques one by one into an existing model PoseNet always leads to better results. Together these techniques reduce the orientation error by 15.9% and the translation error by 38.3% compared to the state-of-the-art model Bayesian PoseNet. We implement BranchNet on an Intel NUC mobile platform and reach a speed of 43 fps, which meets the real-time requirement of many robotic applications. Liwei Ma, Xiaolin Hu 0001 |
ICRA | 3 |
| 2017 | Accelerating convolutional neural networks by group-wise 2D-filter pruningabstractNetwork pruning is an effective way to accelerate Convolutional Neural Networks (CNNs). In recent years, structured pruning methods are proposed in favor of unstructured methods as they have shown greater speedup in practical use. Existing structured methods does pruning along two main dimensions: 3D-filter wise, i.e., remove a 3D-fllter as a whole, and filter-shape wise, i.e., remove a same position from all 3D-filters. In this work, we propose a new group-wise 2D-fllter pruning approach that is orthogonal and complementary to the existing methods. The proposed approach removes a portion of 2D-fllters from each 3D-filter according to the pruning patterns learned from the data, and leads to compressed models that do not require sophisticated implementation of convolution operations. A fine-tuning process is followed to recover the accuracy. The knowledge distillation (KD) framework is explored in the fine-tuning process to improve the performance. We present our method for learning the pruning pattens as well as the fine-tuning strategy based on knowledge distillation. The proposed approach is validated on two representative CNN models - ZF and VGG16, pre-trained on ILSVRC12. Experimental results demonstrate the effectiveness of our approach. In VGG16, we get even higher accuracy after speeding-up the network by 4 times. Niange Yu, Xiaolin Hu 0001, Jianmin Li 0001 |
IJCNN | 3 |
| 2017 | Reservoir Computing with a Small-World Network for Discriminating Two Sequential Stimuli
Fangzhou Liao, Xiaolin Hu 0001 |
ISNN (1) | 3 |
| 2017 | Gated Recurrent Convolution Neural Network for OCRabstractOptical Character Recognition (OCR) aims to recognize text in natural images. Inspired by a recently proposed model for general image classification, Recurrent Convolution Neural Network (RCNN), we propose a new architecture named Gated RCNN (GRCNN) for solving this problem. Its critical component, Gated Recurrent Convolution Layer (GRCL), is constructed by adding a gate to the Recurrent Convolution Layer (RCL), the critical component of RCNN. The gate controls the context modulation in RCL and balances the feed-forward information and the recurrent information. In addition, an efficient Bidirectional Long Short-Term Memory (BLSTM) is built for sequence modeling. The GRCNN is combined with BLSTM to recognize text in natural images. The entire GRCNN-BLSTM model can be trained end-to-end. Experiments show that the proposed model outperforms existing methods on several benchmark datasets including the IIIT-5K, Street View Text (SVT) and ICDAR. Xiaolin Hu 0001 |
NIPS | 2 |
| 2017 | Convolution Neural Networks With Two Pathways for Image Style RecognitionabstractAutomatic recognition of an image's style is important for many applications, including artwork analysis, photo organization, and image retrieval. Traditional convolution neural network (CNN) approach uses only object features for image style recognition. This approach may not be optimal, because the same object in two images may have different styles. We propose a CNN architecture with two pathways extracting object features and texture features, respectively. The object pathway represents the standard CNN architecture and the texture pathway intermixes the object pathway by outputting the gram matrices of intermediate features in the object pathway. The two pathways are jointly trained. In experiments, two deep CNNs, AlexNet and VGG-19, pretrained on the ImageNet classification data set are fine-tuned for this task. For any model, the two-pathway architecture performs much better than individual pathways, which indicates that the two pathways contain complementary information of an image's style. In particular, the model based on VGG-19 achieves the state-of-the-art results on three benchmark data sets, WikiPaintings, Flickr Style, and AVA Style. Tiancheng Sun, Jian Yang 0003, Xiaolin Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Joint Training of Cascaded CNN for Face DetectionabstractCascade has been widely used in face detection, where classifier with low computation cost can be firstly used to shrink most of the background while keeping the recall. The cascade in detection is popularized by seminal Viola-Jones framework and then widely used in other pipelines, such as DPM and CNN. However, to our best knowledge, most of the previous detection methods use cascade in a greedy manner, where previous stages in cascade are fixed when training a new stage. So optimizations of different CNNs are isolated. In this paper, we propose joint training to achieve end-to-end optimization for CNN cascade. We show that the back propagation algorithm used in training CNN can be naturally used in training CNN cascade. We present how jointly training can be conducted on naive CNN cascade and more sophisticated region proposal network (RPN) and fast R-CNN. Experiments on face detection benchmarks verify the advantages of the joint training. Hongwei Qin, Xiu Li 0001, Xiaolin Hu 0001 |
CVPR | 4 |
| 2016 | Fuzzy String Matching Using Sentence Embedding Algorithms
Yu Rong 0003, Xiaolin Hu 0001 |
ICONIP (3) | 2 |
| 2016 | Guest Editorial Special Issue on Neurodynamic Systems for Optimization and ApplicationsabstractRecurrent neural networks, as neurodynamic systems, are a class of connectionist models that capture the dynamics of sequences via cycles in artificial neurons. Since the invention of Hopfield neural network, recurrent neural networks have attracted considerable attention, which marks the beginning of the modern age of neural network studies. Thanks to their inherent nature of parallel and distributed information processing, many computationally intensive applications can be solved by recurrent neural networks in the real-time environment. Zhigang Zeng, Andrzej Cichocki, Long Cheng 0001, Youshen Xia, Xiaolin Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2015 | Interlinked Convolutional Neural Networks for Face ParsingabstractFace parsing is a basic task in face image analysis. It amounts to labeling each pixel with appropriate facial parts such as eyes and nose. In the paper, we present a interlinked convolutional neural network (iCNN) for solving this problem in an end-to-end fashion. It consists of multiple convolutional neural networks (CNNs) taking input in different scales. A special interlinking layer is designed to allow the CNNs to exchange information, enabling them to integrate local and contextual information efficiently. The hallmark of iCNN is the extensive use of downsampling and upsampling in the interlinking layers, while traditional CNNs usually uses downsampling only. A two-stage pipeline is proposed for face parsing and both stages use iCNN. The first stage localizes facial parts in the size-reduced image and the second stage labels the pixels in the identified facial parts in the original image. On a benchmark dataset we have obtained better results than the state-of-the-art methods. Yisu Zhou, Xiaolin Hu 0001, Bo Zhang 0010 |
ISNN | 2 |
| 2015 | Convolutional Neural Networks with Intra-Layer Recurrent Connections for Scene LabelingabstractScene labeling is a challenging computer vision task. It requires the use of both local discriminative features and global context information. We adopt a deep recurrent convolutional neural network (RCNN) for this task, which is originally proposed for object recognition. Different from traditional convolutional neural networks (CNN), this model has intra-layer recurrent connections in the convolutional layers. Therefore each convolutional layer becomes a two-dimensional recurrent neural network. The units receive constant feed-forward inputs from the previous layer and recurrent inputs from their neighborhoods. While recurrent iterations proceed, the region of context captured by each unit expands. In this way, feature extraction and context modulation are seamlessly integrated, which is different from typical methods that entail separate modules for the two steps. To further utilize the context, a multi-scale RCNN is proposed. Over two benchmark datasets, Standford Background and Sift Flow, the model outperforms many state-of-the-art models in accuracy and efficiency. Xiaolin Hu 0001, Bo Zhang 0010 |
NIPS | 2 |
| 2015 | Comparison of ℓ1-Norm SVR and Sparse Coding Algorithms for Linear RegressionabstractSupport vector regression (SVR) is a popular function estimation technique based on Vapnik's concept of support vector machine. Among many variants, the l1-norm SVR is known to be good at selecting useful features when the features are redundant. Sparse coding (SC) is a technique widely used in many areas and a number of efficient algorithms are available. Both l1-norm SVR and SC can be used for linear regression. In this brief, the close connection between the l1-norm SVR and SC is revealed and some typical algorithms are compared for linear regression. The results show that the SC algorithms outperform the Newton linear programming algorithm, an efficient l1-norm SVR algorithm, in efficiency. The algorithms are then used to design the radial basis function (RBF) neural networks. Experiments on some benchmark data sets demonstrate the high efficiency of the SC algorithms. In particular, one of the SC algorithms, the orthogonal matching pursuit is two orders of magnitude faster than a well-known RBF network designing algorithm, the orthogonal least squares algorithm. Qingtian Zhang, Xiaolin Hu 0001, Bo Zhang 0010 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Modeling response properties of V2 neurons using a hierarchical K-means model
Xiaolin Hu 0001, Jianwei Zhang 0001, Peng Qi 0004, Bo Zhang 0010 |
Neurocomputing | 1 |
| 2014 | Learning Nonlinear Statistical Regularities in Natural Images by Modeling the Outer Product of Image IntensitiesabstractIt is well known that there exist nonlinear statistical regularities in natural images. Existing approaches for capturing such regularities always model the image intensities by assuming a parameterized distribution for the intensities and learn the parameters. In the letter, we propose to model the outer product of image intensities by assuming a gaussian distribution for it. A two-layer structure is presented, where the first layer is nonlinear and the second layer is linear. Trained on natural images, the first-layer bases resemble the receptive fields of simple cells in the primary visual cortex (V1), while the second-layer units exhibit some properties of the complex cells in V1, including phase invariance and masking effect. The model can be seen as an approximation of the covariance model proposed in Karklin and Lewicki (2009) but has more robust and efficient learning algorithms. Peng Qi 0004, Xiaolin Hu 0001 |
Neural Comput. | 2 |
| 2013 | Traffic sign detection by ROI extraction and histogram features-based recognitionabstractWe present a traffic sign detection model consisting of two modules. The first module is for ROI (region of interest) extraction. By supervised learning, it transforms the color images to gray images such that the characteristic colors for the traffic signs are more distinguishable in the gray images. It follows shape template matching, where a set of templates for each target category of signs are designed. After that, a set of ROIs are generated. The second module is for recognition. It validates if an ROI belongs to a target category of traffic signs by supervised learning. Local shape and color features are extracted. The supervised learning methods used in the model are SVMs. The overall model is applied on the GTSDB benchmark and achieves 100%, 98.85% and 92.00% AUC (area under the precision-recall curve) for Prohibitory, Danger and Mandatory signs, respectively. The testing speed is 0.4-1.0 second per image on a mainstream PC, which demonstrates the great potential of the proposed model in real-time applications. Mingyi Yuan, Xiaolin Hu 0001, Jianmin Li 0001, Huaping Liu 0001 |
IJCNN | 3 |
| 2013 | Traffic sign detection based on convolutional neural networksabstractWe propose an approach for traffic sign detection based on Convolutional Neural Networks (CNN). We first transform the original image into the gray scale image by using support vector machines, then use convolutional neural networks with fixed and learnable layers for detection and recognition. The fixed layer can reduce the amount of interest areas to detect, and crop the boundaries very close to the borders of traffic signs. The learnable layers can increase the accuracy of detection significantly. Besides, we use bootstrap methods to improve the accuracy and avoid overfitting problem. In the German Traffic Sign Detection Benchmark, we obtained competitive results, with an area under the precision-recall curve(AUC) of 99.73% in the category “Danger”, and an AUC of 97.62% in the category “Mandatory”. Yihui Wu, Jianmin Li 0001, Huaping Liu 0001, Xiaolin Hu 0001 |
IJCNN | 5 |
| 2013 | Data-based control, optimization, modeling and applications
Dongbin Zhao, Yi Shen 0002, Xiaolin Hu 0001 |
Neural Comput. Appl. | 4 |
| 2012 | Hierarchical K-Means Algorithm for Modeling Visual Area V2 Neurons
Xiaolin Hu 0001, Peng Qi 0004, Bo Zhang 0010 |
ICONIP (3) | 1 |
| 2012 | Solving the Assignment Problem Using Continuous-Time and Discrete-Time Improved Dual NetworksabstractThe assignment problem is an archetypal combinatorial optimization problem. In this brief, we present a continuous-time version and a discrete-time version of the improved dual neural network (IDNN) for solving the assignment problem. Compared with most assignment networks in the literature, the two versions of IDNNs are advantageous in circuit implementation due to their simple structures. Both of them are theoretically guaranteed to be globally convergent to a solution of the assignment problem if only the solution is unique. Xiaolin Hu 0001, Jun Wang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | Solving the Assignment Problem with the Improved Dual Neural Network
Xiaolin Hu 0001, Jun Wang 0002 |
ISNN (1) | 1 |
| 2010 | A Gaussian Attractor Network for Memory and Recognition with Experience-Dependent LearningabstractAttractor networks are widely believed to underlie the memory systems of animals across different species. Existing models have succeeded in qualitatively modeling properties of attractor dynamics, but their computational abilities often suffer from poor representations for realistic complex patterns, spurious attractors, low storage capacity, and difficulty in identifying attractive fields of attractors. We propose a simple two-layer architecture, gaussian attractor network, which has no spurious attractors if patterns to be stored are uncorrelated and can store as many patterns as the number of neurons in the output layer. Meanwhile the attractive fields can be precisely quantified and manipulated. Equipped with experience-dependent unsupervised learning strategies, the network can exhibit both discrete and continuous attractor dynamics. A testable prediction based on numerical simulations is that there exist neurons in the brain that can discriminate two similar stimuli at first but cannot after extensive exposure to physically intermediate stimuli. Inspired by this network, we found that adding some local feedbacks to a well-known hierarchical visual recognition model, HMAX, can enable the model to reproduce some recent experimental results related to high-level visual perception. Xiaolin Hu 0001, Bo Zhang 0010 |
Neural Comput. | 1 |
| 2010 | Design of recurrent neural networks for solving constrained least absolute deviation problemsabstractRecurrent neural networks for solving constrained least absolute deviation (LAD) problems or L(1)-norm optimization problems have attracted much interest in recent years. But so far most neural networks can only deal with some special linear constraints efficiently. In this paper, two neural networks are proposed for solving LAD problems with various linear constraints including equality, two-sided inequality and bound constraints. When tailored to solve some special cases of LAD problems in which not all types of constraints are present, the two networks can yield simpler architectures than most existing ones in the literature. In particular, for solving problems with both equality and one-sided inequality constraints, another network is invented. All of the networks proposed in this paper are rigorously shown to be capable of solving the corresponding problems. The different networks designed for solving the same types of problems possess the same structural complexity, which is due to the fact these architectures share the same computing blocks and only differ in connections between some blocks. By this means, some flexibility for circuits realization is provided. Numerical simulations are carried out to illustrate the theoretical results and compare the convergence rates of the networks. Xiaolin Hu 0001, Changyin Sun 0001, Bo Zhang 0010 |
IEEE Trans. Neural Networks | 1 |
| 2009 | Another Simple Recurrent Neural Network for Quadratic and Linear Programming
Xiaolin Hu 0001, Bo Zhang 0010 |
ISNN (3) | 1 |
| 2009 | Motion Planning with Obstacle Avoidance for Kinematically Redundant Manipulators Based on Two Recurrent Neural NetworksabstractInverse kinematic motion planning of redundant manipulators by using recurrent neural networks in the presence of obstacles and uncertainties is a real-time nonlinear optimization problem. To tackle this problem, two subproblems should be resolved in real time. One is the determination of critical points on a given manipulator closest to obstacles, and the other is the computation of joint velocities of the manipulator which can direct the manipulator following a desired trajectory and away from obstacles if it is getting close to them. Different from our previous approaches where the critical points on the manipulator were assumed to be known, these points are to be computed by using a recurrent neural network in the paper. A time-varying quadratic programming problem is formulated for avoiding polyhedral obstacles. In view that the problem is not strictly convex, an existing recurrent neural network, general projection neural network, is applied for solving it. By introducing a velocity smoothing technique into our previous quadratic programming formulation of the joint velocity assignment problem, a recently developed recurrent neural network, improved dual neural network, is proposed to solve it, which features lower structural complexity compared with existing neural networks. Moreover, The effectiveness of the proposed neural networks is demonstrated by simulations on the Mitsubishi PA10-7C manipulator. Xiaolin Hu 0001, Jun Wang 0002, Bo Zhang 0010 |
SMC | 1 |
| 2009 | A New Recurrent Neural Network for Solving Convex Quadratic Programming Problems With an Application to the k -Winners-Take-All ProblemabstractIn this paper, a new recurrent neural network is proposed for solving convex quadratic programming (QP) problems. Compared with existing neural networks, the proposed one features global convergence property under weak conditions, low structural complexity, and no calculation of matrix inverse. It serves as a competitive alternative in the neural network family for solving linear or quadratic programming problems. In addition, it is found that by some variable substitution, the proposed network turns out to be an existing model for solving minimax problems. In this sense, it can be also viewed as a special case of the minimax neural network. Based on this scheme, a k-winners-take-all ( k-WTA) network with O(n) complexity is designed, which is characterized by simple structure, global convergence, and capability to deal with some ill cases. Numerical simulations are provided to validate the theoretical results obtained. More importantly, the network design method proposed in this paper has great potential to inspire other competitive inventions along the same line. Xiaolin Hu 0001, Bo Zhang 0010 |
IEEE Trans. Neural Networks | 1 |
| 2009 | An Alternative Recurrent Neural Network for Solving Variational Inequalities and Related Optimization ProblemsabstractThere exist many recurrent neural networks for solving optimization-related problems. In this paper, we present a method for deriving such networks from existing ones by changing connections between computing blocks. Although the dynamic systems may become much different, some distinguished properties may be retained. One example is discussed to solve variational inequalities and related optimization problems with mixed linear and nonlinear constraints. A new network is obtained from two classical models by this means, and its performance is comparable to its predecessors. Thus, an alternative choice for circuits implementation is offered to accomplish such computing tasks. Xiaolin Hu 0001, Bo Zhang 0010 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2008 | Three Global Exponential Convergence Results of the GPNN for Solving Generalized Linear Variational Inequalities
Xiaolin Hu 0001, Zhigang Zeng, Bo Zhang 0010 |
ISNN (1) | 1 |
| 2008 | An Improved Dual Neural Network for Solving a Class of Quadratic Programming Problems and Its k-Winners-Take-All ApplicationabstractThis paper presents a novel recurrent neural network for solving a class of convex quadratic programming (QP) problems, in which the quadratic term in the objective function is the square of the Euclidean norm of the variable. This special structure leads to a set of simple optimality conditions for the problem, based on which the neural network model is formulated. Compared with existing neural networks for general convex QP, the new model is simpler in structure and easier to implement. The new model can be regarded as an improved version of the dual neural network in the literature. Based on the new model, a simple neural network capable of solving the k-winners-take-all ( k-WTA) problem is formulated. The stability and global convergence of the proposed neural network is proved rigorously and substantiated by simulation results. Xiaolin Hu 0001, Jun Wang 0002 |
IEEE Trans. Neural Networks | 1 |
| 2007 | Solving the k-Winners-Take-All Problem and the Oligopoly Cournot-Nash Equilibrium Problem Using the General Projection Neural Networks
Xiaolin Hu 0001, Jun Wang 0002 |
ICONIP (1) | 1 |
| 2007 | Convergence of a Recurrent Neural Network for Nonconvex Optimization Based on an Augmented Lagrangian Function
Xiaolin Hu 0001, Jun Wang 0002 |
ISNN (3) | 1 |
| 2007 | Solving Generally Constrained Generalized Linear Variational Inequalities Using the General Projection Neural NetworksabstractGeneralized linear variational inequality (GLVI) is an extension of the canonical linear variational inequality. In recent years, a recurrent neural network (NN) called general projection neural network (GPNN) was developed for solving GLVIs with simple bound (often box-type or sphere-type) constraints. The aim of this paper is twofold. First, some further stability results of the GPNN are presented. Second, the GPNN is extended for solving GLVIs with general linear equality and inequality constraints. A new design methodology for the GPNN is then proposed. Furthermore, in view of different types of constraints, approaches for reducing the number of neurons of the GPNN are discussed, which results in two specific GPNNs. Moreover, some distinct properties of the resulting GPNNs are also explored based on their particular structures. Numerical simulation results are provided to validate the results. Xiaolin Hu 0001, Jun Wang 0002 |
IEEE Trans. Neural Networks | 1 |
| 2007 | A Recurrent Neural Network for Solving a Class of General Variational InequalitiesabstractThis paper presents a recurrent neural-network model for solving a special class of general variational inequalities (GVIs), which includes classical VIs as special cases. It is proved that the proposed neural network (NN) for solving this class of GVIs can be globally convergent, globally asymptotically stable, and globally exponentially stable under different conditions. The proposed NN can be viewed as a modified version of the general projection NN existing in the literature. Several numerical examples are provided to demonstrate the effectiveness and performance of the proposed NN. Xiaolin Hu 0001, Jun Wang 0002 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | Design of General Projection Neural Networks for Solving Monotone Linear Variational Inequalities and Linear and Quadratic Optimization ProblemsabstractMost existing neural networks for solving linear variational inequalities (LVIs) with the mapping Mx + p require positive definiteness (or positive semidefiniteness) of M. In this correspondence, it is revealed that this condition is sufficient but not necessary for an LVI being strictly monotone (or monotone) on its constrained set where equality constraints are present. Then, it is proposed to reformulate monotone LVIs with equality constraints into LVIs with inequality constraints only, which are then possible to be solved by using some existing neural networks. General projection neural networks are designed in this correspondence for solving the transformed LVIs. Compared with existing neural networks, the designed neural networks feature lower model complexity. Moreover, the neural networks are guaranteed to be globally convergent to solutions of the LVI under the condition that the linear mapping Mx + p is monotone on the constrained set. Because quadratic and linear programming problems are special cases of LVI in terms of solutions, the designed neural networks can solve them efficiently as well. In addition, it is discovered that the designed neural network in a specific case turns out to be the primal-dual network for solving quadratic or linear programming problems. The effectiveness of the neural networks is illustrated by several numerical examples. Xiaolin Hu 0001, Jun Wang 0002 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2006 | Solving Extended Linear Programming Problems Using a Class of Recurrent Neural Networks
Xiaolin Hu 0001, Jun Wang 0002 |
ICONIP (2) | 1 |
| 2006 | A Recurrent Neural Network for Solving Nonconvex Optimization ProblemsabstractAn existing recurrent neural network for convex optimization is extended to solve nonconvex optimization problems. One of the prominent features of this neural network is the one-to-one correspondence between its equilibria and the Karush-Kuhn-Tucker (KKT) points of the nonconvex optimization problem. The conditions are derived under which the neural network (locally) converges to the KKT points. It is desired that the neural network is stable at minimum solutions, and unstable at maximum solutions or saddle solutions. It is found in the paper that most likely the neural network is unstable at the maximum solutions. Moreover, we found that if the derived conditions are not satisfied at minimum solutions, by transforming the original problem into an equivalent one with the p-power (or partial p-power) method, these conditions can be satisfied. As a result, the neural network will locally converge to a minimum solution. Finally, two illustrative examples are provided to demonstrate the performance of the recurrent neural network. Xiaolin Hu 0001, Jun Wang 0002 |
IJCNN | 1 |
| 2006 | Global stability of a recurrent neural network for solving pseudomonotone variational inequalitiesabstractSolving variational inequality problems by using neural networks are of great interest in recent years. To date, most work in this direction focus on solving monotone variational inequalities. In this paper, we show that an existing recurrent neural network proposed originally for solving monotone variational inequalities can be used to solve pseudomonotone variational inequalities with proper choice of a system parameter. The global convergence, global asymptotic stability and global exponential stability of the neural network are discussed under various conditions. The existing stability results are thus extended in view of the fact that pseudomonotonicity is a weaker condition than monotonicity Xiaolin Hu 0001, Jun Wang 0002 |
ISCAS | 1 |
| 2006 | Solving Pseudomonotone Variational Inequalities and Pseudoconvex Optimization Problems Using the Projection Neural NetworkabstractIn recent years, a recurrent neural network called projection neural network was proposed for solving monotone variational inequalities and related convex optimization problems. In this paper, we show that the projection neural network can also be used to solve pseudomonotone variational inequalities and related pseudoconvex optimization problems. Under various pseudomonotonicity conditions and other conditions, the projection neural network is proved to be stable in the sense of Lyapunov and globally convergent, globally asymptotically stable, and globally exponentially stable. Since monotonicity is a special case of pseudomononicity, the projection neural network can be applied to solve a broader class of constrained optimization problems related to variational inequalities. Moreover, a new concept, called componentwise pseudomononicity, different from pseudomononicity in general, is introduced. Under this new concept, two stability results of the projection neural network for solving variational inequalities are also obtained. Finally, numerical examples show the effectiveness and performance of the projection neural network. Xiaolin Hu 0001, Jun Wang 0002 |
IEEE Trans. Neural Networks | 1 |