VLDB 2026 Research / reviewers in the wild / expert
Xuhui Huang
dblp:01/8502
· DBLP profile ↗
64ranked-venue papers
0as first author
29since 2021 · last 2026
0000-0002-7119-9358ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 19 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 since 2021Human-computer interaction and ubiquitous computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HLML-SNN:Fast Continual Learning in Spiking Neural Networks Achieved via Hebbian Learning-Driven Meta-LearningabstractCatastrophic forgetting remains a fundamental barrier to artificial continual learning (CL) - a capability innate to humans. Existing CL methods often incur prohibitive computational costs in resource-constrained scenarios. Spiking neural networks (SNNs), with their biological plausibility and energy efficiency, offer distinct advantages for CL. Inspired by cortico-hippocampal memory mechanisms, we propose a spiking neural network framework integrating Hebbian plasticity with meta-learning, named HLML-SNN. This architecture emulates a dual-phase CL process: (1) In the short-term phase, sample-level Hebbian learning rapidly adapts to new inputs through local synaptic updates; (2) In the long-term phase, task-level meta-learning optimizes cross-task parameters using consolidated synaptic weights, mimicking cortical memory integration to refine shared representations and initialize subsequent Hebbian learning. HLML-SNN incrementally transforms short-term adaptations into stable long-term knowledge, where the synergy of rapid synaptic updates and meta-driven global optimization enables efficient continual learning while balancing stability and plasticity. Empirical results establish HLML-SNN's state-of-the-art performance across split-MNIST/CIFAR10/CIFAR100/TinyImageNet while markedly reducing training time compared to existing methods, demonstrating substantial practical potential for rapid deployment scenarios. The code and appendix are available on https://github.com/JiangshuaiXu/HLML SNN. Jiangshuai Xu, Peiyun Xue, Jiacheng Song, Xuhui Huang, Qingshan Hou |
AAAI | 4 |
| 2026 | DMV - CLIP : Disentangled Multimodal Visual Adaptation for Text-Driven Face EditingabstractABSTRACT Text‐driven face editing has attracted widespread interest due to its intuitive control and user‐friendly interaction. However, current state‐of‐the‐art (SOTA) methods face two main challenges: (1) they utilize unfinetuned general image‐text encoders for modality fusion, making it difficult to comprehend domain‐specific knowledge in facial attribute editing (dozens of fine‐grained facial attributes such as moustache and lipsticks); (2) they roughly optimize all attributes simultaneously using a cross‐entropy loss, leading to severe mutual interference among attributes. To this end, we propose Disentangled Multimodal Visual Adaptation for CLIP (DMV‐CLIP). First, DMV‐CLIP incorporates learnable context tokens to inject facial domain knowledge into the CLIP model via multimodal prompt learning (MPL). Second, it employs directional contrastive learning (DCL) to disentangle facial attributes and enable precise editing. Finally, DMV‐CLIP utilizes a vision‐language consistency model (VLCM) to maintain identity consistency while ensuring that the generated images strictly adhere to the semantic instructions. Xin Wei 0002, Huan Wan, Haoruo Zhang, Xuhui Huang |
Expert Syst. J. Knowl. Eng. | 6 |
| 2026 | S- CNN : A Dual-Region Feature Convolutional Network for Fish Freshness Assessment Based on Eyes and Gills CharacteristicsabstractABSTRACT Freshness is a core quality indicator that determines the utilisation and commercial value of fish products. Traditional fish freshness detection methods are highly subjective and destructive, while existing neural network models suffer from low detection accuracy, unsatisfactory recall and confidence scores and limited generalisation ability. To address these limitations, this paper proposes a novel Dual‐Sensitive Convolutional Neural Network (S‐CNN), where the letter ‘S’ stands for Sensitive. The model simultaneously extracts and fuses discriminative features from fish eye and gill images, capturing subtle freshness differences through a dual‐sensitive feature extraction mechanism. In data preprocessing, all image pixels are normalised to the range [0, 1] to unify numerical scales, stabilise gradient descent and mitigate overfitting. The proposed S‐CNN is composed of seven convolutional blocks, each equipped with batch normalisation, L2 regularisation and a pooling layer; the pooling operation is omitted in the last block to avoid excessive dimensionality reduction. After the flatten layer, a Dropout regularisation module is adopted, and L2 regularisation is applied to all convolutional and fully connected layers. The network uses categorical crossentropy as the loss function. Experimental results demonstrate that the S‐CNN achieves a detection accuracy of 98.70% and an average confidence score of 99.18% on the fish freshness dataset, outperforming other comparative models. The results confirm that the fusion of fish eye and gill features can effectively evaluate fish freshness, providing a reliable method for nondestructive detection and quality assessment of fish products. Boqi Suzhang, Xiaozhou He, Xuhui Huang, Suo Gao, Jun Mou, Ahmed A. Abd El-Latif 0001, Basma Abd El-Rahiem |
Expert Syst. J. Knowl. Eng. | 4 |
| 2026 | The gift of clinical knowledge: annotation-free liver tumor segmentation via knowledge-driven synthesis
Keyi Zhong, Feng Ouyang, Peter Xiaoping Liu, Xuhui Huang, Huan Wan, Xin Wei 0002 |
Multim. Syst. | 4 |
| 2026 | Hierarchical Concept Bottleneck With Compensation Concept LearningabstractConcept Bottleneck Models (CBMs) enhance the interpretability of deep neural networks by mapping images to human-understandable concepts and then using the concepts to make predictions. While they improve transparency, existing CBMs primarily explain only the final layer's features, limiting the interpretability of intermediate layers. Additionally, constructing a comprehensive concept set remains a challenging task, further constraining model performance. In this paper, we investigate the assignment of concept granularity across model layers and propose theHierarchicalConceptBottleneckModel (HCBM) to enhance interpretability. HCBM introduces a Hybrid Concept Bottleneck Layer (HCBL) at each layer, consisting of a Predefined Concept Bottleneck (PCB) that maps visual features to concepts of corresponding granularity and a Compensation Concept Bottleneck (CCB) which incorporates the concept frequency loss and the concept semantic loss to capture compensation concepts for improving performance. Extensive experiments demonstrate that HCBM outperforms state-of-the-art methods. It is worth noting that the HCBM with CLIP RN50 as the backbone outperforms the black-box model. Miao Shang, Kaige Mao, Xiaopeng Hong, Xuhui Huang |
IEEE Trans. Multim. | 6 |
| 2025 | Improving Transformer Based Line Segment Detection with Matched Predicting and Re-rankingabstractClassical Transformer-based line segment detection methods have delivered impressive results. However, we observe that some accurately detected line segments are assigned low confidence scores during prediction, causing them to be ranked lower and potentially suppressed. Additionally, these models often require prolonged training periods to achieve strong performance, largely due to the necessity of bipartite matching. In this paper, we introduce RANK-LETR, a novel Transformer-based line segment detection method. Our approach leverages learnable geometric information to refine the ranking of predicted line segments by enhancing the confidence scores of high-quality predictions in a posterior verification step. We also propose a new line segment proposal method, wherein the feature point nearest to the centroid of the line segment directly predicts the location, significantly improving training efficiency and stability. Moreover, we introduce a line segment ranking loss to stabilize rankings during training, thereby enhancing the generalization capability of the model. Experimental results demonstrate that our method outperforms other Transformer-based and CNN-based approaches in prediction accuracy while requiring fewer training epochs than previous Transformer-based models. Shi Peng, Baojie Tian, Yufei Guo 0001, Xuhui Huang, Zhe Ma 0001 |
AAAI | 5 |
| 2025 | AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-ScientistsabstractYifei Li, Hanane Nour Moussa, Ziru Chen, Shijie Chen, Botao Yu, Mingyi Xue, Benjamin Burns, Tzu-Yao Chiu, Vishal Dey, Zitong Lu, Chen Wei, Qianheng Zhang, Tianyu Zhang, Song Gao, Xuhui Huang, Xia Ning, Nesreen K. Ahmed, Ali Payani, Huan Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yifei Li 0005, Hanane Nour Moussa, Ziru Chen, Botao Yu, Mingyi Xue 0001, Benjamin Burns, Tzu-Yao Chiu, Vishal Dey, Zitong Lu, Qianheng Zhang, Song Gao 0001, Xuhui Huang, Xia Ning, Nesreen K. Ahmed, Ali Payani, Huan Sun 0001 |
EMNLP | 15 |
| 2025 | ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryabstractThe advancements of language language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about the true capabilities of such agents. In this work, we argue that for an agent to fully automate scientific discovery, it must be able to complete all essential tasks in the workflow. Thus, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for evaluating language agents for data-driven scientific discovery. To ensure the scientific authenticity and real-world relevance of our benchmark, we extract 102 tasks from 44 peer-reviewed publications in four disciplines and engage nine subject matter experts to validate them. We unify the target output for every task to a self-contained Python program file and employ an array of evaluation metrics to examine the generated programs, execution results, and costs. Each task goes through multiple rounds of manual validation by annotators and subject matter experts to ensure its annotation quality and scientific plausibility. We also propose two effective strategies to mitigate data contamination concerns. Using our benchmark, we evaluate five open-weight and proprietary LLMs, each with three frameworks: direct prompting, OpenHands, and self-debug. Given three attempts for each task, the best-performing agent can only solve 32.4% of the tasks independently and 34.3% with expert-provided knowledge. These results underscore the limited capacities of current language agents in generating code for data-driven discovery, let alone end-to-end automation for scientific research. Ziru Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li 0005, Zeyi Liao, Zitong Lu, Vishal Dey, Mingyi Xue 0001, Frazier N. Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao 0001, Yu Su 0001, Huan Sun 0001 |
ICLR | 16 |
| 2025 | Dynamic dual consistency distillation for imbalanced facial attribute editing
Xuhui Huang, Huan Wan, Xin Wei 0002 |
Knowl. Based Syst. | 2 |
| 2025 | Chaos-MLP: Chaotic Transform MLP-Like Architecture for Medical Images Multi-Label Recognition TaskabstractThe theory of "three-stage prevention" in view of the body constitution is the key technology of modern Chinese medicine for "Preventive Treatment of Diseases". In particular, automated body constitution recognition (BCR) is an integral part of intelligent Traditional Chinese Medicine (TCM), which is extremely valuable for disease prevention and diagnosis. Actually, BCR is a challenging multi-label recognition task by the TCM composite constitution theory. First, two new databases are constructed, one is a multi-label facial body constitution (MFBC), and another is a multi-label tongue body constitution (MTBC). Second, a novel MLP-like architecture, named Chaos-MLP, is designed for the BCR task, which interacts with the channel chaotic features of extracted medical images and fuses them with the width and height channel direction features, respectively. Notably, the chaotic transform can enhance the distinguishability of extracted features from the medical images. Moreover, we propose a binary center cognitive gravity loss (BCCGL) to enhance the learning ability of the Chaos-MLP for unbalanced body constitution labels. Our proposed method shows superior performance on both MFBC and MTBC datasets than other state-of-the-art (SOTA) MLP-like networks and a vision graph-based neural network (VGNN), which include Wave-MLP, Cycle-MLP, Vip, and Active-MLP. Mengjian Zhang, Guihua Wen, Pei Yang 0001, Changjun Wang, Xuhui Huang, Chuyun Chen |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Ternary Spike: Learning Ternary Spikes for Spiking Neural NetworksabstractThe Spiking Neural Network (SNN), as one of the biologically inspired neural network infrastructures, has drawn increasing attention recently. It adopts binary spike activations to transmit information, thus the multiplications of activations and weights can be substituted by additions, which brings high energy efficiency. However, in the paper, we theoretically and experimentally prove that the binary spike activation map cannot carry enough information, thus causing information loss and resulting in accuracy decreasing. To handle the problem, we propose a ternary spike neuron to transmit information. The ternary spike neuron can also enjoy the event-driven and multiplication-free operation advantages of the binary spike neuron but will boost the information capacity. Furthermore, we also embed a trainable factor in the ternary spike neuron to learn the suitable spike amplitude, thus our SNN will adopt different spike amplitudes along layers, which can better suit the phenomenon that the membrane potential distributions are different along layers. To retain the efficiency of the vanilla ternary spike, the trainable ternary spike SNN will be converted to a standard one again via a re-parameterization technique in the inference. Extensive experiments with several popular network structures over static and dynamic datasets show that the ternary spike can consistently outperform state-of-the-art methods. Our code is open-sourced at https://github.com/yfguo91/Ternary-Spike. Yufei Guo 0001, Yuanpei Chen, Xiaode Liu, Weihang Peng 0001, Yuhan Zhang 0006, Xuhui Huang, Zhe Ma 0001 |
AAAI | 6 |
| 2024 | End-to-End Real-Time Vanishing Point Detection with TransformerabstractIn this paper, we propose a novel transformer-based end-to-end real-time vanishing point detection method, which is named Vanishing Point TRansformer (VPTR). The proposed method can directly regress the locations of vanishing points from given images. To achieve this goal, we pose vanishing point detection as a point object detection task on the Gaussian hemisphere with region division. Considering low-level features always provide more geometric information which can contribute to accurate vanishing point prediction, we propose a clear architecture where vanishing point queries in the decoder can directly gather multi-level features from CNN backbone with deformable attention in VPTR. Our method does not rely on line detection or Manhattan world assumption, which makes it more flexible to use. VPTR runs at an inferring speed of 140 FPS on one NVIDIA 3090 card. Experimental results on synthetic and real-world datasets demonstrate that our method can be used in both natural and structural scenes, and is superior to other state-of-the-art methods on the balance of accuracy and efficiency. Shi Peng, Yufei Guo 0001, Xuhui Huang |
AAAI | 4 |
| 2024 | Enhancing Representation of Spiking Neural Networks via Similarity-Sensitive Contrastive LearningabstractSpiking neural networks (SNNs) have attracted intensive attention as a promising energy-efficient alternative to conventional artificial neural networks (ANNs) recently, which could transmit information in form of binary spikes rather than continuous activations thus the multiplication of activation and weight could be replaced by addition to save energy. However, the binary spike representation form will sacrifice the expression performance of SNNs and lead to accuracy degradation compared with ANNs. Considering improving feature representation is beneficial to training an accurate SNN model, this paper focuses on enhancing the feature representation of the SNN. To this end, we establish a similarity-sensitive contrastive learning framework, where SNN could capture significantly more information from its ANN counterpart to improve representation by Mutual Information (MI) maximization with layer-wise sensitivity to similarity. In specific, it enriches the SNN’s feature representation by pulling the positive pairs of SNN's and ANN's feature representation of each layer from the same input samples closer together while pushing the negative pairs from different samples further apart. Experimental results show that our method consistently outperforms the current state-of-the-art algorithms on both popular non-spiking static and neuromorphic datasets. Yuhan Zhang 0006, Xiaode Liu, Yuanpei Chen, Weihang Peng 0001, Yufei Guo 0001, Xuhui Huang, Zhe Ma 0001 |
AAAI | 6 |
| 2023 | Deep Dive into Gradients: Better Optimization for 3D Object Detection with Gradient-Corrected IoU SupervisionabstractIntersection-over-Union (IoU) is the most popular metric to evaluate regression performance in 3D object detection. Recently, there are also some methods applying IoU to the optimization of 3D bounding box regression. However, we demonstrate through experiments and mathematical proof that the 3D IoU loss suffers from abnormal gradient w.r.t. angular error and object scale, which further leads to slow convergence and suboptimal regression process, respectively. In this paper, we propose a Gradient-Corrected IoU (GCIoU) loss to achieve fast and accurate 3D bounding box regression. Specifically, a gradient correction strategy is designed to endow 3D IoU loss with a reasonable gradient. It ensures that the model converges quickly in the early stage of training, and helps to achieve fine-grained refinement of bounding boxes in the later stage. To solve suboptimal regression of 3D IoU loss for objects at different scales, we introduce a gradient rescaling strategy to adaptively optimize the step size. Finally, we integrate GCIoU Loss into multiple models to achieve stable performance gains and faster model convergence. Experiments on KITTI dataset demonstrate superiority of the proposed method. The code is available at https://github.com/ming71/GCIoU-loss. Qi Ming, Lingjuan Miao, Zhe Ma 0001, Zhiqiang Zhou 0001, Xuhui Huang, Yuanpei Chen, Yufei Guo 0001 |
CVPR | 6 |
| 2023 | PeakConv: Learning Peak Receptive Field for Radar Semantic SegmentationabstractThe modern machine learning-based technologies have shown considerable potential in automatic radar scene understanding. Among these efforts, radar semantic segmentation (RSS) can provide more refined and detailed information including the moving objects and background clutters within the effective receptive field of the radar. Motivated by the success of convolutional networks in various visual computing tasks, these networks have also been introduced to solve RSS task. However, neither the regular convolution operation nor the modified ones are specific to interpret radar signals. The receptive fields of existing convolutions are defined by the object presentation in optical signals, but these two signals have different perception mechanisms. In classic radar signal processing, the object signature is detected according to a local peak response, i.e., CFAR detection. Inspired by this idea, we redefine the receptive field of the convolution operation as the peak receptive field (PRF) and propose the peak convolution operation (PeakConv) to learn the object signatures in an end-to-end network. By incorporating the proposed PeakConv layers into the encoders, our RSS network can achieve better segmentation results compared with other SoTA methods on a multi-view real-measured dataset collected from an FMCW radar. Our code for PeakConv is available at https://github.com/zlw9161/PKC. Liwen Zhang 0001, Youcheng Zhang, Yufei Guo 0001, Yuanpei Chen, Xuhui Huang, Zhe Ma 0001 |
CVPR | 6 |
| 2023 | RMP-Loss: Regularizing Membrane Potential Distribution for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) as one of the biology-inspired models have received much attention recently. It can significantly reduce energy consumption since they quantize the real-valued membrane potentials to 0/1 spikes to transmit information thus the multiplications of activations and weights can be replaced by additions when implemented on hardware. However, this quantization mechanism will inevitably introduce quantization error, thus causing catastrophic information loss. To address the quantization error problem, we propose a regularizing membrane potential loss (RMP-Loss) to adjust the distribution which is directly related to quantization error to a range close to the spikes. Our method is extremely simple to implement and straightforward to train an SNN. Furthermore, it is shown to consistently outperform previous state-of-the-art methods over different network architectures and datasets. Yufei Guo 0001, Xiaode Liu, Yuanpei Chen, Liwen Zhang 0001, Weihang Peng 0001, Yuhan Zhang 0006, Xuhui Huang, Zhe Ma 0001 |
ICCV | 7 |
| 2023 | Membrane Potential Batch Normalization for Spiking Neural NetworksabstractAs one of the energy-efficient alternatives of conventional neural networks (CNNs), spiking neural networks (SNNs) have gained more and more interest recently. To train the deep models, some effective batch normalization (BN) techniques are proposed in SNNs. All these BNs are suggested to be used after the convolution layer as usually doing in CNNs. However, the spiking neuron is much more complex with the spatio-temporal dynamics. The regulated data flow after the BN layer will be disturbed again by the membrane potential updating operation before the firing function, i.e., the nonlinear activation. Therefore, we advocate adding another BN layer before the firing function to normalize the membrane potential again, called MPBN. To eliminate the induced time cost of MPBN, we also propose a training-inference-decoupled re-parameterization technique to fold the trained MPBN into the firing threshold. With the re-parameterization technique, the MPBN will not introduce any extra time burden in the inference. Furthermore, the MPBN can also adopt the element-wised form, while these BNs after the convolution layer can only use the channel-wised form. Experimental results show that the proposed MPBN performs well on both popular non-spiking static and neuromorphic datasets. Yufei Guo 0001, Yuhan Zhang 0006, Yuanpei Chen, Weihang Peng 0001, Xiaode Liu, Liwen Zhang 0001, Xuhui Huang, Zhe Ma 0001 |
ICCV | 7 |
| 2023 | Alleviating Catastrophic Forgetting of Incremental Object Detection via Within-Class and Between-Class Knowledge DistillationabstractIncremental object detection (IOD) task requires a model to learn continually from newly added data. However, directly fine-tuning a well-trained detection model on a new task will sharply decrease the performance on old tasks, which is known as catastrophic forgetting. Knowledge distillation, including feature distillation and response distillation, has been proven to be an effective way to alleviate catastrophic forgetting. However, previous works on feature distillation heavily rely on low-level feature information, while under-exploring the importance of high-level semantic information. In this paper, we discuss the cause of catastrophic forgetting in IOD task as destruction of semantic feature space. We propose a method that dynamically distills both semantic and feature information with consideration of both between-class discriminativeness and within-class consistency on Transformer-based detector. Between-class discriminativeness is preserved by distilling class-level semantic distance and feature distance among various categories, while within-class consistency is preserved by distilling instance-level semantic information and feature information within each category. Extensive experiments are conducted on both Pascal VOC and MS COCO benchmarks. Our method outperforms all the previous CNN-based SOTA methods under various experimental scenarios, with a remarkable mAP improvement from 36.90% to 39.80% under one-step IOD task. Mengxue Kang, Xiashuang Wang, Zhe Ma 0001, Xuhui Huang |
ICCV | 7 |
| 2023 | Joint A-SNN: Joint training of artificial and spiking neural networks via self-Distillation and weight factorization
Yufei Guo 0001, Weihang Peng 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, Xuhui Huang, Zhe Ma 0001 |
Pattern Recognit. | 6 |
| 2023 | MLP-Like Model With Convolution Complex Transformation for Auxiliary Diagnosis Through Medical ImagesabstractMedical images such as facial and tongue images have been widely used for intelligence-assisted diagnosis, which can be regarded as the multi-label classification task for disease location (DL) and disease nature (DN) of biomedical images. Compared with complicated convolutional neural networks and Transformers for this task, recent MLP-like architectures are not only simple and less computationally expensive, but also have stronger generalization capabilities. However, MLP-like models require better input features from the image. Thus, this study proposes a novel convolution complex transformation MLP-like (CCT-MLP) model for the multi-label DL and DN recognition task for facial and tongue images. Notably, the convolutional Tokenizer and multiple convolutional layers are first used to extract the better shallow features from input biomedical images to make up for the loss of spatial information obtained by the simple MLP structure. Subsequently, the Channel-MLP architecture with complex transformations is used to extract deep-level contextual features. In this way, multi-channel features are extracted and mixed to perform the multi-label classification of the input biomedical images. Experimental results on our constructed multi-label facial and tongue image datasets demonstrate that our method outperforms existing methods in terms of both accuracy (Acc) and mean average precision (mAP). Mengjian Zhang, Guihua Wen, Jiahui Zhong, Changjun Wang, Xuhui Huang |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Zero-Shot Predicate Prediction for Scene Graph ParsingabstractThe scene graph is a structured semantic representation of an image, which represents objects and relationships with vertices and edges, respectively. Since it is impossible to manually label all potential relationships in the real world, some previous methods try to apply the zero-shot method for scene graph generation. However, existing methods take triplet (i.e., hsubject-predicate-objecti) as the basic unit of a relationship. Each element (i.e., subject, predicate, or object) of the unseen relationship is actually seen in the training data. Therefore, they ignore the unseen predicate. To predict the unseen predicate, we introduce a novel task named zero-shot predicate prediction, which is crucial to extending existing scene graph generation methods to recognize more relationship classes. The new task is challenging and cannot be simply resolved through conventional zero-shot learning methods because there is a large intra-class variation of each predicate. Firstly, the large intra-class variation leads to the difficulty of computing the discriminative instancelevel feature of the predicate class. Secondly, the large intraclass variation also brings more difficulties when knowledge is transferred from seen classes to unseen classes. For the first challenge, we propose distilling lexical knowledge of different objects and construct multi-modal representations of pairwise objects to reduce the intra-class variation of the predicate. To respond to the second challenge, we build a compact semantic space where the representations of unseen classes are reconstructed based on the seen classes for zero-shot predicate classification. We evaluate the proposed method on the public dataset Visual Genome. The extensive experiment results under the zeroshot/few-shot/supervised settings demonstrate the effectiveness of the proposed method. Yiming Li 0008, Xiaoshan Yang, Xuhui Huang, Zhe Ma 0001, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2022 | RecDis-SNN: Rectifying Membrane Potential Distribution for Directly Training Spiking Neural NetworksabstractThe brain-inspired and event-driven Spiking Neural Network (SNN) aiming at mimicking the synaptic activity of biological neurons has received increasing attention. It transmits binary spike signals between network units when the membrane potential exceeds the firing threshold. This biomimetic mechanism of SNN appears energy-efficiency with its power sparsity and asynchronous operations on spike events. Unfortunately, with the propagation of binary spikes, the distribution of membrane potential will shift, leading to degeneration, saturation, and gradient mismatch problems, which would be disadvantageous to the network optimization and convergence. Such undesired shifts would prevent the SNN from performing well and going deep. To tackle these problems, we attempt to rectify the membrane potential distribution (MPD) by designing a novel distribution loss, MPD-Loss, which can explicitly penalize the un-desired shifts without introducing any additional operations in the inference phase. Moreover, the proposed method can also mitigate the quantization error in SNNs, which is usually ignored in other works. Experimental results demonstrate that the proposed method can directly train a deeper, larger, and better-performing SNN within fewer timesteps. Yufei Guo 0001, Xinyi Tong 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, Zhe Ma 0001, Xuhui Huang |
CVPR | 7 |
| 2022 | Reducing Information Loss for Spiking Neural Networks
Yufei Guo 0001, Yuanpei Chen, Liwen Zhang 0001, YingLei Wang, Xiaode Liu, Xinyi Tong 0001, Yuanyuan Ou, Xuhui Huang, Zhe Ma 0001 |
ECCV (11) | 8 |
| 2022 | Real Spike: Learning Real-Valued Spikes for Spiking Neural Networks
Yufei Guo 0001, Liwen Zhang 0001, Yuanpei Chen, Xinyi Tong 0001, Xiaode Liu, YingLei Wang, Xuhui Huang, Zhe Ma 0001 |
ECCV (12) | 7 |
| 2022 | Balanced Ranking and Sorting For Class Incremental Object DetectionabstractClass incremental learning has drawn much attention recently. Although many algorithms have been proposed for class incremental image classification, developing object detectors which can learn incrementally is still a challenge. Existing methods rely on knowledge distillation to achieve class incremental object detection (CIOD), which suffer from performance tradeoff between old and new classes. In this paper, we propose balanced ranking and sorting (BRS), to tackle the catastrophic forgetting and data imbalance problems for CIOD. Specifically, ranking & sorting with pseudo ground truths (RSP) and ranking & sorting transfer (RST) are developed to preserve the learned knowledge from the old model while learning new classes, in an unified framework. To mitigate the data imbalance problem, gradient rebalancing is performed with specific sample pairs. We demonstrate the effectiveness of our approach with extensive experiments on PASCAL VOC and COCO datasets, in which significant improvement over state-of-the-art methods is achieved. Bo Cui 0004, Xuhui Huang |
ICASSP | 3 |
| 2022 | IM-Loss: Information Maximization Loss for Spiking Neural NetworksabstractSpiking Neural Network (SNN), recognized as a type of biologically plausible architecture, has recently drawn much research attention. It transmits information by $0/1$ spikes. This bio-mimetic mechanism of SNN demonstrates extreme energy efficiency since it avoids any multiplications on neuromorphic hardware. However, the forward-passing $0/1$ spike quantization will cause information loss and accuracy degradation. To deal with this problem, the Information maximization loss (IM-Loss) that aims at maximizing the information flow in the SNN is proposed in the paper. The IM-Loss not only enhances the information expressiveness of an SNN directly but also plays a part of the role of normalization without introducing any additional operations (\textit{e.g.}, bias and scaling) in the inference phase. Additionally, we introduce a novel differentiable spike activity estimation, Evolutionary Surrogate Gradients (ESG) in SNNs. By appointing automatic evolvable surrogate gradients for spike activity function, ESG can ensure sufficient model updates at the beginning and accurate gradients at the end of the training, resulting in both easy convergence and high task performance. Experimental results on both popular non-spiking static and neuromorphic datasets show that the SNN models trained by our method outperform the current state-of-the-art algorithms. Yufei Guo 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, YingLei Wang, Xuhui Huang, Zhe Ma 0001 |
NeurIPS | 6 |
| 2022 | Clustering by centroid drift and boundary shrinkage
Hui Qv, Tao Ma 0008, Xinyi Tong 0001, Xuhui Huang, Zhe Ma 0001, Jiehong Feng |
Pattern Recognit. | 4 |
| 2021 | An Improved Advancing-front-Delaunay Method for Triangular Mesh Generation
Yufei Guo 0001, Xuhui Huang, Zhe Ma 0001, Yongqing Hai, Rongli Zhao, Kewu Sun |
CGI | 2 |
| 2021 | ECKPN: Explicit Class Knowledge Propagation Network for Transductive Few-Shot LearningabstractRecently, the transductive graph-based methods have achieved great success in the few-shot classification task. However, most existing methods ignore exploring the class-level knowledge that can be easily learned by humans from just a handful of samples. In this paper, we propose an Explicit Class Knowledge Propagation Network (ECKPN), which is composed of the comparison, squeeze and calibration modules, to address this problem. Specifically, we first employ the comparison module to explore the pairwise sample relations to learn rich sample representations in the instance-level graph. Then, we squeeze the instance-level graph to generate the class-level graph, which can help obtain the class-level visual knowledge and facilitate modeling the relations of different classes. Next, the calibration module is adopted to characterize the relations of the classes explicitly to obtain the more discriminative class-level knowledge representations. Finally, we combine the class-level knowledge with the instance-level sample representations to guide the inference of the query samples. We conduct extensive experiments on four few-shot classification benchmarks, and the experimental results show that the proposed ECKPN significantly outperforms the state-of-the art methods. Xiaoshan Yang, Changsheng Xu, Xuhui Huang, Zhe Ma 0001 |
CVPR | 4 |
| 2020 | A biologically plausible supervised learning method for spiking neural networks using the symmetric STDP rule
Yunzhe Hao, Xuhui Huang, Meng Dong, Bo Xu 0002 |
Neural Networks | 2 |
| 2019 | Attentional Residual Dense Factorized Network for Real-Time Semantic Segmentation
Long Lan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
ICANN (3) | 4 |
| 2019 | Short-term synaptic plasticity expands the operational range of long-term synaptic changes in neural networks
Guanxiong Zeng, Xuhui Huang, Tianzi Jiang |
Neural Networks | 2 |
| 2019 | Stacked Marginal Time Warping for Temporal Alignment
Xiang Zhang 0008, Liquan Nie, Long Lan, Xuhui Huang, Zhigang Luo |
Neural Process. Lett. | 4 |
| 2019 | Nonnegative Constrained Graph Based Canonical Correlation Analysis for Multi-view Feature Learning
Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
Neural Process. Lett. | 4 |
| 2018 | Margin-Embedding Canonical Correlation Analysis with Feature Selection for Person Re-IdentificationabstractCanonical correlation analysis (CCA) is a classical subspace learning method of capturing the common semantic information underlying multi-view data. It has been used in person re-identification (re-ID) task by treating the task of matching identical individuals across non-overlapping multi-cameras as a multi-view learning problem. However, CCA-based reID methods still achieve unsatisfactory results because few jointly consider discriminative margin information and selecting importantly relevant features. To address this issue, we propose a novel l2,1-norm regularized margin-embedding CCA ( l2,1-MCCA), which learns a generalized discriminative subspace by employing more discriminative margin information. Moreover, the new method enforces the l2,1-norm regularization term over the learned subspace to identify the relevant features. Both lightweight and effective schemes can benefit from each other and endeavor to enlarge the interclass variations whilst reducing the intra-class variations. Experiments on three popular datasets show the efficacy of l2,1-MCCA as compared with recently representative re-ID methods. Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICIP | 4 |
| 2018 | Graph-Laplacian Correlated Low-Rank Representation for Subspace ClusteringabstractSubspace clustering seeks to segment a given unlabeled data into clusters with the hope of each cluster corresponding to a union of low-dimensional subspaces. Among them, low-rank representation (LRR) is a promising potential method which intends to build a good affinity matrix by using the self-expression of inputs. However, it completely ignores the important data locality. Although several works in this regard have considered the local geometric structure through the Laplacian regularizer, they also neglect the correlation of the data. In this paper, we propose a graph-Laplacian correlated low-rank representation model (GCLRR) to address such an issue. Particularly, GCLRR factorizes the self-expression as the product of two low-dimensional matrices, of which one is the latent representation of the self-expression. On the basis of the latent representation, the Laplacian regularizer is integrated with the orthogonal constraint together and behaves like the spectral clustering. Moreover, we devise a Frobenius norm based trace loss and use it to constrain both the latent representation and the self-expression to capture the correlation of the data. Our improved trace loss is more efficient than the original one. More importantly, GCLRR provides an effective unified framework to seamlessly integrate both aspects above. Then, we optimize GCLRR in the frame of alternating direction method (ADM) and fortunately derive the analytical solution to each subproblem. Experiments of motion segmentation and image clustering confirm the efficacy of the proposed GCLRR. Huayue Cai, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICIP | 5 |
| 2018 | Discrete Graph Hashing via Affine TransformationabstractIn unsupervised graph-based hashing for large-scale image retrieval, many efforts have been made to bridge the gap between the learned graph embedding and the corresponding binary codes. Relatively, few studies focus on the issue of the discrimination of graph embedding. In this paper, we firstly devise a discrete graph hashing model (DGH) that smooths graph embedding and simultaneously solving binary codes under the balanced discrete constraint, which equals a novel method of jointly learning graph embedding and spectral rotation, theoretically. To further induce discriminant graph embedding, we substitute affine transformation for spectral rotation in our DGH (abbreviated as ADGH). This is because affine transformation can accommodate both rotational angle and distance of graph embedding, while respecting the neighborhood structure among most samples. Besides, each subproblem of ADGH can yield the closed-form solution. Experiments of image retrieval on three benchmark datasets show that ADGH outperforms the representative hashing methods in quantity. Guohua Dong, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICME | 4 |
| 2018 | Cross-Layer Convolutional Siamese Network for Visual Tracking
Yanyin Chen, Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICONIP (2) | 6 |
| 2018 | Box-constrained Discriminant Projective Non-negative Matrix Factorization through Augmented Lagrangian Multiplier MethodabstractProjective non-negative matrix factorization (PNMF) learns a non-negative projection matrix to project high-dimensional examples onto a lower-dimensional space spanned by the transpose of the learned projection matrix. Since PNMF can learn parts-based representation, it has attracted ample attention from computer vision community. However, existing PNMF methods either completely ignore labels of the dataset or endure the slow convergent optimization algorithm. In this paper, we propose a box-constrained discriminant PNMF (BDPNMF) method to address these issues. Specifically, BDPNMF jointly exploits the Fisher's criterion and the augmented Lagrangian multiplier (ALM) method into PNMF to boost discriminative capacity of the learned subspace and its efficiency. Experimental results on four popular face image datasets confirm the efficacy of BDPNMF compared to previous PNMF methods in quantity. Huayue Cai, Xiang Zhang 0008, Zhigang Luo, Xuhui Huang |
IJCNN | 4 |
| 2018 | Multi-granularity Hierarchical Attention Siamese Network for Visual TrackingabstractSpeed and accuracy are the two most important focuses for many visual tracking methods. Recently, siamese networks based trackers have shown very promising potentials in both aspects, which develop a twin network to measure the responses between target and hypotheses with a fully convolutional operation. However, the learned response maps are vulnerable to background clutters and scale changes as they ignore priori knowledge such as the object salience and multi-granularity cues. To explore the benefits of priori, this paper devises a multi-granularity hierarchical attention siamese network tracker (MHA-Siam) to further enhance the tracking stability without sacrificing real-time speed. Particularly, the channel-wise attention mechanism is exploited here to filter out the background while remain the salient object region; then, the response maps of the coarse-to-finer multi-layer features are fused to capture multi-granularity location information helpful for improvement in tracking stability. To make full use of them, MHA-Siam imposes the element-wise max-and-sum operation on them to induce a reliable response map for accurate location. Experiments of visual tracking on OTB benchmark shows the superiority of MHA-Siam with the competitive efficiency to its counterpart trackers. Xiang Zhang 0008, Huibin Tan, Long Lan, Zhigang Luo, Xuhui Huang |
IJCNN | 6 |
| 2018 | Ranking-Embedded Transfer Canonical Correlation Analysis for Person Re-IdentificationabstractPerson re-identification (re-ID) seeks to match the identical individuals across different cameras and is still a challenging visual task due to substantial variances of person appearance in complex scenarios. Different from most of conventional person re-ID methods, which generally reduce person re-ID task to either a multi-view learning problem or a multi- domain learning problem alone, this paper treats such a task as a multi-view multi-domain (MVMD) learning problem to exploit the both benefits by refreshing canonical correlation analysis (CCA) with two improvements, termed as ranking-embedded transfer CCA (RTCCA). Specifically, to bridge the semantic gap between different views, we first embed a ranking weight matrix into CCA to strength the correlations among the multi-view images of the same identity and simultaneously to weaken that of different identities. Furthermore, we utilize the well-known distribution metric maximum mean discrepancy (MMD) as a regularization term to reduce the domain shift between training set and testing set. More importantly, the two improvements benefit from each other and the joint merit can further boost the re-ID performance. Experiments on three benchmarks verify the efficacy of the proposed RTCCA when compared with the recently representative baseline person re-ID methods. Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
IJCNN | 4 |
| 2018 | Low-Rank Matrix Recovery via Continuation-Based Approximate Low-Rank Minimization
Xiang Zhang 0008, Yongqiang Gao, Long Lan, Xuhui Huang, Zhigang Luo |
PRICAI (1) | 5 |
| 2017 | Attention Focused Spatial Pyramid Pooling for Boxless Action Recognition in Still Images
Weijiang Feng, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
ICANN (2) | 3 |
| 2017 | GNMF Revisited: Joint Robust k-NN Graph and Reconstruction-Based Graph Regularization for Image Clustering
Wenju Zhang, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
ICANN (2) | 5 |
| 2017 | Unsupervised domain adaptation with joint supervised sparse coding and discriminative regularization termabstractDomain adaptation (DA) attempts to enhance the generalization capability of classifier through narrowing the gap of the distributions across domains. This paper focuses on unsupervised domain adaptation where labels are not available in target domain. Most existing approaches explore the domain-invariant features shared by domains but ignore the discriminative information of source domain. To address this issue, we propose a discriminative domain adaptation method (DDA) to reduce domain shift by seeking a common latent subspace jointly using supervised sparse coding (SSC) and discriminative regularization term. Particularly, DDA adapts SSC to yield discriminative coefficients of target data and further unites with discriminative regularization term to induce a common latent subspace across domains. We show that both strategies can boost the ability of transferring knowledge from source to target domain. Experiments on two real world datasets demonstrate the effectiveness of our proposed method over several existing state-of-the-art domain adaptation methods. Xiang Zhang 0008, Wenju Zhang, Xuhui Huang, Naiyang Guan, Zhigang Luo |
ICIP | 4 |
| 2017 | Boxless Action Recognition in Still Images via Recurrent Visual Attention
Weijiang Feng, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
ICONIP (2) | 3 |
| 2016 | Enhancing temporal alignment with autoencoder regularizationabstractTemporal alignment aligns two temporal sequences and is quite challenging due to drastic differences among temporal sequences and source data from different views. Canonical time warping (CTW) has shown great potential in temporal alignment tasks because it can reduce data redundancy by transforming high-dimensional data to a lower-dimensional subspace via canonical correlation analysis (CCA). However, CTW cannot uncover the underlying nonlinear structure embedded in the dataset. In this paper, we propose an autoencoder regularized canonical time warping method (AECTW) to overcome this drawback. Specifically, AECTW enhances lower-dimensional representation of each sequence by incorporating an autoencoder regularization, meanwhile reveals the nonlinear structure of features by explicit nonlinear transformation. By these strategies, AECTW significantly boosts CTW in temporal alignment tasks. Experiments on both synthetic data and two practical human action datasets demonstrate that AECTW outperforms the representative DTW-based methods. Liquan Nie, Yuanyuan Wang 0004, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
IJCNN | 4 |
| 2016 | Distributed graph regularized non-negative matrix factorization with greedy coordinate descentabstractGraph regularized non-negative matrix factorization (GNMF) decomposes a high-dimensional non-negative data matrix into two low-dimensional matrices with the non-negativity property kept and the geometric structure preserved. Due to its effectiveness, GNMF has been widely used in many fields such as computer vision and data mining. However, GNMF cannot process large-scale datasets on distributed system because the gradient of the graph regularization term costs huge amount of communication overheads among computing nodes. In this paper, we proposed a distributed GNMF (DGNMF) algorithm to overcome this deficiency. Particularly, DGNMF reformulates the graph regularization term to avoid multiplying graph Laplacian by factor matrix through introducing an auxiliary variable and incorporating an equality constraint over it. We optimize DGNMF by using greedy coordinate descent method in the frame of augmented Lagrange method and implement this algorithm on a distributed system. Since DGNMF requires quite few communication overheads among computing nodes, it can be applied to large scale dataset. The preliminary results illustrate efficiency, scalability, and effectiveness of DGNMF. Ziheng Gao, Naiyang Guan, Xuhui Huang, Xuefeng Peng, Zhigang Luo, Yuhua Tang |
SMC | 3 |
| 2015 | Labelwalking nonnegative matrix factorizationabstractSemi-supervised learning (SSL) utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples. Due to its great discriminant power, SSL has been widely applied to various real-world tasks such as information retrieval, pattern recognition, and speech separa- tion. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, LP assumes nearby examples should share the same label, thus, it unavoidably pushes the labels to the wrong examples, especially when different la- beled examples are not strictly separated. Seed K-means uses labeled examples to initialize class centers, and avoid getting stuck in poor local optima comparing to traditional K-means, however the hard constraint of each example's membership makes Seed K-means failed in many real world applications. This paper proposes a novel label walking nonnegative matrix factorization method (LWNMF) to handle labeled examples in SSL based on the framework of NMF. LWNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix, and to travel labels to unlabeled examples, LWNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. Since LWNMF learns comprehensive class centroids, labels iteratively walk to unlabeled examples through these significant centroids. Long Lan, Naiyang Guan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
ICASSP | 4 |
| 2015 | Semi-supervised Non-negative Local Coordinate Factorization
Cherong Zhou, Xiang Zhang 0008, Naiyang Guan, Xuhui Huang, Zhigang Luo |
ICONIP (2) | 4 |
| 2015 | Correntropy supervised non-negative matrix factorizationabstractNon-negative matrix factorization (NMF) is a powerful dimension reduction method and has been widely used in many pattern recognition and computer vision problems. However, conventional NMF methods are neither robust enough as their loss functions are sensitive to outliers, nor discriminative because they completely ignore labels in a dataset. In this paper, we proposed a correntropy supervised NMF (CSNMF) to simultaneously overcome aforementioned deficiencies. In particular, CSNMF maximizes the correntropy between the data matrix and its reconstruction in low-dimensional space to inhibit outliers during learning the subspace, and narrows the minimizes the distances between coefficients of any two samples with the same class labels to enhance the subsequent classification performance. To solve CSNMF, we developed a multiplicative update rules and theoretically proved its convergence. Experimental results on popular face image datasets verify the effectiveness of CSNMF comparing with NMF, its supervised variants, and its robustified variants. Wenju Zhang, Naiyang Guan, Dacheng Tao, Bin Mao, Xuhui Huang, Zhigang Luo |
IJCNN | 5 |
| 2015 | Two-Dimensional Euler PCA for Face Recognition
Huibin Tan, Xiang Zhang 0008, Naiyang Guan, Dacheng Tao, Xuhui Huang, Zhigang Luo |
MMM (2) | 5 |
| 2015 | Non-negative Low-Rank and Group-Sparse Matrix Factorization
Shuyi Wu, Xiang Zhang 0008, Naiyang Guan, Dacheng Tao, Xuhui Huang, Zhigang Luo |
MMM (2) | 5 |
| 2015 | Constraint-Relaxation Approach for Nonnegative Matrix Factorization: A Case StudyabstractNonnegative matrix factorization (NMF) is a powerful technique for dimensionality reduction. Conventional NMF algorithms usually keep the matrices W and H nonnegative while iterating. However, to get the NMF of a matrix, it's unnecessary to force the temporary solutions in iterations nonnegative. In this paper, we propose a two-staged approach for NMF. At the relaxation stage, the nonnegative constraint of temporary solutions is relaxed and a real valued matrix factorization is generated. At the constraint stage, the real valued matrix factorization is transformed to a nonnegative matrix factorization by an invertible linear transformation. Based on this approach, we study on exact nonnegative matrix factorization when rank=2. We proved that, given two real valued matrices of rank=2, there exists an invertible linear transformation which can transform the real valued matrices to nonnegative matrices with their product stable. We propose an algorithm to find out the transformation. When rank is higher than 2, this kind of transformation may not exist. In the experiments, it's showed that this approach can reach a nonnegative matrix factorization with lower reconstruction error than conventional methods, and the technique for rank=2 exact NMF works well. Naiyang Guan, Xuhui Huang, Zhigang Luo |
SMC | 3 |
| 2015 | Symmetric Non-negative Matrix Factorization Based Link Partition Method for Overlapping Community DetectionabstractPartitioning links rather than nodes is effective in overlapping community detection (OCD) on complex networks. However, it consumes high CPU and memory overheads because the volume of links is huge especially when the network is rather complex. In this paper, we proposes a symmetric non-negative matrix factorization (SNMF) based link partition method called SNMF-Link to overcome this deficiency. In particular, SNMF-Link represents data in a lower-dimensional space spanned by the node-link incidence matrix. By solving a lighter SNMF problem, SNMF-Link learns the clustering indicators of each links. Since traditional multiplicative update rule (MUR) based optimization algorithm for SNMF suffers from slow convergence, we applied the augmented Lagrangian method (ALM) to efficiently optimize SNMF. Experimental results show that SNMF-Link is much more efficient than the representative clustering algorithms without reducing the OCD performance. Xiang Zhang 0008, Naiyang Guan, Wenju Zhang, Xuhui Huang, Shuyi Wu, Zhigang Luo |
SMC | 4 |
| 2015 | Markov State Models Reveal a Two-Step Mechanism of miRNA Loading into the Human Argonaute Protein: Selective Binding followed by Structural Re-arrangementabstractArgonaute (Ago) proteins and microRNAs (miRNAs) are central components in RNA interference, which is a key cellular mechanism for sequence-specific gene silencing. Despite intensive studies, molecular mechanisms of how Ago recognizes miRNA remain largely elusive. In this study, we propose a two-step mechanism for this molecular recognition: selective binding followed by structural re-arrangement. Our model is based on the results of a combination of Markov State Models (MSMs), large-scale protein-RNA docking, and molecular dynamics (MD) simulations. Using MSMs, we identify an open state of apo human Ago-2 in fast equilibrium with partially open and closed states. Conformations in this open state are distinguished by their largely exposed binding grooves that can geometrically accommodate miRNA as indicated in our protein-RNA docking studies. miRNA may then selectively bind to these open conformations. Upon the initial binding, the complex may perform further structural re-arrangement as shown in our MD simulations and eventually reach the stable binary complex structure. Our results provide novel insights in Ago-miRNA recognition mechanisms and our methodology holds great potential to be widely applied in the studies of other important molecular recognition systems. Hanlun Jiang, Fu Kit Sheong, Lizhe Zhu, Xin Gao 0001, Julie Bernauer, Xuhui Huang |
PLoS Comput. Biol. | 6 |
| 2015 | Structural Model of RNA Polymerase II Elongation Complex with Complete Transcription Bubble Reveals NTP Entry RoutesabstractThe RNA polymerase II (Pol II) is a eukaryotic enzyme that catalyzes the synthesis of the messenger RNA using a DNA template. Despite numerous biochemical and biophysical studies, it remains elusive whether the "secondary channel" is the only route for NTP to reach the active site of the enzyme or if the "main channel" could be an alternative. On this regard, crystallographic structures of Pol II have been extremely useful to understand the structural basis of transcription, however, the conformation of the unpaired non-template DNA part of the full transcription bubble (TB) is still unknown. Since diffusion routes of the nucleoside triphosphate (NTP) substrate through the main channel might overlap with the TB region, gaining structural information of the full TB is critical for a complete understanding of Pol II transcription process. In this study, we have built a structural model of Pol II with a complete transcription bubble based on multiple sources of existing structural data and used Molecular Dynamics (MD) simulations together with structural analysis to shed light on NTP entry pathways. Interestingly, we found that although both channels have enough space to allow NTP loading, the percentage of MD conformations containing enough space for NTP loading through the secondary channel is twice higher than that of the main channel. Further energetic study based on MD simulations with NTP loaded in the channels has revealed that the diffusion of the NTP through the main channel is greatly disfavored by electrostatic repulsion between the NTP and the highly negatively charged backbones of nucleotides in the non-template DNA strand. Taken together, our results suggest that the secondary channel is the major route for NTP entry during Pol II transcription. Lu Zhang 0079, Daniel-Adriano Silva, Fátima Pardo Avila, Xuhui Huang |
PLoS Comput. Biol. | 5 |
| 2014 | Translation non-negative matrix factorization with fast optimizationabstractNon-negative matrix factorization (NMF) reconstructs the original samples in a lower dimensional space and has been widely used in pattern recognition and data mining because it usually yields sparse representation. Since NMF leads to unsatisfactory reconstruction for the datasets that contain translations of large magnitude, it is required to develop translation NMF (TNMF) to first remove the translation and then conduct a decomposition. However, existing multiplicative update rule based algorithm for TNMF is not efficient enough. In this paper, we reformulate TNMF and show that it can be efficiently solved by using the state-of-the-art solvers such as NeNMF. Experimental results on face image datasets confirm both efficiency and effectiveness of the reformulated TNMF. Yuanyuan Wang 0004, Naiyang Guan, Bin Mao, Xuhui Huang, Zhigang Luo |
SMC | 4 |
| 2014 | Quantitatively Characterizing the Ligand Binding Mechanisms of Choline Binding Protein Using Markov State Model AnalysisabstractProtein-ligand recognition plays key roles in many biological processes. One of the most fascinating questions about protein-ligand recognition is to understand its underlying mechanism, which often results from a combination of induced fit and conformational selection. In this study, we have developed a three-pronged approach of Markov State Models, Molecular Dynamics simulations, and flux analysis to determine the contribution of each model. Using this approach, we have quantified the recognition mechanism of the choline binding protein (ChoX) to be ∼90% conformational selection dominant under experimental conditions. This is achieved by recovering all the necessary parameters for the flux analysis in combination with available experimental data. Our results also suggest that ChoX has several metastable conformational states, of which an apo-closed state is dominant, consistent with previous experimental findings. Our methodology holds great potential to be widely applied to understand recognition mechanisms underlining many fundamental biological processes. Shuo Gu, Daniel-Adriano Silva, Luming Meng, Alexander Yue, Xuhui Huang |
PLoS Comput. Biol. | 5 |
| 2013 | A Two-State Model for the Dynamics of the Pyrophosphate Ion Release in Bacterial RNA PolymeraseabstractThe dynamics of the PPi release during the transcription elongation of bacterial RNA polymerase and its effects on the Trigger Loop (TL) opening motion are still elusive. Here, we built a Markov State Model (MSM) from extensive all-atom molecular dynamics (MD) simulations to investigate the mechanism of the PPi release. Our MSM has identified a simple two-state mechanism for the PPi release instead of a more complex four-state mechanism observed in RNA polymerase II (Pol II). We observed that the PPi release in bacterial RNA polymerase occurs at sub-microsecond timescale, which is ∼3-fold faster than that in Pol II. After escaping from the active site, the (Mg-PPi)(2-) group passes through a single elongated metastable region where several positively charged residues on the secondary channel provide favorable interactions. Surprisingly, we found that the PPi release is not coupled with the TL unfolding but correlates tightly with the side-chain rotation of the TL residue R1239. Our work sheds light on the dynamics underlying the transcription elongation of the bacterial RNA polymerase. Fátima Pardo Avila, Xuhui Huang |
PLoS Comput. Biol. | 4 |
| 2012 | Graph Based Semi-supervised Non-negative Matrix Factorization for Document ClusteringabstractNon-negative matrix factorization (NMF) approximates a non-negative matrix by the product of two low-rank matrices and achieves good performance in clustering. Recently, semi-supervised NMF (SS-NMF) further improves the performance by incorporating part of the labels of few samples into NMF. In this paper, we proposed a novel graph based SS-NMF (GSS-NMF). For each sample, GSS-NMF minimizes its distances to the same labeled samples and maximizes the distances against different labeled samples to incorporate the discriminative information. Since both labeled and unlabeled samples are embedded in the same reduced dimensional space, the discriminative information from the labeled samples is successfully transferred to the unlabeled samples, and thus it greatly improves the clustering performance. Since the traditional multiplicative update rule converges slowly, we applied the well-known projected gradient method to optimizing GSS-NMF and the proposed algorithm can be applied to optimizing other manifold regularized NMF efficiently. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that GSS-NMF outperforms the representative SS-NMF algorithms. Naiyang Guan, Xuhui Huang, Long Lan, Zhigang Luo, Xiang Zhang 0008 |
ICMLA (1) | 2 |
| 2012 | Semi-supervised Non-negative Patch Alignment FrameworkabstractNon-negative matrix factorization (NMF) learns the latent semantic space more direct and reliable than the latent semantic indexing (LSI) and the spectral clustering methods, thus performs well in document clustering. Recently, semi-supervised NMF such as N2S2L, CNMF and unsupervised method such as GNMF significantly improve the face recognition performance, but they are designed for classification. In this paper, we combine both geometric structure and label information with NMF under the non-negative patch alignment framework (NPAF) to form SS-NPAF. Due to this combination, it greatly improves the clustering performance. To optimize SS-NPAF, we apply the well-known projected gradient method to overcome the slow convergence problem of the mostly used multiplicative update rule. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that SS-NPAF outperforms the representative SS-NMF algorithms. Long Lan, Xuhui Huang, Naiyang Guan, Zhigang Luo, Xiang Zhang 0008 |
ICMLA (1) | 2 |
| 2012 | Multiscale modeling of macromolecular biosystemsabstractIn this article, we review the recent progress in multiresolution modeling of structure and dynamics of protein, RNA and their complexes. Many approaches using both physics-based and knowledge-based potentials have been developed at multiple granularities to model both protein and RNA. Coarse graining can be achieved not only in the length, but also in the time domain using discrete time and discrete state kinetic network models. Models with different resolutions can be combined either in a sequential or parallel fashion. Similarly, the modeling of assemblies is also often achieved using multiple granularities. The progress shows that a multiresolution approach has considerable potential to continue extending the length and time scales of macromolecular modeling. Samuel Flores, Julie Bernauer, Seokmin Shin, Ruhong Zhou, Xuhui Huang |
Briefings Bioinform. | 5 |
| 2011 | A Role for Both Conformational Selection and Induced Fit in Ligand Binding by the LAO ProteinabstractMolecular recognition is determined by the structure and dynamics of both a protein and its ligand, but it is difficult to directly assess the role of each of these players. In this study, we use Markov State Models (MSMs) built from atomistic simulations to elucidate the mechanism by which the Lysine-, Arginine-, Ornithine-binding (LAO) protein binds to its ligand. We show that our model can predict the bound state, binding free energy, and association rate with reasonable accuracy and then use the model to dissect the binding mechanism. In the past, this binding event has often been assumed to occur via an induced fit mechanism because the protein's binding site is completely closed in the bound state, making it impossible for the ligand to enter the binding site after the protein has adopted the closed conformation. More complex mechanisms have also been hypothesized, but these have remained controversial. Here, we are able to directly observe roles for both the conformational selection and induced fit mechanisms in LAO binding. First, the LAO protein tends to form a partially closed encounter complex via conformational selection (that is, the apo protein can sample this state), though the induced fit mechanism can also play a role here. Then, interactions with the ligand can induce a transition to the bound state. Based on these results, we propose that MSMs built from atomistic simulations may be a powerful way of dissecting ligand-binding mechanisms and may eventually facilitate a deeper understanding of allostery as well as the prediction of new protein-ligand interactions, an important step in drug discovery. Daniel-Adriano Silva, Gregory R. Bowman, Alejandro Sosa-Peinado, Xuhui Huang |
PLoS Comput. Biol. | 4 |