Kaizhu Huang

dblp:99/3390 · DBLP profile ↗
← Back
240ranked-venue papers
21as first author
119since 2021 · last 2026
0000-0002-3034-9639ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 193 · 19 first-author · 90 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 1 first-author · 36 since 2021Databases, data management, data science and information retrieval · 32 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Computer networks · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Logic Matters in Lightweight Hallucination Classification for RAG System
abstract
We propose a lightweight, modular framework for hallucination detection in Retrieval-Augmented Generation (RAG) systems, addressing the critical challenge where logical dependencies span across fragmented retrieval results.Through graph-based semantic evidence aggregation, which captures the implicit logical structure by clustering semantically coherent segments across retrieved documents via betweenness centrality, our approach enables small NLI models to handle multihop reasoning without task-specific training.We present two deployment configurations: a resource-efficient variant (≈0.5B parameters) achieving 82.4% accuracy on HotPotQA-Derived at 85 ms latency, outperforming all sub-1B baselines by over 30%; and a higheraccuracy variant (≈1.5B parameters) reaching 85.6%, surpassing 11B TrueTeacher while being 7× smaller and 1.7× faster.Experiments with six NLI discriminator models show consistent gains of +6.7%-+29.9%,confirming that graph-based evidence aggregation is NLIagnostic and the primary performance driver.We also contribute HotPotQA-Derived, a new multi-hop hallucination benchmark preserving separate retrieved documents for systematic evaluation.
Ningyuan Yang, Kaizhu Huang
ACL (1)2
2026 Rethinking Real Image Editing: Unleashing Diverse Editing Operators via Multi-Objective Optimization
abstract
Text-conditioned diffusion models have revolutionized the field of controllable real image editing, enabling high-fidelity and precise image manipulation. Recent methods target specific editing tasks, using internal representations from reconstruction to ensure consistency. Although effective for single tasks, they fail to balance precision and consistency across diverse image editing tasks. In this work, we propose a novel inference-time real-image editing framework that enables executing multiple editing tasks by tuning editing operators. Our key insight is to treat real image editing as a multi-objective optimization problem, optimizing editing operators for a Pareto optimal solution that balances editing accuracy and consistency at each denoising iteration. Additionally, we design a benchmark for operator-guided real-image editing that covers various local and global editing tasks. Extensive experimental evaluations demonstrate the method’s effectiveness in executing precise edits while preserving image fidelity across all tasks, thereby establishing it as the new state-of-the-art.
Xi Yang 0008, Huiru Shao, Rui Zhang 0012, Kaizhu Huang
WACV8
2026 Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Chaolong Yang, Yuyao Yan, Chenru Jiang, Weiguang Zhao, Jie Sun 0024, Bin Dong 0003, Kaizhu Huang
Int. J. Comput. Vis.10
2026 Lena-TRNN: Exploring energy flow for time series prediction
Penglei Gao, Rui Zhang 0012, Xi Yang 0008, Zhuang Qian, Kaizhu Huang
Neural Networks5
2026 You look from old classes: Towards accurate few shot class-incremental learning
Yijie Hu, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
Pattern Recognit.2
2026 A benchmark and method for photographed table reasoning
Xiaoqiang Kang, Xiaochen Zi, Xiao-Bo Jin, Kaizhu Huang, Qiufeng Wang 0001
Pattern Recognit.5
2026 A comprehensive survey of oracle character recognition: Challenges, datasets, methodology, and beyond
Xueke Chi, Qiufeng Wang 0001, Kaizhu Huang, Dahan Wang, Yongge Liu
Pattern Recognit.4
2026 Point2pix-Zero: Point-driven refined diffusion for multi-object image editing
Yuyao Yan, Kaizhu Huang
Pattern Recognit.7
2026 IDEA: Image description enhanced CLIP-adapter for image classification
Zhipeng Ye, Qiufeng Wang 0001, Kaizhu Huang
Pattern Recognit.4
2026 Switch, Reason, and Revise: Enhancing Reasoning Capability of Video Game AI by Large Language Models
abstract
Attributing to the strong reasoning capability, behavior models in artificial intelligence for games play a crucial role in creating gaming experiences. For further enhancing the reasoning capability of behavior models, we propose a novel Switch, Reason, and Revise (SRR) framework, which integrates them with Large Language Model (LLM). The SRR framework contains three core components. The component of Dual-Track Experiential Reasoning fully utilizes the agent experiences for reasoning. The component of Block-Retrieval-Augmented Thoughts adaptively determines the granularity of information retrieval for external sources. The component of Self-Reliant Thinking System Switch increases the reasoning speed by model switching and performs the model switching automatically upon the LLM. Together, these three components can strengthen the agent reasoning capability in complex tasks. Experimental results in the Pokémon battle environment demonstrate the effectiveness and efficiency superiority of SRR over the rival methods. Furthermore, we conduct an exploratory study to reveal the potential of the SRR-empowered agent for guiding new players in Pokémon battle games.
Wei Li 0049, Jiali Lv, Kaizhu Huang, Aiguo Song, Zhen Lei 0001
IEEE Trans. Games4
2026 TPGCA: Transferable Policy Generation and Credit Assignment Network for Cooperative Multiagent Reinforcement Learning
abstract
Multiagent reinforcement learning (MARL) methods have good application performances and prospects in cooperative tasks. To improve the capability of agent policy learning in new scenarios, some methods transfer the learned policy knowledge to new scenarios. However, most methods only focus on the knowledge transfer of individual agent policies, neglecting the credit assignment among agents in cooperative tasks, which results in a transfer bias of cooperative policies. In this paper, we propose a novel method, transferable policy generation and credit assignment (TPGCA) network for cooperative MARL. TPGCA can transfer the entire MARL model by the constructed transferable$Q$-value network and mixing network. Specifically, in TPGCA, to enhance the effectivity and transferability of agent policies, we design the correspondence network between observations and actions (COA) on the basis of transformer and gated recurrent unit (GRU). To implement the reliable credit assignment and diminish the transfer bias, we devise the role-based joint$Q$-value decomposition network (RVD) that can evaluate the contributions of agents from different observation perspectives. Experimental results in various micro-management scenarios on StarCraft multiagent challenge (SMAC) and multiagent particle environment (MPE) sufficiently demonstrate the effectiveness and transferability of TPGCA.
Wei Li 0049, Jiali Lv, Kaizhu Huang, Aiguo Song
IEEE Trans. Comput. Soc. Syst.4
2026 MedMAP: Promoting Incomplete Multi-Modal Brain Tumor Segmentation With Alignment
abstract
Brain tumor segmentation is often based on multiple magnetic resonance imaging (MRI). However, in clinical practice, certain modalities of MRI may be missing, which presents a more difficult scenario. To cope with this challenge, Knowledge Distillation, Domain Adaption, and Shared Latent Space have emerged as commonly promising strategies. However, recent efforts to address the missing modality problem in brain tumor segmentation typically overlook the modality gaps and thus fail to learn important invariant feature representations across different modalities. Such drawback consequently leads to limited performance for missing modality models. To ameliorate these problems, pre-trained models are used in natural visual segmentation tasks to minimize the gaps. However, promising pre-trained models are difficult to obtain in the brain tumor segmentation task due to the lack of sufficient data. Along this line, in this paper, we propose a novel paradigm that aligns latent features of involved modalities to a well-defined distribution anchor as the substitution of the pre-trained model. As a major contribution, we prove that our novel training paradigm ensures a tight evidence lower bound, thus theoretically certifying its effectiveness. Extensive experiments on different backbones validate that the proposed paradigm can enable invariant feature representations and produce models with narrowed modality gaps. Models with our alignment paradigm show their superior performance on both BraTS2018, BraTS2020 and Brain Metastasis datasets.
Zhaorui Tan, Muyin Chen, Xi Yang 0008, Haochuan Jiang, Kaizhu Huang
IEEE J. Biomed. Health Informatics6
2026 Diff-Oracle: Learning Styles and Contents to Augment Realistic Oracle Characters in Diffusion Model
abstract
Recognizing oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains because of the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel multi-modal conditional diffusion model that generates a diverse range of controllable oracle characters by inputting random combinations of references. Given the challenge of accurately describing oracle character styles using natural language, Diff-Oracle departs from traditional diffusion models that rely primarily on text prompts by introducing a style encoder. This encoder extracts style prompts from existing oracle character images, where style details are converted into a text embedding format via a pre-trained language-vision model. Additionally, given the lack of explicit content information for oracle characters, ensuring that generated characters accurately represent the intended glyphs is challenging. Therefore, we pre-generate pixel-level paired oracle character images (i.e., style and content images) by an image-to-image translation model, providing content information for the generation process. Meanwhile, Diff-Oracle integrates a content encoder designed to capture specific content details from content reference images. Extensive experiments on Oracle-241 and OBC306 datasets demonstrate that Diff-Oracle significantly outperforms existing generative methods in image quality and diversity. Moreover, Diff-Oracle substantially benefits downstream recognition tasks, outperforming all existing state-of-the-art methods by a large margin. In particular, on the challenging OBC306 dataset, Diff-Oracle achieves a 7.70% accuracy gain in the zero-shot setting and reaches 84.62% accuracy for unseen oracle characters, setting a new benchmark for oracle character recognition. The code is available at https://github.com/JJJingLi/Diff-Oracle .
Jing Li 0049, Qiufeng Wang 0001, Siyuan Wang 0017, Rui Zhang 0012, Kaizhu Huang, Erik Cambria
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation
abstract
Solving tabular math word problems (TMWPs) has become a critical role in evaluating the mathematical reasoning ability of large language models (LLMs), where large-scale TMWP samples are commonly required for fine-tuning. Since the collection of high-quality TMWP datasets is costly and time-consuming, recent research has concentrated on automatic TMWP generation. However, current generated samples usually suffer from issues of either correctness or diversity. In this paper, we propose a Template-driven LLM-paraphrased (TeLL) framework for generating high-quality TMWP samples with diverse backgrounds and accurate tables, questions, answers, and solutions. To this end, we first extract templates from existing real samples to generate initial problems, ensuring correctness. Then, we adopt an LLM to extend templates and paraphrase problems, obtaining diverse TMWP samples. Furthermore, we find the reasoning annotation is important for solving TMWPs. Therefore, we propose to enrich each solution with illustrative reasoning steps. Through the proposed framework, we construct a high-quality dataset TabMWP-TeLL by adhering to the question types in the TabMWP dataset, and we conduct extensive experiments on a variety of LLMs to demonstrate the effectiveness of TabMWP-TeLL in improving TMWP-solving performance.
Xiaoqiang Kang, Xiao-Bo Jin, Wei Wang 0042, Kaizhu Huang, Qiufeng Wang 0001
AAAI5
2025 GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMs
abstract
With the outstanding capabilities of Large Language Models (LLMs), solving math word problems (MWP) has greatly progressed, achieving higher performance on several benchmark datasets. However, it is more challenging to solve plane geometry problems (PGPs) due to the necessity of understanding, reasoning and computation on two modality data including both geometry diagrams and textual questions, where Multi-Modal Large Language Models (MLLMs) have not been extensively explored. Previous works simply regarded a plane geometry problem as multi-modal QA task, which ignored the importance of explicit parsing geometric elements from problems. To tackle this limitation, we propose to solve plane Geometry problems by Neural-Symbolic reasoning with MLLMs (GNS). We first leverage an MLLM to understand PGPs through knowledge prediction and symbolic parsing, next perform mathematical reasoning to obtain solutions, last adopt a symbolic solver to compute answers. Correspondingly, we introduce the largest PGPs dataset GNS-260K with multiple annotations including symbolic parsing, understanding, reasoning and computation. In experiments, our Phi3-Vision-based MLLM wins the first place on the PGPs solving task of MathVista benchmark, outperforming GPT-4o, Gemini Ultra and other much larger MLLMs. While LLaVA-13B-based MLLM markedly exceeded other close-source and open-source MLLMs on the MathVerse benchmark and also achieved the new SOTA on GeoQA dataset.
Maizhen Ning, Qiufeng Wang 0001, Xiaowei Huang 0001, Kaizhu Huang
AAAI5
2025 Towards Better Robustness Against Natural Corruptions in Document Tampering Localization
abstract
Marvelous advances have been exhibited in recent document tampering localization (DTL) systems. However, confronted with corrupted tampered document images, their vulnerability is fatal in real-world scenarios. While robustness against adversarial attack has been extensively studied by adversarial training (AT), the robustness on natural corruptions remains under-explored for DTL. In this paper, to overcome forensic dependency, we propose the adversarial forensic regularization (AFR) based on min-max optimization to improve robustness. Specifically, we adopt mutual information (MI) to represent forensic dependency between two random variable over tampered and authentic pixels spaces, where the MI can be approximated by Jensen-Shannon-Divergence (JSD) with empirical sampling. To further enable a trade-off between predictive representations in clean tampered document pixels and robust ones in corrupted pixels, an additional regularization term is formulated with divergence between clean and perturbed pixels distribution (DDR). Following min-max optimization framework, our method can also work well against adversarial attacks. To evaluate our proposed method, we collect a dataset (i.e., TSorie-CRP) for evaluating robustness against natural corruptions in real scenarios. Extensive experiments demonstrate the effectiveness of our method against natural corruptions. Without any surprise, our method also achieves good performance against adversarial attack on DTL benchmark datasets.
Huiru Shao, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
AAAI2
2025 Disentangling Tabular Data Towards Better One-Class Anomaly Detection
abstract
Tabular anomaly detection under the one-class classification setting poses a significant challenge, as it involves accurately conceptualizing "normal" derived exclusively from a single category to discern anomalies from normal data variations. Capturing the intrinsic correlation among attributes within normal samples presents one promising method for learning the concept. To do so, the most recent effort relies on a learnable mask strategy with a reconstruction task. However, this wisdom may suffer from the risk of producing uniform masks, i.e., essentially nothing is masked, leading to less effective correlation learning. To address this issue, we presume that attributes related to others in normal samples can be divided into two non-overlapping and correlated subsets, defined as CorrSets, to capture the intrinsic correlation effectively. Accordingly, we introduce an innovative method that disentangles CorrSets from normal tabular data. To our knowledge, this is a pioneering effort to apply the concept of disentanglement for one-class anomaly detection on tabular data. Extensive experiments on 20 tabular datasets show that our method substantially outperforms the state-of-the-art methods and leads to an average performance improvement of 6.1% on AUC-PR and 2.1% on AUC-ROC.
Jianan Ye, Zhaorui Tan, Yijie Hu, Xi Yang 0008, Kaizhu Huang
AAAI6
2025 KMD: Koopman Multi-modality Decomposition for Generalized Brain Tumor Segmentation under Incomplete Modalities
abstract
Magnetic resonance imaging (MRI), with modalities including T1, T2, T1ce, and Flair, providing complementary information critical for sub-region analysis, is widely used for brain tumor diagnosis. However, clinical practice often suffers from varying degrees of incompleteness of necessary modalities due to reasons such as susceptibility to artifacts. It significantly impairs segmentation model performance. Given the limited available modalities at hand, existing approaches attempt to project them into a shared latent space. However, they ignore decomposing the modality-shared and modality-specific information and failed to construct the relationship among different modalities. Such deficiency limits the effectiveness of the segmentation performance, particularly at the time when the amount of data in each modality is different. In this paper, we propose the plug-and-play Koopman Multi-modality Decomposition (KMD) module, leveraging the Koopman Invariant Subspace to disentangle modality-common and modality-specific information. It is capable of constructing modality relationships that minimize bias toward modalities across various modality-incomplete scenarios. More importantly, it can be integrated into several existing backbones feasibility. Through theoretical deductions and extensive empirical experiences on the BraTS2018 and BraTS2020 datasets, we have sufficiently demonstrated the effectiveness of the proposed KMD to promote generalization performance. Code is available at: https://github.com/T-Y-Liu/KMD.
Haochuan Jiang, Kaizhu Huang
CVPR3
2025 PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly Detection
abstract
Point cloud anomaly detection under the anomaly-free setting poses significant challenges as it requires accurately capturing the features of 3D normal data to identify deviations indicative of anomalies. Current efforts focus on devising reconstruction tasks, such as acquiring normal data representations by restoring normal samples from altered, pseudo-anomalous counterparts. Our findings reveal that distributing attention equally across normal and pseudo-anomalous data tends to dilute the model’s focus on anomalous deviations. The challenge is further compounded by the inherently disordered and sparse nature of 3D point cloud data. In response to those predicaments, we introduce an innovative approach that emphasizes learning point offsets, targeting more informative pseudo-abnormal points, thus fostering more effective distillation of normal data representations. We also have crafted an augmentation technique that is steered by normal vectors, facilitating the creation of credible pseudo anomalies that enhance the efficiency of the training process. Our comprehensive experimental evaluation on the Anomaly-ShapeNet and Real3DAD datasets evidences that our proposed method outperforms existing state-of-the-art approaches, achieving an average enhancement of 9.0% and 1.4% in the AUC-ROC detection metric across these datasets, respectively. Code is available at https://github.com/yjnanan/PO3AD.
Jianan Ye, Weiguang Zhao, Xi Yang 0008, Kaizhu Huang
CVPR5
2025 BFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysis
abstract
3D semantic segmentation plays a fundamental and crucial role to understand 3D scenes. While contemporary state-of-the-art techniques predominantly concentrate on elevating the overall performance of 3D semantic segmentation based on general metrics (e.g. mIoU, mAcc, and oAcc), they unfortunately leave the exploration of challenging regions for segmentation mostly neglected. In this paper, we revisit 3D semantic segmentation through a more granular lens, shedding light on subtle complexities that are typically overshadowed by broader performance metrics. Concretely, we have delineated 3D semantic segmentation errors into four comprehensive categories as well as corresponding evaluation metrics tailored to each. Building upon this categorical framework, we introduce an innovative 3D semantic segmentation network called BFANet that incorporates detailed analysis of semantic boundary features. First, we design the boundary-semantic module to decouple point cloud features into semantic and boundary features, and fuse their query queue to enhance semantic features with attention. Second, we introduce a more concise and accelerated boundary pseudo-label calculation algorithm, which is 3.9 times faster than the state-of-the-art, offering compatibility with data augmentation and enabling efficient computation in training. Extensive experiments on benchmark data indicate the superiority of our BFANet model, confirming the significance of emphasizing the four uniquely designed metrics. Code is available at https://github.com/weiguangzhao/BFANet.
Weiguang Zhao, Rui Zhang 0012, Qiufeng Wang 0001, Kaizhu Huang
CVPR5
2025 Can GRPO Boost Complex Multimodal Table Understanding?
abstract
Xiaoqiang Kang, Shengen Wu, Zimu Wang, Yilin Liu, Xiaobo Jin, Kaizhu Huang, Wei Wang, Yutao Yue, Xiaowei Huang, Qiufeng Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xiaoqiang Kang, Shengen Wu, Xiao-Bo Jin, Kaizhu Huang, Wei Wang 0042, Yutao Yue
EMNLP6
2025 Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant Representation
Zhaorui Tan, Xi Yang 0008, Tan Pan, Chen Jiang 0006, Xin Guo 0010, Qiufeng Wang 0001, Anh Nguyen 0003, Yuan Qi 0001, Kaizhu Huang
ICCV10
2025 The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
Kaizhu Huang, Qiufeng Wang 0001, Xiao-Bo Jin
ICDM4
2025 ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25.
Zihan Ye, Shreyank N. Gowda, Shiming Chen 0002, Xiaowei Huang 0001, Fahad Shahbaz Khan, Yaochu Jin, Kaizhu Huang, Xiao-Bo Jin
ICLR8
2025 Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
abstract
Exceptional mathematical reasoning ability is one of the key features that demonstrate the power of large language models (LLMs). How to comprehensively define and evaluate the mathematical abilities of LLMs, and even reflect the user experience in real-world scenarios, has emerged as a critical issue. Current benchmarks predominantly concentrate on problem-solving capabilities, presenting a substantial risk of model overfitting and fails to accurately measure the genuine mathematical reasoning abilities. In this paper, we argue that if a model really understands a problem, it should be robustly and readily applied across a diverse array of tasks. To this end, we introduce MathCheck, a well-designed checklist for testing task generalization and reasoning robustness, as well as an automatic tool to generate checklists efficiently. MathCheck includes multiple mathematical reasoning tasks and robustness tests to facilitate a comprehensive evaluation of both mathematical reasoning ability and behavior testing. Utilizing MathCheck, we develop MathCheck-GSM and MathCheck-GEO to assess mathematical textual reasoning and multi-modal reasoning capabilities, respectively, serving as upgraded versions of benchmarks including GSM8k, GeoQA, UniGeo, and Geometry3K. We adopt MathCheck-GSM and MathCheck-GEO to evaluate over 26 LLMs and 17 multi-modal LLMs, assessing their comprehensive mathematical reasoning abilities. Our results demonstrate that while frontier LLMs like GPT-4o continue to excel in various abilities on the checklist, many other model families exhibit a significant decline. Further experiments indicate that, compared to traditional math benchmarks, MathCheck better reflects true mathematical abilities and represents mathematical intelligence more linearly, thereby supporting our design. Using MathCheck, we can also efficiently conduct informative behavior analysis to deeply investigate models. Finally, we show that our proposed checklist paradigm can easily extend to other reasoning tasks for their comprehensive evaluation.
Shudong Liu 0004, Maizhen Ning, Wei Liu 0131, Jindong Wang 0001, Derek F. Wong, Xiaowei Huang 0001, Qiufeng Wang 0001, Kaizhu Huang
ICLR9
2025 From 2D Images to 3D Model: Weakly Supervised Multi-View Face Reconstruction with Deep Fusion
abstract
While weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models. In this regard, we propose a novel pipeline called Deep Fusion MVR (DF-MVR) to explore the feature correspondences between multi-view images and reconstruct high-precision 3D faces. Specifically, we present a novel multi-view feature fusion backbone that utilizes face masks to align features from multiple encoders and integrates one multi-layer attention mechanism to enhance feature interaction and fusion, resulting in one unified facial representation. Additionally, we develop one concise face mask mechanism that facilitates multi-view feature fusion and facial reconstruction by identifying common areas and guiding the network’s focus on critical facial features (e.g., eyes, brows, nose, and mouth). Experiments on Pixel-Face and Bosphorus datasets indicate the superiority of the proposed method. Without the 3D annotation, DF-MVR achieves relative 5.2% and 3.0% RMSE improvement over the existing weakly supervised MVRs, respectively, on Pixel-Face and Bosphorus datasets. Our code is available at https://github.com/weiguangzhao/DF_MVR.
Weiguang Zhao, Chaolong Yang, Jianan Ye, Rui Zhang 0012, Yuyao Yan, Xi Yang 0008, Bin Dong 0003, Amir Hussain 0001, Kaizhu Huang
ICME9
2025 Entropy-Guided Distillation for Medical Image Segmentation Under Missing Modalities
Yuyao Yan, Xi Yang 0008, Kaizhu Huang
ICONIP (5)4
2025 Defending Against Jailbreak Through Early Exit Generation of Large Language Models
Chongwen Zhao, Zhihao Dou, Kaizhu Huang
ICONIP (5)3
2025 Towards Training-Free Open-World Classification with 3D Generative Models
abstract
3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring robust subsequent knowledge adaptation capabilities. While current approaches predominantly rely on 2D pre-trained models through 3D-to-2D projection, their performance degrades severely under arbitrary object orientations. Unlike these present efforts, this work makes a pioneering exploration of 3D generative models for 3D open-world classification-specifically, leverageing the accumulated prior knowledge from these models to provide anchors for novel categories, while integrating a rotation-invariant feature extractor. This innovative synergy endows our pipeline with the advantages of being training-free and pose-invariant, thus well suited to adapt novel categories in 3D open-world classification. Extensive experiments on benchmark datasets demonstrate the potential of this pipeline, achieving state-of-the-art performance on ModelNet10‡ and McGill‡ with 32.7% and 8.7% overall accuracy improvement, respectively. The code is available in the supplementary materials.
Xinzhe Xia, Weiguang Zhao, Yuyao Yan, Guanyu Yang 0002, Rui Zhang 0012, Kaizhu Huang, Xi Yang 0008
ACM Multimedia6
2025 KDTalker++: Controllable Talking Portrait Generation with Audio, Text, and Expression Editing
abstract
This work presents KDTalker++, a real-time system for generating talking portrait videos from a single image using audio or text input. Built on a keypoint-based spatiotemporal diffusion model, it adds voice cloning, background editing, and fine-grained expression control. The demo is available at https://kdtalker.com. A live presentation video is available at https://drive.google.com/file/d/1N4Ggu0Y32DTsV3mbKhXS4kGY6l2ZYKSp/view.
Chaolong Yang, Yinuo Guo, Yuyao Yan, Jie Sun 0024, Kaizhu Huang
ACM Multimedia6
2025 Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
abstract
Math reasoning has been one crucial ability of large language models (LLMs), where significant advancements have been achieved in recent years. However, most efforts focus on LLMs by curating high-quality annotation data and intricate training (or inference) paradigms, while the math reasoning performance of multi-modal LLMs (MLLMs) remains lagging behind. Since the MLLM typically consists of an LLM and vision block, we wonder: \textit{Can MLLMs directly absorb math reasoning abilities from off-the-shelf math LLMs without tuning?} Recent model-merging approaches may offer insights into this question. However, they overlook the alignment between the MLLM and LLM, where we find that there is a large gap between their parameter spaces, resulting in lower performance. Our empirical evidence reveals two key factors behind this issue: the identification of crucial reasoning-associated layers in the model and the mitigation of the gaps in parameter space. Based on the empirical insights, we propose \textbf{IP-Merging} that first \textbf{I}dentifies the reasoning-associated parameters in both MLLM and Math LLM, then \textbf{P}rojects them into the subspace of MLLM aiming to maintain the alignment, finally merges parameters in this subspace. IP-Merging is a tuning-free approach since parameters are directly adjusted. Extensive experiments demonstrate that our IP-Merging method can enhance the math reasoning ability of MLLMs directly from Math LLMs without compromising their other capabilities.
Yijie Hu, Kaizhu Huang, Xiaowei Huang 0001
NeurIPS3
2025 DvD: Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion Model
abstract
Document dewarping aims to rectify deformations in photographic document images, thus improving text readability, which has attracted much attention and made great progress, but it is still challenging to preserve document structures. Given recent advances in diffusion models, it is natural for us to consider their potential applicability to document dewarping. However, it is far from straightforward to adopt diffusion models in document dewarping due to their unfaithful control on highly complex document images (e.g., 2000 × 3000 resolution). In this paper, we propose DvD, the first generative model to tackle document Dewarping via a Diffusion framework. To be specific, DvD introduces a coordinate-level denoising instead of typical pixel-level denoising, generating a mapping for deformation rectification. In addition, we further propose a time-variant condition refinement mechanism to enhance the preservation of document structures. In experiments, we find that current document dewarping benchmarks can not evaluate dewarping models comprehensively. To this end, we present AnyPhotoDoc6300, a rigorously designed large-scale document dewarping benchmark comprising 6,300 real image pairs across three distinct domains, enabling fine-grained evaluation of dewarping models. Comprehensive experiments demonstrate that our proposed DvD can achieve state-of-the-art performance with acceptable computational efficiency on multiple metrics across various benchmarks, including DocUNet, DIR300, and AnyPhotoDoc6300. The new benchmark and code will be publicly available at https://github.com/hanquansanren/DvD.
Huangcheng Lu, Maizhen Ning, Xiaowei Huang 0001, Wei Wang 0042, Kaizhu Huang, Qiufeng Wang 0001
SIGGRAPH Asia6
2025 Covariance-Based Space Regularization for Few-Shot Class Incremental Learning
Yijie Hu, Guanyu Yang 0002, Zhaorui Tan, Xiaowei Huang 0001, Kaizhu Huang, Qiufeng Wang 0001
WACV5
2025 Prompt-Enhanced: Leveraging language representation for prompt continual learning
Wei Li 0049, Shitong Shao, Kaizhu Huang, Zhen Lei 0001
Neural Networks4
2025 Stagger Network: Rethinking information loss in medical image segmentation with various-sized targets
Zhaorui Tan, Haochuan Jiang, Kaizhu Huang
Neural Networks4
2025 Open-Pose 3D zero-shot learning: Benchmark and challenges
Weiguang Zhao, Guanyu Yang 0002, Rui Zhang 0012, Chenru Jiang, Chaolong Yang, Yuyao Yan, Amir Hussain 0001, Kaizhu Huang
Neural Networks8
2025 Revisiting 3D point cloud analysis with Markov process
Chenru Jiang, Wuwei Ma, Kaizhu Huang, Qiufeng Wang 0001, Xi Yang 0008, Weiguang Zhao, Junwei Wu 0001, Xinheng Wang 0001, Jimin Xiao, Zhenxing Niu
Pattern Recognit.3
2025 SCMix: Stochastic Compound Mixing for Open Compound Domain Adaptation in Semantic Segmentation
abstract
Open compound domain adaptation (OCDA) aims to transfer knowledge from a labeled source domain to a mix of unlabeled homogeneous compound target domains while generalizing to open unseen domains. Existing OCDA methods solve the intradomain gaps by a divide-and-conquer strategy, which decomposes the problem into several individual and parallel domain adaptation (DA) tasks. In this work, starting from the general DA theory, we establish a novel generalization bound for the setting of OCDA. Built upon this, we argue that conventional OCDA approaches may substantially underestimate the inherent variance inside the compound target domains for model generalization, constraining the model's performance. We subsequently present stochastic compound mixing (SCMix), an augmentation strategy with the primary objective of mitigating the divergence between the source and mixed target distributions. Theoretical analyses are conducted to substantiate the superiority of SCMix, proving that single-target mixing is a subgroup of our method. Extensive experiments show that our method attains a lower empirical risk on OCDA semantic segmentation tasks, thus supporting our theories. In particular, combining the transformer architecture, SCMix achieves a notable performance boost compared to SoTA results.
Zhaorui Tan, Zixian Su, Xi Yang 0008, Jie Sun 0024, Kaizhu Huang
IEEE Trans. Neural Networks Learn. Syst.6
2024 Unraveling Batch Normalization for Realistic Test-Time Adaptation
abstract
While recent test-time adaptations exhibit efficacy by adjusting batch normalization to narrow domain disparities, their effectiveness diminishes with realistic mini-batches due to inaccurate target estimation. As previous attempts merely introduce source statistics to mitigate this issue, the fundamental problem of inaccurate target estimation still persists, leaving the intrinsic test-time domain shifts unresolved. This paper delves into the problem of mini-batch degradation. By unraveling batch normalization, we discover that the inexact target statistics largely stem from the substantially reduced class diversity in batch. Drawing upon this insight, we introduce a straightforward tool, Test-time Exponential Moving Average (TEMA), to bridge the class diversity gap between training and testing batches. Importantly, our TEMA adaptively extends the scope of typical methods beyond the current batch to incorporate a diverse set of class information, which in turn boosts an accurate target estimation. Built upon this foundation, we further design a novel layer-wise rectification strategy to consistently promote test-time performance. Our proposed method enjoys a unique advantage as it requires neither training nor tuning parameters, offering a truly hassle-free solution. It significantly enhances model robustness against shifted domains and maintains resilience in diverse real-world scenarios with various batch sizes, achieving state-of-the-art performance on several major benchmarks. Code is available at https://github.com/kiwi12138/RealisticTTA.
Zixian Su, Jingwei Guo 0001, Xi Yang 0008, Qiufeng Wang 0001, Kaizhu Huang
AAAI6
2024 Semantic-Aware Data Augmentation for Text-to-Image Synthesis
abstract
Data augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the semantic mismatch between augmented paired data. Even worse, semantic collapse may occur when generated images are less semantically constrained. In this paper, we develop a novel Semantic-aware Data Augmentation (SADA) framework dedicated to T2Isyn. In particular, we propose to augment texts in the semantic space via an Implicit Textual Semantic Preserving Augmentation, in conjunction with a specifically designed Image Semantic Regularization Loss as Generated Image Semantic Conservation, to cope well with semantic mismatch and collapse. As one major contribution, we theoretically show that Implicit Textual Semantic Preserving Augmentation can certify better text-image consistency while Image Semantic Regularization Loss regularizing the semantics of generated images would avoid semantic collapse and enhance image quality. Extensive experiments validate that SADA enhances text-image consistency and improves image quality significantly in T2Isyn models across various backbones. Especially, incorporating SADA during the tuning process of Stable Diffusion models also yields performance improvements.
Zhaorui Tan, Xi Yang 0008, Kaizhu Huang
AAAI3
2024 MathAttack: Attacking Large Language Models towards Math Solving Ability
abstract
With the boom of Large Language Models (LLMs), the research of solving Math Word Problem (MWP) has recently made great progress. However, there are few studies to examine the robustness of LLMs in math solving ability. Instead of attacking prompts in the use of LLMs, we propose a MathAttack model to attack MWP samples which are closer to the essence of robustness in solving math problems. Compared to traditional text adversarial attack, it is essential to preserve the mathematical logic of original MWPs during the attacking. To this end, we propose logical entity recognition to identify logical entries which are then frozen. Subsequently, the remaining text are attacked by adopting a word-level attacker. Furthermore, we propose a new dataset RobustMath to evaluate the robustness of LLMs in math solving ability. Extensive experiments on our RobustMath and two another math benchmark datasets GSM8K and MultiAirth show that MathAttack could effectively attack the math solving ability of LLMs. In the experiments, we observe that (1) Our adversarial samples from higher-accuracy LLMs are also effective for attacking LLMs with lower accuracy (e.g., transfer from larger to smaller-size LLMs, or from few-shot to zero-shot prompts); (2) Complex MWPs (such as more solving steps, longer text, more numbers) are more vulnerable to attack; (3) We can improve the robustness of LLMs by using our adversarial samples in few-shot prompts. Finally, we hope our practice and observation can serve as an important attempt towards enhancing the robustness of LLMs in math solving ability. The code and dataset is available at: https://github.com/zhouzihao501/MathAttack.
Qiufeng Wang 0001, Mingyu Jin, Jianan Ye, Wei Liu 0131, Wei Wang 0042, Xiaowei Huang 0001, Kaizhu Huang
AAAI9
2024 Mind the Gap: Promoting Missing Modality Brain Tumor Segmentation with Alignment
abstract
Brain tumor segmentation is often based on multiple magnetic resonance imaging (MRI). However, in clinical practice, certain modalities of MRI may be missing, which presents an even more difficult scenario. To cope with this challenge, knowledge distillation has emerged as one promising strategy. However, recent efforts typically overlook the modality gaps and thus fail to learn invariant feature representations across different modalities. Such drawback consequently leads to limited performance for both teachers and students. To ameliorate these problems, in this paper, we propose a novel paradigm that aligns latent features of involved modalities to a well-defined distribution anchor. As a major contribution, we prove that our novel training paradigm ensures a tight evidence lower bound, thus theoretically certifying its effectiveness. Extensive experiments on different backbones validate that the proposed paradigm can enable invariant feature representations and produce a teacher with narrowed modality gaps. This further offers superior guidance for missing modality students, achieving an average improvement of 1.75 on dice score.
Zhaorui Tan, Haochuan Jiang, Xi Yang 0008, Kaizhu Huang
BIBM5
2024 Rethinking Multi-Domain Generalization with A General Learning Objective
abstract
Multi-domain generalization$(mDG)$is universally aimed to minimize the discrepancy between training and testing distributions to enhance marginal-to-label distribution mapping. However, existing$mDG$literature lacks a general learning objective paradigm and often imposes constraints on static target marginal distributions. In this paper, we propose to leverage a Y-mapping to relax the constraint. We rethink the learning objective for$mDG$and design a new general learning objective to interpret and analyze most existing$mDG$wisdom. This general objective is bifurcated into two synergistic amis: learning domain-independent conditional features and maximizing a posterior. Explorations also extend to two effective regularization terms that incorporate prior information and suppress invalid causality, alleviating the issues that come with relaxed constraints. We theoretically contribute an upper bound for the domain alignment of domain-independent conditional features, disclosing that many previous$mDG$endeavors actually optimize partially the objective and thus lead to limited performance. As such, our study distills a general learning objective into four practical components, providing a general, robust, and flexible mechanism to handle complex domain shifts. Extensive empirical results indicate that the proposed objective with Y -mapping leads to substantially better$mDG$performance in various downstream tasks, including regression, segmentation, and classification. Code is available at htttps://github.com/zhaorui-t.an/GMDG/tree/main.
Zhaorui Tan, Xi Yang 0008, Kaizhu Huang
CVPR3
2024 Delving into Adversarial Robustness on Document Tampering Localization
Huiru Shao, Zhuang Qian, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
ECCV (65)3
2024 Class Incremental Learning for Character String Recognition
Yijie Hu, Yan-Ming Zhang 0001, Kaizhu Huang, Qiufeng Wang 0001
ICDAR (5)3
2024 Coarse-to-Fine Document Image Registration for Dewarping
Qiufeng Wang 0001, Kaizhu Huang, Xiaomeng Gu, Fengjun Guo
ICDAR (4)3
2024 Lite-SVO: Towards A Lightweight Self-Supervised Semantic Visual Odometry Exploiting Multi-Feature Sharing Architecture
abstract
Not relying on ground-truth data for training, self-supervised semantic visual odometry (SVO) has recently gained considerable attention. Within self-supervised SVO, feature representation inconsistency between semantic/depth and pose tasks presents a significant challenge, as it may disrupt cross-task feature representations and lead to notable performance degradation. Regrettably, existing self-supervised SVO lacks an effective solution to address this obstacle, for either overlooking this issue or exploiting a too heavy architecture. In response to this challenge, we propose a groundbreaking solution within the Single-Stream architecture, known as Lite-SVO, which is a lightweight yet efficient multi-feature sharing architecture. Lite-SVO is designed to bolster self-supervised SVO, facilitating its adoption on edge devices without compromising accuracy and performance. The crucial innovation lies in the multi-feature sharing architecture, which fuses the semantic and depth maps as pose features, thus significantly reducing the model complexity and boosting the speed in edge devices. Built upon the novel feature sharing framework, Lite-SVO further optimizes the feature sharing representation to improve the performance. Specifically, a cross-feature sharing module alleviates the impact of object boundary in depth estimation, while a multi-feature sharing module focuses on extracting and fusing spatial features to enhance pose estimation. Experimental results demonstrate that our method is at least 84.46% faster than the state-of-the-art Single-Stream approaches, and excitingly, our method’s pose accuracy is about 79.83% higher than theirs.
Wenhui Wei, Kaizhu Huang, Jiadong Li, Xin Liu 0102, Yangfan Zhou 0004
ICRA3
2024 Document Registration: Towards Automated Labeling of Pixel-Level Alignment Between Warped-Flat Documents
Qiufeng Wang 0001, Kaizhu Huang, Xiaowei Huang 0001, Fengjun Guo, Xiaomeng Gu
ACM Multimedia3
2024 Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
abstract
Vision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A logical regularization termed L-Reg is derived which bridges a logical analysis framework to image classification. Our work reveals that L-Reg reduces the complexity of the model in terms of the feature distribution and classifier weights. Specifically, we unveil the interpretability brought by L-Reg, as it enables the model to extract the salient features, such as faces to persons, for classification. Theoretical analysis and experiments demonstrate that L-Reg enhances generalization across various scenarios, including multi-domain generalization and generalized category discovery. In complex real-world scenarios where images span unknown classes and unseen domains, L-Reg consistently improves generalization, highlighting its practical efficacy.
Zhaorui Tan, Xi Yang 0008, Qiufeng Wang 0001, Anh Nguyen 0003, Kaizhu Huang
NeurIPS5
2024 Inter-feature Relationship Certifies Robust Generalization of Adversarial Training
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Bin Gu 0001, Huan Xiong, Xinping Yi
Int. J. Comput. Vis.3
2024 Zero-shot text classification with knowledge resources under label-fully-unseen setting
Wei Wang 0042, Qi Chen 0026, Kaizhu Huang, Anh Nguyen 0003, Suparna De
Neurocomputing4
2024 Instance-Specific Model Perturbation Improves Generalized Zero-Shot Learning
abstract
Zero-shot learning (ZSL) refers to the design of predictive functions on new classes (unseen classes) of data that have never been seen during training. In a more practical scenario, generalized zero-shot learning (GZSL) requires predicting both seen and unseen classes accurately. In the absence of target samples, many GZSL models may overfit training data and are inclined to predict individuals as categories that have been seen in training. To alleviate this problem, we develop a parameter-wise adversarial training process that promotes robust recognition of seen classes while designing during the test a novel model perturbation mechanism to ensure sufficient sensitivity to unseen classes. Concretely, adversarial perturbation is conducted on the model to obtain instance-specific parameters so that predictions can be biased to unseen classes in the test. Meanwhile, the robust training encourages the model robustness, leading to nearly unaffected prediction for seen classes. Moreover, perturbations in the parameter space, computed from multiple individuals simultaneously, can be used to avoid the effect of perturbations that are too extreme and ruin the predictions. Comparison results on four benchmark ZSL data sets show the effective improvement that the proposed framework made on zero-shot methods with learned metrics.
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, Xi Yang 0008
Neural Comput.2
2024 Perturbation diversity certificates robust generalization
Zhuang Qian, Shufei Zhang, Kaizhu Huang, Qiufeng Wang 0001, Xinping Yi, Bin Gu 0001, Huan Xiong
Neural Networks3
2024 ES-GNN: Generalizing Graph Neural Networks Beyond Homophily With Edge Splitting
abstract
While Graph Neural Networks (GNNs) have achieved enormous success in multiple graph analytical tasks, modern variants mostly rely on the strong inductive bias of homophily. However, real-world networks typically exhibit both homophilic and heterophilic linking patterns, wherein adjacent nodes may share dissimilar attributes and distinct labels. Therefore, GNNs smoothing node proximity holistically may aggregate both task-relevant and irrelevant (even harmful) information, limiting their ability to generalize to heterophilic graphs and potentially causing non-robustness. In this work, we propose a novel Edge Splitting GNN (ES-GNN) framework to adaptively distinguish between graph edges either relevant or irrelevant to learning tasks. This essentially transfers the original graph into two subgraphs with the same node set but complementary edge sets dynamically. Given that, information propagation separately on these subgraphs and edge splitting are alternatively conducted, thus disentangling the task-relevant and irrelevant features. Theoretically, we show that our ES-GNN can be regarded as a solution to a disentangled graph denoising problem, which further illustrates our motivations and interprets the improved generalization beyond homophily. Extensive experiments over 11 benchmark and 1 synthetic datasets not only demonstrate the effective performance of ES-GNN but also highlight its robustness to adversarial graphs and mitigation of the over-smoothing problem.
Jingwei Guo 0001, Kaizhu Huang, Rui Zhang 0012, Xinping Yi
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 SaliencyCut: Augmenting plausible anomalies for anomaly detection
Jianan Ye, Yijie Hu, Xi Yang 0008, Qiufeng Wang 0001, Kaizhu Huang
Pattern Recognit.6
2024 A generalizable framework for low-rank tensor completion with numerical priors
Shiran Yuan, Kaizhu Huang
Pattern Recognit.2
2024 EgPDE-Net: Building Continuous Neural Networks for Time Series Prediction With Exogenous Variables
abstract
While exogenous variables have a major impact on performance improvement in time series analysis, interseries correlation and time dependence among them are rarely considered in the present continuous methods. The dynamical systems of multivariate time series could be modeled with complex unknown partial differential equations (PDEs) which play a prominent role in many disciplines of science and engineering. In this article, we propose a continuous-time model for arbitrary-step prediction to learn an unknown PDE system in multivariate time series whose governing equations are parameterized by self-attention and gated recurrent neural networks. The proposed model, exogenous-guided PDE network (EgPDE-Net), takes account of the relationships among the exogenous variables and their effects on the target series. Importantly, the model can be reduced into a regularized ordinary differential equation (ODE) problem with specially designed regularization guidance, which makes the PDE problem tractable to obtain numerical solutions and feasible to predict multiple future values of the target series at arbitrary time points. Extensive experiments demonstrate that our proposed model could achieve competitive accuracy over strong baselines: on average, it outperforms the best baseline by reducing 9.85% on RMSE and 13.98% on MAE for arbitrary-step prediction.
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Ping Guo 0002, John Yannis Goulermas, Kaizhu Huang
IEEE Trans. Cybern.6
2024 Learning Disentangled Graph Convolutional Networks Locally and Globally
abstract
Graph convolutional networks (GCNs) emerge as the most successful learning models for graph-structured data. Despite their success, existing GCNs usually ignore the entangled latent factors typically arising in real-world graphs, which results in nonexplainable node representations. Even worse, while the emphasis has been placed on local graph information, the global knowledge of the entire graph is lost to a certain extent. In this work, to address these issues, we propose a novel framework for GCNs, termed LGD-GCN, taking advantage of both local and global information for disentangling node representations in the latent space. Specifically, we propose to represent a disentangled latent continuous space with a statistical mixture model, by leveraging neighborhood routing mechanism locally. From the latent space, various new graphs can then be disentangled and learned, to overall reflect the hidden structures with respect to different factors. On the one hand, a novel regularizer is designed to encourage interfactor diversity for model expressivity in the latent space. On the other hand, the factor-specific information is encoded globally via employing a message passing along these new graphs, in order to strengthen intrafactor consistency. Extensive evaluations on both synthetic and five benchmark datasets show that LGD-GCN brings significant performance gains over the recent competitive models in both disentangling and node classification. Particularly, LGD-GCN is able to outperform averagely the disentangled state-of-the-arts by 7.4% on social network datasets.
Jingwei Guo 0001, Kaizhu Huang, Xinping Yi, Rui Zhang 0012
IEEE Trans. Neural Networks Learn. Syst.2
2024 Can Perturbations Help Reduce Investment Risks? Risk-aware Stock Recommendation via Split Variational Adversarial Training
abstract
In the stock market, a successful investment requires a good balance between profits and risks. Based on the learning to rank paradigm, stock recommendation has been widely studied in quantitative finance to recommend stocks with higher return ratios for investors. Despite the efforts to make profits, many existing recommendation approaches still have some limitations in risk control, which may lead to intolerable paper losses in practical stock investing. To effectively reduce risks, we draw inspiration from adversarial learning and propose a novel Split Variational Adversarial Training (SVAT) method for risk-aware stock recommendation. Essentially, SVAT encourages the stock model to be sensitive to adversarial perturbations of risky stock examples and enhances the model’s risk awareness by learning from perturbations. To generate representative adversarial examples as risk indicators, we devise a variational perturbation generator to model diverse risk factors. Particularly, the variational architecture enables our method to provide a rough risk quantification for investors, showing an additional advantage of interpretability. Experiments on several real-world stock market datasets demonstrate the superiority of our SVAT method. By lowering the volatility of the stock-recommendation model, SVAT effectively reduces investment risks and outperforms state-of-the-art baselines by more than 30% in terms of risk-adjusted profits. All the experimental data and source code are available at https://drive.google.com/drive/folders/14AdM7WENEvIp5x5bV3zV_i4Aev21C9g6?usp=sharing .
Jiezhu Cheng, Kaizhu Huang, Zibin Zheng
ACM Trans. Inf. Syst.2
2024 Continuous Image Outpainting with Neural ODE
abstract
Generalised image outpainting is an important and active research topic in computer vision, which aims to extend appealing content all-side around a given image. Existing state-of-the-art outpainting methods often rely on discrete extrapolation to extend the feature map in the bottleneck. They thus suffer from content unsmoothness, especially in circumstances where the outlines of objects in the extrapolated regions are incoherent with the input sub-images. To mitigate this issue, we design a novel bottleneck with Neural ODEs to make continuous extrapolation in latent space, which could be a plug-in for many deep learning frameworks. Our ODE-based network continuously transforms the state and makes accurate predictions by learning the incremental relationship among latent points, leading to both smooth and structured feature representation. Experimental results on three real-world datasets both applied on transformer-based and CNN-based frameworks show that our methods could generate more realistic and coherent images against the state-of-the-art image outpainting approaches. Our code is available at https://github.com/PengleiGao/Continuous-Image-Outpainting-with-Neural-ODE .
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Kaizhu Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Scene Text Recognition via Dual-path Network with Shape-driven Attention Alignment
abstract
Scene text recognition (STR), one typical sequence-to-sequence problem, has drawn much attention recently in multimedia applications. To guarantee good performance, it is essential for STR to obtain aligned character-wise features from the whole-image feature maps. While most present works adopt fully data-driven attention-based alignment, such practice ignores specific character geometric information. In this article, built upon a group of learnable geometric points, we propose a novel shape-driven attention alignment method that is able to obtain character-wise features. Concretely, we first design a corner detector to generate a shape map to guide the attention alignments explicitly, where a series of points can be learned to represent character-wise features flexibly. We then propose a dual-path network with a mutual learning and cooperating strategy that successfully combines CNN with a ViT-based model, leading to further accuracy improvement. We conduct extensive experiments to evaluate the proposed method on various scene text benchmarks, including six popular regular and irregular datasets, two more challenging datasets (i.e., WordArt and OST), and three Chinese datasets. Experimental results indicate that our method can achieve superior performance with a comparable model size against many state-of-the-art models.
Yijie Hu, Bin Dong 0003, Kaizhu Huang, Lei Ding 0012, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation
abstract
Single-source domain generalization (SDG) in medical image segmentation is a challenging yet essential task as domain shifts are quite common among clinical image datasets. Previous attempts most conduct global-only/random augmentation. Their augmented samples are usually insufficient in diversity and informativeness, thus failing to cover the possible target domain distribution. In this paper, we rethink the data augmentation strategy for SDG in medical image segmentation. Motivated by the class-level representation invariance and style mutability of medical images, we hypothesize that unseen target data can be sampled from a linear combination of C (the class number) random variables, where each variable follows a location-scale distribution at the class level. Accordingly, data augmented can be readily made by sampling the random variables through a general form. On the empirical front, we implement such strategy with constrained Bezier transformation on both global and local (i.e. class-level) regions, which can largely increase the augmentation diversity. A Saliency-balancing Fusion mechanism is further proposed to enrich the informativeness by engaging the gradient information, guiding augmentation with proper orientation and magnitude. As an important contribution, we prove theoretically that our proposed augmentation can lead to an upper bound of the generalization risk on the unseen target domain, thus confirming our hypothesis. Combining the two strategies, our Saliency-balancing Location-scale Augmentation (SLAug) exceeds the state-of-the-art works by a large margin in two challenging SDG tasks. Code is available at https://github.com/Kaiseem/SLAug.
Zixian Su, Xi Yang 0008, Kaizhu Huang, Qiufeng Wang 0001, Jie Sun 0024
AAAI4
2023 Explore Epistemic Uncertainty in Domain Adaptive Semantic Segmentation
abstract
In domain adaptive segmentation, domain shift may cause erroneous high-confidence predictions on the target domain, resulting in poor self-training. To alleviate the potential error, most previous works mainly consider aleatoric uncertainty arising from the inherit data noise. This may however lead to overconfidence in incorrect predictions and thus limit the performance. In this paper, we take advantage of Deterministic Uncertainty Methods (DUM) to explore the epistemic uncertainty, which reflects accurately the domain gap depending on the model choice and parameter fitting trained on source domain. The epistemic uncertainty on target domain is evaluated on-the-fly to facilitate online reweighting and correction in the self-training process. Meanwhile, to tackle the class-wise quantity and learning difficulty imbalance problem, we introduce a novel data resampling strategy to promote simultaneous convergence across different categories. This strategy prevents the class-level over-fitting in source domain and further boosts the adaptation performance by better quantifying the uncertainty in target domain. We illustrate the superiority of our method compared with the state-of-the-art methods.
Zixian Su, Xi Yang 0008, Jie Sun 0024, Kaizhu Huang
CIKM5
2023 Multi-agent Reinforcement Learning Based Collaborative Multi-task Scheduling for Vehicular Edge Computing
Peisong Li, Ziren Xiao, Xinheng Wang 0001, Kaizhu Huang, Yi Huang 0001, Andrei Tchernykh
CollaborateCom (3)4
2023 Towards Better Robustness against Common Corruptions for Unsupervised Domain Adaptation
abstract
Recent studies have investigated how to achieve robustness for unsupervised domain adaptation (UDA). While most efforts focus on adversarial robustness, i.e. how the model performs against unseen malicious adversarial perturbations, robustness against benign common corruption (RaCC) surprisingly remains under-explored for UDA. Towards improving RaCC for UDA methods in an unsupervised manner, we propose a novel Distributionally and Discretely Adversarial Regularization (DDAR) framework in this paper. Formulated as a min-max optimization with a distribution distance, DDAR1is theoretically well-founded to ensure generalization over unknown common corruptions. Meanwhile, we show that our regularization scheme effectively reduces a surrogate of RaCC, i.e., the perceptual distance between natural data and common corruption. To enable a abetter adversarial regularization, the design of the optimization pipeline relies on an image discretization scheme that can transform "out-of-distribution" adversarial data into "in-distribution" data augmentation. Through extensive experiments, in terms of RaCC, our method is superior to conventional unsupervised regularization mechanisms, widely improves the robustness of existing UDA methods, and achieves state-of-the-art performance.
Kaizhu Huang, Rui Zhang 0012, Dawei Liu 0001, Jieming Ma
ICCV2
2023 Divide and Conquer: 3D Point Cloud Instance Segmentation With Point-Wise Binarization
abstract
Instance segmentation on point clouds is crucially important for 3D scene understanding. Most SOTAs adopt distance clustering, which is typically effective but does not perform well in segmenting adjacent objects with the same semantic label (especially when they share neighboring points). Due to the uneven distribution of offset points, these existing methods can hardly cluster all instance points. To this end, we design a novel divide-and-conquer strategy named PBNet that binarizes each point and clusters them separately to segment instances. Our binary clustering divides offset instance points into two categories: high and low density points (HPs vs. LPs). Adjacent objects can be clearly separated by removing LPs, and then be completed and refined by assigning LPs via a neighbor voting method. To suppress potential over-segmentation, we propose to construct local scenes with the weight mask for each instance. As a plug-in, the proposed binary clustering can replace the traditional distance clustering and lead to consistent performance gains on many mainstream baselines. A series of experiments on ScanNetV2 and S3DIS datasets indicate the superiority of our model. In particular, PBNet ranks first on the ScanNetV2 official benchmark challenge, achieving the highest mAP. Code will be available publicly at https://github.com/weiguangzhao/PBNet.
Weiguang Zhao, Yuyao Yan, Chaolong Yang, Jianan Ye, Xi Yang 0008, Kaizhu Huang
ICCV6
2023 Decoupled Learning for Long-Tailed Oracle Character Recognition
Jing Li 0049, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Rui Zhang 0012, Kaizhu Huang
ICDAR (4)6
2023 Context Does Matter: End-to-end Panoptic Narrative Grounding with Deformable Attention Refined Matching Network
abstract
Panoramic Narrative Grounding (PNG) is an emerging visual grounding task that aims to segment visual objects in images based on dense narrative captions. The current state-of-the-art methods first refine the representation of phrase by aggregating the most similar k image pixels, and then match the refined text representations with the pixels of the image feature map to generate segmentation results. However, simply aggregating sampled image features ignores the contextual information, which can lead to phrase-to-pixel mis-match. In this paper, we propose a novel learning framework called Deformable Attention Refined Matching Network (DRMN), whose main idea is to bring deformable attention in the iterative process of feature learning to incorporate essential context information of different scales of pixels. DRMN iteratively re-encodes pixels with the deformable attention network after updating the feature representation of the top-k most similar pixels. As such, DRMN can lead to accurate yet discriminative pixel representations, purify the top-k most similar pixels, and consequently alleviate the phrase-to-pixel mis-match substantially. Experimental results show that our novel design significantly improves the matching results between text phrases and image pixels. Concretely, DRMN achieves new state-of-the-art performance on the PNG benchmark with an average recall improvement 3.5%. The codes are available in: https://github.com/JaMesLiMers/DRMN.
Xiao-Bo Jin, Qiufeng Wang 0001, Kaizhu Huang
ICDM4
2023 Structure First Detail Next: Image Inpainting with Pyramid Generator
abstract
Recent deep generative models have achieved promising performance in image inpainting. However, it is still challenging for a neural network to generate realistic image details and textures due to its inherent spectral bias. We suggest adopting a ‘structure first detail next’ workflow for image inpainting by knowing how artists work. Thus, we propose to build a Pyramid Generator by stacking several sub-generators, where lower-layer sub-generators focus on restoring image structures. In contrast, the higher-layer sub-generators emphasize image details. Our model progressively restores the input through the entire pyramid in a bottom-up fashion. Notably, our approach has a learning scheme of progressively increasing hole size, which allows it to restore large-hole images. In addition, our method could fully exploit the benefits of learning with high-resolution images and hence is suitable for high-resolution image inpainting. Extensive experimental results on benchmark datasets have validated the effectiveness of our approach compared with state-of-the-arts.
Shuyi Qu, Zhenxing Niu, Jianke Zhu, Bin Dong 0003, Kaizhu Huang
ICME5
2023 Improving Handwritten Mathematical Expression Recognition via an Attention Refinement Network
Qiufeng Wang 0001, Jianghan Chen, Kaizhu Huang
ICONIP (13)5
2023 Progressive Supervision for Tampering Localization in Document Images
Huiru Shao, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
ICONIP (15)2
2023 PAG: Protecting Artworks from Personalizing Image Generative Models
Zhaorui Tan, Siyuan Wang 0017, Xi Yang 0008, Kaizhu Huang
ICONIP (4)4
2023 Towards Deeper and Better Multi-view Feature Fusion for 3D Semantic Segmentation
Chaolong Yang, Yuyao Yan, Weiguang Zhao, Jianan Ye, Xi Yang 0008, Amir Hussain 0001, Bin Dong 0003, Kaizhu Huang
ICONIP (15)8
2023 A Symbolic Characters Aware Model for Solving Geometry Problems
abstract
AI has made significant progress in solving math problems, but geometry problems remain challenging due to their reliance on both text and diagrams. In the text description, symbolic characters such as "ABC" often serve as a bridge to connect the corresponding diagram. However, by simply tokenizing symbolic characters into individual letters (e.g., 'A', 'B' and 'C'), existing works fail to study them explicitly and thus lose the semantic relationship with the diagram. In this paper, we develop a symbolic character-aware model to fully explore the role of these characters in both text and diagram understanding and optimize the model under a multi-modal reasoning framework. In the text encoder, we propose merging individual symbolic characters to form one semantic unit along with geometric information from the corresponding diagram. For the diagram encoder, we pre-train it under a multi-label classification framework with the symbolic characters as labels. In addition, we enhance the geometry diagram understanding ability via a self-supervised learning method under the masked image modeling auxiliary task. By integrating the proposed model into a general encoder-decoder pipeline for solving geometry problems, we demonstrate its superiority on two benchmark datasets, including GeoQA and Geometry3K, with extensive experiments. Specifically, on GeoQA, the question-solving accuracy is increased from 60.0% to 64.1%, achieving a new state-of-the-art accuracy; on Geometry3K, we reduce the question average solving steps from 6.9 down to 6.0 with marginally higher solving accuracy.
Maizhen Ning, Qiufeng Wang 0001, Kaizhu Huang, Xiaowei Huang 0001
ACM Multimedia3
2023 Graph Neural Networks with Diverse Spectral Filtering
abstract
Spectral Graph Neural Networks (GNNs) have achieved tremendous success in graph machine learning, with polynomial filters applied for graph convolutions, where all nodes share the identical filter weights to mine their local contexts. Despite the success, existing spectral GNNs usually fail to deal with complex networks (e.g., WWW) due to such homogeneous spectral filtering setting that ignores the regional heterogeneity as typically seen in real-world networks. To tackle this issue, we propose a novel diverse spectral filtering (DSF) framework, which automatically learns node-specific filter weights to exploit the varying local structure properly. Particularly, the diverse filter weights consist of two components — A global one shared among all nodes, and a local one that varies along network edges to reflect node difference arising from distinct graph parts — to balance between local and global information. As such, not only can the global graph characteristics be captured, but also the diverse local patterns can be mined with awareness of different node positions. Interestingly, we formulate a novel optimization problem to assist in learning diverse filters, which also enables us to enhance any spectral GNNs with our DSF framework. We showcase the proposed framework on three state-of-the-arts including GPR-GNN, BernNet, and JacobiConv. Extensive experiments over 10 benchmark datasets demonstrate that our framework can consistently boost model performance by up to 4.92% in node classification tasks, producing diverse filters with enhanced interpretability.
Jingwei Guo 0001, Kaizhu Huang, Xinping Yi, Rui Zhang 0012
WWW2
2023 Self-supervised generative learning for sequential data prediction
Guoqiang Zhong 0001, Zhaoyang Deng, Kang Zhang 0007, Kaizhu Huang
Appl. Intell.5
2023 Randomized block-coordinate adaptive algorithms for nonconvex optimization problems
Yangfan Zhou 0004, Kaizhu Huang, Amir Hussain 0001, Xin Liu 0102
Eng. Appl. Artif. Intell.2
2023 Retrieval-based language model adaptation for handwritten Chinese text recognition
Shuying Hu, Qiufeng Wang 0001, Kaizhu Huang, Frans Coenen
Int. J. Document Anal. Recognit.3
2023 Robust generative adversarial network
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Rui Zhang 0012, Jimin Xiao, Canyi Lu
Mach. Learn.3
2023 Generalized image outpainting with U-transformer
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas, Yujie Geng, Yuyao Yan, Kaizhu Huang
Neural Networks7
2023 Multi-semantic hypergraph neural network for effective few-shot learning
Hao Chen 0011, Fuyuan Hu, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Zhenping Xia
Pattern Recognit.6
2023 Aggregated pyramid gating network for human pose estimation without pre-training
Chenru Jiang, Kaizhu Huang, Shufei Zhang, Xinheng Wang 0001, Jimin Xiao, John Yannis Goulermas
Pattern Recognit.2
2023 Towards better long-tailed oracle character recognition with adversarial data augmentation
abstract
Deciphering oracle bone script is of great significance to the study of ancient Chinese culture as well as archaeology. Although recent studies on oracle character recognition have made substantial progress, they still suffer from the long-tailed data situation that results in a noticeable performance drop on the tail classes. To mitigate this issue, we propose a generative adversarial framework to augment oracle characters in the problematic classes. In this framework, the generator produces synthetic data through convex combinations of all the available samples in the corresponding classes, and is further optimized through adversarial learning with the classifier and simultaneously the discriminator . Meanwhile, we introduce Repatch to generalize samples in the generator. Since tail classes do not have sufficient data for convex combinations , we propose the TailMix mechanism to generate suitable tail class samples from other classes. Experimental results show that our proposed algorithm obtains remarkable performance in oracle character recognition and achieves new state-of-the-art average (total) accuracy with 86.03% (89.46%), 86.54% (93.86%), 95.22% (96.17%) on the three datasets Oracle-AYNU, OBC306 and Oracle-20K, respectively.
Jing Li 0049, Qiufeng Wang 0001, Kaizhu Huang, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas
Pattern Recognit.3
2023 Semantic Similarity Distance: Towards better text-image consistency metric in text-to-image generation
Zhaorui Tan, Xi Yang 0008, Zihan Ye, Qiufeng Wang 0001, Yuyao Yan, Anh Nguyen 0003, Kaizhu Huang
Pattern Recognit.7
2023 Rebalanced Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to identify unseen classes with zero samples during training. Broadly speaking, present ZSL methods usually adopt class-level semantic labels and compare them with instance-level semantic predictions to infer unseen classes. However, we find that such existing models mostly produce imbalanced semantic predictions, i.e. these models could perform precisely for some semantics, but may not for others. To address the drawback, we aim to introduce an imbalanced learning framework into ZSL. However, we find that imbalanced ZSL has two unique challenges: (1) Its imbalanced predictions are highly correlated with the value of semantic labels rather than the number of samples as typically considered in the traditional imbalanced learning; (2) Different semantics follow quite different error distributions between classes. To mitigate these issues, we first formalize ZSL as an imbalanced regression problem which offers empirical evidences to interpret how semantic labels lead to imbalanced semantic predictions. We then propose a re-weighted loss termed Re-balanced Mean-Squared Error (ReMSE), which tracks the mean and variance of error distributions, thus ensuring rebalanced learning across classes. As a major contribution, we conduct a series of analyses showing that ReMSE is theoretically well established. Extensive experiments demonstrate that the proposed method effectively alleviates the imbalance in semantic prediction and outperforms many state-of-the-art ZSL methods.
Zihan Ye, Guanyu Yang 0002, Xiao-Bo Jin, Youfa Liu, Kaizhu Huang
IEEE Trans. Image Process.5
2023 Mind the Gap: Alleviating Local Imbalance for Unsupervised Cross-Modality Medical Image Segmentation
abstract
Unsupervised cross-modality medical image adaptation aims to alleviate the severe domain gap between different imaging modalities without using the target domain label. A key in this campaign relies upon aligning the distributions of source and target domain. One common attempt is to enforce the global alignment between two domains, which, however, ignores the fatal local-imbalance domain gap problem, i.e., some local features with larger domain gap are harder to transfer. Recently, some methods conduct alignment focusing on local regions to improve the efficiency of model learning. While this operation may cause a deficiency of critical information from contexts. To tackle this limitation, we propose a novel strategy to alleviate the domain gap imbalance considering the characteristics of medical images, namely Global-Local Union Alignment. Specifically, a feature-disentanglement style-transfer module first synthesizes the target-like source images to reduce the global domain gap. Then, a local feature mask is integrated to reduce the 'inter-gap' for local features by prioritizing those discriminative features with larger domain gap. This combination of global and local alignment can precisely localize the crucial regions in segmentation target while preserving the overall semantic consistency. We conduct a series of experiments with two cross-modality adaptation tasks, i,e. cardiac substructure and abdominal multi-organ segmentation. Experimental results indicate that our method achieves state-of-the-art performance in both tasks.
Zixian Su, Xi Yang 0008, Qiufeng Wang 0001, Yuyao Yan, Jie Sun 0024, Kaizhu Huang
IEEE J. Biomed. Health Informatics7
2023 Fitting Imbalanced Uncertainties in Multi-output Time Series Forecasting
abstract
We focus on multi-step ahead time series forecasting with the multi-output strategy. From the perspective of multi-task learning (MTL), we recognize imbalanced uncertainties between prediction tasks of different future time steps. Unexpectedly, trained by the standard summed Mean Squared Error (MSE) loss, existing multi-output forecasting models may suffer from performance drops due to the inconsistency between the loss function and the imbalance structure. To address this problem, we reformulate each prediction task as a distinct Gaussian Mixture Model (GMM) and derive a multi-level Gaussian mixture loss function to better fit imbalanced uncertainties in multi-output time series forecasting. Instead of using the two-step Expectation-Maximization (EM) algorithm, we apply the self-attention mechanism on the task-specific parameters to learn the correlations between different prediction tasks and generate the weight distribution for each GMM component. In this way, our method jointly optimizes the parameters of the forecasting model and the mixture model simultaneously in an end-to-end fashion, avoiding the need of two-step optimization. Experiments on three real-world datasets demonstrate the effectiveness of our multi-level Gaussian mixture loss compared to models trained with the standard summed MSE loss function. All the experimental data and source code are available at https://github.com/smallGum/GMM-FNN .
Jiezhu Cheng, Kaizhu Huang, Zibin Zheng
ACM Trans. Knowl. Discov. Data2
2023 Explainable Tensorized Neural Ordinary Differential Equations for Arbitrary-Step Time Series Prediction
abstract
In this work, we propose a continuous neural network architecture, referred to as Explainable Tensorized Neural - Ordinary Differential Equations (ETN-ODE) network for multi-step time series prediction at arbitrary time points. Unlike existing approaches which mainly handle univariate time series for multi-step prediction, or multivariate time series for single-step predictions, ETN-ODE is capable of handling multivariate time series with arbitrary-step predictions. An additional benefit is its tandem attention mechanism, with respect to temporal and variable attention, which enable it to greatly facilitate data interpretability. Specifically, the proposed model combines an explainable tensorized gated recurrent unit with ordinary differential equations, with the derivatives of the latent states parameterized through a neural network. We quantitatively and qualitatively demonstrate the effectiveness and interpretability of ETN-ODE on one arbitrary-step prediction task and five standard multi-step prediction tasks. Extensive experiments show that the proposed method achieves very accurate predictions at arbitrary time points while attaining very competitive performance against the baseline methods in standard multi-step time series prediction.
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Kaizhu Huang, John Yannis Goulermas
IEEE Trans. Knowl. Data Eng.4
2023 FastAdaBelief: Improving Convergence Rate for Belief-Based Adaptive Optimizers by Exploiting Strong Convexity
abstract
AdaBelief, one of the current best optimizers, demonstrates superior generalization ability over the popular Adam algorithm by viewing the exponential moving average of observed gradients. AdaBelief is theoretically appealing in which it has a data-dependent O(√T) regret bound when objective functions are convex, where T is a time horizon. It remains, however, an open problem whether the convergence rate can be further improved without sacrificing its generalization ability. To this end, we make the first attempt in this work and design a novel optimization algorithm called FastAdaBelief that aims to exploit its strong convexity in order to achieve an even faster convergence rate. In particular, by adjusting the step size that better considers strong convexity and prevents fluctuation, our proposed FastAdaBelief demonstrates excellent generalization ability and superior convergence. As an important theoretical contribution, we prove that FastAdaBelief attains a data-dependent O(logT) regret bound, which is substantially lower than AdaBelief in strongly convex cases. On the empirical side, we validate our theoretical analysis with extensive experiments in scenarios of strong convexity and nonconvexity using three popular baseline models. Experimental results are very encouraging: FastAdaBelief converges the quickest in comparison to all mainstream algorithms while maintaining an excellent generalization ability, in cases of both strong convexity or nonconvexity. FastAdaBelief is, thus, posited as a new benchmark model for the research community.
Yangfan Zhou 0004, Kaizhu Huang, Amir Hussain 0001, Xin Liu 0102
IEEE Trans. Neural Networks Learn. Syst.2
2022 Outpainting by Queries
Penglei Gao, Xi Yang 0008, Jie Sun 0024, Rui Zhang 0012, Kaizhu Huang
ECCV (23)6
2022 Towards Accurate Alignment and Sufficient Context in Scene Text Recognition
Yijie Hu, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Xiao-Bo Jin, Kaizhu Huang
ICONIP (3)6
2022 Rethinking Image Inpainting with Attention Feature Fusion
Shuyi Qu, Kaizhu Huang, Qiufeng Wang 0001, Bin Dong 0003
ICONIP (3)2
2022 Certifying Better Robust Generalization for Unsupervised Domain Adaptation
abstract
Recent studies explore how to obtain adversarial robustness for unsupervised domain adaptation (UDA). These efforts are however dedicated to achieving an optimal trade-off between accuracy and robustness on a given or seen target domain but ignore the robust generalization issue over unseen adversarial data. Consequently, degraded performance will be often observed when existing robust UDAs are applied to future adversarial data. In this work, we make a first attempt to address the robust generalization issue of UDA. We conjecture that the poor robust generalization of present robust UDAs may be caused by the large distribution gap among adversarial examples. We then provide an empirical and theoretical analysis showing that this large distribution gap is mainly owing to the discrepancy between feature-shift distributions. To reduce such discrepancy, a novel Anchored Feature-Shift Regularization (AFSR) method is designed with a certificated robust generalization bound. We conduct a series of experiments on benchmark UDA datasets. Experimental results validate the effectiveness of our proposed AFSR over many existing robust UDA methods.
Shufei Zhang, Kaizhu Huang, Qiufeng Wang 0001, Rui Zhang 0012, Chaoliang Zhong
ACM Multimedia3
2022 Harnessing Multi-Semantic Hypergraph for Few-Shot Learning
Hao Chen 0011, Zhenping Xia, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Fuyuan Hu
PRCV (1)6
2022 Generalised Zero-shot Learning for Entailment-based Text Classification with External Knowledge
abstract
Text classification techniques have been substantially important to many smart computing applications, e.g. topic extraction and event detection. However, classification is always challenging when only insufficient amount of labelled data for model training is available. To mitigate this issue, zero-shot learning (ZSL) has been introduced for models to recognise new classes that have not been observed during the training stage. We propose an entailment-based zero-shot text classification model, named as S-BERT-CAM, to better capture the relationship between the premise and hypothesis in the BERT embedding space. Two widely used textual datasets are utilised to conduct the experiments. We fine-tune our model using 50% of the labels for each dataset and evaluate it on the label space containing all labels (including both seen and unseen labels). The experimental results demonstrate that our model is more robust to the generalised ZSL and significantly improves the overall performance against baselines.
Wei Wang 0042, Qi Chen 0026, Kaizhu Huang, Anh Nguyen 0003, Suparna De
SMARTCOMP4
2022 Zero-Shot Text Classification via Knowledge Graph Embedding for Social Media Data
abstract
The idea of “citizen sensing” and “human as sensors” is crucial for social Internet of Things, an integral part of cyber–physical–social systems (CPSSs). Social media data, which can be easily collected from the social world, has become a valuable resource for research in many different disciplines, e.g., crisis/disaster assessment, social event detection, or the recent COVID-19 analysis. Useful information, or knowledge derived from social data, could better serve the public if it could be processed and analyzed in more efficient and reliable ways. Advances in deep neural networks have significantly improved the performance of many social media analysis tasks. However, deep learning models typically require a large amount of labeled data for model training, while most CPSS data is not labeled, making it impractical to build effective learning models using traditional approaches. In addition, the current state-of-the-art, pretrained natural language processing (NLP) models do not make use of existing knowledge graphs, thus often leading to unsatisfactory performance in real-world applications. To address the issues, we propose a new zero-shot learning method which makes effective use of existing knowledge graphs for the classification of very large amounts of social text data. Experiments were performed on a large, real-world tweet data set related to COVID-19, the evaluation results show that the proposed method significantly outperforms six baseline models implemented with state-of-the-art deep learning models for NLP.
Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Frans Coenen
IEEE Internet Things J.3
2022 Re-thinking model robustness from stability: a new insight to defend adversarial examples
Shufei Zhang, Kaizhu Huang, Zenglin Xu
Mach. Learn.2
2022 Sparse matrix factorization with L2, 1 norm for matrix completion
Xiao-Bo Jin, Jianyu Miao, Qiufeng Wang 0001, Guanggang Geng, Kaizhu Huang
Pattern Recognit.5
2022 A survey of robust adversarial training in pattern recognition: Fundamental, theory, and methodologies
Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Xu-Yao Zhang
Pattern Recognit.2
2022 End-to-end weakly supervised semantic segmentation with reliable region mining
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Kaizhu Huang, Shan Luo 0001, Yao Zhao 0001
Pattern Recognit.4
2022 Unsupervised domain adaptation in homogeneous distance space for person re-identification
Dingyuan Zheng, Jimin Xiao, Yunchao Wei, Qiufeng Wang 0001, Kaizhu Huang, Yao Zhao 0001
Pattern Recognit.5
2022 A Novel 3D Unsupervised Domain Adaptation Framework for Cross-Modality Medical Image Segmentation
abstract
We consider the problem of volumetric (3D) unsupervised domain adaptation (UDA) in cross-modality medical image segmentation, aiming to perform segmentation on the unannotated target domain (e.g. MRI) with the help of labeled source domain (e.g. CT). Previous UDA methods in medical image analysis usually suffer from two challenges: 1) they focus on processing and analyzing data at 2D level only, thus missing semantic information from the depth level; 2) one-to-one mapping is adopted during the style-transfer process, leading to insufficient alignment in the target domain. Different from the existing methods, in our work, we conduct a first of its kind investigation on multi-style image translation for complete image alignment to alleviate the domain shift problem, and also introduce 3D segmentation in domain adaptation tasks to maintain semantic consistency at the depth level. In particular, we develop an unsupervised domain adaptation framework incorporating a novel quartet self-attention module to efficiently enhance relationships between widely separated features in spatial regions on a higher dimension, leading to a substantial improvement in segmentation accuracy in the unlabeled target domain. In two challenging cross-modality tasks, specifically brain structures and multi-organ abdominal segmentation, our model is shown to outperform current state-of-the-art methods by a significant margin, demonstrating its potential as a benchmark resource for the biomedical and health informatics research community.
Zixian Su, Kaizhu Huang, Xi Yang 0008, Jie Sun 0024, Amir Hussain 0001, Frans Coenen
IEEE J. Biomed. Health Informatics3
2022 Disentangling Semantic-to-Visual Confusion for Zero-Shot Learning
abstract
Using generative models to synthesize visual features from semantic distribution is one of the most popular solutions to ZSL image classification in recent years. The triplet loss (TL) is popularly used to generate realistic visual distributions from semantics by automatically searching discriminative representations. However, the traditional TL cannot search reliable unseen disentangled representations due to the unavailability of unseen classes in ZSL. To alleviate this drawback, we propose in this work a multi-modal triplet loss (MMTL) which utilizes multi-modal information to search adisentangledrepresentation space. As such, all classes can interplay which can benefit learning disentangled class representations in the searched space. Furthermore, we develop a novel model called Disentangling Class Representation Generative Adversarial Network (DCR-GAN) focusing on exploiting the disentangled representations in training, feature synthesis, and final recognition stages. Benefiting from the disentangled representations, DCR-GAN could fit a more realistic distribution over both seen and unseen features. Extensive experiments show that our proposed model can lead to superior performance to the state-of-the-arts on four benchmark datasets.
Zihan Ye, Fuyuan Hu, Fan Lyu, Kaizhu Huang
IEEE Trans. Multim.5
2022 Exploiting Attention-Consistency Loss For Spatial-Temporal Stream Action Recognition
abstract
Currently, many action recognition methods mostly consider the information from spatial streams. We propose a new perspective inspired by the human visual system to combine both spatial and temporal streams to measure their attention consistency. Specifically, a branch-independent convolutional neural network (CNN) based algorithm is developed with a novel attention-consistency loss metric, enabling the temporal stream to concentrate on consistent discriminative regions with the spatial stream in the same period. The consistency loss is further combined with the cross-entropy loss to enhance the visual attention consistency. We evaluate the proposed method for action recognition on two benchmark datasets: Kinetics400 and UCF101. Despite its apparent simplicity, our proposed framework with the attention consistency achieves better performance than most of the two-stream networks, i.e., 75.7% top-1 accuracy on Kinetics400 and 95.7% on UCF101, while reducing 7.1% computational cost compared with our baseline. Particularly, our proposed method can attain remarkable improvements on complex action classes, showing that our proposed network can act as a potential benchmark to handle complicated scenarios in industry 4.0 applications.
Xiao-Bo Jin, Qiufeng Wang 0001, Amir Hussain 0001, Kaizhu Huang
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Each Attribute Matters: Contrastive Attention for Sentence-based Image Editing
Liuqing Zhao, Fan Lyu, Fuyuan Hu, Kaizhu Huang, Fenglei Xu
BMVC4
2021 Gradient Distribution Alignment Certificates Better Adversarial Domain Adaptation
abstract
The latest heuristic for handling the domain shift in un-supervised domain adaptation tasks is to reduce the data distribution discrepancy using adversarial learning. Recent studies improve the conventional adversarial domain adaptation methods with discriminative information by integrating the classifier’s outputs into distribution divergence measurement. However, they still suffer from the equilibrium problem of adversarial learning in which even if the discriminator is fully confused, sufficient similarity between two distributions cannot be guaranteed. To overcome this problem, we propose a novel approach named feature gradient distribution alignment (FGDA)1. We demonstrate the rationale of our method both theoretically and empirically. In particular, we show that the distribution discrepancy can be reduced by constraining feature gradients of two domains to have similar distributions. Meanwhile, our method enjoys a theoretical guarantee that a tighter error upper bound for target samples can be obtained than that of conventional adversarial domain adaptation methods. By integrating the proposed method with existing adversarial domain adaptation models, we achieve state-of-the-art performance on two real-world benchmark datasets.
Shufei Zhang, Kaizhu Huang, Qiufeng Wang 0001, Chaoliang Zhong
ICCV3
2021 Mix-Up Augmentation for Oracle Character Recognition with Imbalanced Data Distribution
Jing Li 0049, Qiufeng Wang 0001, Rui Zhang 0012, Kaizhu Huang
ICDAR (1)4
2021 Towards Better Robust Generalization with Shift Consistency Regularization
abstract
While adversarial training becomes one of the most promising defending approaches against adversarial attacks for deep neural networks, the conventional wisdom through robust optimization may usually not guarantee good generalization for robustness. Concerning with robust generalization over unseen adversarial data, this paper investigates adversarial training from a novel perspective of shift consistency in latent space. We argue that the poor robust generalization of adversarial training is owing to the significantly dispersed latent representations generated by training and test adversarial data, as the adversarial perturbations push the latent features of natural examples in the same class towards diverse directions. This is underpinned by the theoretical analysis of the robust generalization gap, which is upper-bounded by the standard one over the natural data and a term of feature inconsistent shift caused by adversarial perturbation {–} a measure of latent dispersion. Towards better robust generalization, we propose a new regularization method {–} shift consistency regularization (SCR) {–} to steer the same-class latent features of both natural and adversarial data into a common direction during adversarial training. The effectiveness of SCR in adversarial training is evaluated through extensive experiments over different datasets, such as CIFAR-10, CIFAR-100, and SVHN, against several competitive methods.
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Rui Zhang 0012, Xinping Yi
ICML3
2021 A Segment-Based Layout Aware Model for Information Extraction on Document Images
Maizhen Ning, Qiufeng Wang 0001, Kaizhu Huang, Xiaowei Huang 0001
ICONIP (5)3
2021 Global-aware Beam Search for Neural Abstractive Summarization
abstract
This study develops a calibrated beam-based algorithm with awareness of the global attention distribution for neural abstractive summarization, aiming to improve the local optimality problem of the original beam search in a rigorous way. Specifically, a novel global protocol is proposed based on the attention distribution to stipulate how a global optimal hypothesis should attend to the source. A global scoring mechanism is then developed to regulate beam search to generate summaries in a near-global optimal fashion. This novel design enjoys a distinctive property, i.e., the global attention distribution could be predicted before inference, enabling step-wise improvements on the beam search through the global scoring mechanism. Extensive experiments on nine datasets show that the global (attention)-aware inference significantly improves state-of-the-art summarization models even using empirical hyper-parameters. The algorithm is also proven robust as it remains to generate meaningful texts with corrupted attention distributions. The codes and a comprehensive set of examples are available.
Zixun Lan, Lu Zong, Kaizhu Huang
NeurIPS4
2021 Multi-modal generative adversarial networks for traffic event detection in smart cities
Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Suparna De, Frans Coenen
Expert Syst. Appl.3
2021 Novel Artificial Immune Networks-based optimization of shallow machine learning (ML) classifiers
Summrina Kanwal, Amir Hussain 0001, Kaizhu Huang
Expert Syst. Appl.3
2021 Residual attention-based multi-scale script identification in scene text images
Mengkai Ma, Qiufeng Wang 0001, Shen Huang, John Yannis Goulermas, Kaizhu Huang
Neurocomputing6
2021 Coarse-grained generalized zero-shot learning with efficient self-focus mechanism
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
Neurocomputing2
2021 Domain adaptation with feature and label adversarial networks
Wenhua Zang, Bin Liu 0022, Zhao Kang 0001, Kaizhu Huang, Zenglin Xu
Neurocomputing6
2021 Artificial Intelligence in Collaborative Computing
Xinheng Wang 0001, Honghao Gao, Kaizhu Huang
Mob. Networks Appl.3
2021 Improving generative adversarial networks with simple latent distributions
Shufei Zhang, Kaizhu Huang, Zhuang Qian, Rui Zhang 0012, Amir Hussain 0001
Neural Comput. Appl.2
2021 Manifold adversarial training for supervised and semi-supervised learning
Shufei Zhang, Kaizhu Huang, Jianke Zhu
Neural Networks2
2021 Automated Social Text Annotation With Joint Multilabel Attention Networks
abstract
Automated social text annotation is the task of suggesting a set of tags for shared documents on social media platforms. The automated annotation process can reduce users' cognitive overhead in tagging and improve tag management for better search, browsing, and recommendation of documents. It can be formulated as a multilabel classification problem. We propose a novel deep learning-based method for this problem and design an attention-based neural network with semantic-based regularization, which can mimic users' reading and annotation behavior to formulate better document representation, leveraging the semantic relations among labels. The network separately models the title and the content of each document and injects an explicit, title-guided attention mechanism into each sentence. To exploit the correlation among labels, we propose two semantic-based loss regularizers, i.e., similarity and subsumption, which enforce the output of the network to conform to label semantics. The model with the semantic-based loss regularizers is referred to as the joint multilabel attention network (JMAN). We conducted a comprehensive evaluation study and compared JMAN to the state-of-the-art baseline models, using four large, real-world social media data sets. In terms of F1, JMAN significantly outperformed bidirectional gated recurrent unit (Bi-GRU) relatively by around 12.8%-78.6% and the hierarchical attention network (HAN) by around 3.9%-23.8%. The JMAN model demonstrates advantages in convergence and training speed. Further improvement of performance was observed against latent Dirichlet allocation (LDA) and support vector machine (SVM). When applying the semantic-based loss regularizers, the performance of HAN and Bi-GRU in terms of F1was also boosted. It is also found that dynamic update of the label semantic matrices (JMANd) has the potential to further improve the performance of JMAN but at the cost of substantial memory and warrants further study.
Hang Dong 0002, Wei Wang 0042, Kaizhu Huang, Frans Coenen
IEEE Trans. Neural Networks Learn. Syst.3
2020 Towards Better Forecasting by Fusing Near and Distant Future Visions
abstract
Multivariate time series forecasting is an important yet challenging problem in machine learning. Most existing approaches only forecast the series value of one future moment, ignoring the interactions between predictions of future moments with different temporal distance. Such a deficiency probably prevents the model from getting enough information about the future, thus limiting the forecasting accuracy. To address this problem, we propose Multi-Level Construal Neural Network (MLCNN), a novel multi-task deep learning framework. Inspired by the Construal Level Theory of psychology, this model aims to improve the predictive performance by fusing forecasting information (i.e., future visions) of different future time. We first use the Convolution Neural Network to extract multi-level abstract representations of the raw data for near and distant future predictions. We then model the interplay between multiple predictive tasks and fuse their future visions through a modified Encoder-Decoder architecture. Finally, we combine traditional Autoregression model with the neural network to solve the scale insensitive problem. Experiments on three real-world datasets show that our method achieves statistically significant improvements compared to the most state-of-the-art baseline methods, with average 4.59% reduction on RMSE metric and average 6.87% reduction on MAE metric.
Jiezhu Cheng, Kaizhu Huang, Zibin Zheng
AAAI2
2020 Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation Approach
abstract
Weakly supervised semantic segmentation is a challenging task as it only takes image-level information as supervision for training but produces pixel-level predictions for testing. To address such a challenging task, most recent state-of-the-art approaches propose to adopt two-step solutions, i.e. 1) learn to generate pseudo pixel-level masks, and 2) engage FCNs to train the semantic segmentation networks with the pseudo masks. However, the two-step solutions usually employ many bells and whistles in producing high-quality pseudo masks, making this kind of methods complicated and inelegant. In this work, we harness the image-level labels to produce reliable pixel-level annotations and design a fully end-to-end network to learn to predict segmentation maps. Concretely, we firstly leverage an image classification branch to generate class activation maps for the annotated categories, which are further pruned into confident yet tiny object/background regions. Such reliable regions are then directly served as ground-truth labels for the parallel segmentation branch, where a newly designed dense energy loss function is adopted for optimization. Despite its apparent simplicity, our one-step solution achieves competitive mIoU scores (val: 62.6, test: 62.9) on Pascal VOC compared with those two-step state-of-the-arts. By extending our one-step method to two-step, we get a new state-of-the-art performance on the Pascal VOC (val: 66.3, test: 66.5).
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, Kaizhu Huang
AAAI5
2020 A Covert Ultrasonic Phone-to-Phone Communication Scheme
Liming Shi, Limin Yu, Kaizhu Huang, Xu Zhu 0001, Zhi Wang 0003, Xiaofei Li 0001, Wenwu Wang 0001, Xinheng Wang 0001
CollaborateCom (1)3
2020 Feature Representation Matters: End-to-End Learning for Reference-Based Image Super-Resolution
Yanchun Xie, Jimin Xiao, Mingjie Sun, Kaizhu Huang
ECCV (4)5
2020 Adversarial Rectification Network for Scene Text Regularization
Jing Li 0049, Qiufeng Wang 0001, Rui Zhang 0012, Kaizhu Huang
ICONIP (2)4
2020 MCRN: A New Content-Based Music Classification and Recommendation Network
Yuxu Mao, Guoqiang Zhong 0001, Haizhen Wang, Kaizhu Huang
ICONIP (4)4
2020 CDMC'19 - The 10th International Cybersecurity Data Mining Competition
Shaoning Pang 0001, Tao Ban, Youki Kadobayashi, Kaizhu Huang, Geongsen Poh, Iqbal Gondal, Kitsuchart Pasupa, Fadi A. Aloul
ICONIP (2)5
2020 Feature Redirection Network for Few-Shot Classification
Guoqiang Zhong 0001, Yuxu Mao, Kaizhu Huang
ICONIP (4)4
2020 Multi-scale Attention Consistency for Multi-label Image Classification
Xiao-Bo Jin, Qiufeng Wang 0001, Kaizhu Huang
ICONIP (4)4
2020 Maximum Power Point Tracking of Photovoltaic Systems Using Deep Q-networks
abstract
A photovoltaic (PV) generator exhibits nonlinear current-voltage characteristics and its maximum power point varies with incident atmospheric conditions. Therefore, maximum power point tracking (MPPT) control is required to maximize the output power of the PV generator. In this paper, deep Q-network based reinforcement learning strategy is proposed to optimize MPPT process for the photovoltaic system. The proposed system uses a novel control method which introduces agent to interface with the environment and finally gets the strategy of maximum reward accordingly. Simulations and experiments show the feasibility and effectiveness of the proposed system. Compared with the traditional perturb and observe (P&O) and incremental conductance (InC) methods, this method prominently saves tracking steps.
Kangshi Wang, Dou Hong, Jieming Ma, Ka Lok Man, Kaizhu Huang, Xiaowei Huang 0001
INDIN5
2020 Pay Attention Selectively and Comprehensively: Pyramid Gating Network for Human Pose Estimation without Pre-training
abstract
Deep neural network with multi-scale feature fusion has achieved great success in human pose estimation. However, drawbacks still exist in these methods: 1) they consider multi-scale features equally, which may over-emphasize redundant features; 2) preferring deeper structures, they can learn features with the strong semantic representation, but tend to lose natural discriminative information; 3) to attain good performance, they rely heavily on pretraining, which is time-consuming, or even unavailable practically. To mitigate these problems, we propose a novel comprehensive recalibration model called Pyramid GAting Network (PGA-Net) that is capable of distillating, selecting, and fusing the discriminative and attention-aware features at different scales and different levels (i.e., both semantic and natural levels). Meanwhile, focusing on fusing features both selectively and comprehensively, PGA-Net can demonstrate remarkable stability and encouraging performance even without pre-training, making the model can be trained truly from scratch. We demonstrate the effectiveness of PGA-Net through validating on COCO and MPII benchmarks, attaining new state-of-the-art performance. https://github.com/ssr0512/PGA-Net
Chenru Jiang, Kaizhu Huang, Shufei Zhang, Xinheng Wang 0001, Jimin Xiao
ACM Multimedia2
2020 Inductive Generalized Zero-Shot Learning with Adversarial Relation Network
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
ECML/PKDD (2)2
2020 Multi-modal Adversarial Training for Crisis-related Data Classification on Social Media
abstract
Social media platforms such as Twitter are increasingly used to collect data of all kinds. During natural disasters, users may post text and image data on social media platforms to report information about infrastructure damage, injured people, cautions and warnings. Effective processing and analysing tweets in real time can help city organisations gain situational awareness of the affected citizens and take timely operations. With the advances in deep learning techniques, recent studies have significantly improved the performance in classifying crisis-related tweets. However, deep learning models are vulnerable to adversarial examples, which may be imperceptible to the human, but can lead to model's misclassification. To process multi-modal data as well as improve the robustness of deep learning models, we propose a multi-modal adversarial training method for crisis-related tweets classification in this paper. The evaluation results clearly demonstrate the advantages of the proposed model in improving the robustness of tweet classification.
Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Suparna De, Frans Coenen
SMARTCOMP3
2020 Triple loss for hard face detection
Zhenyu Fang, Jinchang Ren, Stephen Marshall, Huimin Zhao 0001, Zheng Wang 0008, Kaizhu Huang, Bing Xiao 0005
Neurocomputing6
2020 Hybrid channel based pedestrian detection
Fiseha B. Tesema, Junpeng Lin, William Zhu 0001, Kaizhu Huang
Neurocomputing6
2020 Improving deep neural network performance by integrating kernelized Min-Max objective
Qiufeng Wang 0001, Rui Zhang 0012, Amir Hussain 0001, Kaizhu Huang
Neurocomputing5
2020 Knowledge base enrichment by relation learning from social tagging data
Hang Dong 0002, Wei Wang 0042, Frans Coenen, Kaizhu Huang
Inf. Sci.4
2020 Editorial: Collaborative Computing for Data-Driven Systems
Xinheng Wang 0001, Muddesar Iqbal, Honghao Gao, Kaizhu Huang, Andrei Tchernykh
Mob. Networks Appl.4
2020 Novel deep neural network based pattern field classification architectures
Kaizhu Huang, Shufei Zhang, Rui Zhang 0012, Amir Hussain 0001
Neural Networks1
2020 Generative adversarial networks with mixture of t-distributions noise for diverse image generation
Jinxuan Sun, Guoqiang Zhong 0001, Yang Chen 0036, Tao Li 0031, Kaizhu Huang
Neural Networks6
2020 Encoding primitives generation policy learning for robotic arm to overcome catastrophic forgetting in sequential multi-tasks learning
Fangzhou Xiong, Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Hong Qiao, Amir Hussain 0001
Neural Networks3
2020 Generative adversarial networks with decoder-encoder output noises
Guoqiang Zhong 0001, Youzhao Yang, Dahan Wang, Kaizhu Huang
Neural Networks6
2020 Generative adversarial classifier for handwriting characters super-resolution
Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Jimin Xiao, Rui Zhang 0012
Pattern Recognit.2
2020 Segmentation mask guided end-to-end person search
Dingyuan Zheng, Jimin Xiao, Kaizhu Huang, Yao Zhao 0001
Signal Process. Image Commun.3
2020 Correlation Filter Selection for Visual Tracking Using Reinforcement Learning
abstract
Correlation filter has been proven to be an effective tool for a number of approaches in visual tracking, particularly for seeking a good balance between tracking accuracy and speed. However, correlation filter-based models are susceptible to wrong updates stemming from inaccurate tracking results. To date, very little effort has been devoted towards handling the correlation filter update problem. In this paper, we propose a novel approach to address the correlation filter update problem. In our approach, we update and maintain multiple correlation filter models in parallel, and we use deep reinforcement learning for the selection of an optimal correlation filter model among them. To facilitate the decision process in an efficient manner, we propose a decision-net to deal with target appearance modeling, which is trained through hundreds of challenging videos using proximal policy optimization and a lightweight learning network. An exhaustive evaluation of the proposed approach on the OTB100 and OTB2013 benchmarks shows that the approach is effective enough to achieve the average success rate of 62.3% and the average precision score of 81.2%, both exceeding the performance of traditional correlation filter-based trackers.
Yanchun Xie, Jimin Xiao, Kaizhu Huang, Jeyan Thiyagalingam, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 An Interactive and Generative Approach for Chinese Shanshui Painting Document
abstract
Chinese Shanshui is a landscape painting document mainly drawing mountain and water, which is popular in Chinese culture. However, it is very challenging to create this by general people. In this paper, we propose an interactive and generative approach to automatically generate the Chinese Shanshui painting documents based on users' input, where the users only need to sketch simple lines to represent their ideal landscape without any professional Shanshui painting skills. This sketch-to-Shanshui translation is optimized by the model of cycle Generative Adversarial Networks (GAN). To evaluate the proposed approach, we collected a large set of both sketch data and Chinese Shanshui painting data to train the model of cycle-GAN, and developed an interactive system called Shanshui-DaDA (i.e., Design and Draw with AI) to generate Chinese Shanshui painting documents in real-time. The experimental results show that this system can generate satisfied Chinese Shanshui painting documents by general users.
Aven-Le Zhou, Qiufeng Wang 0001, Kaizhu Huang, Cheng-Hung Lo
ICDAR3
2019 VSB-DVM: An End-to-End Bayesian Nonparametric Generalization of Deep Variational Mixture Model
abstract
Mixture of factor analyzers is a fundamental model in unsupervised learning, which is particularly useful for high dimensional data. Recent efforts on deep auto-encoding mixture models made a fruitful progress in clustering. However, in most cases, their performance depends highly on the results of pre-training. Moreover, they tend to ignore the prior information when making clustering assignment, leading to a less strict inference and consequently limiting the performance. In this paper, we propose an end-to-end Bayesian nonparametric generalization of deep mixture model with a Variational Auto-Encoder (VAE) framework. Specifically, we develop a novel model called VSB-DVM exploiting the Variational Stick-Breaking Process to design a Deep Variational Mixture Model. Distinct from the existing deep auto-encoding mixture models, this novel unsupervised deep generative model can learn low-dimensional representations and clustering simultaneously without pre-training. Importantly, a strict inference is proposed using weights of stick-breaking process in a variational way. Furthermore, able to capture the richer statistical structure of the data, VSB-DVM can also generate highly realistic samples for any specified cluster. A series of experiments are carried out, both qualitatively and quantitatively, on benchmark clustering and generation tasks. Comparative results show that the proposed model is able to generate diverse and high-quality samples of data, and also achieves encouraging clustering results outperforming the state-of-the-art algorithms on four real-world datasets.
Xi Yang 0008, Yuyao Yan, Kaizhu Huang, Rui Zhang 0012
ICDM3
2019 Generalized Adversarial Training in Riemannian Space
abstract
Adversarial examples, referred to as augmented data points generated by imperceptible perturbations of input samples, have recently drawn much attention. Well-crafted adversarial examples may even mislead state-of-the-art deep neural network (DNN) models to make wrong predictions easily. To alleviate this problem, many studies have focused on investigating how adversarial examples can be generated and/or effectively handled. All existing works tackle this problem in the Euclidean space. In this paper, we extend the learning of adversarial examples to the more general Riemannian space over DNNs. The proposed work is important in that (1) it is a generalized learning methodology since Riemmanian space will be degraded to the Euclidean space in a special case; (2) it is the first work to tackle the adversarial example problem tractably through the perspective of Riemannian geometry; (3) from the perspective of geometry, our method leads to the steepest direction of the loss function, by considering the second order information of the loss function. We also provide a theoretical study showing that our proposed method can truly find the descent direction for the loss function, with a comparable computational time against traditional adversarial methods. Finally, the proposed framework demonstrates superior performance over traditional counterpart methods, using benchmark data including MNIST, CIFAR-10 and SVHN.
Shufei Zhang, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICDM2
2019 Enhanced LSTM with Batch Normalization
Li-Na Wang, Guoqiang Zhong 0001, Shoujun Yan, Junyu Dong, Kaizhu Huang
ICONIP (1)5
2019 MPSSD: Multi-Path Fusion Single Shot Detector
abstract
Recent prevalent one stage detectors, such as single shot detector (SSD) and RetinaNet, are able to detect objects faster than two stage ones while maintaining comparable accuracy. To further boost the accuracy, many studies focus on enhancing the multi-scale feature pyramid. Most of these current proposals focus on strengthening features on one pyramid, ignoring the rich connection among different scale features. In contrast, we propose a novel multi-path design to fully utilize the localization and semantics information. First, we exploit the original SSD multi-scale features as our base pyramid. Then we fuse these features in different groups to generate multi-path feature pyramids. Finally, we combine these pyramids through a novel and effective aggregation module, to obtain the final informative pyramid for detection. Comparative experiments on benchmark PASCAL VOC and MS COCO datasets have shown that our proposed method outperforms many state-of-the-art detectors. As an illustrative example, for input image with size 512×512, we can achieve a mean Average Precision (mAP) of 81.8% on VOC2007 test and 33.1% mAP on COCO test-dev2015.
Shuyi Qu, Kaizhu Huang, Amir Hussain 0001, John Yannis Goulermas
IJCNN2
2019 Special issue on advances in graph algorithm and applications
Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Cheng-Lin Liu 0001
Neurocomputing2
2019 IAN: The Individual Aggregation Network for Person Search
Jimin Xiao, Yanchun Xie, Tammam Tillo, Kaizhu Huang, Yunchao Wei, Jiashi Feng
Pattern Recognit.4
2019 Stochastic Conjugate Gradient Algorithm With Variance Reduction
abstract
Conjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction1and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functions. We experimentally demonstrate that the CG with variance reduction algorithm converges faster than its counterparts for four learning models, which may be convex, nonconvex or nonsmooth. In addition, its area under the curve performance on six large-scale data sets is comparable to that of the LIBLINEAR solver for the L2-regularized L2-loss but with a significant improvement in computational efficiency.
Xiao-Bo Jin, Xu-Yao Zhang, Kaizhu Huang, Guanggang Geng
IEEE Trans. Neural Networks Learn. Syst.3
2019 Guided Policy Search for Sequential Multitask Learning
abstract
Policy search in reinforcement learning (RL) is a practical approach to interact directly with environments in parameter spaces, that often deal with dilemmas of local optima and real-time sample collection. A promising algorithm, known as guided policy search (GPS), is capable of handling the challenge of training samples using trajectory-centric methods. It can also provide asymptotic local convergence guarantees. However, in its current form, the GPS algorithm cannot operate in sequential multitask learning scenarios. This is due to its batch-style training requirement, where all training samples are collectively provided at the start of the learning process. The algorithm’s adaptation is thus hindered for real-time applications, where training samples or tasks can arrive randomly. In this paper, the GPS approach is reformulated, by adapting a recently proposed, lifelong-learning method, and elastic weight consolidation. Specifically, Fisher information is incorporated to impart knowledge from previously learned tasks. The proposed algorithm, termed sequential multitask learning-GPS, is able to operate in sequential multitask learning settings and ensuring continuous policy learning, without catastrophic forgetting. Pendulum and robotic manipulation experiments demonstrate the new algorithms efficacy to learn control policies for handling sequentially arriving training samples, delivering comparable performance to the traditional, and batch-based GPS algorithm. In conclusion, the proposed algorithm is posited as a new benchmark for the real-time RL and robotics research community.
Fangzhou Xiong, Biao Sun 0005, Xu Yang 0004, Hong Qiao, Kaizhu Huang, Amir Hussain 0001, Zhiyong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2018 W-Net: One-Shot Arbitrary-Style Chinese Character Generation with Deep Neural Networks
Haochuan Jiang, Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012
ICONIP (5)3
2018 Improving Deep Neural Network Performance with Kernelized Min-Max Objective
Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)2
2018 Three-Dimensional Local Energy-Based Shape Histogram (3D-LESH): A Novel Feature Extraction Technique
abstract
In this paper, we present a novel feature extraction technique, termed Three-Dimensional Local Energy-Based Shape Histogram (3D-LESH), and exploit it to detect breast cancer in volumetric medical images. The technique is incorporated as part of an intelligent expert system that can aid medical practitioners making diagnostic decisions. Analysis of volumetric images, slice by slice, is cumbersome and inefficient. Hence, 3D-LESH is designed to compute a histogram-based feature set from a local energy map, calculated using a phase congruency (PC) measure of volumetric Magnetic Resonance Imaging (MRI) scans in 3D space. 3D-LESH features are invariant to contrast intensity variations within different slices of the MRI scan and are thus suitable for medical image analysis. The contribution of this article is manifold. First, we formulate a novel 3D-LESH feature extraction technique for 3D medical images to analyse volumetric images. Further, the proposed 3D-LESH algorithmis, for the first time, applied to medical MRI images. The final contribution is the design of an intelligent clinical decision support system (CDSS) as a multi-stage approach, combining novel 3D-LESH feature extraction with machine learning classifiers, to detect cancer from breast MRI scans. The proposed system applies contrast-limited adaptive histogram equalisation (CLAHE) to the MRI images before extracting 3D-LESH features. Furthermore, a selected subset of these features is fed into a machine-learning classifier, namely, a support vector machine (SVM), an extreme learning machine (ELM) or an echo state network (ESN) classifier, to detect abnormalities and distinguish between different stages of abnormality. We demonstrate the performance of the proposed technique by its application to benchmark breast cancer MRI images. The results indicate high-performance accuracy of the proposed system (98%±0.0050, with an area under a receiver operating charactertistic curve value of 0.9900 ± 0.0050) with multiple classifiers. When compared with the state-of-the-art wavelet-based feature extraction technique, statistical analysis provides conclusive evidence of the significance of our proposed 3D-LESH algorithm.
Summrina Kanwal Wajid, Amir Hussain 0001, Kaizhu Huang
Expert Syst. Appl.3
2018 Siamese network ensemble for visual tracking
Chenru Jiang, Jimin Xiao, Yanchun Xie, Tammam Tillo, Kaizhu Huang
Neurocomputing5
2018 A new two-layer mixture of factor analyzers with joint factor loading model for the classification of small dataset problems
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
Neurocomputing2
2018 Approximately optimizing NDCG using pair-wise loss
Xiao-Bo Jin, Guanggang Geng, Guosen Xie, Kaizhu Huang
Inf. Sci.4
2018 Banzhaf random forests: Cooperative game theory based random forests with consistency
Jianyuan Sun, Guoqiang Zhong 0001, Kaizhu Huang, Junyu Dong
Neural Networks3
2018 Zero-Shot Learning via Attribute Regression and Class Prototype Rectification
abstract
Zero-shot learning (ZSL) aims at classifying examples for unseen classes (with no training examples) given some other seen classes (with training examples). Most existing approaches exploit intermedia-level information (e.g., attributes) to transfer knowledge from seen classes to unseen classes. A common practice is to first learn projections from samples to attributes on seen classes via a regression method, and then apply such projections to unseen classes directly. However, it turns out that such a manner of learning strategy easily causes projection domain shift problem and hubness problem, which hinder the performance of ZSL task. In this paper, we also formulate ZSL as an attribute regression problem. However, different from general regression-based solutions, the proposed approach is novel in three aspects. First, a class prototype rectification method is proposed to connect the unseen classes to the seen classes. Here, a class prototype refers to a vector representation of a class, and it is also known as a class center, class signature, or class exemplar. Second, an alternating learning scheme is proposed for jointly performing attribute regression and rectifying the class prototypes. Finally, a new objective function which takes into consideration both the attribute regression accuracy and the class prototype discrimination is proposed. By introducing such a solution, domain shift problem and hubness problem can be mitigated. Experimental results on three public datasets (i.e., CUB200-2011, SUN Attribute, and aPaY) well demonstrate the effectiveness of our approach.
Changzhi Luo, Zhetao Li, Kaizhu Huang, Jiashi Feng, Meng Wang 0001
IEEE Trans. Image Process.3
2017 Field Support Vector Regression
Haochuan Jiang, Kaizhu Huang, Rui Zhang 0012
ICONIP (1)2
2017 Deep Mixtures of Factor Analyzers with Common Loadings: A Novel Deep Generative Approach to Clustering
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012
ICONIP (1)2
2017 Improve Deep Learning with Unsupervised Objective
Shufei Zhang, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)2
2017 Customer churn prediction in the telecommunication sector using a rough set approach
Adnan Amin, Sajid Anwar 0001, Awais Adnan, Khalid Alawfi, Amir Hussain 0001, Kaizhu Huang
Neurocomputing7
2017 Joint Learning of Unsupervised Dimensionality Reduction and Gaussian Mixture Model
Xi Yang 0008, Kaizhu Huang, John Yannis Goulermas, Rui Zhang 0012
Neural Process. Lett.2
2016 Learning Latent Features with Infinite Non-negative Binary Matrix Tri-factorization
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)2
2016 Learning from Few Samples with Memory Network
Shufei Zhang, Kaizhu Huang
ICONIP (1)2
2016 A fast projected fixed-point algorithm for large graph matching
Kaizhu Huang, Cheng-Lin Liu 0001
Pattern Recognit.2
2015 A Unified Gradient Regularization Family for Adversarial Examples
abstract
Adversarial examples are augmented data points generated by imperceptible perturbation of input samples. They have recently drawn much attention with the machine learning and data mining community. Being difficult to distinguish from real examples, such adversarial examples could change the prediction of many of the best learning models including the state-of-the-art deep learning models. Recent attempts have been made to build robust models that take into account adversarial examples. However, these methods can either lead to performance drops or lack mathematical motivations. In this paper, we propose a unified framework to build robust machine learning models against adversarial examples. More specifically, using the unified framework, we develop a family of gradient regularization methods that effectively penalize the gradient of loss function w.r.t. inputs. Our proposed framework is appealing in that it offers a unified view to deal with adversarial examples. It incorporates another recently-proposed perturbation based approach as a special case. In addition, we present some visual effects that reveals semantic meaning in those perturbations, and thus support our regularization method and provide another explanation for generalizability of adversarial examples. By applying this technique to Maxout networks, we conduct a series of experiments and achieve encouraging results on two benchmark datasets. In particular, we attain the best accuracy on MNIST data (without data augmentation) and competitive performance on CIFAR-10 data.
Chunchuan Lyu, Kaizhu Huang, Hai-Ning Liang
ICDM2
2015 Is DeCAF Good Enough for Accurate Image Classification?
Yajuan Cai, Guoqiang Zhong 0001, Yuchen Zheng 0001, Kaizhu Huang, Junyu Dong
ICONIP (2)4
2015 Two-layer Mixture of Factor Analyzers with Joint Factor Loading
abstract
Dimensionality Reduction (DR) is a fundamental yet active research topic in pattern recognition and machine learning. When used in classification, previous research usually performs DR separately, and then inputs the reduced features to other available models, e.g., Gaussian Mixture Model (GMM). Such independent learning could however significantly limit the classification performance, since the optimal subspace given by a particular DR approach may not be appropriate for the following classification model. More seriously, for high-dimensional data classification in the face of a limited number of samples (called small sample size or S3 problem), independent learning of DR and classification model may even deteriorate the classification accuracy. To solve this problem, we propose a joint learning model, called Two-layer Mixture of Factor Analyzers with Joint Factor Loading (2L-MJFA) for classification. More specifically, our proposed model enjoys a two-layer mixture structure, or a mixture of mixtures structure, with each component (representing each specific class) as another mixture model of Factor Analyzer (MFA). Importantly, all the involved factor analyzers are intentionally designed to share the same loading matrix. On one hand, such joint loading matrix can be considered as the dimensionality reduction matrix; on the other hand, a joint common matrix would largely reduce the parameters, making the proposed algorithm very suitable for S3 problems. We describe our model definition and propose a modified EM algorithm to optimize the model. A series of experiments demonstrates that our proposed model significantly outperforms the other three competitive algorithms on five data sets.
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas
IJCNN2
2015 WSDM'15 Workshop Summary / Scalable Data Analytics: Theory and Applications
abstract
The SDA workshop at WSDM 2015 is the fifth International Workshop on Scalable Data Analytics, following the previous four workshops of SDA respectively held at IEEE Big Data 2013, PAKDD 2014, IEEE Big Data 2014, and IEEE ICDM 2014. This series of workshops aims to provide professionals, researchers, and technologists with a single forum where they can discuss and share the state-of-the-art theories and applications of scalable data analytics technologies. In particular, in the era of information explosion, the scientific, biomedical, and engineering research communities are undergoing a profound transformation where discoveries and innovations increasingly rely on massive amounts of data. The characteristics of volume, velocity, variety and veracity originated in the massive big data then bring challenges to current data analytics techniques. The focus of the fifth SDA is to discuss how we can scale up data analytics techniques for modeling and analyzing big data from various domains.
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu
WSDM1
2015 DE2: Dynamic ensemble of ensembles for learning nonstationary data
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao
Neurocomputing2
2015 Maximum margin semi-supervised learning with irrelevant data
Haiqin Yang, Kaizhu Huang, Irwin King, Michael R. Lyu
Neural Networks2
2015 Learning Imbalanced Classifiers Locally and Globally with One-Side Probability Machine
Kaizhu Huang, Rui Zhang 0012, Xu-Cheng Yin
Neural Process. Lett.1
2015 MTC: A Fast and Robust Graph-Based Transductive Learning Method
abstract
Despite the great success of graph-based transductive learning methods, most of them have serious problems in scalability and robustness. In this paper, we propose an efficient and robust graph-based transductive classification method, called minimum tree cut (MTC), which is suitable for large-scale data. Motivated from the sparse representation of graph, we approximate a graph by a spanning tree. Exploiting the simple structure, we develop a linear-time algorithm to label the tree such that the cut size of the tree is minimized. This significantly improves graph-based methods, which typically have a polynomial time complexity. Moreover, we theoretically and empirically show that the performance of MTC is robust to the graph construction, overcoming another big problem of traditional graph-based methods. Extensive experiments on public data sets and applications on web-spam detection and interactive image segmentation demonstrate our method's advantages in aspect of accuracy, speed, and robustness.
Yan-Ming Zhang 0001, Kaizhu Huang, Guanggang Geng, Cheng-Lin Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2014 Unsupervised Dimensionality Reduction for Gaussian Mixture Model
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012
ICONIP (2)2
2014 Text Categorization with Diversity Random Forests
Xu-Cheng Yin, Kaizhu Huang
ICONIP (3)3
2014 A Novel Hybrid Approach for Combining Deep and Traditional Neural Networks
Rui Zhang 0012, Shufei Zhang, Kaizhu Huang
ICONIP (3)3
2014 Graphical lasso quadratic discriminant function and its application to character recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi
Neurocomputing2
2014 A novel classifier ensemble method with sparsity and diversity
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao, Khalid Iqbal, Zhi-Bin Wang
Neurocomputing2
2014 Robust Text Detection in Natural Scene Images
abstract
Text detection in natural scene images is an important prerequisite for many content-based image analysis tasks. In this paper, we propose an accurate and robust method for detecting texts in natural scene images. A fast and effective pruning algorithm is designed to extract Maximally Stable Extremal Regions (MSERs) as character candidates using the strategy of minimizing regularized variations. Character candidates are grouped into text candidates by the single-link clustering algorithm, where distance weights and clustering threshold are learned automatically by a novel self-training distance metric learning algorithm. The posterior probabilities of text candidates corresponding to non-text are estimated with a character classifier; text candidates with high non-text probabilities are eliminated and texts are identified with a text classifier. The proposed system is evaluated on the ICDAR 2011 Robust Reading Competition database; the f-measure is over 76%, much better than the state-of-the-art performance of 71%. Experiments on multilingual, street view, multi-orientation and even born-digital databases also demonstrate the effectiveness of the proposed method.
Xu-Cheng Yin, Xuwang Yin, Kaizhu Huang, Hongwei Hao
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Combination of Classification and Clustering Results with Label Propagation
abstract
This letter considers the combination of multiple classification and clustering results to improve the prediction accuracy. First, an object-similarity graph is constructed from multiple clustering results. The labels predicted by the classification models are then propagated on this graph to adaptively satisfy the smoothness of the prediction over the graph. The convex learning problem is efficiently solved by the label propagation algorithm. A semi-supervised extension is also provided to further improve the performance. Experiments on 11 tasks identify the validity of the proposed models.
Xu-Yao Zhang, Peipei Yang, Yan-Ming Zhang 0001, Kaizhu Huang, Cheng-Lin Liu 0001
IEEE Signal Process. Lett.4
2014 Learning Locality Preserving Graph from Data
abstract
Machine learning based on graph representation, or manifold learning, has attracted great interest in recent years. As the discrete approximation of data manifold, the graph plays a crucial role in these kinds of learning approaches. In this paper, we propose a novel learning method for graph construction, which is distinct from previous methods in that it solves an optimization problem with the aim of directly preserving the local information of the original data set. We show that the proposed objective has close connections with the popular Laplacian Eigenmap problem, and is hence well justified. The optimization turns out to be a quadratic programming problem with n(n-1)/2 variables (n is the number of data points). Exploiting the sparsity of the graph, we further propose a more efficient cutting plane algorithm to solve the problem, making the method better scalable in practice. In the context of clustering and semi-supervised learning, we demonstrated the advantages of our proposed method by experiments.
Yan-Ming Zhang 0001, Kaizhu Huang, Xinwen Hou, Cheng-Lin Liu 0001
IEEE Trans. Cybern.2
2013 Feature Transformation with Class Conditional Decorrelation
abstract
The well-known feature transformation model of Fisher linear discriminant analysis (FDA) can be decomposed into an equivalent two-step approach: whitening followed by principal component analysis (PCA) in the whitened space. By proving that whitening is the optimal linear transformation to the Euclidean space in the sense of minimum log-determinant divergence, we propose a transformation model called class conditional decor relation (CCD). The objective of CCD is to diagonalize the covariance matrices of different classes simultaneously, which is efficiently optimized using a modified Jacobi method. CCD is effective to find the common principal components among multiple classes. After CCD, the variables become class conditionally uncorrelated, which will benefit the subsequent classification tasks. Combining CCD with the nearest class mean (NCM) classification model can significantly improve the classification accuracy. Experiments on 15 small-scale datasets and one large-scale dataset (with 3755 classes) demonstrate the scalability of CCD for different applications. We also discuss the potential applications of CCD for other problems such as Gaussian mixture models and classifier ensemble learning.
Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001
ICDM2
2013 Dynamic Ensemble of Ensembles in Nonstationary Environments
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao
ICONIP (2)2
2013 One-Side Probability Machine: Learning Imbalanced Classifiers Locally and Globally
Rui Zhang 0012, Kaizhu Huang
ICONIP (2)2
2013 Fast kNN Graph Construction with Locality Sensitive Hashing
Yan-Ming Zhang 0001, Kaizhu Huang, Guanggang Geng, Cheng-Lin Liu 0001
ECML/PKDD (2)2
2013 Accurate and robust text detection: a step-in for text retrieval in natural scene images
abstract
We propose and implement a robust text detection system, which is a prominent step-in for text retrieval in natural scene images or videos. Our system includes several key components: (1) A fast and effective pruning algorithm is designed to extract Maximally Stable Extremal Regions as character candidates using the strategy of minimizing regularized variations. (2) Character candidates are grouped into text candidates by the single-link clustering algorithm, where distance weights and threshold of clustering are learned automatically by a novel self-training distance metric learning algorithm. (3) The posterior probabilities of text candidates corresponding to non-text are estimated with an character classifier; text candidates with high probabilities are then eliminated and finally texts are identified with a text classifier. The proposed system is evaluated on the ICDAR 2011 Robust Reading Competition dataset and a publicly available multilingual dataset; the f measures are over 76% and 74% which are significantly better than the state-of-the-art performances of 71% and 65%, respectively.
Xu-Cheng Yin, Xuwang Yin, Kaizhu Huang, Hongwei Hao
SIGIR3
2013 Geometry preserving multi-task metric learning
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001
Mach. Learn.2
2013 A multi-task framework for metric learning with common subspace
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001
Neural Comput. Appl.2
2012 Multiple Outlooks Learning with Support Vector Machines
Yinglu Liu, Xu-Yao Zhang, Kaizhu Huang, Xinwen Hou, Cheng-Lin Liu 0001
ICONIP (3)3
2012 Manifold Regularized Multi-Task Learning
Peipei Yang, Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001
ICONIP (3)3
2012 Classifier Ensemble Using a Heuristic Learning with Sparsity and Diversity
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao, Khalid Iqbal, Zhi-Bin Wang
ICONIP (2)2
2012 Geometry Preserving Multi-task Metric Learning
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001
ECML/PKDD (1)2
2012 Joint learning of error-correcting output codes and dichotomizers from data
Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001
Neural Comput. Appl.2
2012 Maxi-Min discriminant analysis via online learning
Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001
Neural Networks2
2011 Fast and Robust Graph-based Transductive Learning via Minimum Tree Cut
abstract
In this paper, we propose an efficient and robust algorithm for graph-based transductive classification. After approximating a graph with a spanning tree, we develop a linear-time algorithm to label the tree such that the cut size of the tree is minimized. This significantly improves typical graph-based methods, which either have a cubic time complexity (for a dense graph) or O(kn2) (for a sparse graph with k denoting the node degree). Furthermore, our method shows great robustness to the graph construction both theoretically and empirically; this overcomes another big problem of traditional graph-based methods. In addition to its good scalability and robustness, the proposed algorithm demonstrates high accuracy. In particular, on a graph with 400,000 nodes (in which 10,000 nodes are labeled) and 10,455,545 edges, our algorithm achieves the highest accuracy of 99.6% but takes less than 10 seconds to label all the unlabeled data.
Yan-Ming Zhang 0001, Kaizhu Huang, Cheng-Lin Liu 0001
ICDM2
2011 Low Rank Metric Learning with Manifold Regularization
abstract
In this paper, we present a semi-supervised method to learn a low rank Mahalanobis distance function. Based on an approximation to the projection distance from a manifold, we propose a novel parametric manifold regularizer. In contrast to previous approaches that usually exploit side information only, our proposed method can further take advantages of the intrinsic manifold information from data. In addition, we focus on learning a metric of low rank directly, this is different from traditional approaches that often enforce the l1norm on the metric. The resulting configuration is convex with respect to the manifold structure and the distance function, respectively. We solve it with an alternating optimization algorithm, which proves effective to find a satisfactory solution. For efficient implementation, we even present a fast algorithm, in which the manifold structure and the distance function are learned independently without alternating minimization. Experimental results over 12 standard UCI data sets demonstrate the advantages of our method.
Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001
ICDM2
2011 Graphical Lasso Quadratic Discriminant Function for Character Recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi
ICONIP (3)2
2011 Multi-Task Low-Rank Metric Learning Based on Common Subspace
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001
ICONIP (2)2
2011 Pattern Field Classification with Style Normalized Transformation
Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001
IJCAI2
2011 Generalized sparse metric learning with relative comparisons
Kaizhu Huang, Yiming Ying, Colin Campbell
Knowl. Inf. Syst.1
2011 Exchange rate prediction with non-numerical information
Zhi-Bin Wang, Hongwei Hao, Xu-Cheng Yin, Kaizhu Huang
Neural Comput. Appl.5
2011 FMI image based rock structure classification using classifier combination
Xu-Cheng Yin, Hongwei Hao, Zhi-Bin Wang, Kaizhu Huang
Neural Comput. Appl.5
2011 m-SNE: Multiview Stochastic Neighbor Embedding
abstract
Dimension reduction has been widely used in real-world applications such as image retrieval and document classification. In many scenarios, different features (or multiview data) can be obtained, and how to duly utilize them is a challenge. It is not appropriate for the conventional concatenating strategy to arrange features of different views into a long vector. That is because each view has its specific statistical property and physical interpretation. Even worse, the performance of the concatenating strategy will deteriorate if some views are corrupted by noise. In this paper, we propose a multiview stochastic neighbor embedding (m-SNE) that systematically integrates heterogeneous features into a unified representation for subsequent processing based on a probabilistic framework. Compared with conventional strategies, our approach can automatically learn a combination coefficient for each view adapted to its contribution to the data embedding. This combination coefficient plays an important role in utilizing the complementary information in multiview data. Also, our algorithm for learning the combination coefficient converges at a rate of O(1/k(2)), which is the optimal rate for smooth problems. Experiments on synthetic and real data sets suggest the effectiveness and robustness of m-SNE for data visualization, image retrieval, object categorization, and scene recognition.
Bo Xie 0002, Yang Mu, Dacheng Tao, Kaizhu Huang
IEEE Trans. Syst. Man Cybern. Part B4
2010 Similar Handwritten Chinese Characters Recognition by Critical Region Selection Based on Average Symmetric Uncertainty
abstract
We consider the problem of similar Chinese character recognition in this paper. Engaging the Average Symmetric Uncertainty (ASU) criterion to measure the correlation between different image regions and the class label, we manage to detect the most critical regions for each pair of similar characters. These critical regions are proved to contain more discriminative information and hence can largely benefit the classification accuracy for similar characters. We conduct a series of experiments on the CASIA Chinese character data set. Experimental results show that our proposed method is superior to three competitive approaches in terms of both accuracy and efficiency.
Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001
ICFHR2
2010 Ellipse Detection with an Improved Randomized Hough Transform for Intellectual Phacoemulsification Surgery Systems
Hongwei Hao, Xu-Cheng Yin, Zhi-Bin Wang, Kaizhu Huang
ICIP5
2010 Learning ECOC and Dichotomizers Jointly from Data
Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001
ICONIP (1)2
2010 Dimensionality Reduction by Minimal Distance Maximization
abstract
In this paper, we propose a novel discriminant analysis method, called Minimal Distance Maximization (MDM). In contrast to the traditional LDA, which actually maximizes the average divergence among classes, MDM attempts to find a low-dimensional subspace that maximizes the minimal (worst-case) divergence among classes. This ``minimal" setting solves the problem caused by the ``average" setting of LDA that tends to merge similar classes with smaller divergence when used for multi-class data. Furthermore, we elegantly formulate the worst-case problem as a convex problem, making the algorithm solvable for larger data sets. Experimental results demonstrate the advantages of our proposed method against five other competitive approaches on one synthetic and six real-life data sets.
Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001
ICPR2
2010 Robust Metric Learning by Smooth Optimization
Kaizhu Huang, Rong Jin 0001, Zenglin Xu, Cheng-Lin Liu 0001
UAI1
2010 Sparse learning for support vector classification
Kaizhu Huang, Danian Zheng, Jun Sun 0004, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi
Pattern Recognit. Lett.1
2009 GSML: A Unified Framework for Sparse Metric Learning
abstract
There has been significant recent interest in sparse metric learning (SML) in which we simultaneously learn both a good distance metric and a low-dimensional representation. Unfortunately, the performance of existing sparse metric learning approaches is usually limited because the authors assumed certain problem relaxations or they target the SML objective indirectly. In this paper, we propose a Generalized Sparse Metric Learning method (GSML). This novel framework offers a unified view for understanding many of the popular sparse metric learning algorithms including the Sparse Metric Learning framework proposed, the Large Margin Nearest Neighbor (LMNN), and the D-ranking Vector Machine (D-ranking VM). Moreover, GSML also establishes a close relationship with the Pairwise Support Vector Machine. Furthermore, the proposed framework is capable of extending many current non-sparse metric learning models such as Relevant Vector Machine (RCA) and a state-of-the-art method proposed into their sparse versions. We present the detailed framework, provide theoretical justifications, build various connections with other models, and propose a practical iterative optimization method, making the framework both theoretically important and practically scalable for medium or large datasets. A series of experiments show that the proposed approach can outperform previous methods in terms of both test accuracy and dimension reduction, on six real-world benchmark datasets.
Kaizhu Huang, Yiming Ying, Colin Campbell
ICDM1
2009 Exchange Rate Forecasting Using Classifier Ensemble
Zhi-Bin Wang, Hongwei Hao, Xu-Cheng Yin, Kaizhu Huang
ICONIP (1)5
2009 A Rock Structure Recognition System Using FMI Images
Xu-Cheng Yin, Hongwei Hao, Zhi-Bin Wang, Kaizhu Huang
ICONIP (1)5
2009 Supervised Self-taught Learning: Actively transferring knowledge from unlabeled data
abstract
We consider the task of Self-taught Learning (STL) from unlabeled data. In contrast to semi-supervised learning, which requires unlabeled data to have the same set of class labels as labeled data, STL can transfer knowledge from different types of unlabeled data. STL uses a three-step strategy: (1) learning high-level representations from unlabeled data only, (2) re-constructing the labeled data via such representations and (3) building a classifier over the re-constructed labeled data. However, the high-level representations which are exclusively determined by the unlabeled data, may be inappropriate or even misleading for the latter classifier training process. In this paper, we propose a novel Supervised Self-taught Learning (SSTL) framework that successfully integrates the three isolated steps of STL into a single optimization problem. Benefiting from the interaction between the classifier optimization and the process of choosing high-level representations, the proposed model is able to select those discriminative representations which are more appropriate for classification. One important feature of our novel framework is that the final optimization can be iteratively solved with convergence guaranteed. We evaluate our novel framework on various data sets. The experimental results show that the proposed SSTL can outperform STL and traditional supervised learning methods in certain instances.
Kaizhu Huang, Zenglin Xu, Irwin King, Michael R. Lyu, Colin Campbell
IJCNN1
2009 Sparse Metric Learning via Smooth Optimization
abstract
In this paper we study the problem of learning a low-dimensional (sparse) distance matrix. We propose a novel metric learning model which can simultaneously conduct dimension reduction and learn a distance matrix. The sparse representation involves a mixed-norm regularization which is non-convex. We then show that it can be equivalently formulated as a convex saddle (min-max) problem. From this saddle representation, we develop an efficient smooth optimization approach for sparse metric learning although the learning model is based on a non-differential loss function. This smooth optimization approach has an optimal convergence rate of $O(1 /\ell^2)$ for smooth problems where $\ell$ is the iteration number. Finally, we run experiments to validate the effectiveness and efficiency of our sparse metric learning model on various datasets.
Yiming Ying, Kaizhu Huang, Colin Campbell
NIPS2
2009 Enhanced protein fold recognition through a novel data integration approach
abstract
BACKGROUND: Protein fold recognition is a key step in protein three-dimensional (3D) structure discovery. There are multiple fold discriminatory data sources which use physicochemical and structural properties as well as further data sources derived from local sequence alignments. This raises the issue of finding the most efficient method for combining these different informative data sources and exploring their relative significance for protein fold classification. Kernel methods have been extensively used for biological data analysis. They can incorporate separate fold discriminatory features into kernel matrices which encode the similarity between samples in their respective data sources. RESULTS: In this paper we consider the problem of integrating multiple data sources using a kernel-based approach. We propose a novel information-theoretic approach based on a Kullback-Leibler (KL) divergence between the output kernel matrix and the input kernel matrix so as to integrate heterogeneous data sources. One of the most appealing properties of this approach is that it can easily cope with multi-class classification and multi-task learning by an appropriate choice of the output kernel matrix. Based on the position of the output and input kernel matrices in the KL-divergence objective, there are two formulations which we respectively refer to as MKLdiv-dc and MKLdiv-conv. We propose to efficiently solve MKLdiv-dc by a difference of convex (DC) programming method and MKLdiv-conv by a projected gradient descent algorithm. The effectiveness of the proposed approaches is evaluated on a benchmark dataset for protein fold recognition and a yeast protein function prediction problem. CONCLUSION: Our proposed methods MKLdiv-dc and MKLdiv-conv are able to achieve state-of-the-art performance on the SCOP PDB-40D benchmark dataset for protein fold prediction and provide useful insights into the relative significance of informative data sources. In particular, MKLdiv-dc further improves the fold discrimination accuracy to 75.19% which is a more than 5% improvement over competitive Bayesian probabilistic and SVM margin-based kernel learning methods. Furthermore, we report a competitive performance on the yeast protein function prediction problem.
Yiming Ying, Kaizhu Huang, Colin Campbell
BMC Bioinform.2
2009 Localized support vector regression for time series prediction
Haiqin Yang, Kaizhu Huang, Irwin King, Michael R. Lyu
Neurocomputing2
2009 Arbitrary Norm Support Vector Machines
abstract
Support vector machines (SVM) are state-of-the-art classifiers. Typically L2-norm or L1-norm is adopted as a regularization term in SVMs, while other norm-based SVMs, for example, the L0-norm SVM or even the L(infinity)-norm SVM, are rarely seen in the literature. The major reason is that L0-norm describes a discontinuous and nonconvex term, leading to a combinatorially NP-hard optimization problem. In this letter, motivated by Bayesian learning, we propose a novel framework that can implement arbitrary norm-based SVMs in polynomial time. One significant feature of this framework is that only a sequence of sequential minimal optimization problems needs to be solved, thus making it practical in many real applications. The proposed framework is important in the sense that Bayesian priors can be efficiently plugged into most learning methods without knowing the explicit form. Hence, this builds a connection between Bayesian learning and the kernel machines. We derive the theoretical framework, demonstrate how our approach works on the L0-norm SVM as a typical example, and perform a series of experiments to validate its advantages. Experimental results on nine benchmark data sets are very encouraging. The implemented L0-norm is competitive with or even better than the standard L2-norm SVM in terms of accuracy but with a reduced number of support vectors, -9.46% of the number on average. When compared with another sparse model, the relevance vector machine, our proposed algorithm also demonstrates better sparse properties with a training speed over seven times faster.
Kaizhu Huang, Danian Zheng, Irwin King, Michael R. Lyu
Neural Comput.1
2009 A novel kernel-based maximum a posteriori classification method
Zenglin Xu, Kaizhu Huang, Jianke Zhu, Irwin King, Michael R. Lyu
Neural Networks2
2008 Semi-supervised text categorization by active search
abstract
In automated text categorization, given a small number of labeled documents, it is very challenging, if not impossible, to build a reliable classifier that is able to achieve high classification accuracy. To address this problem, a novel web-assisted text categorization framework is proposed in this paper. Important keywords are first automatically identified from the available labeled documents to form the queries. Search engines are then utilized to retrieve from the Web a multitude of relevant documents, which are then exploited by a semi-supervised framework. To our best knowledge, this work is the first study of this kind. Extensive experimental study shows the encouraging results of the proposed text categorization framework: using Google as the web search engine, the proposed framework is able to reduce the classification error by 30% when compared with the state-of-the-art supervised text categorization method.
Zenglin Xu, Rong Jin 0001, Kaizhu Huang, Michael R. Lyu, Irwin King
CIKM3
2008 Direct Zero-Norm Optimization for Feature Selection
abstract
Zero-norm, defined as the number of non-zero elements in a vector, is an ideal quantity for feature selection. However, minimization of zero-norm is generally regarded as a combinatorially difficult optimization problem. In contrast to previous methods that usually optimize a surrogate of zero-norm, we propose a direct optimization method to achieve zero-norm for feature selection in this paper. Based on Expectation Maximization (EM), this method boils down to solving a sequence of Quadratic Programming problems and hence can be practically optimized in polynomial time. We show that the proposed optimization technique has a nice Bayesian interpretation and converges to the true zero norm asymptotically, provided that a good starting point is given. Following the scheme of our proposed zero-norm, we even show that an arbitrary-norm based Support Vector Machine can be achieved in polynomial time. A series of experiments demonstrate that our proposed EM based zero-norm outperforms other state-of-the-art methods for feature selection on biological microarray data and UCI data, in terms of both the accuracy and the learning efficiency.
Kaizhu Huang, Irwin King, Michael R. Lyu
ICDM1
2008 Semi-supervised Learning from General Unlabeled Data
abstract
We consider the problem of semi-supervised learning (SSL) from general unlabeled data, which may contain irrelevant samples. Within the binary setting, our model manages to better utilize the information from unlabeled data by formulating them as a three-class (-1,+1, 0) mixture, where class 0 represents the irrelevant data. This distinguishes our work from the traditional SSL problem where unlabeled data are assumed to contain relevant samples only, either +1 or -1, which are forced to be the same as the given labeled samples. This work is also different from another family of popular models, universum learning (universum means "irrelevant" data), in that the universum need not to be specified beforehand. One significant contribution of our proposed framework is that such irrelevant samples can be automatically detected from the available unlabeled data, even though they are mixed with relevant data. This hence presents a general SSL framework that does not force "clean" unlabeled data.More importantly, we formulate this general learning framework as a Semi-definite Programming problem, making it solvable in polynomial time. A series of experiments demonstrate that the proposed framework can outperform the traditional SSL on both synthetic and real data.
Kaizhu Huang, Zenglin Xu, Irwin King, Michael R. Lyu
ICDM1
2008 Efficient Minimax Clustering Probability Machine by Generalized Probability Product Kernel
abstract
Minimax Probability Machine (MPM), learning a decision function by minimizing the maximum probability of misclassification, has demonstrated very promising performance in classification and regression. However, MPM is often challenged for its slow training and test procedures. Aiming to solve this problem, we propose an efficient model named Minimax Clustering Probability Machine (MCPM). Following many traditional methods, we represent training data points by several clusters. Different from these methods, a Generalized Probability Product Kernel is appropriately defined to grasp the inner distributional information over the clusters. Incorporating clustering information via a non-linear kernel, MCPM can fast train and test in classification problem with promising performance. Another appealing property of the proposed approach is that MCPM can still derive an explicit worst-case accuracy bound for the decision boundary. Experimental results on synthetic and real data validate the effectiveness of MCPM for classification while attaining high accuracy.
Haiqin Yang, Kaizhu Huang, Irwin King, Michael R. Lyu
IJCNN2
2008 Maxi-Min Margin Machine: Learning Large Margin Classifiers Locally and Globally
abstract
In this paper, we propose a novel large margin classifier, called the maxi-min margin machine M(4). This model learns the decision boundary both locally and globally. In comparison, other large margin classifiers construct separating hyperplanes only either locally or globally. For example, a state-of-the-art large margin classifier, the support vector machine (SVM), considers data only locally, while another significant model, the minimax probability machine (MPM), focuses on building the decision hyperplane exclusively based on the global information. As a major contribution, we show that SVM yields the same solution as M(4) when data satisfy certain conditions, and MPM can be regarded as a relaxation model of M(4). Moreover, based on our proposed local and global view of data, another popular model, the linear discriminant analysis, can easily be interpreted and extended as well. We describe the M(4) model definition, provide a geometrical interpretation, present theoretical justifications, and propose a practical sequential conic programming method to solve the optimization problem. We also show how to exploit Mercer kernels to extend M(4) for nonlinear classifications. Furthermore, we perform a series of evaluations on both synthetic data sets and real-world benchmark data sets. Comparison with SVM and MPM demonstrates the advantages of our new model.
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu
IEEE Trans. Neural Networks1
2007 An SVM-Based High-accurate Recognition Approach for Handwritten Numerals by Using Difference Features
abstract
Handwritten numeral recognition is an important pattern recognition task. It can be widely used in various domains, e.g., bank money recognition, which requires a very high recognition rate. As a state-of-the-art classifier, support vector machine (SVM), has been extensively used in this area. Typically, SVM is trained in a batch model, i.e., all data points are simultaneously input for training the classification boundary. However, some slightly exceptional data, only accounting for a small proportion, are critical for the recognition rates. Training a classifier among all the data may possibly treat such legal but slightly exceptional samples as "noise ". In this paper, we propose a novel approach to attack this problem. This approach exploits a two-stage framework by using difference features. In the first stage, a regular SVM is trained on all the training data; in the second stage, only the samples misclassified in the first stage are specially considered. Therefore, the performance can be lifted. The number of misclassifications is often small because of the good performance of SVM. This will present difficulties in training an accurate SVM engine only for these misclassified samples. We then further propose a multi-way to binary approach using difference features. This approach successfully transforms multi-category classification to binary classification and expands the training samples greatly. In order to evaluate the proposed method, experiments are performed on 10,000 handwritten numeral samples extracted from real banks forms. This new algorithm achieves 99.0% accuracy. In comparison, the traditional SVM only gets 98.4%.
Kaizhu Huang, Jun Sun 0004, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi
ICDAR1
2007 Degraded Character Recognition by Complementary Classifiers Combination
abstract
Character degradation is a big problem for machine printed character recognition. Two main reasons for degradation are extrinsic image degradation such as blurring and low image dimension, and intrinsic degradation caused by font variations. A recognition method that combines two complementary classifiers is proposed in this paper. The local feature based classifier extracts the local contour direction changes, which is effective for character patterns with less structure deterioration. The global feature based classifier extracts the texture distribution of the character image, which is effective when the character structure is hard to discriminate. The two complementary classifiers are combined by candidate fusion in a coarse-to-fine style. Experiments are carried on degraded Chinese character recognition. The results prove the effectiveness of our method.
Jun Sun 0004, Kaizhu Huang, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi
ICDAR2
2007 Kernel Maximum a Posteriori Classification with Error Bound Analysis
Zenglin Xu, Kaizhu Huang, Jianke Zhu, Irwin King, Michael R. Lyu
ICONIP (1)2
2006 A Hybrid Handwritten Chinese Address Recognition Approach
Kaizhu Huang, Jun Sun 0004, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi, Chong Long, Li Zhuang, Xiaoyan Zhu 0001
ICONIP (2)1
2006 Local Support Vector Regression for Financial Time Series Prediction
abstract
We consider the regression problem for financial time series. Typically, financial time series are non-stationary and volatile in nature. Because of its good generalization power and the tractability of the problem, the Support Vector Regression (SVR) has been extensively applied in financial time series prediction. The standard SVR adopts the lp-norm (p = 1 or 2) to model the functional complexity of the whole data set and employs a fixed ε-tube to tolerate noise. Although this approach has proved successful both theoretically and empirically, it considers data in a global fashion only. Therefore it may lack the flexibility to capture the local trend of data; this is a critical aspect of volatile data, especially financial time series data. Aiming to address this issue, we propose the Local Support Vector Regression (LSVR) model. This novel model is demonstrated to provide a systematic and automatic scheme to adapt the margin locally and flexibly; the margin is fixed globally in the standard SVR. Therefore, the LSVR can tolerate noise adaptively. We provide both theoretical justifications and empirical evaluations for this novel model. The experimental results on synthetic data and real financial data demonstrate its advantages over the standard SVR.
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu
IJCNN1
2006 Imbalanced learning with a biased minimax probability machine
abstract
Imbalanced learning is a challenged task in machine learning. In this context, the data associated with one class are far fewer than those associated with the other class. Traditional machine learning methods seeking classification accuracy over a full range of instances are not suitable to deal with this problem, since they tend to classify all the data into a majority class, usually the less important class. In this correspondence, the authors describe a new approach named the biased minimax probability machine (BMPM) to deal with the problem of imbalanced learning. This BMPM model is demonstrated to provide an elegant and systematic way for imbalanced learning. More specifically, by controlling the accuracy of the majority class under all possible choices of class-conditional densities with a given mean and covariance matrix, this model can quantitatively and systematically incorporate a bias for the minority class. By establishing an explicit connection between the classification accuracy and the bias, this approach distinguishes itself from the many current imbalanced-learning methods; these methods often impose a certain bias on the minority data by adapting intermediate factors via the trial-and-error procedure. The authors detail the theoretical foundation, prove its solvability, propose an efficient optimization algorithm, and perform a series of experiments to evaluate the novel model. The comparison with other competitive methods demonstrates the effectiveness of this new model.
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu
IEEE Trans. Syst. Man Cybern. Part B1
2004 Learning Classifiers from Imbalanced Data Based on Biased Minimax Probability Machine
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu
CVPR (2)1
2004 Learning large margin classifiers locally and globally
abstract
A new large margin classifier, named Maxi-Min Margin Machine (M4) is proposed in this paper. This new classifier is constructed based on both a "local: and a "global" view of data, while the most popular large margin classifier, Support Vector Machine (SVM) and the recently-proposed important model, Minimax Probability Machine (MPM) consider data only either locally or globally. This new model is theoretically important in the sense that SVM and MPM can both be considered as its special case. Furthermore, the optimization of M4 can be cast as a sequential conic programming problem, which can be solved efficiently. We describe the M4 model definition, provide a clear geometrical interpretation, present theoretical justifications, propose efficient solving methods, and perform a series of evaluations on both synthetic data sets and real world benchmark data sets. Its comparison with SVM and MPM also demonstrates the advantages of our new model.
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu
ICML1
2004 Outliers Treatment in Support Vector Regression for Financial Time Series Prediction
Haiqin Yang, Kaizhu Huang, Lai-Wan Chan, Irwin King, Michael R. Lyu
ICONIP2
2004 Biased support vector machine for relevance feedback in image retrieval
abstract
Recently, support vector machines (SVMs) have been engaged on relevance feedback tasks in content-based image retrieval. Typical approaches by SVMs treat the relevance feedback as a strict binary classification problem. However, these approaches do not consider an important issue of relevance feedback, i.e. the unbalanced dataset problem, in which the negative instances largely outnumber the positive instances. For solving this problem, we propose a novel technique to formulate the relevance feedback based on a modified SVM called biased support vector machine (Biased SVM or BSVM). Mathematical formulation and explanations are provided for showing the advantages. Experiments are conducted to evaluate the performance of our algorithms, in which promising results demonstrate the effectiveness of our techniques.
Steven C. H. Hoi, Chi-Hang Chan, Kaizhu Huang, Michael R. Lyu, Irwin King
IJCNN3
2004 The Minimum Error Minimax Probability Machine
Kaizhu Huang, Haiqin Yang, Irwin King, Michael R. Lyu, Lai-Wan Chan
J. Mach. Learn. Res.1
2003 Finite Mixture Model of Bounded Semi-naive Bayesian Networks Classifier
Kaizhu Huang, Irwin King, Michael R. Lyu
ICANN1
2003 Discriminative training of Bayesian Chow-Liu multinet classifiers
abstract
Discriminative classifiers such as support vector machines directly learn a discriminant function or a posterior probability model to perform classification. On the other hand, generative classifiers often learn a joint probability model and then use Bayes rules to construct a posterior classifier from this model. In general, generative classifiers are not as accurate as discriminant classifier. However generative classifiers provide a principled way to handle the missing information problems, which discriminant classifiers cannot easily deal with. To achieve good performances in various classification tasks, it is better to combine these two strategies. In this paper, we develop a novel method to iteratively train a kind of generative Bayesian classifier: Bayesian Chow-Liu multinet classifier in a discriminative way. Different with the traditional Bayesian multinet classifiers, our discriminative method adds into the optimization function a penalty item, which represents the divergence between classes as big as possible. We state the theoretical justification, outline of the algorithm and also perform a series of experiments to demonstrate the advantages of our method. The experiments results are promising and encouraging.
Kaizhu Huang, Irwun King, Michael R. Lyu
IJCNN1