Zhi-Hong Deng 0001

dblp:161/4814-1 · also Zhihong Deng 0001 · DBLP profile ↗
← Back
100ranked-venue papers
22as first author
30since 2021 · last 2026
0000-0003-2285-9169ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 78 · 14 first-author · 30 since 2021Databases, data management, data science and information retrieval · 25 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Efficient Thought Space Exploration Through Strategic Intervention
abstract
While large language models (LLMs) demonstrate emerging reasoning capabilities, current inference-time expansion methods incur prohibitive computational costs through exhaustive sampling. Through analyzing decoding trajectories, we observe that most next-token predictions align well with the golden output, except for a few critical tokens that lead to deviations. Inspired by this phenomenon, we propose a novel Hint-Practice Reasoning (HPR) framework that operationalizes this insight through two synergistic components: 1) a hinter (powerful LLM) that provides probabilistic guidance at critical decision points, and 2) a practitioner (efficient smaller model) that executes major reasoning steps. The framework's core innovation lies in Distributional Inconsistency Reduction (DIR), a theoretically-grounded metric that dynamically identifies intervention points by quantifying the divergence between practitioner's reasoning trajectory and hinter's expected distribution in a tree-structured probabilistic space. Through iterative tree updates guided by DIR, HPR reweights promising reasoning paths while deprioritizing low-probability branches. Experiments across arithmetic and commonsense reasoning benchmarks demonstrate HPR's state-of-the-art efficiency-accuracy tradeoffs: it achieves comparable performance to self-consistency and MCTS baselines while decoding only 1/5 tokens, and outperforms existing methods by at most 5.1% absolute accuracy while maintaining similar or lower FLOPs.
Ziheng Li 0003, Hengyi Cai, Xiaochi Wei, Yuchen Li 0006, Shuaiqiang Wang, Zhi-Hong Deng 0001, Dawei Yin 0001
AAAI6
2026 SDA: Steering-Driven Distribution Alignment for Open LLMs Without Fine-Tuning
abstract
With the rapid advancement of large language models (LLMs), their deployment in real-world applications has become increasingly widespread. LLMs are expected to deliver robust performance across diverse tasks, user preferences, and practical scenarios. However, as demands grow, ensuring that LLMs produce responses aligned with human intent remains a foundational challenge. In particular, aligning model behavior effectively and efficiently during inference, without costly retraining or extensive supervision, is both a critical requirement and a non-trivial technical endeavor. To address the challenge, we propose SDA (Steering-Driven Distribution Alignment), a training-free and model-agnostic alignment framework designed for open-source LLMs. SDA dynamically redistributes model output probabilities based on user-defined alignment instructions, enhancing alignment between model behavior and human intents without fine-tuning. The method is lightweight, resource-efficient, and compatible with a wide range of open-source LLMs. It can function independently during inference or be integrated with training-based alignment strategies. Moreover, SDA supports personalized preference alignment, enabling flexible control over the model’s response behavior. Empirical results demonstrate that SDA consistently improves alignment performance across 8 open-source LLMs with varying scales and diverse origins, evaluated on three key alignment dimensions, helpfulness, harmlessness, and honesty (3H). Specifically, SDA achieves average gains of 64.4% in helpfulness, 30% in honesty and 11.5% in harmlessness across the tested models, indicating its effectiveness and generalization across diverse models and application scenarios.
Zhi-Hong Deng 0001
AAAI2
2025 Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
abstract
Although the LLM-based in-context learning (ICL) paradigm has demonstrated considerable success across various natural language processing tasks, it encounters challenges in event detection. This is because LLMs lack an accurate understanding of event triggers and tend to make over-interpretation, which cannot be effectively corrected through in-context examples alone. In this paper, we focus on the most challenging one-shot setting and propose KeyCP++, a keyword-centric chain-of-thought prompting approach. KeyCP++ addresses the weaknesses of conventional ICL by automatically annotating the logical gaps between input text and detection results for the demonstrations. Specifically, to generate in-depth and meaningful rationale, KeyCP++ constructs a trigger discrimination prompting template. It incorporates the exemplary triggers (a.k.a keywords) into the prompt as the anchor to simply trigger profiling, let LLM propose candidate triggers, and justify each candidate. These propose-and-judge rationales help LLMs mitigate over-reliance on the keywords and promote detection rule learning. Extensive experiments demonstrate the effectiveness of our approach, showcasing significant advancements in one-shot event detection.
Ziheng Li 0003, Zhi-Hong Deng 0001
ECAI2
2025 Instruction-Based Molecular Graph Generation with Unified Text-Graph Diffusion Model
abstract
Recent advancements in computational chemistry have increasingly emphasized generating and editing molecules from textual instructions. However, integrating graph generation with instruction understanding remains challenging, as most existing approaches either rely on molecular sequences in text modality with limited structural information, or struggle with multimodal alignment in graph diffusion methods. To address these limitations, we propose UTGDiff (Unified Text-Graph Diffusion Model), a novel framework that utilizes pre-trained language models for discrete graph diffusion, enabling the generation of molecular graphs from instructions. UTGDiff introduces a unified text-graph transformer as a denoising network, adapted with minimal modifications from language models to process graph data via attention bias. Experimental results show that UTGDiff consistently outperforms both sequence-based and conditional graph-diffusion baselines on instruction-based molecule generation and editing tasks with fewer parameters, covering instructions specifying molecular structures or properties.
Yuran Xiang, Haiteng Zhao, Zhi-Hong Deng 0001
ECAI4
2025 SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
abstract
Transformer-based large language models (LLMs) have already achieved remarkable results on long-text tasks, but the limited GPU memory (VRAM) resources struggle to accommodate the linearly growing demand for key-value (KV) cache as the sequence length increases, which has become a bottleneck for the application of LLMs on long sequences. Existing KV cache compression methods include eviction, merging, or quantization of the KV cache to reduce its size. However, compression results in irreversible information forgetting, potentially affecting the accuracy of subsequent decoding. In this paper, we propose SpeCache, which takes full advantage of the large and easily expandable CPU memory to offload the complete KV cache, and dynamically fetches KV pairs back in each decoding step based on their importance measured by low-precision KV cache copy in VRAM. To avoid inference latency caused by CPU-GPU communication, SpeCache speculatively predicts the KV pairs that the next token might attend to, allowing us to prefetch them before the next decoding step which enables parallelization of prefetching and computation. Experiments on LongBench and Needle-in-a-Haystack benchmarks verify that SpeCache effectively reduces VRAM usage while avoiding information forgetting for long sequences without re-training, even with a 10x high KV cache compression ratio.
Shibo Jie, Yehui Tang 0001, Kai Han 0002, Zhi-Hong Deng 0001
ICML4
2025 Mixture of Lookup Experts
abstract
Mixture-of-Experts (MoE) activates only a subset of experts during inference, allowing the model to maintain low inference FLOPs and latency even as the parameter count scales up. However, since MoE dynamically selects the experts, all the experts need to be loaded into VRAM. Their large parameter size still limits deployment, and offloading, which load experts into VRAM only when needed, significantly increase inference latency. To address this, we propose Mixture of Lookup Experts (MoLE), a new MoE architecture that is efficient in both communication and VRAM usage. In MoLE, the experts are Feed-Forward Networks (FFNs) during training, taking the output of the embedding layer as input. Before inference, these experts can be re-parameterized as lookup tables (LUTs) that retrieves expert outputs based on input ids, and offloaded to storage devices. Therefore, we do not need to perform expert computations during inference. Instead, we directly retrieve the expert’s computation results based on input ids and load them into VRAM, and thus the resulting communication overhead is negligible. Experiments show that, with the same FLOPs and VRAM usage, MoLE achieves inference speeds comparable to dense models and significantly faster than MoE with experts offloading, while maintaining performance on par with MoE. Code: https://github.com/JieShibo/MoLE.
Shibo Jie, Yehui Tang 0001, Kai Han 0002, Duyu Tang, Zhi-Hong Deng 0001, Yunhe Wang 0001
ICML6
2025 Iterative Vectors: In-Context Gradient Steering without Backpropagation
abstract
In-context learning has become a standard approach for utilizing language models. However, selecting and processing suitable demonstration examples can be challenging and time-consuming, especially when dealing with large numbers of them. We propose Iterative Vectors (IVs), a technique that explores activation space to enhance in-context performance by simulating gradient updates during inference. IVs extract and iteratively refine activation-based meta-gradients, applying them during inference without requiring backpropagation at any stage. We evaluate IVs across various tasks using four popular models and observe significant improvements. Our findings suggest that in-context activation steering is a promising direction, opening new avenues for future research.
Zhi-Hong Deng 0001
ICML2
2024 Convolutional Bypasses Are Better Vision Transformer Adapters
abstract
The pretrain-then-finetune paradigm has been widely adopted in computer vision. But as the size of Vision Transformer (ViT) grows exponentially, the full finetuning becomes prohibitive in view of the heavier storage overhead. Motivated by parameter-efficient transfer learning (PETL) on language transformers, recent studies attempt to insert lightweight adaptation modules (e.g., adapter layers or prompt tokens) to pretrained ViT and only finetune these modules while the pretrained weights are frozen. However, these modules were originally proposed to finetune language models and did not take into account the prior knowledge specifically for visual tasks. In this paper, we propose to construct Convolutional Bypasses (Convpass) in ViT as adaptation modules, introducing only a small amount (less than 0.5% of model parameters) of trainable parameters to adapt the large ViT. Different from other PETL methods, Convpass benefits from the hard-coded inductive bias of convolutional layers and thus is more suitable for visual tasks, especially in the low-data regime. Experimental results on VTAB-1K benchmark and few-shot learning datasets show that Convpass outperforms current language-oriented adaptation modules, demonstrating the necessity to tailor vision-oriented adaptation modules for adapting vision models.
Shibo Jie, Zhi-Hong Deng 0001, Shixuan Chen, Zhijuan Jin
ECAI2
2024 Token Compensator: Altering Inference Cost of Vision Transformer Without Re-tuning
Shibo Jie, Yehui Tang 0001, Jianyuan Guo, Zhi-Hong Deng 0001, Kai Han 0002, Yunhe Wang 0001
ECCV (16)4
2024 Retrieved Sequence Augmentation for Protein Representation Learning
abstract
Chang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin, Qintong Li, Lijun Wu, Zhihong Deng, Yang Young Lu, Qi Liu, Sheng Wang, Lingpeng Kong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haiteng Zhao, Jiayi Xin, Qintong Li, Zhi-Hong Deng 0001, Qi Liu 0049, Lingpeng Kong
EMNLP7
2024 Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
abstract
Current solutions for efficiently constructing large vision-language (VL) models follow a two-step paradigm: projecting the output of pre-trained vision encoders to the input space of pre-trained language models as visual prompts; and then transferring the models to downstream VL tasks via end-to-end parameter-efficient fine-tuning (PEFT). However, this paradigm still exhibits inefficiency since it significantly increases the input length of the language models. In this paper, in contrast to integrating visual prompts into inputs, we regard visual prompts as additional knowledge that facilitates language models in addressing tasks associated with visual information. Motivated by the finding that Feed-Forward Network (FFN) of language models acts as "key-value memory", we introduce a novel approach termed memory-space visual prompting (MemVP), wherein visual prompts are concatenated with the weights of FFN for visual knowledge injection. Experimental results across various VL tasks and language models reveal that MemVP significantly reduces the training time and inference latency of the finetuned VL models and surpasses the performance of previous PEFT methods.
Shibo Jie, Yehui Tang 0001, Zhi-Hong Deng 0001, Kai Han 0002, Yunhe Wang 0001
ICML4
2023 FacT: Factor-Tuning for Lightweight Adaptation on Vision Transformer
abstract
Recent work has explored the potential to adapt a pre-trained vision transformer (ViT) by updating only a few parameters so as to improve storage efficiency, called parameter-efficient transfer learning (PETL). Current PETL methods have shown that by tuning only 0.5% of the parameters, ViT can be adapted to downstream tasks with even better performance than full fine-tuning. In this paper, we aim to further promote the efficiency of PETL to meet the extreme storage constraint in real-world applications. To this end, we propose a tensorization-decomposition framework to store the weight increments, in which the weights of each ViT are tensorized into a single 3D tensor, and their increments are then decomposed into lightweight factors. In the fine-tuning process, only the factors need to be updated and stored, termed Factor-Tuning (FacT). On VTAB-1K benchmark, our method performs on par with NOAH, the state-of-the-art PETL method, while being 5x more parameter-efficient. We also present a tiny version that only uses 8K (0.01% of ViT's parameters) trainable parameters but outperforms full fine-tuning and many other PETL methods such as VPT and BitFit. In few-shot settings, FacT also beats all PETL baselines using the fewest parameters, demonstrating its strong capability in the low-data regime.
Shibo Jie, Zhi-Hong Deng 0001
AAAI2
2023 Dual-Alignment Pre-training for Cross-lingual Sentence Embedding
abstract
Ziheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng, Qiang Lou, Haizhen Huang, Jian Jiao, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ziheng Li 0003, Shaohan Huang, Zhi-Hong Deng 0001, Qiang Lou, Haizhen Huang, Jian Jiao 0007, Furu Wei, Qi Zhang 0066
ACL (1)4
2023 Masked Image Modeling with Local Multi-Scale Reconstruction
abstract
Masked Image Modeling (MIM) achieves outstanding success in self-supervised representation learning. Unfortunately, MIM models typically have huge computational burden and slow learning process, which is an inevitable obstacle for their industrial applications. Although the lower layers play the key role in MIM, existing MIM models conduct reconstruction task only at the top layer of encoder. The lower layers are not explicitly guided and the interaction among their patches is only used for calculating new activations. Considering the reconstruction task requires non-trivial inter-patch interactions to reason target signals, we apply it to multiple local layers including lower and upper layers. Further, since the multiple layers expect to learn the information of different scales, we design local multi-scale reconstruction, where the lower and upper layers reconstruct fine-scale and coarse-scale supervision signals respectively. This design not only accelerates the representation learning process by explicitly guiding multiple layers, but also facilitates multi-scale semantical understanding to the input. Extensive experiments show that with significantly less pre-training burden, our model achieves comparable or better performance on classification, detection and segmentation tasks than existing MIM models. Code is available with both MindSpore and PyTorch.
Haoqing Wang, Yehui Tang 0001, Yunhe Wang 0001, Jianyuan Guo, Zhi-Hong Deng 0001, Kai Han 0002
CVPR5
2023 Revisiting the Parameter Efficiency of Adapters from the Perspective of Precision Redundancy
abstract
Current state-of-the-art results in computer vision depend in part on fine-tuning large pre-trained vision models. However, with the exponential growth of model sizes, the conventional full fine-tuning, which needs to store a individual network copy for each tasks, leads to increasingly huge storage and transmission overhead. Adapterbased Parameter-Efficient Tuning (PET) methods address this challenge by tuning lightweight adapters inserted into the frozen pre-trained models. In this paper, we investigate how to make adapters even more efficient, reaching a new minimum size required to store a task-specific fine-tuned network. Inspired by the observation that the parameters of adapters converge at flat local minima, we find that adapters are resistant to noise in parameter space, which means they are also resistant to low numerical precision. To train low-precision adapters, we propose a computational-efficient quantization method which minimizes the quantization error. Through extensive experiments, we find that low-precision adapters exhibit minimal performance degradation, and even 1-bit precision is sufficient for adapters. The experimental results demonstrate that 1-bit adapters outperform all other PET methods on both the VTAB-1K benchmark and few-shot FGVC tasks, while requiring the smallest storage size. Our findings show, for the first time, the significant potential of quantization techniques in PET, providing a general solution to enhance the parameter efficiency of adapter-based PET methods. Code: https://github.com/JieShibo/PETL-ViT
Shibo Jie, Haoqing Wang, Zhi-Hong Deng 0001
ICCV3
2023 Are More Layers Beneficial to Graph Transformers?
Haiteng Zhao, Shuming Ma, Dongdong Zhang 0001, Zhi-Hong Deng 0001, Furu Wei
ICLR4
2023 Focus Your Attention when Few-Shot Classification
abstract
Since many pre-trained vision transformers emerge and provide strong representation for various downstream tasks, we aim to adapt them to few-shot image classification tasks in this work. The input images typically contain multiple entities. The model may not focus on the class-related entities for the current few-shot task, even with fine-tuning on support samples, and the noise information from the class-independent ones harms performance. To this end, we first propose a method that uses the attention and gradient information to automatically locate the positions of key entities, denoted as position prompts, in the support images. Then we employ the cross-entropy loss between their many-hot presentation and the attention logits to optimize the model to focus its attention on the key entities during fine-tuning. This ability then can generalize to the query samples. Our method is applicable to different vision transformers (e.g., columnar or pyramidal ones), and also to different pre-training ways (e.g., single-modal or vision-language pre-training). Extensive experiments show that our method can improve the performance of full or parameter-efficient fine-tuning methods on few-shot tasks. Code is available at https://github.com/Haoqing-Wang/FORT.
Haoqing Wang, Shibo Jie, Zhi-Hong Deng 0001
NeurIPS3
2023 GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning
abstract
Molecule property prediction has gained significant attention in recent years. The main bottleneck is the label insufficiency caused by expensive lab experiments. In order to alleviate this issue and to better leverage textual knowledge for tasks, this study investigates the feasibility of employing natural language instructions to accomplish molecule-related tasks in a zero-shot setting. We discover that existing molecule-text models perform poorly in this setting due to inadequate treatment of instructions and limited capacity for graphs. To overcome these issues, we propose GIMLET, which unifies language models for both graph and text data. By adopting generalized position embedding, our model is extended to encode both graph structures and instruction text without additional graph encoding modules. GIMLET also decouples encoding of the graph from tasks instructions in the attention mechanism, enhancing the generalization of graph features across novel tasks. We construct a dataset consisting of more than two thousand molecule tasks with corresponding instructions derived from task descriptions. We pretrain GIMLET on the molecule tasks along with instructions, enabling the model to transfer effectively to a broad range of tasks. Experimental results demonstrate that GIMLET significantly outperforms molecule-text baselines in instruction-based zero-shot learning, even achieving closed results to supervised GNN models on tasks such as toxcast and muv.
Haiteng Zhao, Shengchao Liu, Hannan Xu, Jie Fu 0001, Zhi-Hong Deng 0001, Lingpeng Kong, Qi Liu 0049
NeurIPS6
2023 Towards well-generalizing meta-learning via adversarial task augmentation
Haoqing Wang, Huiyu Mai, Yuhang Gong, Zhi-Hong Deng 0001
Artif. Intell.4
2022 Switch-GPT: An Effective Method for Constrained Text Generation under Few-Shot Settings (Student Abstract)
abstract
In real-world applications of natural language generation, target sentences are often required to satisfy some lexical constraints. However, the success of most neural-based models relies heavily on data, which is infeasible for data-scarce new domains. In this work, we present FewShotAmazon, the first benchmark for the task of Constrained Text Generation under few-shot settings on multiple domains. Further, we propose the Switch-GPT model, in which we utilize the strong language modeling capacity of GPT-2 to generate fluent and well-formulated sentences, while using a light attention module to decide which constraint to attend to at each step. Experiments show that the proposed Switch-GPT model is effective and remarkably outperforms the baselines. Codes will be available at https://github.com/chang-github-00/Switch-GPT.
Gehui Shen, Zhi-Hong Deng 0001
AAAI4
2022 Rethinking Minimal Sufficient Representation in Contrastive Learning
abstract
Contrastive learning between different views of the data achieves outstanding success in the field of self-supervised representation learning and the learned representations are useful in broad downstream tasks. Since all supervision information for one view comes from the other view, contrastive learning approximately obtains the minimal sufficient representation which contains the shared information and eliminates the non-shared information between views. Considering the diversity of the downstream tasks, it cannot be guaranteed that all task-relevant information is shared between views. Therefore, we assume the non-shared task-relevant information cannot be ignored and theoretically prove that the minimal sufficient representation in contrastive learning is not sufficient for the downstream tasks, which causes performance degradation. This reveals a new problem that the contrastive learning models have the risk of overfitting to the shared information between views. To alleviate this problem, we propose to increase the mutual information between the representation and input as regularization to approximately introduce more task-relevant information, since we cannot utilize any downstream task information during training. Extensive experiments verify the rationality of our analysis and the effectiveness of our method. It significantly improves the performance of several classic contrastive learning models in downstream tasks. Our code is available at https://github.com/Haoqing-Wang/InfoCL.
Haoqing Wang, Xun Guo 0002, Zhi-Hong Deng 0001, Yan Lu 0001
CVPR3
2022 Contrastive Prototypical Network with Wasserstein Confidence Penalty
Haoqing Wang, Zhi-Hong Deng 0001
ECCV (19)2
2022 Certified Robustness Against Natural Language Attacks by Causal Intervention
abstract
Deep learning models have achieved great success in many fields, yet they are vulnerable to adversarial examples. This paper follows a causal perspective to look into the adversarial vulnerability and proposes Causal Intervention by Semantic Smoothing (CISS), a novel framework towards robustness against natural language attacks. Instead of merely fitting observational data, CISS learns causal effects p(y|do(x)) by smoothing in the latent semantic space to make robust predictions, which scales to deep architectures and avoids tedious construction of noise customized for specific attacks. CISS is provably robust against word substitution attacks, as well as empirically robust even when perturbations are strengthened by unknown attack algorithms. For example, on YELP, CISS surpasses the runner-up by 6.8% in terms of certified robustness against word substitutions, and achieves 80.7% empirical robustness when syntactic attacks are integrated.
Haiteng Zhao, Xinshuai Dong, Anh Tuan Luu, Zhi-Hong Deng 0001, Hanwang Zhang
ICML5
2022 Domain Adaptation via Maximizing Surrogate Mutual Information
abstract
Unsupervised domain adaptation (UDA), which is an important topic in transfer learning, aims to predict unlabeled data from target domain with access to labeled data from the source domain. In this work, we propose a novel framework called SIDA (Surrogate Mutual Information Maximization Domain Adaptation) with strong theoretical guarantees. To be specific, SIDA implements adaptation by maximizing mutual information (MI) between features. In the framework, a surrogate joint distribution models the underlying joint distribution of the unlabeled target domain. Our theoretical analysis validates SIDA by bounding the expected risk on target domain with MI and surrogate distribution bias. Experiments show that our approach is comparable with state-of-the-art unsupervised adaptation methods on standard UDA tasks.
Haiteng Zhao, Qinyu Chen, Zhi-Hong Deng 0001
IJCAI4
2021 Cross-Domain Few-Shot Classification via Adversarial Task Augmentation
abstract
Few-shot classification aims to recognize unseen classes with few labeled samples from each class. Many meta-learning models for few-shot classification elaborately design various task-shared inductive bias (meta-knowledge) to solve such tasks, and achieve impressive performance. However, when there exists the domain shift between the training tasks and the test tasks, the obtained inductive bias fails to generalize across domains, which degrades the performance of the meta-learning models. In this work, we aim to improve the robustness of the inductive bias through task augmentation. Concretely, we consider the worst-case problem around the source task distribution, and propose the adversarial task augmentation method which can generate the inductive bias-adaptive 'challenging' tasks. Our method can be used as a simple plug-and-play module for various meta-learning models, and improve their cross-domain generalization capability. We conduct extensive experiments under the cross-domain setting, using nine few-shot classification datasets: mini-ImageNet, CUB, Cars, Places, Plantae, CropDiseases, EuroSAT, ISIC and ChestX. Experimental results show that our method can effectively improve the few-shot classification performance of the meta-learning models under domain shift, and outperforms the existing works. Our code is available at https://github.com/Haoqing-Wang/CDFSL-ATA.
Haoqing Wang, Zhi-Hong Deng 0001
IJCAI2
2021 Generative Feature Replay with Orthogonal Weight Modification for Continual Learning
abstract
The ability of intelligent agents to learn and remember multiple tasks sequentially is crucial to achieving artificial general intelligence. Many continual learning (CL) methods have been proposed to overcome catastrophic forgetting which results from non i.i.d data in the sequential learning of neural networks. In this paper we focus on class incremental learning, a challenging CL scenario. For this scenario, generative replay is a promising strategy which generates and replays pseudo data for previous tasks to alleviate catastrophic forgetting. However, it is hard to train a generative model continually for relatively complex data. Based on recently proposed orthogonal weight modification (OWM) algorithm which can approximately keep previously learned feature invariant when learning new tasks, we propose to 1) replay penultimate layer feature with a generative model; 2) leverage a self-supervised auxiliary task to further enhance the stability of feature. Empirical results on several datasets show our method always achieves substantial improvement over powerful OWM while conventional generative replay always results in a negative effect. Meanwhile our method beats several strong baselines including one based on real data storage. In addition, we conduct experiments to study why our method is effective.
Gehui Shen, Zhi-Hong Deng 0001
IJCNN4
2021 Distributed representations of diseases based on co-occurrence relationship
Haoqing Wang, Huiyu Mai, Zhi-Hong Deng 0001, Luxia Zhang, Huai-Yu Wang
Expert Syst. Appl.3
2021 Towards unsupervised text multi-style transfer with parameter-sharing scheme
Gehui Shen, Zhi-Hong Deng 0001, Unil Yun
Neurocomputing4
2021 Sequence generative adversarial nets with a conditional discriminator
Yongfei Yan, Gehui Shen, Zhi-Hong Deng 0001, Unil Yun
Neurocomputing5
2021 Approximate high utility itemset mining in noisy environments
Yoonji Baek, Unil Yun, Heonho Kim, Jongseong Kim, Bay Vo, Tin Truong 0001, Zhi-Hong Deng 0001
Knowl. Based Syst.7
2020 Dynamically Pruned Message Passing Networks for Large-scale Knowledge Graph Reasoning
Yunsheng Jiang, Xiaohui Xie, Zhiqing Sun, Zhi-Hong Deng 0001
ICLR6
2020 Variational Learning of Bayesian Neural Networks via Bayesian Dark Knowledge
abstract
Bayesian neural networks (BNNs) have received more and more attention because they are capable of modeling epistemic uncertainty which is hard for conventional neural networks. Markov chain Monte Carlo (MCMC) methods and variational inference (VI) are two mainstream methods for Bayesian deep learning. The former is effective but its storage cost is prohibitive since it has to save many samples of neural network parameters. The latter method is more time and space efficient, however the approximate variational posterior limits its performance. In this paper, we aim to combine the advantages of above two methods by distilling MCMC samples into an approximate variational posterior. On the basis of an existing distillation technique we first propose variational Bayesian dark knowledge method. Moreover, we propose Bayesian dark prior knowledge, a novel distillation method which considers MCMC posterior as the prior of a variational BNN. Two proposed methods both not only can reduce the space overhead of the teacher model so that are scalable, but also maintain a distilled posterior distribution capable of modeling epistemic uncertainty. Experimental results manifest our methods outperform existing distillation method in terms of predictive accuracy and uncertainty modeling.
Gehui Shen, Zhi-Hong Deng 0001
IJCAI3
2020 Learning to compose over tree structures via POS tags for sentence representation
Gehui Shen, Zhi-Hong Deng 0001
Expert Syst. Appl.2
2020 Mining high occupancy itemsets
Zhi-Hong Deng 0001
Future Gener. Comput. Syst.1
2020 A Window-Based Self-Attention approach for sentence encoding
Zhi-Hong Deng 0001, Gehui Shen
Neurocomputing2
2019 RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Space
Zhiqing Sun, Zhi-Hong Deng 0001, Jian-Yun Nie, Jian Tang 0005
ICLR (Poster)2
2019 Leap-LSTM: Enhancing Long Short-Term Memory for Text Categorization
abstract
Recurrent Neural Networks (RNNs) are widely used in the field of natural language processing (NLP), ranging from text categorization to question answering and machine translation. However, RNNs generally read the whole text from beginning to end or vice versa sometimes, which makes it inefficient to process long texts. When reading a long document for a categorization task, such as topic categorization, large quantities of words are irrelevant and can be skipped. To this end, we propose Leap-LSTM, an LSTM-enhanced model which dynamically leaps between words while reading texts. At each step, we utilize several feature encoders to extract messages from preceding texts, following texts and the current word, and then determine whether to skip the current word. We evaluate Leap-LSTM on several text categorization tasks: sentiment analysis, news categorization, ontology classification and topic classification, with five benchmark data sets. The experimental results show that our model reads faster and predicts better than standard LSTM. Compared to previous models which can also skip words, our model achieves better trade-offs between performance and efficiency.
Gehui Shen, Zhi-Hong Deng 0001
IJCAI3
2019 Fast Structured Decoding for Sequence Models
abstract
Autoregressive sequence models achieve state-of-the-art performance in domains like machine translation. However, due to the autoregressive factorization nature, these models suffer from heavy latency during inference. Recently, non-autoregressive sequence models were proposed to speed up the inference time. However, these models assume that the decoding process of each token is conditionally independent of others. Such a generation process sometimes makes the output sentence inconsistent, and thus the learned non-autoregressive models could only achieve inferior accuracy compared to their autoregressive counterparts. To improve then decoding consistency and reduce the inference cost at the same time, we propose to incorporate a structured inference module into the non-autoregressive models. Specifically, we design an efficient approximation for Conditional Random Fields (CRF) for non-autoregressive sequence models, and further propose a dynamic transition technique to model positional contexts in the CRF. Experiments in machine translation show that while increasing little latency (8~14ms, our model could achieve significantly better translation performance than previous non-autoregressive models on different translation datasets. In particular, for the WMT14 En-De dataset, our model obtains a BLEU score of 26.80, which largely outperforms the previous non-autoregressive baselines and is only 0.61 lower in BLEU than purely autoregressive models.
Zhiqing Sun, Zhuohan Li 0001, Haoqing Wang, Di He 0001, Zi Lin, Zhi-Hong Deng 0001
NeurIPS6
2019 DivGraphPointer: A Graph Pointer Network for Extracting Diverse Keyphrases
abstract
Keyphrase extraction from documents is useful to a variety of applications such as information retrieval and document summarization. This paper presents an end-to-end method called DivGraphPointer for extracting a set of diversified keyphrases from a document. DivGraphPointer combines the advantages of traditional graph-based ranking methods and recent neural network-based approaches. Specifically, given a document, a word graph is constructed from the document based on word proximity and is encoded with graph convolutional networks, which effectively capture document-level word salience by modeling long-range dependency between words in the document and aggregating multiple appearances of identical words into one node. Furthermore, we propose a diversified point network to generate a set of diverse keyphrases out of the word graph in the decoding process. Experimental results on five benchmark data sets show that our proposed method significantly outperforms the existing state-of-the-art approaches.
Zhiqing Sun, Jian Tang 0005, Pan Du 0001, Zhi-Hong Deng 0001, Jian-Yun Nie
SIGIR4
2018 MEMD: A Diversity-Promoting Learning Framework for Short-Text Conversation
abstract
Neural encoder-decoder models have been widely applied to conversational response generation, which is a research hot spot in recent years. However, conventional neural encoder-decoder models tend to generate commonplace responses like “I don’t know” regardless of what the input is. In this paper, we analyze this problem from a new perspective: latent vectors. Based on it, we propose an easy-to-extend learning framework named MEMD (Multi-Encoder to Multi-Decoder), in which an auxiliary encoder and an auxiliary decoder are introduced to provide necessary training guidance without resorting to extra data or complicating network’s inner structure. Experimental results demonstrate that our method effectively improve the quality of generated responses according to automatic metrics and human evaluations, yielding more diverse and smooth replies.
Meng Zou, Xihan Li 0001, Haokun Liu, Zhi-Hong Deng 0001
COLING4
2018 Unsupervised Neural Word Segmentation for Chinese via Segmental Language Modeling
abstract
Previous traditional approaches to unsupervised Chinese word segmentation (CWS) can be roughly classified into discriminative and generative models.The former uses the carefully designed goodness measures for candidate segmentation, while the latter focuses on finding the optimal segmentation of the highest generative probability.However, while there exists a trivial way to extend the discriminative models into neural version by using neural language models, those of generative ones are non-trivial.In this paper, we propose the segmental language models (SLMs) for CWS.Our approach explicitly focuses on the segmental nature of Chinese, as well as preserves several properties of language models.In SLMs, a context encoder encodes the previous context and a segment decoder generates each segment incrementally.As far as we know, we are the first to propose a neural model for unsupervised CWS and achieve competitive performance to the state-of-theart statistical models on four different datasets from SIGHAN 2005 bakeoff.
Zhiqing Sun, Zhi-Hong Deng 0001
EMNLP2
2018 An efficient structure for fast mining high utility itemsets
Zhi-Hong Deng 0001
Appl. Intell.1
2017 Inter-Weighted Alignment Network for Sentence Pair Modeling
abstract
Sentence pair modeling is a crucial problem in the field of natural language processing.In this paper, we propose a model to measure the similarity of a sentence pair focusing on the interaction information.We utilize the word level similarity matrix to discover fine-grained alignment of two sentences.It should be emphasized that each word in a sentence has a different importance from the perspective of semantic composition, so we exploit two novel and efficient strategies to explicitly calculate a weight for each word.Although the proposed model only use a sequential LSTM for sentence modeling without any external resource such as syntactic parser tree and additional lexicon features, experimental results show that our model achieves state-of-the-art performance on three datasets of two tasks.
Gehui Shen, Yunlun Yang, Zhi-Hong Deng 0001
EMNLP3
2017 A Variational Autoencoding Approach for Inducing Cross-lingual Word Embeddings
abstract
Cross-language learning allows one to use training data from one language to build models for another language. Many traditional approaches require word-level alignment sentences from parallel corpora, in this paper we define a general bilingual training objective function requiring sentence level parallel corpus only. We propose a variational autoencoding approach for training bilingual word embeddings. The variational model introduces a continuous latent variable to explicitly model the underlying semantics of the parallel sentence pairs and to guide the generation of the sentence pairs. Our model restricts the bilingual word embeddings to represent words in exactly the same continuous vector space. Empirical results on the task of cross lingual document classification has shown that our method is effective.
Liang-Chen Wei, Zhi-Hong Deng 0001
IJCAI2
2017 Mining top-k co-occurrence items with sequential pattern
Tung Kieu, Bay Vo, Tuong Le, Zhi-Hong Deng 0001, Bac Le
Expert Syst. Appl.4
2017 A novel approach for mining maximal frequent patterns
Bay Vo, Sang Pham, Tuong Le, Zhi-Hong Deng 0001
Expert Syst. Appl.4
2017 An adaptive Kalman filter estimating process noise covariance
Zhi-Hong Deng 0001, Hongbin Ma, Yuanqing Xia
Neurocomputing2
2016 Identifying Sentiment Words Using an Optimization Model with L1 Regularization
abstract
Sentiment word identification is a fundamental work in numerous applications of sentiment analysis and opinion mining, such as review mining, opinion holder finding, and twitter classification. In this paper, we propose an optimization model with L1 regularization, called ISOMER, for identifying the sentiment words from the corpus. Our model can employ both seed words and documents with sentiment labels, different from most existing researches adopting seed words only. The L1 penalty in the objective function yields a sparse solution since most candidate words have no sentiment. The experiments on the real datasets show that ISOMER outperforms the classic approaches, and that the lexicon learned by ISOMER can be effectively adapted to document-level sentiment analysis.
Zhi-Hong Deng 0001, Yunlun Yang
AAAI1
2016 An Unsupervised Multi-Document Summarization Framework Based on Neural Document Model
abstract
In the age of information exploding, multi-document summarization is attracting particular attention for the ability to help people get the main ideas in a short time. Traditional extractive methods simply treat the document set as a group of sentences while ignoring the global semantics of the documents. Meanwhile, neural document model is effective on representing the semantic content of documents in low-dimensional vectors. In this paper, we propose a document-level reconstruction framework named DocRebuild, which reconstructs the documents with summary sentences through a neural document model and selects summary sentences to minimize the reconstruction error. We also apply two strategies, sentence filtering and beamsearch, to improve the performance of our method. Experimental results on the benchmark datasets DUC 2006 and DUC 2007 show that DocRebuild is effective and outperforms the state-of-the-art unsupervised algorithms.
Shulei Ma, Zhi-Hong Deng 0001, Yunlun Yang
COLING2
2016 A Position Encoding Convolutional Neural Network Based on Dependency Tree for Relation Classification
abstract
With the renaissance of neural network in recent years, relation classification has again become a research hotspot in natural language processing, and leveraging parse trees is a common and effective method of tackling this problem.In this work, we offer a new perspective on utilizing syntactic information of dependency parse tree and present a position encoding convolutional neural network (PECNN) based on dependency parse tree for relation classification.First, treebased position features are proposed to encode the relative positions of words in dependency trees and help enhance the word representations.Then, based on a redefinition of "context", we design two kinds of tree-based convolution kernels for capturing the semantic and structural information provided by dependency trees.Finally, the features extracted by convolution module are fed to a classifier for labelling the semantic relations.Experiments on the benchmark dataset show that PECNN outperforms state-of-the-art approaches.We also compare the effect of different position features and visualize the influence of treebased position feature by tracing back the convolution process.
Yunlun Yang, Yunhai Tong, Shulei Ma, Zhi-Hong Deng 0001
EMNLP4
2016 The SPMF Open-Source Data Mining Library Version 2
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Antonio Gomariz, Ted Gueniche, Azadeh Soltani, Zhi-Hong Deng 0001, Hoang Thanh Lam
ECML/PKDD (3)6
2016 Update with out-of-sequence measurements
Haomiao Zhou, Zhi-Hong Deng 0001, Yuanqing Xia, Mengyin Fu
Neurocomputing2
2016 A new sampling method in particle filter based on Pearson correlation coefficient
Haomiao Zhou, Zhi-Hong Deng 0001, Yuanqing Xia, Mengyin Fu
Neurocomputing2
2016 A novel probabilistic clustering model for heterogeneous networks
Zhi-Hong Deng 0001
Mach. Learn.1
2015 Effectively Predicting Whether and When a Topic Will Become Prevalent in a Social Network
abstract
Effective forecasting of future prevalent topics plays animportant role in social network business development.It involves two challenging aspects: predicting whethera topic will become prevalent, and when. This cannotbe directly handled by the existing algorithms in topicmodeling, item recommendation and action forecasting.The classic forecasting framework based on time seriesmodels may be able to predict a hot topic when a seriesof periodical changes to user-addressed frequency in asystematic way. However, the frequency of topics discussedby users often changes irregularly in social networks.In this paper, a generic probabilistic frameworkis proposed for hot topic prediction, and machine learningmethods are explored to predict hot topic patterns.Two effective models, PreWHether and PreWHen, areintroduced to predict whether and when a topic will becomeprevalent. In the PreWHether model, we simulatethe constructed features of previously observed frequencychanges for better prediction. In the PreWHen model,distributions of time intervals associated with the emergenceto prevalence of a topic are modeled. Extensiveexperiments on real datasets demonstrate that ourmethod outperforms the baselines and generates moreeffective predictions.
Weiwei Liu 0003, Zhi-Hong Deng 0001, Xiuwen Gong, Frank Jiang 0001, Ivor W. Tsang
AAAI2
2015 Multi-Document Summarization Based on Two-Level Sparse Representation Model
abstract
Multi-document summarization is of great value to many real world applications since it can help people get the main ideas within a short time.In this paper, we tackle the problem of extracting summary sentences from multi-document sets by applying sparse coding techniques and present a novel framework to this challenging problem. Based on the data reconstruction and sentence denoising assumption, we present a two-level sparse representation model to depict the process of multi-document summarization. Three requisite properties is proposed to form an ideal reconstructable summary: Coverage, Sparsity and Diversity. We then formalize the task of multi-document summarization as an optimization problem according to the above properties, and use simulated annealing algorithm to solve it.Extensive experiments on summarization benchmark data sets DUC2006 and DUC2007 show that our proposed model is effective and outperforms the state-of-the-art algorithms.
Zhi-Hong Deng 0001
AAAI3
2015 JEAM: A Novel Model for Cross-Domain Sentiment Classification Based on Emotion Analysis
abstract
Cross-domain sentiment classification (CSC) aims at learning a sentiment classifier for unlabeled data in the target domain based on the labeled data from a different source domain.Due to the differences of data distribution of two domains in terms of the raw features, the CSC problem is difficult and challenging.Previous researches mainly focused on concepts mining by clustering words across data domains, which ignored the importance of authors' emotion contained in data, or the different representations of the emotion between domains.In this paper, we propose a novel framework to solve the CSC problem, by modelling the emotion across domains.We first develop a probabilistic model named JEAM to model author's emotion state when writing.Then, an EM algorithm is introduced to solve the likelihood maximum problem and to obtain the latent emotion distribution of the author.Finally, a supervised learning method is utilized to assign the sentiment polarity to a given online review.Experiments show that our approach is effective and outperforms state-of-the-art approaches.
Kun-Hu Luo, Zhi-Hong Deng 0001, Liang-Chen Wei
EMNLP2
2015 Image Tagging via Cross-Modal Semantic Mapping
abstract
Images without annotations are ubiquitous on the Internet, and recommending tags for them has become a challenging open task in image understanding. A common bottleneck of related work is the semantic gap between the image and text representations. In this paper, we bridge the gap by introducing a semantic layer, the space of word embeddings that represents the image tags as the word vectors. Our model first learns the optimal mapping from the visual space to the semantic space using training sources. Then we annotate test images by decoding the semantic representations of the visual features. Extensive experiments demonstrate that our model outperforms the state-of-the-art approaches in predicting the image tags.
Zhi-Hong Deng 0001, Yunlun Yang
ACM Multimedia1
2015 PrePost+: An efficient N-lists-based algorithm for mining frequent itemsets via Children-Parent Equivalence pruning
Zhi-Hong Deng 0001, Sheng-Long Lv
Expert Syst. Appl.1
2015 CLE_LMNN: A novel framework of LMNN based on clustering labeled examples
Zhi-Hong Deng 0001, Kun-Hu Luo
Expert Syst. Appl.1
2015 Automatic Identification and Recognition of Sentiment Words Using an Optimization-Based Model with Propagation
abstract
Sentiment word identification, or SWI, is one of the most basic and important techniques in sentiment analysis. Many existing methods depend on the seed word, and such dependence leads to low robustness. In this paper, we propose a novel method utilizing propagation and optimization model, PRopagation-based Constrained Optimization Model (PR-COM) for SWI. Unlike the previous research, we exploit an iterative algorithm to expand the seed word set from the candidate word set, which brings higher robustness. Experimental results on several data sets show that our PR-COM method is effective and outperforms the state-of-art methods.
Kun-Hu Luo, Zhi-Hong Deng 0001, Shiyingxue Li
Int. J. Intell. Syst.2
2015 Optimal linear estimation with square-based sampling
Haomiao Zhou, Zhi-Hong Deng 0001, Yuanqing Xia, Mengyin Fu
Inf. Sci.2
2015 Mining summarization of high utility itemsets
Zhi-Hong Deng 0001
Knowl. Based Syst.2
2015 Mining Top K Spread Sources for a Specific Topic and a Given Node
abstract
In social networks, nodes (or users) interested in specific topics are often influenced by others. The influence is usually associated with a set of nodes rather than a single one. An interesting but challenging task for any given topic and node is to find the set of nodes that represents the source or trigger for the topic and thus identify those nodes that have the greatest influence on the given node as the topic spreads. We find that it is an NP-hard problem. This paper proposes an effective framework to deal with this problem. First, the topic propagation is represented as the Bayesian network. We then construct the propagation model by a variant of the voter model. The probability transition matrix (PTM) algorithm is presented to conduct the probability inference with the complexity O(θ(3)log2θ), while θ is the number nodes in the given graph. To evaluate the PTM algorithm, we conduct extensive experiments on real datasets. The experimental results show that the PTM algorithm is both effective and efficient.
Weiwei Liu 0003, Zhi-Hong Deng 0001, Longbing Cao, Xiuwen Gong
IEEE Trans. Cybern.2
2014 A Joint Optimization Model for Image Summarization Based on Image Content and Tags
abstract
As an effective technology for navigating a large number of images, image summarization is becoming a promising task with the rapid development of image sharing sites and social networks. Most existing summarization approaches use the visual-based features for image representation without considering tag information.In this paper, we propose a novel framework, named JOINT, which employs both image content and tag information to summarize images. Our model generates the summary images which can best reconstruct the original collection. Based on the assumption that an image with representative content should also have typical tags, we introduce a similarity-inducing regularizer to our model. Furthermore, we impose the lasso penalty on the objective function to yield a concise summary set. Extensive experiments demonstrate our model outperforms the state-of-the-art approaches.
Zhi-Hong Deng 0001, Yunlun Yang
AAAI2
2014 XDist: an effective XML keyword search system with re-ranking model based on keyword distribution
Ning Gao 0006, Zhi-Hong Deng 0001, ShengLong Lü
Sci. China Inf. Sci.2
2014 Fast mining Top-Rank-k frequent patterns by using Node-lists
Zhi-Hong Deng 0001
Expert Syst. Appl.1
2014 Fast mining frequent itemsets using Nodesets
Zhi-Hong Deng 0001, Sheng-Long Lv
Expert Syst. Appl.1
2014 A study of supervised term weighting scheme for sentiment analysis
Zhi-Hong Deng 0001, Kun-Hu Luo
Expert Syst. Appl.1
2013 Semantic Inversion in XML Keyword Search with General Conditional Random Fields
Shu-Han Wang, Zhi-Hong Deng 0001
WISE (1)2
2013 A clique-superposition model for social networks
Shaowei Cai 0001, Ming Zhang 0004, Zhi-Hong Deng 0001
Sci. China Inf. Sci.5
2013 LAF: a new XML encoding and indexing strategy for keyword-based XML search
abstract
ABSTRACT As a large number of corpuses are represented, stored and published in XML format, how to find useful information from XML databases has become an increasingly important issue. Keyword search enables web users to easily access XML data without the need to learn a structured query language or to study complex data schemas. Most existing indexing strategies for XML keyword search are based upon Dewey encoding. In this paper, we proposed a new encoding method called Level Order and Father (LAF) for XML documents. With LAF encoding, we devised a new index structure, called two‐layer LAF inverted index, which can greatly decrease the space complexity compared with Dewey encoding‐based inverted index. Furthermore, with two‐layer LAF inverted index, we proposed a new keyword query algorithm called Algorithm based on Binary Search (ABS) that can quickly find all Smallest Lowest Common Ancestor. We experimentally evaluate two‐layer LAF inverted index and ABS algorithm on four real XML data sets selected from Wikipedia. The experimental results prove the advantages of our index method and querying algorithm. The space consumed by two‐layer LAF index is less than half of that consumed by Dewey inverted index. Moreover, ABS is about one to two orders of magnitude faster than the classic Stack algorithm. Concurrency and Computation: Practice and Experience, 2012.© 2012 Wiley Periodicals, Inc.
Zhi-Hong Deng 0001, Yong-Qing Xiang, Ning Gao 0006
Concurr. Comput. Pract. Exp.1
2013 ROBIN: A novel personal recommendation model based on information propagation
Zhi-Hong Deng 0001, Zhonghui Wang
Expert Syst. Appl.1
2013 Mining Top-Rank-k Erasable Itemsets by PID_lists
abstract
Mining erasable itemsets are one of new emerging data mining tasks. In this paper, we present a new data representation called a PID_list, which keeps track of the id_nums (identification number) of products that include an itemset. On the basis of the PID_list, we propose a new algorithm called VM for mining top-rank-k erasable itemsets efficiently. The VM algorithm can avoid the time-consuming process of calculating the gain of the candidate itemsets and lots of scans of the databases. Therefore, it can accelerate the task of mining greatly. For evaluating the VM algorithm, we have conducted experiments on six synthetic product databases. Our performance study shows that the VM algorithm is efficient and much faster than the MIKE algorithm, which is the first algorithm for dealing with the problem of mining top-rank-k erasable itemsets.
Zhi-Hong Deng 0001
Int. J. Intell. Syst.1
2013 A new continuous-discrete particle filter for continuous-discrete nonlinear systems
Yuanqing Xia, Zhi-Hong Deng 0001, Li Li 0050, Xiumei Geng
Inf. Sci.2
2013 Personalized search in digital libraries via spreading activation model
abstract
With the tremendous development of information technology, the volume of data in digital libraries is increasing enormously, and the magnitude is putting users at risk of information overload. Personalized search aims at solving the problem by tailor
Ming Zhang 0004, Zhi-Hong Deng 0001
Web Intell. Agent Syst.4
2012 Guess What I Want: Inferring the Semantics of Keyword Queries Using Evidence Theory
Jia-Jian Jiang, Zhi-Hong Deng 0001, Ning Gao 0006, Sheng-Long Lv
APWeb2
2012 A novel local patch framework for fixing supervised learning models
abstract
In the past decades, machine learning models, especially supervised learning algorithms, have been widely used in various real world applications. However, no matter how strong a learning model is, it will suffer from the prediction errors when it is applied to real world problems. Due to the black box nature of supervised learning models, it is a challenging problem to fix the supervised learning models by further learning from the failure cases it generates. In this paper, we propose a novel Local Patch Framework (LPF) to locally fix supervised learning models by learning from its predicted failure cases. Since the learning models are generally globally optimized during training process, our proposed LPF assumes that most of the learning errors are led by local errors in the model. Thus we aim to break the black boxes of learning models by identifying and fixing the local errors of various models automatically. The proposed LPF has two key steps, which are local error region subspace learning and local patch model learning. Through this way, we aim to fix the errors of learning models locally and automatically with certain generalization ability on unseen testing data. Experiments on both classification and ranking problems show that the proposed LPF is effective and outperforms the original algorithms and the incremental learning model.
Bingzheng Wei, Jun Yan 0001, Zhi-Hong Deng 0001, Zheng Chen 0001
CIKM5
2012 A new algorithm for fast mining frequent itemsets using N-lists
Zhi-Hong Deng 0001, Zhonghui Wang, Jia-Jian Jiang
Sci. China Inf. Sci.1
2012 PAV: A novel model for ranking heterogeneous objects in bibliographic information networks
Zhi-Hong Deng 0001, Bo-Yan Lai, Zhonghui Wang, Guo-Dong Fang
Expert Syst. Appl.1
2012 Fast mining erasable itemsets using NC_sets
Zhi-Hong Deng 0001
Expert Syst. Appl.1
2012 MAXLCA: A New Query Semantic Model for XML Keyword Search
Ning Gao 0006, Zhi-Hong Deng 0001, Jia-Jian Jiang
J. Web Eng.2
2011 Fully Utilize Feedbacks: Language Model Based Relevance Feedback in Information Retrieval
Sheng-Long Lv, Zhi-Hong Deng 0001, Ning Gao 0006, Jia-Jian Jiang
ADMA (1)2
2011 BibClus: A Clustering Algorithm of Bibliographic Networks by Message Passing on Center Linkage Structure
abstract
Multi-type objects with multi-type relations are ubiquitous in real-world networks, e.g. bibliographic networks. Such networks are also called heterogeneous information networks. However, the research on clustering for heterogeneous information networks is little. A new algorithm, called NetClus, has been proposed in recent two years. Although NetClus is applied on a heterogeneous information network with a star network schema, considering the relations between center objects and all attribute objects linking to them, it ignores the relations between center objects such as citation relations, which also contain rich information. Hence, we think the star network schema cannot be used to characterize all possible relations without integrating the linkage structure among center objects, which we call the Center Linkage Structure, and there has been no practical way good enough to solve it. In this paper, we present a novel algorithm, BibClus, for clustering heterogeneous objects with center linkage structure by taking a bibliographic information network as an example. In BibClus, we build a probabilistic model of pair wise hidden Markov random field (P-HMRF) to characterize the center linkage structure, and convert it to a factor graph. We further combine EM algorithm with factor graph theory, and design an efficient way based on message passing algorithm to inference marginal probabilities and estimate parameters at each iteration of EM. We also study how factor functions affect clustering performance with different function forms and constraints. For evaluating our proposed method, we have conducted thorough experiments on a real dataset that we had crawled from ACM Digital Library. The experimental results show that BibClus is effective and has a much higher quantity than the recently proposed algorithm, NetClus, in both recall and precision.
Zhi-Hong Deng 0001
ICDM2
2011 ListOPT: Learning to Optimize for XML Ranking
Ning Gao 0006, Zhi-Hong Deng 0001, Jia-Jian Jiang
PAKDD (2)2
2011 Mop: An Efficient Algorithm for Mining Frequent Pattern with Subtree Traversing
abstract
Mining frequent patterns in database has emerged as an important task in knowledge discovery and data mining. In this paper, we present an efficient algorithm called Mop for fast frequent pattern discovery. Mop utilizes a new kind of data structure called OP_tree (ordered pattern tree) and some particular properties of frequent patterns to facilitate the process of mining frequent patterns. An OP_tree is a special frequent pattern tree, where the children of any node are sorted according to the supports of corresponding items. Efficiency of Mop is achieved with three techniques: (1) it adopts OP_tree to store a large database to avoid repetitive database scans, (2) it finds all frequent 2-patterns in the construction of OP_tree to avoid the costly generation of a large number of candidate 2-patterns, (3) the supports of candidate k-patterns (k>2) can be obtained by traversing a few of specific subtrees of the OP_tree, which greatly reduces the search space and avoid multi-scans of a database. We experimentally compare our algorithm with the Apriori algorithm and the FP-growth algorithm on one real database and one synthetical database. The experimental results show that Mop is about an order of magnitude faster than the Apriori algorithm. Mop also outperforms the FP-growth algorithm, especially when support threshold is very low and databases are quite large.
Zhi-Hong Deng 0001, Ning Gao 0006
Fundam. Informaticae1
2010 An Efficient Algorithm for Mining Erasable Itemsets
Zhi-Hong Deng 0001
ADMA (1)1
2010 Tag Recommendation Based on Bayesian Principle
Zhonghui Wang, Zhi-Hong Deng 0001
ADMA (2)2
2010 Recommended or Not Recommended? Review Classification through Opinion Extraction
abstract
With the rapid growth of web 2.0, online product reviews generated by users are becoming increasingly useful for customers to make purchase decisions. In this paper, we focus on the problem of classifying user reviews as recommended the product or not. The proposed method first mines the product features and relevant opinions, and then determines the overall sentiment orientation of the review based on the polarity and strength of these opinions. The evaluation results show the effectiveness of our proposed method in product feature mining and review classification.
Sheng Feng, Ming Zhang 0004, Yanxing Zhang, Zhi-Hong Deng 0001
APWeb4
2010 Adaptive Top-k Algorithm in SLCA-Based XML Keyword Search
abstract
Computing top-k results matching XML queries is gaining importance due to the increasing of large XML repositories. In this paper, we propose a novel two-layer-based index construction and associated algorithms for efficiently computing top-k results for SLCA-based XML keyword search. We have conducted expensive experiments and the results show great advantage on efficiency compared with existing approaches.
Zhi-Hong Deng 0001, Yong-Qing Xiang, Ning Gao 0006, Ming Zhang 0004, Shiwei Tang
APWeb2
2010 Users' Book-Loan Behaviors Analysis and Knowledge Dependency Mining
Ming Zhang 0004, Jian Tang 0005, Zhi-Hong Deng 0001, Long Xiao
WAIM5
2010 Mining frequent patterns from network flows for monitoring network
Zhi-Hong Deng 0001
Expert Syst. Appl.2
2009 Mining Frequent Patterns from Network Data Flow
Zhi-Hong Deng 0001, Shiwei Tang, Bei Zhang 0002
ADMA2
2007 A novel clustering-based RSS aggregator
abstract
In recent years, different commercial Weblog subscribing systems have been proposed to return stories from users. subscribed feeds. In this paper, we propose a novel clustering-based RSS aggregator called as RSS Clusgator System (RCS) for Weblog reading. Note that an RSS feed may have several different topics. A user may only be interested in a subset of these topics. In addition there could be many different stories from multiple RSS feeds, which discuss similar topic from different perspectives. A user may be interested in this topic but do not know how to collect all feeds related to this topic. In contrast to many previous works, we cluster all stories in RSS feeds into hierarchical structure to better serve the readers. Through this way, users can easily find all their interested stories. To make the system current, we propose a flexible time window for incremental clustering. RCS utilizes both link information and content information for efficient clustering. Experiments show the effectiveness of RCS.
Jun Yan 0001, Zhi-Hong Deng 0001, Lei Ji 0001, Weiguo Fan, Benyu Zhang, Zheng Chen 0001
WWW3
2006 A Fast Algorithm for Maintenance of Association Rules in Incremental Databases
Zhi-Hong Deng 0001, Shiwei Tang
ADMA2
2005 A Non-VSM kNN Algorithm for Text Classification
Zhi-Hong Deng 0001, Shiwei Tang
ADMA1
2005 Mining Frequent Ordered Patterns
Zhi-Hong Deng 0001, Cong-Rui Ji, Ming Zhang 0004, Shiwei Tang
PAKDD1
2005 An Efficient Approach for Interactive Mining of Frequent Itemsets
Zhi-Hong Deng 0001, Shiwei Tang
WAIM1
2004 A Comparative Study on Feature Weight in Text Categorization
Zhi-Hong Deng 0001, Shiwei Tang, Dongqing Yang, Ming Zhang 0004, Liyu Li, Kunqing Xie
APWeb1
2004 WIEAS: Helping to Discover Web Information Sources and Extract Data from Them
Liyu Li, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Zhi-Hong Deng 0001, Zhihua Su
APWeb5