EDBT 2026 Demo / reviewers in the wild / expert
Huijia Zhu
dblp:50/7121
· DBLP profile ↗
39ranked-venue papers
0as first author
31since 2021 · last 2026
0009-0008-5784-7225ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 19 since 2021Databases, data management, data science and information retrieval · 9 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic ForgettingabstractLarge Vision-Language Models (LVLMs) demonstrate impressive general-purpose capabilities but often suffer from catastrophic forgetting when incorporating specialized knowledge. To address this plasticity-stability dilemma, we introduce Structured Dialogue Fine-Tuning (SDFT), a data-centric approach that injects domain-specific concepts while preserving foundational abilities. Distinct from parameter-constrained continual learning methods, SDFT leverages a three-phase dialogue structure: Foundation Preservation reinforces pre-trained visual-linguistic alignment through captioning tasks; Contrastive Disambiguation uses carefully designed counterfactual examples to establish precise semantic boundaries; and Knowledge Specialization embeds specialized information via chain-of-thought reasoning. Evaluations across personalized entity recognition, abstract concept understanding, and biomedical domains show that SDFT significantly outperforms state-of-the-art baselines, including EWC-LoRA and O-LoRA. The results confirm SDFT’s effectiveness in balancing specialized knowledge acquisition and general capability retention. Yijie Hong, Xiaofei Yin, Xinzhong Wang, Huijia Zhu, Sufeng Duan |
ICMR | 5 |
| 2026 | OpenImplicit: Benchmarking Implicit Reasoning in MLLMs via Open-Ended Evaluation
Jidong Li, Xiaofei Yin, Shuheng Zhou 0001, Haodong Zhao, Sufeng Duan, Gongshen Liu, Huijia Zhu |
ICMR | 9 |
| 2026 | Generalizable and Adaptive Continual Learning Framework for AI-Generated Image DetectionabstractThe malicious misuse and widespread dissemination of AI-generated images pose a significant threat to the authenticity of online information. Current detection methods often struggle to generalize to unseen generative models, and the rapid evolution of generative techniques continuously exacerbates this challenge. Without adaptability, detection models risk becoming ineffective in real-world applications. To address this critical issue, we propose a novel three-stage domain continual learning framework designed for continuous adaptation to evolving generative models. In the first stage, we employ a strategic parameter-efficient fine-tuning approach to develop a transferable offline detection model with strong generalization capabilities. Building upon this foundation, the second stage integrates unseen data streams into a continual learning process. To efficiently learn from limited samples of novel generated models and mitigate overfitting, we design a data augmentation chain with progressively increasing complexity. Furthermore, we leverage the Kronecker-Factored Approximate Curvature (K-FAC) method to approximate the Hessian and alleviate catastrophic forgetting. Finally, the third stage utilizes a linear interpolation strategy based on Linear Mode Connectivity, effectively capturing commonalities across diverse generative models and further enhancing overall performance. We establish a comprehensive benchmark of 27 generative models, including GANs, deepfakes, and diffusion models, chronologically structured up to August 2024 to simulate real-world scenarios. Extensive experiments demonstrate that our initial offline detectors surpass the leading baseline by +5.51% in terms of mean average precision. Our continual learning strategy achieves an average accuracy of 92.20%, outperforming state-of-the-art methods. Jun Lan 0001, Yaoyu Kang, Huijia Zhu, Weiqiang Wang 0002, Zhuosheng Zhang 0001, Shi-Lin Wang |
IEEE Trans. Multim. | 4 |
| 2025 | WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images DetectionabstractThe development of text-to-image generative models has enabled the creation of images so realistic that distinguishing between AI-generated images and real photos is becoming a challenge. This progress offers new possibilities but also raises concerns over privacy, authenticity, and security. Detecting AI-generated images is crucial to prevent misuse. To assess the generalizability and robustness of AI-generated image detection, we present a large-scale dataset, referred to as WildFake. This dataset features cutting-edge image generators, a wide variety of generator categories, and generators for various applications, organized in a hierarchical framework. WildFake collects fake images from the open-source community, enriching its diversity with a broad range of image classes and image styles. Its design significantly improves the effectiveness of detection algorithms, making it a valuable resource for enhancing AI-generated image detection in practical applications. Our evaluations offer insights into the performance of generative models at various levels, showcasing WildFake's unique hierarchical structure's benefits. Yan Hong 0001, Jianming Feng, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003 |
AAAI | 5 |
| 2025 | SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation MethodsabstractAs speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems.However, many existing speech deepfake datasets are limited in scale and diversity, making it challenging to train models that can generalize well to unseen deepfakes.To address these gaps, we introduce SpeechFake, a largescale dataset designed specifically for speech deepfake detection.SpeechFake includes over 3 million deepfake samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools.The dataset encompasses a wide range of generation techniques, including text-to-speech, voice conversion, and neural vocoder, incorporating the latest cuttingedge methods.It also provides multilingual support, spanning 46 languages.In this paper, we offer a detailed overview of the dataset's creation, composition, and statistics.We also present baseline results by training detection models on SpeechFake, demonstrating strong performance on both its own test sets and various unseen test sets.Additionally, we conduct experiments to rigorously explore how generation methods, language diversity, and speaker variation affect detection performance.We believe SpeechFake will be a valuable resource for advancing speech deepfake detection and developing more robust models for evolving generation techniques 1 . Wen Huang 0004, Yanmei Gu, Huijia Zhu, Yanmin Qian |
ACL (1) | 4 |
| 2025 | Sparse Latents Steer Retrieval-Augmented GenerationabstractChunlei Xin, Shuheng Zhou, Huijia Zhu, Weiqiang Wang, Xuanang Chen, Xinyan Guan, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chunlei Xin, Shuheng Zhou 0001, Huijia Zhu, Weiqiang Wang 0002, Xuanang Chen, Xinyan Guan, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 3 |
| 2025 | Advancing Controllable Music Generation with Latent Rectified Flow Guided by Rhythm and HarmonyabstractRectified flow models have shown considerable potential in various generation tasks, but their capability for music generation remains largely unexplored. These models use ordinary differential equations (ODEs) with linear interpolation, allowing more straightforward distribution transportation compared to diffusion models. In this paper, we present a text-to-music generation framework based on latent rectified flow. Additionally, to further improve its controllability and generation quality, we inject rhythmic and harmonic control signals into the generation process. Extensive objective and subjective evaluations demonstrate that the rectified flow model can generate music of comparable quality as existing systems based on diffusion and language models. Furthermore, integrating external controls enables the rectified flow model to achieve improved performance. Samples are available on https://anonymous.4open.science/w/MusicLRF-demo-2DA3/ Huijia Zhu, Yanmin Qian |
ASRU | 5 |
| 2025 | Aligning Retrieval with Reader Needs: Reader-Centered Passage Selection for Open-Domain Question AnsweringabstractOpen-Domain Question Answering (ODQA) systems often struggle with the quality of retrieved passages, which may contain conflicting information and be misaligned with the reader’s needs. Existing retrieval methods aim to gather relevant passages but often fail to prioritize consistent and useful information for the reader. In this paper, we introduce a novel Reader-Centered Passage Selection (R-CPS) method, which enhances the performance of the retrieve-then-read pipeline by re-ranking and clustering passages from the reader’s perspective. Our method re-ranks passages based on the reader’s prediction probability distribution and clusters passages according to the predicted answers, prioritizing more useful and relevant passages to the top and reducing inconsistent information. Experiments on ODQA datasets demonstrate the effectiveness of our approach in improving the quality of evidence passages under zero-shot settings. Chunlei Xin, Shuheng Zhou 0001, Xuanang Chen, Yaojie Lu 0001, Huijia Zhu, Weiqiang Wang 0002, Xianpei Han, Le Sun 0001 |
COLING | 5 |
| 2025 | Generalizable Audio Deepfake Detection via Latent Space Refinement and AugmentationabstractAdvances in speech synthesis technologies, like text-to-speech (TTS) and voice conversion (VC), have made detecting deepfake speech increasingly challenging. Spoofing countermeasures often struggle to generalize effectively, particularly when faced with unseen attacks. To address this, we propose a novel strategy that integrates Latent Space Refinement (LSR) and Latent Space Augmentation (LSA) to improve the generalization of deepfake detection systems. LSR introduces multiple learnable prototypes for the spoof class, refining the latent space to better capture the intricate variations within spoofed data. LSA further diversifies spoofed data representations by applying augmentation techniques directly in the latent space, enabling the model to learn a broader range of spoofing patterns. We evaluated our approach on four representative datasets, i.e. ASVspoof 2019 LA, ASVspoof 2021 LA and DF, and In-The-Wild. The results show that LSR and LSA perform well individually, and their integration achieves competitive results, matching or surpassing current state-of-the-art methods. Wen Huang 0004, Yanmei Gu, Huijia Zhu, Yanmin Qian |
ICASSP | 4 |
| 2025 | Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge Editing
Lingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu, Huijia Zhu, Zhuosheng Zhang 0001, Gongshen Liu |
ICCV | 6 |
| 2025 | Stochastic Layer-Wise Shuffle for Improving Vision Mamba TrainingabstractRecent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies and their potential are still not sufficiently explored. In this paper, we investigate strategies for Vim and propose Stochastic Layer-Wise Shuffle (SLWS), a novel regularization method that can effectively improve the Vim training. Without architectural modifications, this approach enables the non-hierarchical Vim to get leading performance on ImageNet-1K compared with the similar type counterparts. Our method operates through four simple steps per layer: probability allocation to assign layer-dependent shuffle rates, operation sampling via Bernoulli trials, sequence shuffling of input tokens, and order restoration of outputs. SLWS distinguishes itself through three principles: \textit{(1) Plug-and-play:} No architectural modifications are needed, and it is deactivated during inference. \textit{(2) Simple but effective:} The four-step process introduces only random permutations and negligible overhead. \textit{(3) Intuitive design:} Shuffling probabilities grow linearly with layer depth, aligning with the hierarchical semantic abstraction in vision models. Our work underscores the importance of tailored training strategies for Vim models and provides a helpful way to explore their scalability. Code and models are available at https://github.com/huangzizheng01/ShuffleMamba Zizheng Huang, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Limin Wang 0002 |
ICML | 5 |
| 2025 | BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
Xun Gong 0005, Anqi Lv, Wangyou Zhang, Huijia Zhu, Yanmin Qian |
INTERSPEECH | 5 |
| 2025 | Generalizable Audio Deepfake Detection via Hierarchical Structure Learning and Feature Whitening in Poincaré sphere
Mingru Yang, Yanmei Gu, Qianhua He, Yanxiong Li, Peirong Zhang 0001, Huijia Zhu, Weiqiang Wang 0002 |
INTERSPEECH | 8 |
| 2025 | Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsabstractProgress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake. Yikun Ji, Yan Hong 0001, Jiahui Zhan, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Liqing Zhang 0001, Jianfu Zhang 0003 |
ACM Multimedia | 6 |
| 2025 | InterAnimate: Taming Region-Aware Diffusion Model for Realistic Human Interaction Animation
Yukang Lin, Yan Hong 0001, Zunnan Xu, Xindi Li, Chuanbiao Song, Ronghui Li, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Xiu Li 0001 |
ACM Multimedia | 10 |
| 2025 | Generalizable Audio Deepfake Detection via Risk-Aware Style Alignment and Structural Empirical Risk MinimizationabstractWith the rapid advancement of AIGC technologies, audio deepfakes have become increasingly realistic, posing serious threats to information security and biometric authentication. Therefore, audio deepfake detection (ADD) has emerged as a critical and fast-evolving research area, particularly requiring superior generalization in out-of-domain scenarios. However, existing ADD methods suffer from constrained generalization and limited access to target data. To address these challenges, we propose Risk-Aware Style Alignment (RASA), a novel generalizable ADD framework that projects the style of any input feature into a shared style space through similarity-based projection. This alignment reduces both inter-domain and intra-source discrepancies without requiring target data during training. In addition, we adopt Structural Empirical Risk Minimization (SERM) in the Poincaré ball model to capture the hierarchical structure of the data and further minimize source risk. By jointly optimizing RASA and SERM, the proposed method effectively tightens the theoretical upper bound of target risk across three key dimensions: source risk, inter-domain divergence, and intra-source discrepancy. Extensive experiments demonstrate that our approach achieves superior generalization and outperforms existing state-of-the-art methods. Mingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang 0001, Haolin He, Huijia Zhu, Weiqiang Wang 0002 |
ACM Multimedia | 7 |
| 2025 | Conditional Prototype Rectification Prompt LearningabstractPre-trained large-scale vision-language models (VLMs) have acquired profound understanding of general visual concepts. Recent advancements in efficient transfer learning (ETL) have shown remarkable success in fine-tuning VLMs within the scenario of limited data, introducing only a few parameters to harness task-specific insights from VLMs. Despite significant progress, current leading ETL methods tend to overfit the narrow distributions of base classes seen during training and encounter two primary challenges: (i) only utilizing uni-modal information to modeling task-specific knowledge; and (ii) using costly and time-consuming methods to supplement knowledge. To address these issues, we propose a Conditional Prototype Rectification Prompt Learning (CPR) method to correct the bias of the base examples and augment limited data in an effective way. Specifically, we alleviate over-fitting on base classes from two aspects. First, each input image acquires knowledge from both textual and visual prototypes and then generates sample-conditional text tokens. Second, we extract utilizable knowledge from unlabeled data to further refine the prototypes. These two strategies mitigate biases that stem from base classes, yielding a more effective classifier. Extensive experiments on 11 benchmark datasets show that our CPR achieves state-of-the-art performance on few-shot classification, base-to-new generalization, and cross-dataset generalization tasks. Our code is available at https://github.com/chenhaoxing/CPR. Haoxing Chen, Zizheng Huang, Yan Hong 0001, Zhuoer Xu, Zhangxuan Gu, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Beyond Full Fine-tuning: Harnessing the Power of LoRA for Multi-Task Instruction TuningabstractLow-Rank Adaptation (LoRA) is a widespread parameter-efficient fine-tuning algorithm for large-scale language models. It has been commonly accepted that LoRA mostly achieves promising results in single-task, low-resource settings, and struggles to handle multi-task instruction tuning scenarios. In this paper, we conduct a systematic study of LoRA on diverse tasks and rich resources with different learning capacities, examining its performance on seen tasks during training and its cross-task generalization on unseen tasks. Our findings challenge the prevalent assumption that the limited learning capacity will inevitably result in performance decline. In fact, our study reveals that when configured with an appropriate rank, LoRA can achieve remarkable performance in high-resource and multi-task scenarios, even comparable to that achieved through full fine-tuning. It turns out that the constrained learning capacity encourages LoRA to prioritize conforming to instruction requirements rather than memorizing specialized features of particular tasks or instances. This study reveals the underlying connection between learning capacity and generalization capabilities for robust parameter-efficient fine-tuning, highlighting a promising direction for the broader application of LoRA across various tasks and settings. Chunlei Xin, Yaojie Lu 0001, Shuheng Zhou 0001, Huijia Zhu, Weiqiang Wang 0002, Xianpei Han, Le Sun 0001 |
LREC/COLING | 5 |
| 2024 | Probe Then Retrieve and Reason: Distilling Probing and Reasoning Capabilities into Smaller Language ModelsabstractStep-by-step reasoning methods, such as the Chain-of-Thought (CoT), have been demonstrated to be highly effective in harnessing the reasoning capabilities of Large Language Models (LLMs). Recent research efforts have sought to distill LLMs into Small Language Models (SLMs), with a significant focus on transferring the reasoning capabilities of LLMs to SLMs via CoT. However, the outcomes of CoT distillation are inadequate for knowledge-intensive reasoning tasks. This is because generating accurate rationales requires crucial factual knowledge, which SLMs struggle to retain due to their parameter constraints. We propose a retrieval-based CoT distillation framework, named Probe then Retrieve and Reason (PRR), which distills the question probing and reasoning capabilities from LLMs into SLMs. We train two complementary distilled SLMs, a probing model and a reasoning model, in tandem. When presented with a new question, the probing model first identifies the necessary knowledge to answer it, generating queries for retrieval. Subsequently, the reasoning model uses the retrieved knowledge to construct a step-by-step rationale for the answer. In knowledge-intensive reasoning tasks, such as StrategyQA and OpenbookQA, our distillation framework yields superior performance for SLMs compared to conventional methods, including simple CoT distillation and knowledge-augmented distillation using raw questions. Yichun Zhao, Shuheng Zhou 0001, Huijia Zhu |
LREC/COLING | 3 |
| 2024 | ComFusion: Enhancing Personalized Generation by Instance-Scene Compositing and Fusion
Yan Hong 0001, Yuxuan Duan, Bo Zhang 0075, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003 |
ECCV (44) | 6 |
| 2024 | COIN-Matting: Confounder Intervention for Image Matting
Zhaohe Liao, Jiangtong Li, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Li Niu 0002, Liqing Zhang 0001 |
ECCV (19) | 4 |
| 2024 | Modeling Layout Reading Order as Ordering Relations for Visually-rich Document UnderstandingabstractChong Zhang, Yi Tu, Yixi Zhao, Chenshu Yuan, Huan Chen, Yue Zhang, Mingxu Chai, Ya Guo, Huijia Zhu, Qi Zhang, Tao Gui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yixi Zhao, Chenshu Yuan, Huan Chen 0012, Yue Zhang 0073, Mingxu Chai, Huijia Zhu, Qi Zhang 0001, Tao Gui |
EMNLP | 9 |
| 2024 | UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich DocumentsabstractThe recognition of named entities in visually-rich documents (VrD-NER) plays a critical role in various real-world scenarios and applications. However, the research in VrD-NER faces three major challenges: complex document layouts, incorrect reading orders, and unsuitable task formulations. To address these challenges, we propose a query-aware entity extraction head, namely UNER, to collaborate with existing multi-modal document transformers to develop more robust VrD-NER models. The UNER head considers the VrD-NER task as a combination of sequence labeling and reading order prediction, effectively addressing the issues of discontinuous entities in documents. Experimental evaluations on diverse datasets demonstrate the effectiveness of UNER in improving entity extraction performance. Moreover, the UNER head enables a supervised pre-training stage on various VrD-NER datasets to enhance the document transformer backbones and exhibits substantial knowledge transfer from the pre-training stage to the fine-tuning stage. By incorporating universal layout understanding, a pre-trained UNER-based model demonstrates significant advantages in few-shot and cross-linguistic scenarios and exhibits zero-shot entity extraction abilities. Huan Chen 0012, Jinyang Tang, Huijia Zhu, Qi Zhang 0001 |
ACM Multimedia | 6 |
| 2024 | DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric FinetuningabstractThe recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of these models is still limited when we expect to generate images that fall into a specific domain either hard to describe or just unseen to the models. In this work, we propose DomainGallery, a few-shot domain-driven image generation method which aims at finetuning pretrained Stable Diffusion on few-shot target datasets in an attribute-centric manner. Specifically, DomainGallery features prior attribute erasure, attribute disentanglement, regularization and enhancement. These techniques are tailored to few-shot domain-driven generation in order to solve key issues that previous works have failed to settle. Extensive experiments are given to validate the superior performance of DomainGallery on a variety of domain-driven generation scenarios. Yuxuan Duan, Yan Hong 0001, Bo Zhang 0075, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
NeurIPS | 5 |
| 2023 | Acoustics-Text Dual-Modal Joint Representation Learning for Cover Song IdentificationabstractCover Song Identification (CSI) is an important and challenging task in Music Information Retrieval (MIR). This paper focuses on investigating the multi-modal features of audio and text in the music domain and proposes two significant improvements to enhance the model performance for CSI. Firstly, our approach consists of a dual-encoder architecture that learns the embedding between the audio and corresponding song title information of music. Secondly, we propose a multi-modal representation learning strategy by jointly optimizing classification and metric learning losses in the audio modality, and contrastive learning loss in the audio-text modality. Experimental results demonstrate that our method efficiently learns more robust multi-modal representations for cover songs compared to a single audio encoder and achieves state-of-the-art results in CSI tasks. Yanmei Gu, Huijia Zhu |
ASRU | 5 |
| 2023 | Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path PredictionabstractRecent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs), in which named entity recognition (NER) is treated as a sequence-labeling task of predicting the BIO entity tags for tokens, following the typical setting of NLP.However, BIO-tagging scheme relies on the correct order of model inputs, which is not guaranteed in real-world NER on scanned VrDs where text are recognized and arranged by OCR systems.Such reading order issue hinders the accurate marking of entities by BIO-tagging scheme, making it impossible for sequencelabeling methods to predict correct named entities.To address the reading order issue, we introduce Token Path Prediction (TPP), a simple prediction head to predict entity mentions as token sequences within documents.Alternative to token classification, TPP models the document layout as a complete directed graph of tokens, and predicts token paths within the graph as entities.For better evaluation of VrD-NER systems, we also propose two revised benchmark datasets of NER on scanned documents which can reflect real-world scenarios.Experiment results demonstrate the effectiveness of our method, and suggest its potential to be a universal solution to various information extraction tasks on documents. Huan Chen 0012, Jinyang Tang, Huijia Zhu, Qi Zhang 0001, Tao Gui |
EMNLP | 6 |
| 2023 | LayoutGCN: A Lightweight Architecture for Visually Rich Document Understanding
Dengliang Shi, Jintao Du, Huijia Zhu |
ICDAR (3) | 4 |
| 2023 | DiffUTE: Universal Text Editing Diffusion ModelabstractDiffusion model based language-guided image editing has achieved great success recently. However, existing state-of-the-art diffusion models struggle with rendering correct text and text style during generation. To tackle this problem, we propose a universal self-supervised text editing diffusion model (DiffUTE), which aims to replace or modify words in the source image with another one while maintaining its realistic appearance. Specifically, we build our model on a diffusion model and carefully modify the network structure to enable the model for drawing multilingual characters with the help of glyph and position information. Moreover, we design a self-supervised learning framework to leverage large amounts of web data to improve the representation ability of the model. Experimental results show that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. Our code will be avaliable in \url{https://github.com/chenhaoxing/DiffUTE}. Haoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan 0001, Xing Zheng, Changhua Meng, Huijia Zhu, Weiqiang Wang 0002 |
NeurIPS | 8 |
| 2022 | A Multi-Task Dual-Tree Network for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) aims at extracting triplets from a given sentence, where each triplet includes an aspect, its sentiment polarity, and a corresponding opinion explaining the polarity. Existing methods are poor at detecting complicated relations between aspects and opinions as well as classifying multiple sentiment polarities in a sentence. Detecting unclear boundaries of multi-word aspects and opinions is also a challenge. In this paper, we propose a Multi-Task Dual-Tree Network (MTDTN) to address these issues. We employ a constituency tree and a modified dependency tree in two sub-tasks of Aspect Opinion Co-Extraction (AOCE) and ASTE, respectively. To enhance the information interaction between the two sub-tasks, we further design a Transition-Based Inference Strategy (TBIS) that transfers the boundary information from tags of AOCE to ASTE through a transition matrix. Extensive experiments are conducted on four popular datasets, and the results show the effectiveness of our model. Yichun Zhao, Gongshen Liu, Jintao Du, Huijia Zhu |
COLING | 5 |
| 2022 | Ant Multilingual Recognition System for OLR 2021 Challenge
Anqi Lyu, Huijia Zhu |
INTERSPEECH | 3 |
| 2021 | AntVoice Neural Speaker Embedding System for FFSVC 2020
Furong Xu, Kaisheng Yao, Huijia Zhu |
Interspeech | 6 |
| 2016 | Semantic Documents Relatedness using Concept Graph RepresentationabstractWe deal with the problem of document representation for the task of measuring semantic relatedness between documents. A document is represented as a compact concept graph where nodes represent concepts extracted from the document through references to entities in a knowledge base such as DBpedia. Edges represent the semantic and structural relationships among the concepts. Several methods are presented to measure the strength of those relationships. Concepts are weighted through the concept graph using closeness centrality measure which reflects their relevance to the aspects of the document. A novel similarity measure between two concept graphs is presented. The similarity measure first represents concepts as continuous vectors by means of neural networks. Second, the continuous vectors are used to accumulate pairwise similarity between pairs of concepts while considering their assigned weights. We evaluate our method on a standard benchmark for document similarity. Our method outperforms state-of-the-art methods including ESA (Explicit Semantic Annotation) while our concept graphs are much smaller than the concept vectors generated by ESA. Moreover, we show that by combining our concept graph with ESA, we obtain an even further improvement. Yuan Ni, Qiongkai Xu, Yosi Mass, Dafna Sheinwald, Huijia Zhu, Shao Sheng Cao |
WSDM | 6 |
| 2011 | Domain customization for aspect-oriented opinion analysis with multi-level latent sentiment cluesabstractAspect-oriented opinion mining detects the reviewers' sentiment orientation (e.g. positive, negative or neutral) towards different product-features. Domain customization is a big challenge for opinion mining due to the accuracy loss across domains. In this paper, we show our experiences and lessons learned in the domain customization for the aspect-oriented opinion analysis system OpinionIt. We present a customization method for sentiment classification with multi-level latent sentiment clues. We first construct Latent Semantic Association model to capture latent association among product-features from the unlabeled corpus. Meanwhile, we present an unsupervised method to effectively extract various domain-specific sentiment clues from the unlabeled corpus. In the customization, we tune the sentiment classifier on the labeled source domain data by incorporating the multi-level latent sentiment clues (e.g. latent association among product-features, domain-specific and generic sentiment clues). Experimental results show that the proposed method significantly reduces the accuracy loss of sentiment classification without any labeled target domain data. Huijia Zhu, Zhili Guo, Zhong Su |
CIKM | 2 |
| 2010 | OpinionIt: a text mining system for cross-lingual opinion analysisabstractOpinion mining focuses on extracting customers' opinions from the reviews and predicting their sentiment orientation. Reviewers usually praise a product in some aspects and bemoan it in other aspects. With the business globalization, it is very important for enterprises to extract the opinions toward different aspects and find out cross-lingual/cross-culture difference in opinions. Cross-lingual opinion mining is a very challenging task as amounts of opinions are written in different languages, and not well structured. Since people usually use different words to describe the same aspect in the reviews, product-feature (PF) categorization becomes very critical in cross-lingual opinion mining. Manual cross-lingual PF categorization is time consuming, and practically infeasible for the massive amount of data written in different languages. In order to effectively find out cross-lingual difference in opinions, we present an aspect-oriented opinion mining method with Cross-lingual Latent Semantic Association (CLaSA). We first construct CLaSA model to learn the cross-lingual latent semantic association among all the PFs from multi-dimension semantic clues in the review corpus. Then we employ CLaSA model to categorize all the multilingual PFs into semantic aspects, and summarize cross-lingual difference in opinions towards different aspects. Experimental results show that our method achieves better performance compared with the existing approaches. With CLaSA model, our text mining system OpinionIt can effectively discover cross-lingual difference in opinions. Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su |
CIKM | 2 |
| 2010 | CasJoin: a cascade chain for text similarity joinsabstractWe are concerned with the problem of similarity joins of text data, where the task is to find all pairs of documents above an expected similarity. Such a problem often serves as an indispensable step in many web applications. A crucial issue is to preclude unnecessary candidate pairs as many as possible ahead of expensive similarity evaluation. In this paper, we initiate an idea of adopting a cascade structure in text joins for a large speedup, where a latter stage can exclude a considerable number of invalid pairs survived in former stages. The proposed algorithm is shortly referred to as CasJoin. We further adopt a prefix filter to build the stage of CasJoin by introducing a novel vision to the dynamic generation of document vector. Specifically, a vector is partitioned into a chain of multiple prefixes that are appended one by one for cascade joining. We evaluate our CasJoin on a typical web corpus, ODP. Experiments indicate that, comparing to the state-of-the-art prefix algorithms, CasJoin can achieve a drastic reduction of candidates by as much as 98.15% and a dramatic speedup of joining by up to 13.34x. Xiaoxun Zhang, Zhili Guo, Huijia Zhu, Zhong Su |
CIKM | 4 |
| 2009 | Product feature categorization with multilevel latent semantic associationabstractIn recent years, the number of freely available online reviews is increasing at a high speed. Aspect-based opinion mining technique has been employed to find out reviewers' opinions toward different product aspects. Such finer-grained opinion mining is valuable for the potential customers to make their purchase decisions. Product-feature extraction and categorization is very important for better mining aspect-oriented opinions. Since people usually use different words to describe the same aspect in the reviews, product-feature extraction and categorization becomes more challenging. Manually product-feature extraction and categorization is tedious and time consuming, and practically infeasible for the massive amount of products. In this paper, we propose an unsupervised product-feature categorization method with multilevel latent semantic association. After extracting product-features from the semi-structured reviews, we construct the first latent semantic association (LaSA) model to group words into a set of concepts according to their virtual context documents. It generates the latent semantic structure for each product-feature. The second LaSA model is constructed to categorize the product-features according to their latent semantic structures and context snippets in the reviews. Experimental results demonstrate that our method achieves better performance compared with the existing approaches. Moreover, the proposed method is language- and domain-independent. Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su |
CIKM | 2 |
| 2009 | Address standardization with latent semantic associationabstractAddress standardization is a very challenging task in data cleansing. To provide better customer relationship management and business intelligence for customer-oriented cooperates, millions of free-text addresses need to be converted to a standard format for data integration, de-duplication and householding. Existing commercial tools usually employ lots of hand-craft, domain-specific rules and reference data dictionary of cities, states etc. These rules work better for the region they are designed. However, rule-based methods usually require more human efforts to rewrite these rules for each new domain since address data are very irregular and varied with countries and regions. Supervised learning methods usually are more adaptable than rule-based approaches. However, supervised methods need large-scale labeled training data. It is a labor-intensive and time-consuming task to build a large-scale annotated corpus for each target domain. For minimizing human efforts and the size of labeled training data set, we present a free-text address standardization method with latent semantic association (LaSA). LaSA model is constructed to capture latent semantic association among words from the unlabeled corpus. The original term space of the target domain is projected to a concept space using LaSA model at first, then the address standardization model is active learned from LaSA features and informative samples. The proposed method effectively captures the data distribution of the domain. Experimental results on large-scale English and Chinese corpus show that the proposed method significantly enhances the performance of standardization with less efforts and training data. Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su |
KDD | 2 |
| 2009 | Domain Adaptation with Latent Semantic Association for Named Entity Recognition
Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su |
HLT-NAACL | 2 |
| 2006 | Dependency Parsing Based on Dynamic Local Optimization
Ting Liu 0001, Jinshan Ma, Huijia Zhu, Sheng Li 0003 |
CoNLL | 3 |