VLDB 2026 Research / reviewers in the wild / expert
Huixing Jiang
dblp:23/8214
· DBLP profile ↗
18ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0001-8046-0370ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data with High and Consistent Preference Difference Are Better for Reward ModelabstractReinforcement Learning from Human Feedback (RLHF) is a commonly used alignment method for Large Language Models (LLMs). This method relies on a reward model trained on a preference dataset to provide scalar rewards. However, the human-annotated preference data is often sparse, noisy, and costly to obtain, necessitating more efficient utilization. This paper proposes a new metric for better preference data utilization from both theoretical and empirical perspectives. Starting with the Bradley-Terry model, we compute the Mean Square Error (MSE) between the expected loss and empirical loss of the reward model. Our findings reveal that data with higher and more consistent difference result in lower MSE. We therefore propose the Preference Difference (PD), the reward difference between two samples, as a filter for preference data. Experimental results on three open-source models show that reward models trained by filtered data with PD achieve higher calibrated accuracy, as well as better RLHF alignment performance. The conclusion remains consistent when we extend the experiments and theoretical derivations to implicit reward alignment algorithms, such as Direct Preference Optimization (DPO). Hengtong Lu, Caixia Yuan, Huixing Jiang |
AAAI | 5 |
| 2025 | A Systematic Exploration of Knowledge Graph Alignment with Large Language Models in Retrieval Augmented GenerationabstractRetrieval Augmented Generation (RAG) with Knowledge Graphs (KGs) is an effective way to enhance Large Language Models (LLMs). Due to the natural discrepancy between structured KGs and sequential LLMs, KGs must be linearized to text before being inputted into LLMs, leading to the problem of KG Alignment with LLMs (KGA). However, recent KG+RAG methods only consider KGA as a simple step without comprehensive and in-depth explorations, leaving three essential problems unclear: (1) What are the factors and their effects in KGA? (2) How do LLMs understand KGs? (3) How to improve KG+RAG by KGA? To fill this gap, we conduct systematic explorations on KGA, where we first define the problem of KGA and subdivide it into the graph transformation phase (graph-to-graph) and the linearization phase (graph-to-text). In the graph transformation phase, we study graph features at the node, edge, and full graph levels from low to high granularity. In the linearization phase, we study factors on formats, orders, and templates from structural to token levels. We conduct substantial experiments on 15 typical LLMs and three common datasets. Our main findings include: (1) The centrality of the KG affects the final generation; formats have the greatest impact on KGA; orders are model-dependent, without an optimal order adapting for all models; the templates with special token separators are better. (2) LLMs understand KGs by a unique mechanism, different from processing natural sentences, and separators play an important role. (3) We achieved 7.3% average performance improvements on four common LLMs on the KGQA task by combining the optimal factors to enhance KGA. Shiyu Tian, Shuyue Xing, Xingrui Li, Yangyang Luo, Caixia Yuan, Huixing Jiang, Xiaojie Wang 0006 |
AAAI | 7 |
| 2025 | Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language ModelsabstractLarge language models (LLMs) exhibit remarkable capabilities in natural language processing but face catastrophic forgetting when learning new tasks, where adaptation to a new domain leads to a substantial decline in performance on previous tasks. In this paper, we propose Controlled LoRA (CLoRA), a subspace regularization method on LoRA structure. Aiming to reduce the scale of output change while introducing minimal constraint on model capacity, CLoRA imposes constraints on the direction of updating matrix’s null space. Experimental results on one-stage LLM finetuning tasks and continual learning settings highlight the superiority of CLoRA as an effective parameter-efficient finetuning method with catastrophic forgetting mitigating. Further investigation for model parameters indicates that CLoRA effectively balances the trade-off between model capacity and degree of forgetting. The code for implementing CLoRA will be publicly available. Yuheng Lu, Bingshuo Qian, Caixia Yuan, Huixing Jiang |
ACL (1) | 4 |
| 2025 | VL-DynaRefine: A Vision-Language Dynamic Refinement Approach for Visual ReasoningabstractVisual reasoning is a key capability that significantly impacts the performance of multimodal tasks, such as compositional visual question answering and visual grounding. These tasks often require complex, multi-step reasoning processes. In recent years, several training-free methods for Vision-Language Models (VLMs) have emerged, with visual programming methods being proposed to enhance the capability of VLMs in visual reasoning tasks. While these methods have made some progress, they still face two primary challenges due to the lack of verification and refinement mechanisms for each action's output during the reasoning process: error accumulation and feedback delay, as well as insufficient utilization of multimodal contextual information. To address these challenges, we propose VL-DynaRefine, a training-free approach consisting of three modules: a planner, a verifier, and a refiner. The planner generates programmatic actions to solve the problem and executes each action in sequence, which is inspected by a verifier that reassesses the actions via confidence scores and determines whether refinement is necessary based on the evaluation results. In the refiner module, we incorporate a context-aware local refinement mechanism and a global refinement mechanism based on visual and action trajectories to reduce the impact of reasoning errors on the outcome. We evaluate our approach on multiple visual reasoning datasets, and the experimental results show that our method outperforms existing visual programming methods in both reasoning accuracy and efficiency, further validating its effectiveness in visual reasoning tasks. Zeyuan Zang, Fangxiang Feng, Caixia Yuan, Huixing Jiang, Xiaojie Wang 0006 |
ACM Multimedia | 7 |
| 2024 | Q-MoE: Connector for MLLMs with Text-Driven RoutingabstractMultimodal Large Language Models (MLLMs) have showcased remarkable advances in handling various vision-language tasks. These models typically consist of a Large Language Model (LLM), a vision encoder and a connector structure, which is used to bridge the modality gap between vision and language. It is challenging for the connector to filter the right visual information for LLM according to the task in hand. Most of previous connectors, such as light-weight projection and Q-former, treat visual information for diverse tasks uniformly, therefore lacking task-specific visual information extraction capabilities. To address the issue, this paper proposes Q-MoE, a query-based connector with Mixture-of-Experts (MoE) to extract task-specific information with text-driven routing. Furthermore, an optimal path based training strategy is proposed to find an optimal expert combination. Extensive experiments on two popular open-source LLMs and several different visual-language tasks demonstrate the effectiveness of the Q-MoE connecter. Hanzi Wang, Jiamin Ren, Huixing Jiang, Fangxiang Feng, Xiaojie Wang 0006 |
ACM Multimedia | 5 |
| 2024 | Improving Causal Inference of Large Language Models with SCM Tools
Zhenyang Hua, Shuyue Xing, Huixing Jiang, Xiaojie Wang 0006 |
NLPCC (3) | 3 |
| 2022 | A Region-based Document VQAabstractPractical Document Visual Question Answering (DocVQA) needs not only to recognize and extract the document contents, but also reason on them for answering questions. However, previous DocVQA data mainly focuses on in-line questions, where the answers could be directly extracted after locating keywords in the documents, which needs less reasoning. This paper therefore builds a large-scale dataset named Region-based Document VQA (RDVQA), which includes more practical questions for DocVQA. We then propose a novel Reason-over-In-region-Question-answering (ReIQ) model for addressing the problems. It is a pre-training-based model, where a Spatial-Token Pre-trained Model (STPM) is employed as the backbone. Two novel pre-training tasks, Masked Text Box Regression and Shuffled Triplet Reconstruction, are proposed to learn the entailment relationship between text blocks and tokens as well as contextual information, respectively. Moreover, a DocVQA State Tracking Module (DocST) is also proposed to track the DocVQA state in the fine-tuning stage. Experimental results show that our model improves the performance onRDVQA significantly, although more work should be done for practical DocVQA as shown inRDVQA. Xinya Wu, Duo Zheng, Jiashen Sun, Minzhen Hu, Fangxiang Feng, Xiaojie Wang 0006, Huixing Jiang, Fan Yang 0087 |
ACM Multimedia | 8 |
| 2022 | Revisit Overconfidence for OOD Detection: Reassigned Contrastive Learning with Adaptive Class-dependent ThresholdabstractYanan Wu, Keqing He, Yuanmeng Yan, QiXiang Gao, Zhiyuan Zeng, Fujia Zheng, Lulu Zhao, Huixing Jiang, Wei Wu, Weiran Xu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yanan Wu 0002, Keqing He 0001, Yuanmeng Yan, QiXiang Gao, Zhiyuan Zeng 0002, Fujia Zheng, Huixing Jiang, Wei Wu 0014, Weiran Xu |
NAACL-HLT | 8 |
| 2022 | Domain-Oriented Prefix-Tuning: Towards Efficient and Generalizable Fine-tuning for Zero-Shot Dialogue SummarizationabstractLulu Zhao, Fujia Zheng, Weihao Zeng, Keqing He, Weiran Xu, Huixing Jiang, Wei Wu, Yanan Wu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Fujia Zheng, Weihao Zeng 0003, Keqing He 0001, Weiran Xu, Huixing Jiang, Wei Wu 0014, Yanan Wu 0002 |
NAACL-HLT | 6 |
| 2022 | ADPL: Adversarial Prompt-based Domain Adaptation for Dialogue Summarization with Knowledge DisentanglementabstractTraditional dialogue summarization models rely on a large-scale manually-labeled corpus, lacking generalization ability to new domains, and domain adaptation from a labeled source domain to an unlabeled target domain is important in practical summarization scenarios. However, existing domain adaptation works in dialogue summarization generally require large-scale pre-training using extensive external data. To explore the lightweight fine-tuning methods, in this paper, we propose an efficient Adversarial Disentangled Prompt Learning (ADPL) model for domain adaptation in dialogue summarization. We introduce three kinds of prompts including domain-invariant prompt (DIP), domain-specific prompt (DSP), and task-oriented prompt (TOP). DIP aims to disentangle and transfer the shared knowledge from the source domain and target domain in an adversarial way, which improves the accuracy of prediction about domain-invariant information and enhances the ability for generalization to new domains. DSP is designed to guide our model to focus on domain-specific knowledge using domain-related features. TOP is to capture task-oriented knowledge to generate high-quality summaries. Instead of fine-tuning the whole pre-trained language model (PLM), we only update the prompt networks but keep PLM fixed. Experimental results on the zero-shot setting show that the novel design of prompts can yield more coherent, faithful, and relevant summaries than baselines using the prefix-tuning, and perform at par with fine-tuning while being more efficient. Overall, our work introduces a prompt-based perspective to the zero-shot learning for dialogue summarization task and provides valuable findings and insights for future research. Fujia Zheng, Weihao Zeng 0003, Keqing He 0001, Ruotong Geng, Huixing Jiang, Wei Wu 0014, Weiran Xu |
SIGIR | 6 |
| 2021 | Converse, Focus and Guess - Towards Multi-Document Driven Dialogue
Caixia Yuan, Xiaojie Wang 0006, Yushu Yang, Huixing Jiang, Zhongyuan Wang 0006 |
AAAI | 5 |
| 2021 | Novel Slot Detection: A Benchmark for Discovering Unknown Slot Types in the Task-Oriented Dialogue SystemabstractYanan Wu, Zhiyuan Zeng, Keqing He, Hong Xu, Yuanmeng Yan, Huixing Jiang, Weiran Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yanan Wu 0002, Zhiyuan Zeng 0002, Keqing He 0001, Hong Xu 0009, Yuanmeng Yan, Huixing Jiang, Weiran Xu |
ACL/IJCNLP (1) | 6 |
| 2021 | Capturing Event Argument Interaction via A Bi-Directional Entity-Level Recurrent DecoderabstractXi Xiangyu, Wei Ye, Shikun Zhang, Quanxiu Wang, Huixing Jiang, Wei Wu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xiangyu Xi, Wei Ye 0004, Shikun Zhang, Quanxiu Wang, Huixing Jiang, Wei Wu 0014 |
ACL/IJCNLP (1) | 5 |
| 2021 | Improving Event Detection by Exploiting Label HierarchyabstractEvent types are hierarchical, yet most existing methods for event detection classify candidate triggers into fine-grained event types directly, without considering the rich semantic correlations in the hierarchy of event types. To fully utilize such information to improve the detection of fine-grained event types, we propose a three-layer label hierarchy and introduce the detection of two coarser-grained types as auxiliary classification tasks. In particular, we leverage the supplementary supervision information from label hierarchy by a novel Logits Mapping (LM) strategy, which generates logits (the intermediate representations fed into classifier) for coarser-grained types by heuristic mapping of logits for fine-grained types. In this way, training signals provided by auxiliary tasks can help the encoder produce more precise logits via back propagation, thus providing a simple (no extra parameter needed) yet effective way to improve the target task. Results of extensive experiments on the ACE 2005 show that LM can not only be easily integrated into the state-of-the-art methods and achieve significant improvement over them, but also can effectively alleviate the data sparseness problem. Xiangyu Xi, Wei Ye 0004, Tong Zhang 0001, Quanxiu Wang, Shikun Zhang, Huixing Jiang, Wei Wu 0014 |
ICASSP | 6 |
| 2020 | Syntactic Graph Convolutional Network for Spoken Language UnderstandingabstractSlot filling and intent detection are two major tasks for spoken language understanding.In most existing work, these two tasks are built as joint models with multi-task learning with no consideration of prior linguistic knowledge.In this paper, we propose a novel joint model that applies a graph convolutional network over dependency trees to integrate the syntactic structure for learning slot filling and intent detection jointly.Experimental results show that our proposed model achieves state-of-the-art performance on two public benchmark datasets and outperforms existing work.At last, we apply the BERT model to further improve the performance on both slot filling and intent detection. Keqing He 0001, Shuyu Lei, Yushu Yang, Huixing Jiang |
COLING | 4 |
| 2020 | Learning Visual Features from Product Title for Image RetrievalabstractThere is a huge market demand for searching for products by images in e-commerce sites. Visual features play the most important role in solving this content-based image retrieval task. Most existing methods leverage pre-trained models on other large-scale datasets with well-annotated labels, e.g. the ImageNet dataset, to extract visual features. However, due to the large difference between the product images and the images in ImageNet, the feature extractor trained on ImageNet is not efficient in extracting the visual features of product images. And retraining the feature extractor on the product images is faced with the dilemma of lacking the annotated labels. In this paper, we utilize the easily accessible text information, that is, the product title, as a supervised signal to learn the features of the product image. Specifically, we use the n-grams extracted from the product title as the label of the product image to construct a dataset for image classification. This dataset is then used to fine-tuned a pre-trained model. Finally, the basic max-pooling activation of convolutions (MAC) feature is extracted from the fine-tuned model. As a result, we achieve the fourth position in the Grand Challenge of AI Meets Beauty in 2020 ACM Multimedia by using only a single ResNet-50 model without any human annotations and pre-processing or post-processing tricks. Our code is available at: \urlhttps://github.com/FangxiangFeng/AI-Meets-Beauty-2020. Fangxiang Feng, Tianrui Niu, Ruifan Li, Xiaojie Wang 0006, Huixing Jiang |
ACM Multimedia | 5 |
| 2020 | Answer-Driven Visual State Estimator for Goal-Oriented Visual DialogueabstractA goal-oriented visual dialogue involves multi-turn interactions between two agents, Questioner and Oracle. During which, the answer given by Oracle is of great significance, as it provides golden response to what Questioner concerns. Based on the answer, Questioner updates its belief on target visual content and further raises another question. Notably, different answers drive into different visual beliefs and future questions. However, existing methods always indiscriminately encode answers after much longer questions, resulting in a weak utilization of answers. In this paper, we propose an Answer-Driven Visual State Estimator (ADVSE) to impose the effects of different answers on visual states. First, we propose an Answer-Driven Focusing Attention (ADFA) to capture the answer-driven effect on visual attention by sharpening question-related attention and adjusting it by answer-based logical operation at each turn. Then based on the focusing attention, we get the visual state estimation by Conditional Visual Information Fusion (CVIF), where overall information and difference information are fused conditioning on the question-answer state. We evaluate the proposed ADVSE to both question generator and guesser tasks on the large-scale GuessWhat?! dataset and achieve the state-of-the-art performances on both tasks. The qualitative results indicate that the ADVSE boosts the agent to generate highly efficient questions and obtains reliable visual attentions during the reasonable question generation and guess processes. Zipeng Xu, Fangxiang Feng, Xiaojie Wang 0006, Yushu Yang, Huixing Jiang, Zhongyuan Wang 0006 |
ACM Multimedia | 5 |
| 2010 | Second-Order HMM for Event Extraction from Short Message
Huixing Jiang, Xiaojie Wang 0006, Jilei Tian |
NLDB | 1 |