Jinghui Qin

dblp:228/6607 · DBLP profile ↗
← Back
53ranked-venue papers
12as first author
47since 2021 · last 2026
0000-0003-0663-199XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 10 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 18 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 PR-CapsNet: Pseudo-Riemannian Capsule Network with Adaptive Curvature Routing for Graph Learning
abstract
Capsule Networks (CapsNets) show exceptional graph representation capacity via dynamic routing and vectorized hierarchical representations, but they model the complex geometries of real-world graphs (e.g., hierarchies, clusters, cycles) poorly by fixed-curvature space due to the inherent geodesical disconnectedness issues, leading to suboptimal performance. Recent works find that non-Euclidean pseudo-Riemannian manifolds provide specific inductive biases for embedding graph data, but how to leverage them to improve CapsNets is still underexplored. Here, we extend the Euclidean capsule routing into geodesically disconnected pseudo-Riemannian manifolds and derive a Pseudo-Riemannian Capsule Network (PR-CapsNet), which models data in pseudo-Riemannian manifolds of adaptive curvature, for graph representation learning. Specifically, PR-CapsNet enhances the CapsNet with Adaptive Pseudo-Riemannian Tangent Space Routing by utilizing pseudo-Riemannian geometry. Unlike single-curvature or subspace-partitioning methods, PR-CapsNet concurrently models hierarchical and cluster/cyclic graph structures via its versatile pseudo-Riemannian metric. It first deploys Pseudo-Riemannian Tangent Space Routing to decompose capsule states into spherical-temporal and Euclidean-spatial subspaces with diffeomorphic transformations. Then, an Adaptive Curvature Routing is developed to adaptively fuse features from different curvature spaces for complex graphs via a learnable curvature tensor with geometric attention from local manifold properties. Finally, a geometric properties-preserved Pseudo-Riemannian Capsule Classifier is developed to project capsule embeddings to tangent spaces and use curvature-weighted softmax for classification. Extensive experiments on node and graph classification benchmarks show PR-CapsNet outperforms state-of-the-art models, validating PR-CapsNet's strong representation power for complex graph structures.
Ye Qin, Jingchao Wang 0002, Haiying Huang 0005, Junxu Li, Tinghui Chen, Jinghui Qin
WSDM8
2026 Multimodal feature disentangle-fusion network for detecting and grounding multi-modal media manipulation
Jinghui Qin, Tianshui Chen, Zhijing Yang
Expert Syst. Appl.1
2026 Multimodal progressive fusion and enhancement network for multimodal sentiment analysis
Jinghui Qin, Qite Zhou, Lihuang Fang, Zhijing Yang
Expert Syst. Appl.1
2026 Revisiting DIRE: towards universal AI-generated image detection
Huanqi Lin, Jinghui Qin, Xiaoqi Wu, Tianshui Chen, Zhijing Yang
Neural Networks2
2026 4PM: Privacy-Preserving Patient-Provider Matching Service in Digital Healthcare System
abstract
For digital health platforms, the challenge is balancing patient privacy with the ability to match patients to the right providers quickly and accurately. Existing systems often suffer from privacy leakage, insufficient matching precision, and degraded performance when dealing with large-scale data. In this paper, we propose 4PM, a novel privacy-preserving patient-provider matching scheme that leverages secure computation to deliver strong privacy guarantees while ensuring efficient and accurate matching. Our method partitions patient data between two non-colluding servers via secret sharing, employing the optimized Millionaires' Protocol for secure ranking and leveraging oblivious retrieval techniques for privacy-preserving matching. 4PM significantly reduces the computational complexity of high-dimensional data, achieving end-to-end latency within 0.5 seconds in scenarios with 200 doctors and 200-dimensional symptom vectors. Our work contributes to fostering secure and trustworthy healthcare in the digital era.
Jing Lei 0007, Fake Lyu, Jinghui Qin, Qingqi Pei
IEEE J. Biomed. Health Informatics4
2026 CIREC: Causal Intervention-Inspired Policy Learning to Mitigate Exposure Bias for Interactive Recommendation
Yongsen Zheng, Guohua Wang 0005, Jinghui Qin, Ziliang Chen 0001, Junfan Lin, Pengxu Wei, Liang Lin 0004, Kwok-Yan Lam
IEEE Trans. Knowl. Data Eng.3
2025 Mental-Perceiver: Audio-Textual Multi-Modal Learning for Estimating Mental Disorders
abstract
Mental disorders, such as anxiety and depression, have become a global concern that affects people of all ages. Early detection and treatment are crucial to mitigate the negative effects these disorders can have on daily life. Although AI-based detection methods show promise, progress is hindered by the lack of publicly available large-scale datasets. To address this, we introduce the Multi-Modal Psychological assessment corpus (MMPsy), a large-scale dataset containing audio recordings and transcripts from Mandarin-speaking adolescents undergoing automated anxiety/depression assessment interviews. MMPsy also includes self-reported anxiety/depression evaluations using standardized psychological questionnaires. Leveraging this dataset, we propose Mental-Perceiver, a deep learning model for estimating mental disorders from audio and textual data. Extensive experiments on MMPsy and the DAIC-WOZ dataset demonstrate the effectiveness of Mental-Perceiver in anxiety and depression detection.
Jinghui Qin, Changsong Liu, Tianchi Tang, Dahuang Liu, Qianying Huang, Rumin Zhang
AAAI1
2025 Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models
abstract
Recently, Large Language Models (LLMs) with in-context learning have demonstrated remarkable potential in handling neural machine translation. However, existing evidence shows that LLMs are prompt-sensitive and it is sub-optimal to apply the fixed prompt to any input for downstream machine translation tasks. To address this issue, we propose an adaptive few-shot prompting (AFSP) framework to automatically select suitable translation demonstrations for various source input sentences to further elicit the translation capability of an LLM for better machine translation. First, we build a translation demonstration retrieval module based on LLM's embedding to retrieve top-k semantic-similar translation demonstrations from aligned parallel translation corpus. Rather than using other embedding models for semantic demonstration retrieval, we build a hybrid demonstration retrieval module based on the embedding layer of the deployed LLM to build better input representation for retrieving more semantic-related translation demonstrations. Then, to ensure better semantic consistency between source inputs and target outputs, we force the deployed LLM itself to generate multiple output candidates in the target language with the help of translation demonstrations and rerank these candidates. Besides, to better evaluate the effectiveness of our AFSP framework on the latest language and extend the research boundary of neural machine translation, we construct a high-quality diplomatic Chinese-English parallel dataset that consists of 5,528 parallel Chinese-English sentences. Finally, extensive experiments on the proposed diplomatic Chinese-English parallel dataset and the United Nations Parallel Corpus (Chinese-English part) show the effectiveness and superiority of our proposed AFSP.
Jinghui Qin, Wenxuan Ye, Hao Tan 0007, Zhijing Yang
AAAI2
2025 AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
abstract
Recent advancements in multimodal large language models (MLLMs) have garnered significant attention, offering a promising pathway toward artificial general intelligence (AGI).Among the essential capabilities required for AGI, creativity has emerged as a critical trait for MLLMs, with association serving as its foundation.Association reflects a model's ability to think creatively, making it vital to evaluate and understand.While several frameworks have been proposed to assess associative ability, they often overlook the inherent ambiguity in association tasks, which arises from the divergent nature of associations and undermines the reliability of evaluations.To address this issue, we decompose ambiguity into two types-internal ambiguity and external ambiguity-and introduce AssoCiAm, a benchmark designed to evaluate associative ability while circumventing the ambiguity through a hybrid computational method.We then conduct extensive experiments on MLLMs, revealing a strong positive correlation between cognition and association.Additionally, we observe that the presence of ambiguity in the evaluation process causes MLLMs' behavior to become more random-like.Finally, we validate the effectiveness of our method in ensuring more accurate and reliable evaluations.See Project Page for the data and codes.
Wenkuan Zhao, Shanshan Zhong, Jinghui Qin, Mingfu Liang, Zhongzhan Huang, Wushao Wen
EMNLP4
2025 Boundary-Driven Table-Filling with Cross-Granularity Contrastive Learning for Aspect Sentiment Triplet Extraction
abstract
The Aspect Sentiment Triplet Extraction (ASTE) task aims to extract aspect terms, opinion terms, and their corresponding sentiment polarity from a given sentence. It remains one of the most prominent subtasks in fine-grained sentiment analysis. Most existing approaches frame triplet extraction as a 2D table-filling process in an end-to-end manner, focusing primarily on word-level interactions while often overlooking sentence-level representations. This limitation hampers the model’s ability to capture global contextual information, particularly when dealing with multi-word aspect and opinion terms in complex sentences. To address these issues, we propose boundary-driven table-filling with cross-granularity contrastive learning (BTF-CCL) to enhance the semantic consistency between sentence-level representations and word-level representations. By constructing positive and negative sample pairs, the model is forced to learn the associations at both the sentence level and the word level. Additionally, a multi-scale, multi-granularity convolutional method is proposed to capture rich semantic information better. Our approach can capture sentence-level contextual information more effectively while maintaining sensitivity to local details. Experimental results show that the proposed method achieves state-of-the-art performance on public benchmarks according to the F1 score.
Qingling Li, Wushao Wen, Jinghui Qin
ICASSP3
2025 DVIB: Towards Robust Multimodal Recommender Systems via Variational Information Bottleneck Distillation
abstract
In multimodal recommender systems (MRS), integrating various modalities helps to model user preferences and item characteristics more accurately, thereby assisting users in discovering items that match their interests. Although the introduction of multimodal information offers opportunities for performance improvement, it will increase the risks of inherent noise and information redundancy, posing challenges to the robustness of MRS. Many existing methods typically address these two issues separately either by introducing perturbations at the model input for robust training to handle noise or by designing complex network structures to filter out redundant information. In contrast, we propose the DVIB framework to simultaneously address both issues in a simple manner. We found that moving the perturbations from the input layer to the hidden layer, combined with feature self-distillation, can mitigate noise and handle information redundancy without altering the original network architecture. Additionally, we also provide theoretical evidence for the effectiveness of DVIB, demonstrating that the framework not only explicitly enhances the robustness of model training but also implicitly exhibits an information bottleneck effect, which effectively reduces redundant information during multimodal fusion and improves feature extraction quality. Extensive experiments show that DVIB consistently improves the performance of MRS across different datasets and model settings, and it can complement existing robust training methods, representing a promising new paradigm in MRS. See code at https://github.com/MarshmallowLight/DVIB.git.
Wenkuan Zhao, Shanshan Zhong, Wushao Wen, Jinghui Qin, Mingfu Liang, Zhongzhan Huang
WWW5
2025 PSCNet: Long sequence time-series forecasting for photovoltaic power via period selection and cross-variable attention
Hao Tan 0007, Jinghui Qin, Zizheng Li, Weiyan Wu
Appl. Intell.2
2025 CrossGCL: cross-view graph contrastive learning with dual tasks for drug recommendation
Wushao Wen, Lihuang Fang, Qiangpu Chen, Jinghui Qin
Neural Comput. Appl.5
2025 Learning Semantic-aware Representation in Visual-Language Models for Multi-label Recognition with Partial Labels
abstract
Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since collecting large-scale and complete multi-label datasets is difficult in real application scenarios. Recently, vision language models (e.g., CLIP) have demonstrated impressive transferability to downstream tasks in data limited or label limited settings. However, current CLIP-based methods suffer from semantic confusion in MLR task due to the lack of fine-grained information in the single global visual and textual representation for all categories. In this work, we address this problem by introducing a semantic decoupling module and a category-specific prompt optimization method in CLIP-based framework. Specifically, the semantic decoupling module following the visual encoder learns category-specific feature maps by utilizing the semantic-guided spatial attention mechanism. Moreover, the category-specific prompt optimization method is introduced to learn text representations aligned with category semantics. Therefore, the prediction of each category is independent, which alleviate the semantic confusion problem. Extensive experiments on Microsoft COCO 2014 and Pascal VOC 2007 datasets demonstrate that the proposed framework significantly outperforms current state-of-art methods with a simpler model structure. Additionally, visual analysis shows that our method effectively separates information from different categories and achieves better performance compared to CLIP-based baseline method.
Haoxian Ruan, Zhihua Xu, Zhijing Yang, Yongyi Lu, Jinghui Qin, Tianshui Chen
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Adaptive Prompt Routing for Arbitrary Text Style Transfer with Pre-trained Language Models
abstract
Recently, arbitrary text style transfer (TST) has made significant progress with the paradigm of prompt learning. In this paradigm, researchers often design or search for a fixed prompt for any input. However, existing evidence shows that large language models (LLMs) are prompt-sensitive and it is sub-optimal to apply the same prompt to any input for downstream TST tasks. Besides, the prompts obtained by searching are often unreadable and unexplainable to humans. To address these issues, we propose an Adaptive Prompt Routing (APR) framework to adaptively route prompts from a human-readable prompt set for various input texts and given styles. Specifically, we first construct a candidate prompt set of diverse and human-readable prompts for the target style. This set consists of several seed prompts and their variants paraphrased by an LLM. Subsequently, we train a prompt routing model to select the optimal prompts efficiently according to inputs. The adaptively selected prompt can guide the LLMs to perform a precise style transfer for each input sentence while maintaining readability for humans. Extensive experiments on 4 public TST benchmarks over 3 popular LLMs (with parameter sizes ranging from 1.5B to 175B) demonstrate that our APR achieves superior style transfer performances, compared to the state-of-the-art prompt-based and fine-tuning methods. The source code is available at https://github.com/DwyaneLQY/APR
Qingyi Liu, Jinghui Qin, Wenxuan Ye, Hao Mou, Keze Wang
AAAI2
2024 FacetCRS: Multi-Faceted Preference Learning for Pricking Filter Bubbles in Conversational Recommender System
abstract
The filter bubble is a notorious issue in Recommender Systems (RSs), which describes the phenomenon whereby users are exposed to a limited and narrow range of information or content that reinforces their existing dominant preferences and beliefs. This results in a lack of exposure to diverse and varied content. Many existing works have predominantly examined filter bubbles in static or relatively-static recommendation settings. However, filter bubbles will be continuously intensified over time due to the feedback loop between the user and the system in the real-world online recommendation. To address these issues, we propose a novel paradigm, Multi-Facet Preference Learning for Pricking Filter Bubbles in Conversational Recommender System (FacetCRS), which aims to burst filter bubbles in the conversational recommender system (CRS) through timely user-item interactions via natural language conversations. By considering diverse user preferences and intentions, FacetCRS automatically model user preference into multi-facets, including entity-, word-, context-, and review-facet, to capture diverse and dynamic user preferences to prick filter bubbles in the CRS. It is an end-to-end CRS framework to adaptively learn representations of various levels of preference facet and diverse types of external knowledge. Extensive experiments on two publicly available benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance in mitigating filter bubbles and enhancing recommendation quality in CRS.
Yongsen Zheng, Ziliang Chen 0001, Jinghui Qin, Liang Lin 0004
AAAI3
2024 HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
abstract
Yongsen Zheng, Ruilin Xu, Ziliang Chen, Guohua Wang, Mingjie Qian, Jinghui Qin, Liang Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yongsen Zheng, Ruilin Xu 0006, Ziliang Chen 0001, Guohua Wang 0005, Mingjie Qian, Jinghui Qin, Liang Lin 0004
ACL (1)6
2024 Stripe Observation Guided Inference Cost-Free Attention Mechanism
Zhongzhan Huang, Shanshan Zhong, Wushao Wen, Jinghui Qin, Liang Lin 0004
ECCV (24)4
2024 DEEPAM: Toward Deeper Attention Module in Residual Convolutional Neural Networks
Shanshan Zhong, Wushao Wen, Jinghui Qin, Zhongzhan Huang
ICANN (1)3
2024 Exploring Out-of-Distribution Scene Text Recognition for Driving Scenes with Hybrid Test-Time Adaptation
Xiaoyu Xian, Jinghui Qin, Yukai Shi, Daxin Tian, Liang Lin 0004
PRCV (1)2
2024 Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima
abstract
Multimodal recommender systems utilize various types of information to model user preferences and item features, helping users discover items aligned with their interests. The integration of multimodal information mitigates the inherent challenges in recommender systems, e.g., the data sparsity problem and cold-start issues. However, it simultaneously magnifies certain risks from multimodal information inputs, such as information adjustment risk and inherent noise risk. These risks pose crucial challenges to the robustness of recommendation models. In this paper, we analyze multimodal recommender systems from the novel perspective of flat local minima and propose a concise yet effective gradient strategy called Mirror Gradient (MG). This strategy can implicitly enhance the model's robustness during the optimization process, mitigating instability risks arising from multimodal information inputs. We also provide strong theoretical evidence and conduct extensive empirical experiments to show the superiority of MG across various multimodal recommendation models and benchmarks. Furthermore, we find that the proposed MG can complement existing robust training methods and be easily extended to diverse advanced recommendation models, making it a promising new and fundamental paradigm for training multimodal recommender systems. The code is released at https://github.com/Qrange-group/Mirror-Gradient.
Shanshan Zhong, Zhongzhan Huang, Daifeng Li, Wushao Wen, Jinghui Qin, Liang Lin 0004
WWW5
2024 CROSE: Low-light enhancement by CROss-SEnsor interaction for nighttime driving scenes
Xiaoyu Xian, Jinghui Qin, Yin Tian, Yukai Shi, Daxin Tian
Expert Syst. Appl.3
2024 CAT: Continual Adapter Tuning for aspect sentiment classification
Qiangpu Chen, Jiahua Huang, Wushao Wen, Qingling Li, Rumin Zhang, Jinghui Qin
Neurocomputing6
2024 Unsupervised multi-perspective fusing semantic alignment for cross-modal hashing retrieval
Yongfeng Chen, Junpeng Tan, Zhijing Yang, Yukai Shi, Jinghui Qin
Multim. Tools Appl.5
2024 GlobalSR: Global context network for single image super-resolution via deformable convolution attention and fast Fourier convolution
Qiangpu Chen, Wushao Wen, Jinghui Qin
Neural Networks3
2024 Improving span-based Aspect Sentiment Triplet Extraction with part-of-speech filtering and contrastive learning
Qingling Li, Wushao Wen, Jinghui Qin
Neural Networks3
2024 ALAN: Self-Attention Is Not All You Need for Image Super-Resolution
abstract
Vision Transformer (ViT)-based image super-resolution (SR) methods have achieved impressive performance and surpassed CNN-based SR methods by utilizing Multi-Head Self-Attention (MHSA) to model long-range dependencies. However, the quadratic complexity of MHSA and the inefficiency of non-parallelized window partition seriously affect the inference speed, hindering these SR methods from being applied to application scenarios requiring speed and quality. To address this issue, we propose an Asymmetric Large-kernel Attention Network (ALAN) utilizing a stage-to-block design paradigm inspired by ViT. In the ALAN, the core block named Asymmetric Large Kernel Convolution Block (ALKCB) adopts a similar structure to the Swin Transformer Layer but replaces the MHSA with our proposed Asymmetric Depth-Wise Convolution Attention (ADWCA) to enhance both the SR quality and inference speed. The proposed ADWCA, with linear complexity, uses large kernel depth-wise dilation convolution and Hadamard product as the attention map. The structural re-parameterization technique to strengthen the kernel skeletons with asymmetric convolution is also explored. Experimental results demonstrate that ALAN achieves state-of-the-art performance with faster inference speed than ViT-based models and smaller parameters than CNN-based models. Specifically, the tiny size of ALAN (ALAN-T) is$3\times$smaller than ShuffleMixer with similar performance, and ALAN is$4\times$faster than SwinIR-S with 0.1 dB gain in PSNR.
Qiangpu Chen, Jinghui Qin, Wushao Wen
IEEE Signal Process. Lett.2
2024 An Introspective Data Augmentation Method for Training Math Word Problem Solvers
abstract
Though quite challenging, training a deep neural network for automatically solving Math Word Problems (MWPs) has increasingly attracted attention due to its significance in investigating how a machine can understand and reason complex problems like a human. However, the data volume of existing high-quality MWP datasets is far from sufficient to train a robust solver since collecting these datasets would cost a very high price, i.e., they require professional knowledge that accords with the educational standard and massive accessible data. This data bottleneck inspires us to consider using cost-effective data augmentation methods to improve the utilization of the existing data and enhance the performance of an MWP solver. Nevertheless, the traditional input-based data augmentation methods for training natural image or language models are incompetent for training MWP solvers due to the following two reasons. First, MWPs are concise yet comprehensive, so these data augmentation methods are prone to make them more ambiguous. Second, the mathematical dependencies grounded in the problems must be maintained when a batch of augmented examples is generated during the data augmentation process. To address these issues, we propose a simple yet effective data augmentation method called the Introspective Data Augmentation Method (IDAM) that allows the MWP examples to be augmented in latent space during the training of the neural network, instead of making perturbations over the input data. In particular, our IDAM is capable of applying different data augmentation operations on the latent feature representations of MWPs to produce new examples. Moreover, a new training objective is developed to constrain the mathematical dependency consistency between the original MWP and the produced ones. Extensive experiments conducted on standard benchmarks demonstrate the effectiveness of IDAM in generally improving the performance of existing MWP solvers without any elaborated model crafting.
Jinghui Qin, Zhongzhan Huang, Quanshi Zhang, Liang Lin 0004
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 Enhancing Vision and Language Navigation With Prompt-Based Scene Knowledge
abstract
A challenging task in embodied artificial intelligence is enabling the robot to carry out a navigational task following natural language instruction. In the task, the navigator needs to understand objects, directions, as well as room types, which serve as landmarks for navigation. Although it is easy to encode objects and directions with an external encoder like an object detector, current navigators struggle to encode room type information properly due to the low accuracy offered by existing classifiers. This inadequacy poses confusion that navigators find difficult to overcome. Even humans may sometimes fail to determine the exact type of a room since multiple room types may exist in one panorama. To mitigate this problem, we propose to encode room type information in a prompt manner. Specifically, we first establish multi-modal, learnable prompt pools containing knowledge of room types. By querying the prompt pools, the navigator can obtain room-type prompts of the current view, and incorporate them into the navigator using a prompt-based learning method. Experimental results on the REVERIE, R2R and SOON datasets demonstrate the effectiveness of our approach.
Zhaohuan Zhan, Jinghui Qin, Wei Zhuo 0006, Guang Tan
IEEE Trans. Circuits Syst. Video Technol.2
2024 Dynamic Correlation Learning and Regularization for Multi-Label Confidence Calibration
abstract
Modern visual recognition models often display overconfidence due to their reliance on complex deep neural networks and one-hot target supervision, resulting in unreliable confidence scores that necessitate calibration. While current confidence calibration techniques primarily address single-label scenarios, there is a lack of focus on more practical and generalizable multi-label contexts. This paper introduces the Multi-Label Confidence Calibration (MLCC) task, aiming to provide well-calibrated confidence scores in multi-label scenarios. Unlike single-label images, multi-label images contain multiple objects, leading to semantic confusion and further unreliability in confidence scores. Existing single-label calibration methods, based on label smoothing, fail to account for category correlations, which are crucial for addressing semantic confusion, thereby yielding sub-optimal performance. To overcome these limitations, we propose the Dynamic Correlation Learning and Regularization (DCLR) algorithm, which leverages multi-grained semantic correlations to better model semantic confusion for adaptive regularization. DCLR learns dynamic instance-level and prototype-level similarities specific to each category, using these to measure semantic correlations across different categories. With this understanding, we construct adaptive label vectors that assign higher values to categories with strong correlations, thereby facilitating more effective regularization. We establish an evaluation benchmark, re-implementing several advanced confidence calibration algorithms and applying them to leading multi-label recognition (MLR) models for fair comparison. Through extensive experiments, we demonstrate the superior performance of DCLR over existing methods in providing reliable confidence scores in multi-label scenarios.
Tianshui Chen, Weihang Wang 0008, Tao Pu 0002, Jinghui Qin, Zhijing Yang, Jie Liu 0022, Liang Lin 0004
IEEE Trans. Image Process.4
2024 Template-Based Contrastive Distillation Pretraining for Math Word Problem Solving
abstract
Since math word problem (MWP) solving aims to transform natural language problem description into executable solution equations, an MWP solver needs to not only comprehend the real-world narrative described in the problem text but also identify the relationships among the quantifiers and variables implied in the problem and maps them into a reasonable solution equation logic. Recently, although deep learning models have made great progress in MWPs, they ignore the grounding equation logic implied by the problem text. Besides, as we all know, pretrained language models (PLM) have a wealth of knowledge and high-quality semantic representations, which may help solve MWPs, but they have not been explored in the MWP-solving task. To harvest the equation logic and real-world knowledge, we propose a template-based contrastive distillation pretraining (TCDP) approach based on a PLM-based encoder to incorporate mathematical logic knowledge by multiview contrastive learning while retaining rich real-world knowledge and high-quality semantic representation via knowledge distillation. We named the pretrained PLM-based encoder by our approach as MathEncoder. Specifically, the mathematical logic is first summarized by clustering the symbolic solution templates among MWPs and then injected into the deployed PLM-based encoder by conducting supervised contrastive learning based on the symbolic solution templates, which can represent the underlying solving logic in the problems. Meanwhile, the rich knowledge and high-quality semantic representation are retained by distilling them from a well-trained PLM-based teacher encoder into our MathEncoder. To validate the effectiveness of our pretrained MathEncoder, we construct a new solver named MathSolver by replacing the GRU-based encoder with our pretrained MathEncoder in GTS, which is a state-of-the-art MWP solver. The experimental results demonstrate that our method can carry a solver's understanding ability of MWPs to a new stage by outperforming existing state-of-the-art methods on two widely adopted benchmarks Math23K and CM17K. Code will be available at https://github.com/QinJinghui/tcdp.
Jinghui Qin, Xiaodan Liang, Liang Lin 0004
IEEE Trans. Neural Networks Learn. Syst.1
2024 CIPL: Counterfactual Interactive Policy Learning to Eliminate Popularity Bias for Online Recommendation
abstract
Popularity bias, as a long-standing problem in recommender systems (RSs), has been fully considered and explored for offline recommendation systems in most existing relevant researches, but very few studies have paid attention to eliminate such bias in online interactive recommendation scenarios. Bias amplification will become increasingly serious over time due to the existence of feedback loop between the user and the interactive system. However, existing methods have only investigated the causal relations among different factors statically without considering temporal dependencies inherent in the online interactive recommendation system, making them difficult to be adapted to online settings. To address these problems, we propose a novel counterfactual interactive policy learning (CIPL) method to eliminate popularity bias for online recommendation. It first scrutinizes the causal relations in the interactive recommender models and formulates a novel temporal causal graph (TCG) to guide the training and counterfactual inference of the causal interactive recommendation system. Concretely, TCG is used to estimate the causal relations of item popularity on prediction score when the user interacts with the system at each time during model training. Besides, it is also used to remove the negative effect of popularity bias in the test stage. To train the causal interactive recommendation system, we formulated our CIPL by the actor-critic framework with an online interactive environment simulator. We conduct extensive experiments on three public benchmarks and the experimental results demonstrate that our proposed method can achieve the new state-of-the-art performance.
Yongsen Zheng, Jinghui Qin, Pengxu Wei, Ziliang Chen 0001, Liang Lin 0004
IEEE Trans. Neural Networks Learn. Syst.2
2023 HutCRS: Hierarchical User-Interest Tracking for Conversational Recommender System
abstract
Conversational Recommender System (CRS)aims to explicitly acquire user preferences towards items and attributes through natural language conversations.However, existing CRS methods ask users to provide explicit answers (yes/no) for each attribute they require, regardless of users' knowledge or interest, which may significantly reduce the user experience and semantic consistency.Furthermore, these methods assume that users like all attributes of the target item and dislike those unrelated to it, which can introduce bias in attributelevel feedback and impede the system's ability to accurately identify the target item.To address these issues, we propose a more realistic, user-friendly, and explainable CRS framework called Hierarchical User-Interest Tracking for Conversational Recommender System (HutCRS).HutCRS portrays the conversation as a hierarchical interest tree that consists of two stages.In stage I, the system identifies the aspects that the user prefers while the system asks about attributes related to these positive aspects or recommends items in stage II.In addition, we develop a Hierarchical-Interest Policy Learning (HIPL) module to integrate the decision-making process of which aspects to ask and when to ask about attributes or recommend items.Moreover, we classify the attribute-level feedback results to further enhance the system's ability to capture special information, such as attribute instances that are accepted by users but not presented in their historical interactive data.Extensive experiments on four benchmark datasets demonstrate the superiority of our method.The implementation of HutCRS is publicly available at https://github.com/xinle1129/HutCRS.
Mingjie Qian, Yongsen Zheng, Jinghui Qin, Liang Lin 0004
EMNLP3
2023 Understanding Self-attention Mechanism via Dynamical System Perspective
abstract
The self-attention mechanism (SAM) is widely used in various fields of artificial intelligence and has successfully boosted the performance of different models. However, current explanations of this mechanism are mainly based on intuitions and experiences, while there still lacks direct modeling for how the SAM helps performance. To mitigate this issue, in this paper, based on the dynamical system perspective of the residual neural network, we first show that the intrinsic stiffness phenomenon (SP) in the high-precision solution of ordinary differential equations (ODEs) also widely exists in high-performance neural networks (NN). Thus the ability of NN to measure SP at the feature level is necessary to obtain high performance and is an important factor in the difficulty of training NN. Similar to the adaptive step-size method which is effective in solving stiff ODEs, we show that the SAM is also a stiffness-aware step size adaptor that can enhance the model's representational ability to measure intrinsic SP by refining the estimation of stiffness information and generating adaptive attention values, which provides a new understanding about why and how the SAM can benefit the model performance. This novel perspective can also explain the lottery ticket hypothesis in SAM, design new quantitative metrics of representational ability, and inspire a new theoretic-inspired approach, StepNet. Extensive experiments on several popular benchmarks demonstrate that StepNet can extract fine-grained stiffness information and measure SP accurately, leading to significant improvements in various visual tasks.
Zhongzhan Huang, Mingfu Liang, Jinghui Qin, Shanshan Zhong, Liang Lin 0004
ICCV3
2023 LSAS: Lightweight Sub-attention Strategy for Alleviating Attention Bias Problem
abstract
In computer vision, the performance of deep neural networks (DNNs) is highly related to the feature extraction ability, i.e., the ability to recognize and focus on key pixel regions in an image. However, in this paper, we quantitatively and statistically illustrate that DNNs have a serious attention bias problem on many samples from some popular datasets: (1) Position bias: DNNs fully focus on label-independent regions; (2) Range bias: The focused regions from DNN are not completely contained in the ideal region. Moreover, we find that the existing self-attention modules can alleviate these biases to a certain extent, but the biases are still non-negligible. To further mitigate them, we propose a lightweight sub-attention strategy (LSAS), which utilizes high-order sub-attention modules to improve the original self-attention modules. The effectiveness of LSAS is demonstrated by extensive experiments on widely-used benchmark datasets and popular attention networks. We release our code to help other researchers to reproduce the results of LSAS1
Shanshan Zhong, Wushao Wen, Jinghui Qin, Qiangpu Chen, Zhongzhan Huang
ICME3
2023 SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language Models
abstract
Diffusion models, which have emerged to become popular text-to-image generation models, can produce high-quality and content-rich images guided by textual prompts. However, there are limitations to semantic understanding and commonsense reasoning in existing models when the input prompts are concise narrative, resulting in low-quality image generation. To improve the capacities for narrative prompts, we propose a simple-yet-effective parameter-efficient fine-tuning approach called the Semantic Understanding and Reasoning adapter (SUR-adapter) for pre-trained diffusion models. To reach this goal, we first collect and annotate a new dataset SURD which consists of more than 57,000 semantically corrected multi-modal samples. Each sample contains a simple narrative prompt, a complex keyword-based prompt, and a high-quality image. Then, we align the semantic representation of narrative prompts to the complex prompts and transfer knowledge of large language models (LLMs) to our SUR-adapter via knowledge distillation so that it can acquire the powerful semantic understanding and reasoning capabilities to build a high-quality textual semantic representation for text-to-image generation. We conduct experiments by integrating multiple LLMs and popular pre-trained diffusion models to show the effectiveness of our approach in enabling diffusion models to understand and reason concise natural language without image quality degradation. Our approach can make text-to-image diffusion models easier to use with better user experience, which demonstrates our approach has the potential for further advancing the development of user-friendly text-to-image generation models by bridging the semantic gap between simple narrative prompts and complex keyword-based prompts. The code is released at https://github.com/Qrange-group/SUR-adapter.
Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin, Liang Lin 0004
ACM Multimedia4
2023 SPEM: Self-adaptive Pooling Enhanced Attention Module for Image Recognition
Shanshan Zhong, Wushao Wen, Jinghui Qin
MMM (2)3
2023 RecFormer: Recurrent Multi-modal Transformer with History-Aware Contrastive Learning for Visual Dialog
Liucun Lu, Jinghui Qin, Zequn Jie, Lin Ma 0002, Liang Lin 0004, Xiaodan Liang
PRCV (1)2
2023 ESA: Excitation-Switchable Attention for convolutional neural networks
Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Zhijing Yang, Jinghui Qin
Neurocomputing5
2023 Cross-modal hash retrieval based on semantic multiple similarity learning and interactive projection matrix learning
Junpeng Tan, Zhijing Yang, Jielin Ye, Yongqiang Cheng 0001, Jinghui Qin, Yongfeng Chen
Inf. Sci.6
2023 ADASR: An Adversarial Auto-Augmentation Framework for Hyperspectral and Multispectral Data Fusion
abstract
Deep learning-based hyperspectral image (HSI) super-resolution, which aims to generate high spatial resolution HSI (HR-HSI) by fusing hyperspectral image (HSI) and multispectral image (MSI) with deep neural networks (DNNs), has attracted lots of attention. However, neural networks require large amounts of training data, hindering their application in real-world scenarios. In this letter, we propose a novel adversarial automatic data augmentation framework ADASR that automatically optimizes and augments HSI-MSI sample pairs to enrich data diversity for HSI-MSI fusion. Our framework is sample-aware and optimizes an augmentor network and two downsampling networks jointly by adversarial learning so that we can learn more robust downsampling networks for training the upsampling network. Extensive experiments on two public classical hyperspectral datasets demonstrate the effectiveness of our ADASR compared to the state-of-the-art methods.
Jinghui Qin, Lihuang Fang, Ruitao Lu, Liang Lin 0004, Yukai Shi
IEEE Geosci. Remote. Sens. Lett.1
2023 Real-World Image Super-Resolution by Exclusionary Dual-Learning
abstract
Real-world image super-resolution is a practical image restoration problem that aims to obtain high-quality images from in-the-wild input, has recently received considerable attention with regard to its tremendous application potentials. Although deep learning-based methods have achieved promising restoration quality on real-world image super-resolution datasets, they ignore the relationship between L1- and perceptual- minimization and roughly adopt auxiliary large-scale datasets for pre-training. In this paper, we discuss the image types within a corrupted image and the property of perceptual- and Euclidean- based evaluation protocols. Then we propose a method, Real-World image Super-Resolution by Exclusionary Dual-Learning (RWSR-EDL) to address the feature diversity in perceptual- and L1- based cooperative learning. Moreover, a noise-guidance data collection strategy is developed to address the training time consumption in multiple datasets optimization. When an auxiliary dataset is incorporated, RWSR-EDL achieves promising results and repulses any training time increment by adopting the noise-guidance data collection strategy. Extensive experiments show that RWSR-EDL achieves competitive performance over state-of-the-art methods on four in-the-wild image super-resolution datasets.
Hao Li 0058, Jinghui Qin, Zhijing Yang, Pengxu Wei, Jinshan Pan, Liang Lin 0004, Yukai Shi
IEEE Trans. Multim.2
2022 UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
abstract
Geometry problem solving is a well-recognized testbed for evaluating the high-level multimodal reasoning capability of deep models.In most existing works, two main geometry problems: calculation and proving, are usually treated as two specific tasks, hindering a deep model to unify its reasoning capability on multiple math tasks.However, in essence, these two tasks have similar problem representations and overlapped math knowledge which can improve the understanding and reasoning ability of a deep model on both two tasks.Therefore, we construct a large-scale Unified Geometry problem benchmark, UniGeo, which contains 4,998 calculation problems and 9,543 proving problems.Each proving problem is annotated with a multi-step proof with reasons and mathematical expressions.The proof can be easily reformulated as a proving sequence that shares the same formats with the annotated program sequence for calculation problems.Naturally, we also present a unified multitask Geometric Transformer framework, Geoformer, to tackle calculation and proving problems simultaneously in the form of sequence generation, which finally shows the reasoning ability can be improved on both two tasks by unifying formulation.Furthermore, we propose a Mathematical Expression Pretraining (MEP) method that aims to predict the mathematical expressions in the problem solution, thus improving the Geoformer model.Experiments on the UniGeo demonstrate that our proposed Geoformer obtains state-of-the-art performance by outperforming task-specific model NGS with over 5.6% and 3.2% accuracies on calculation and proving problems, respectively.1
Jinghui Qin, Pan Lu, Liang Lin 0004, Chongyu Chen, Xiaodan Liang
EMNLP3
2022 CEM: Machine-Human Chatting Handoff via Causal-Enhance Module
abstract
Aiming to ensure chatbot quality by predicting chatbot failure and enabling human-agent collaboration, Machine-Human Chatting Handoff (MHCH) has attracted lots of attention from both industry and academia in recent years.However, most existing methods mainly focus on the dialogue context or assist with global satisfaction prediction based on multi-task learning, which ignore the grounded relationships among the causal variables, like the user state and labor cost.These variables are significantly associated with handoff decisions, resulting in prediction bias and cost increase.Therefore, we propose Causal-Enhance Module (CEM) by establishing the causal graph of MHCH based on these two variables, which is a simple yet effective module and can be easy to plug into the existing MHCH methods.For the impact of users, we use the user state to correct the prediction bias according to the causal relationship of multi-task.For the labor cost, we train an auxiliary cost simulator to calculate unbiased labor cost through counterfactual learning so that a model becomes cost-aware.Extensive experiments conducted on four real-world benchmarks demonstrate the effectiveness of CEM in generally improving the performance of existing MHCH methods without any elaborated model crafting.
Shanshan Zhong, Jinghui Qin, Zhongzhan Huang, Daifeng Li
EMNLP2
2022 Lightweight single image super-resolution with attentive residual refinement network
Jinghui Qin, Rumin Zhang
Neurocomputing1
2021 Neural-Symbolic Solver for Math Word Problems with Auxiliary Tasks
abstract
Jinghui Qin, Xiaodan Liang, Yining Hong, Jianheng Tang, Liang Lin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jinghui Qin, Xiaodan Liang, Yining Hong, Liang Lin 0004
ACL/IJCNLP (1)1
2021 A Multiblockchain-Oriented Decentralized Market Framework for Frequency Regulation Service
abstract
As an auxiliary to the electricity market, the traditional frequency regulation market (FRM) lacks reliable data storage ways and direct transaction channels for buyers and sellers, and the cost-sharing method of frequency regulation (FR) service is still unfair. To this end, this article proposes a general and decentralized FRM framework based on multiblockchain techniques. According to the scheduling sequence of transactions, the FR services are divided mainly into before-the-fact transactions (BFTs) and after-the-fact transactions (AFTs). We construct a combinatorial double-auction model for BFTs and propose a novel cost-sharing model based on triggers for AFTs. Taking Pennsylvania-New Jersey-Maryland as an example, we design a consensus algorithm for the on-chain transaction process and cost sharing. Numerical results on both types of transactions show that the proposed auction model can meet the different needs of FR buyers. The proposed sharing method highlights that triggers from frequency events should result in higher costs. The consensus algorithm also improves the fault tolerance and throughput of transactions.
Qian Wang 0047, Zhao Luo, Kezhen Liu, Tao Ding 0001, Xi Mo, Jinghui Qin, Leidan Chen
IEEE Trans. Ind. Informatics7
2020 Dynamic Knowledge Routing Network for Target-Guided Open-Domain Conversation
abstract
Target-guided open-domain conversation aims to proactively and naturally guide a dialogue agent or human to achieve specific goals, topics or keywords during open-ended conversations. Existing methods mainly rely on single-turn data-driven learning and simple target-guided strategy without considering semantic or factual knowledge relations among candidate topics/keywords. This results in poor transition smoothness and low success rate. In this work, we adopt a structured approach that controls the intended content of system responses by introducing coarse-grained keywords, attains smooth conversation transition through turn-level supervised learning and knowledge relations between candidate keywords, and drives an conversation towards an specified target with discourse-level guiding strategy. Specially, we propose a novel dynamic knowledge routing network (DRKN) which considers semantic knowledge relations among candidate keywords for accurate next topic prediction of next discourse. With the help of more accurate keyword prediction, our keyword-augmented response retrieval module can achieve better retrieval performance and more meaningful conversations. Besides, we also propose a novel dual discourse-level target-guided strategy to guide conversations to reach their goals smoothly with higher success rate. Furthermore, to push the research boundary of target-guided open-domain conversation to match real-world scenarios better, we introduce a new large-scale Chinese target-guided open-domain conversation dataset (more than 900K conversations) crawled from Sina Weibo. Quantitative and human evaluations show our method can produce meaningful and effective target-guided conversations, significantly improving over other state-of-the-art methods by more than 20% in success rate and more than 0.6 in average smoothness score.
Jinghui Qin, Xiaodan Liang
AAAI1
2020 GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems
abstract
Automatically evaluating dialogue coherence is a challenging but high-demand ability for developing high-quality open-domain dialogue systems.However, current evaluation metrics consider only surface features or utterancelevel semantics, without explicitly considering the fine-grained topic transition dynamics of dialogue flows.Here, we first consider that the graph structure constituted with topics in a dialogue can accurately depict the underlying communication logic, which is a more natural way to produce persuasive metrics.Capitalized on the topic-level dialogue graph, we propose a new evaluation metric GRADE, which stands for Graph-enhanced Representations for Automatic Dialogue Evaluation.Specifically, GRADE incorporates both coarsegrained utterance-level contextualized representations and fine-grained topic-level graph representations to evaluate dialogue coherence.The graph representations are obtained by reasoning over topic-level dialogue graphs enhanced with the evidence from a commonsense graph, including k-hop neighboring representations and hop-attention weights.Experimental results show that our GRADE significantly outperforms other state-of-the-art metrics on measuring diverse dialogue models in terms of the Pearson and Spearman correlations with human judgements.Besides, we release a new large-scale human evaluation benchmark to facilitate future research on automatic metrics.
Lishan Huang, Jinghui Qin, Liang Lin 0004, Xiaodan Liang
EMNLP (1)3
2020 Semantically-Aligned Universal Tree-Structured Solver for Math Word Problems
abstract
A practical automatic textual math word problems (MWPs) solver should be able to solve various textual MWPs while most existing works only focused on one-unknown linear MWPs.Herein, we propose a simple but efficient method called Universal Expression Tree (UET) to make the first attempt to represent the equations of various MWPs uniformly.Then a semantically-aligned universal tree-structured solver (SAU-Solver) based on an encoder-decoder framework is proposed to resolve multiple types of MWPs in a unified model, benefiting from our UET representation.Our SAU-Solver generates a universal expression tree explicitly by deciding which symbol to generate according to the generated symbols' semantic meanings like human solving MWPs.Besides, our SAU-Solver also includes a novel subtree-level semanticallyaligned regularization to further enforce the semantic constraints and rationality of the generated expression tree by aligning with the contextual information.Finally, to validate the universality of our solver and extend the research boundary of MWPs, we introduce a new challenging Hybrid Math Word Problems dataset (HMWP), consisting of three types of MWPs.Experimental results on several MWPs datasets show that our model can solve universal types of MWPs and outperforms several state-of-the-art models 1 .
Jinghui Qin, Lihui Lin, Xiaodan Liang, Rumin Zhang, Liang Lin 0004
EMNLP (1)1
2020 Multi-scale feature fusion residual network for Single Image Super-Resolution
Jinghui Qin, Yongjie Huang, Wushao Wen
Neurocomputing1
2019 Difficulty-Aware Image Super Resolution via Deep Adaptive Dual-Network
abstract
Recently, deep learning based single image super-resolution(SR) approaches have achieved great development. The state-of-the-art SR methods usually adopt a feed-forward pipeline to establish a non-linear mapping between low-res(LR) and high-res(HR) images. However, due to treating all image regions equally without considering the difficulty diversity, these approaches meet an upper bound for optimization. To address this issue, we propose a novel SR approach that discriminately processes each image region within an image by its difficulty. Specifically, we propose a dual-way SR network that one way is trained to focus on easy image regions and another is trained to handle hard image regions. To identify whether a region is easy or hard, we propose a novel image difficulty recognition network based on PSNR prior. Our SR approach that uses the region mask to adaptively enforce the dual-way SR network yields superior results. Extensive experiments on several standard benchmarks (e.g., Set5, Set14, BSD100, and Urban100) show that our approach achieves state-of-the-art performance.
Jinghui Qin, Ziwei Xie, Yukai Shi, Wushao Wen
ICME1
2018 Energy-efficient Offloading Policy for Resource Allocation in Distributed Mobile Edge Computing
abstract
Mobile edge computing (MEC) is a promising paradigm to integrate computing and communication resources in mobile networks. MEC can improve mobile service quality and enhance Quality of Experience (QoE) by offloading computation tasks to MEC servers. However, a MEC server only can provide limited computational resources for users. In this paper, we consider a mobile edge computing system that provides three offloading policies that are: (i) executing tasks in local device, (ii) offloading tasks to servers in a local region, (iii)offloading tasks to servers in a nearby region. In the policy (iii), mobile user equipment can utilize computational resources of MEC servers in nearby regions to solve the problem of insufficient computational resources in local region servers. We formulate the computation offloading problem as a potential game and propose a Distributed Offloading strategy based on Jacobi algorithm (DOJ) for solving the computation offloading problem in a short period. The simulation results show that our proposed algorithm can reduce overall system costs and guarantee the QoE of users.
Chongwu Dong, Jinghui Qin, Xiaoxing Yang, Wushao Wen
ISCC3