Xin Lin 0001

dblp:50/3323-1 · also Xin Alex Lin · DBLP profile ↗
← Back
56ranked-venue papers
12as first author
30since 2021 · last 2026
0009-0008-4110-8989ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 25 · 11 first-author · 7 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Multi-Modal 3D Object Detection in Autonomous Driving: A New Survey
abstract
Recent progress in autonomous driving has underscored the critical importance of multi-modal 3D object detection, which addresses the limitations of single-sensor methods such as cameras or LiDAR. This survey systematically reviews state-of-the-art multi-modal 3D detection techniques for autonomous vehicles, outlining fundamental principles and domain challenges. It categorizes existing approaches through distinct classification schemes and evaluates detection performance and safety across object types. Unlike previous surveys, this work focuses on challenges and solutions in pedestrian detection, long-range detection, road camera integration, and collaborative perception. The robustness of these methods under varying weather and terrain conditions is also analyzed. Using multiple datasets and experiments, the survey compares classification methods and concludes with promising research directions to guide future developments. This work aims to be a definitive reference for advancing 3D object detection in autonomous driving.
Xin Lin 0001
ICMR2
2026 Code LLMs Still Fall Short of Top Programmers: Evaluating Algorithmic Code Generation Through Computational Thinking
abstract
Evaluating the coding capabilities of models through algorithmic code generation is challenging, as it requires deep problem understanding and complex algorithm design. Current benchmarks suffer from a narrow focus on final execution results (such as pass@k), neglecting the crucial reasoning and problem-solving processes inherent in code generation. To address this limitation, we introduce a multi-phase algorithmic code generation benchmark, MUPA, structured around human computational thinking. MUPA dissects the evaluation into four distinct phases: example understanding, algorithm selection, solution description, and code generation. This framework facilitates a comprehensive assessment by providing insights into the model's intermediate problem-solving steps, rather than just the final code. We manually curated 197 high-quality competitive programming problems from Codeforces. Utilizing an LLM-as-a-judge paradigm with specialized prompts, our rigorous evaluation of several existing code generation LLMs reveals significant across-the-board challenges. Notably, we establish a positive correlation, indicating that proficiency in an earlier phase directly impacts performance in subsequent phases, underscoring the interdependency of these algorithmic skills. The benchmark is publicly available at https://github.com/cheniison/MUPA.
Shisong Chen, Ziyu Zhou 0019, Zhixu Li, Yanghua Xiao, Xin Lin 0001, Xiaojun Meng, Jiansheng Wei, Kuien Liu
WSDM7
2026 Large Language Model Judged Self-Training for Named Entity Recognition
abstract
Self-training for Named Entity Recognition (NER) aims at identifying named entities and their types in the text using self-training to fully make use of the limited labeled data and a large amount of unlabeled data. The major challenge in self-training is confirmation bias where incorrect pseudo-labels increase errors. Many efforts have been made to address this challenge, but few labeled data limit their performance. In this paper, we introduce Large Language Model (LLM) into self-training to select high-quality pseudo-labels leveraging its rich knowledge and few-shot learning capability. Specifically, we design a comprehensive prompt to improve the judgment performance of LLM, where the prompt incorporates task rules mined by LLM itself to fully leverage labeled data. In addition, to reduce the impact of LLM's hallucinations, we adopt a collaborative pseudo-label selection based on combined confidence and calibration-guided probability smoothing. Our empirical study conducted on several NER datasets shows that our method outperforms state-of-the-art approaches. The code is available at https://github.com/cheniison/llm-judged-ST.
Shisong Chen, Jiaan Wang, Yanghua Xiao, Zhixu Li, Xin Lin 0001
WSDM6
2025 GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity Recognition
abstract
Large language models (LLMs) demonstrate impressive performance on downstream tasks through in-context learning(ICL). However, there is a significant gap between their performance in Named Entity Recognition (NER) and in fine-tuning methods. We believe this discrepancy is due to inconsistencies in labeling definitions in NER. In addition, recent research indicates that LLMs do not learn the specific input-label mappings from the demonstrations. Therefore, we argue that using examples to implicitly capture the mapping between inputs and labels in in-context learning is not suitable for NER. Instead, it requires explicitly informing the model of the range of entities contained in the labels, such as annotation guidelines. In this paper, we propose GuideNER, which uses LLMs to summarize concise annotation guidelines as contextual information in ICL. We have conducted experiments on widely used NER datasets, and the experimental results indicate that our method can consistently and significantly outperform state-of-the-art methods, while using shorter prompts. Especially on the GENIA dataset, our model outperforms the previous state-of-the-art model by 12.63 F1 scores.
Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001
AAAI5
2025 A Graph Interaction Framework on Relevance for Multimodal Named Entity Recognition with Multiple Images
abstract
Posts containing multiple images have significant research potential in Multimodal Named Entity Recognition nowadays. The previous methods determine whether the images are related to named entities in the text through similarity computation, such as using CLIP. However, it is not effective in some cases and not conducive to task transfer, especially in multi-image scenarios. To address the issue, we propose a graph interaction framework on relevance (GIFR) for Multimodal Named Entity Recognition with multiple images. For humans, they have the abilities to distinguish whether an image is relevant to named entities, but human capabilities are difficult to model. Therefore, we propose using reinforcement learning based on human preference to integrate human abilities into the model to determine whether an image-text pair is relevant, which is referred to as relevance. To better leverage relevance, we construct a heterogeneous graph and introduce graph transformer to enable information interaction. Experiments on benchmark datasets demonstrate that our method achieves the state-of-the-art performance.
Shizhou Huang, Xin Lin 0001
COLING3
2025 Retrieval-Based Multimodal Data Augmentation for Multimodal Information Extraction in Social Media
Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001
DASFAA (4)5
2025 An Interactive Evaluation Framework for Empathetic Response Generation
abstract
Empathetic response generation is a significant domain in Natural Language Processing (NLP). Its development is a critical step toward achieving humanized AI systems. However, current evaluations of empathetic dialogue models are primarily single-turn and static, leading to bias between evaluation results and real-world multi-turn interaction performance. To overcome the longstanding challenge, we propose a novel Interactive Empathy Evaluation Framework (IEEF). It eliminates the bias by facilitating a human-free multi-turn interaction evaluation. Specifically, for human-free interaction, we design a user simulator using reinforcement learning, leveraging a reward model based on LLM scoring. For evalution, we introduce a series of empathy-related metrics based on LLM. The experiments show that IEEF’s evaluation results are highly correlated with real-world multi-turn interaction performance, demonstrating its alignment with human preferences in empathy evaluation.
Xixi Lei, Changqun Li, Liang He 0001, Xin Lin 0001
ICASSP4
2025 Low-Redundancy Knowledge Generation and Modality-Aware Interaction for Multimodal Information Extraction in Social Media
abstract
Multimodal information extraction (MIE) has gained increasing attention, as it helps to accomplish information extraction by adding images as auxiliary information. By acquiring entity-related knowledge, knowledge generation methods can effectively enhance the performance of information extraction models. However, current knowledge generation methods have two weaknesses: (1) they often generate knowledge that includes task-irrelevant information causing redundancy and negatively impacting model performance; (2) they typically concatenate knowledge and text input directly together, ignoring the stylistic and contextual differences arising from their different sources. To address these issues, we propose Low-Redundancy Knowledge Generation and Modality-Aware Interaction (LRKG-MAI). Our approach leverages a large language model to generate task-relevant knowledge with minimal redundancy, while treating knowledge as a distinct modality that interacts with text within its own representation space. Extensive experiments demonstrate the effectiveness of our approach. The source code can be found at https://github.com/JinFish/LRKG-MAI.
Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001
ICME5
2025 Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
abstract
Automatic math correction aims to check students’ solutions to mathematical problems via artificial intelligence technologies. Most existing studies focus on judging the final answer at the problem level, while they ignore detailed feedback on each step in a math problem-solving process, which requires abilities of semantic understanding and reasoning. In this paper, we propose a reinforcement learning (RL)-based method to boost large language model (LLM) for step-level automatic math correction, named StepAMC. Particularly, we convert the step-level automatic math correction within the text classification task into an RL problem to enhance the reasoning capabilities of LLMs. Then, we design a space-constrained policy network to improve the stability of RL. Then, we introduce a fine-grained reward network to convert the binary human feedback into a continuous value. We conduct extensive experiments over two benchmark datasets and the results show that our model outperforms the eleven strong baselines.
Junsong Li, Jie Zhou 0015, Yutao Yang, Bihao Zhan, Qianjun Pan, Yuyang Ding, Qin Chen 0001, Jiang Bo, Xin Lin 0001, Liang He 0001
ICME9
2025 The Social Cognition Ability Evaluation of LLMs: A Dynamic Gamified Assessment and Hierarchical Social Learning Measurement Approach
abstract
Large Language Model (LLM) has shown amazing abilities in reasoning tasks, theory of mind (ToM) has been tested in many studies as part of reasoning tasks, and social learning, which is closely related to ToM, is still lack of investigation. However, the test methods and materials make the test results unconvincing. We propose a dynamic gamified assessment (DGA) and hierarchical social learning measurement to test ToM and social learning capacities in LLMs. The test for ToM consists of five parts. First, we extract ToM tasks from ToM experiments and then design game rules to satisfy the ToM task requirement. After that, we design ToM questions to match the game’s rules and use these to generate test materials. Finally, we go through the above steps to test the model. To assess the social learning ability, we introduce a novel set of social rules (three in total). Experiment results demonstrate that, except GPT-4, LLMs performed poorly on the ToM test but showed a certain level of social learning ability in social learning measurement.
Yangze Yu, Xin Lin 0001, Ciping Deng, Tingjiang Wei, Mo Xuan
ACM Trans. Intell. Syst. Technol.4
2024 BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind
abstract
As a foundational component of cognitive intelligence, theory of mind (ToM) can make AI more closely resemble human thought processes, thereby enhancing their interaction and collaboration with human. In particular, it can significantly improve a model's comprehension of videos in complex scenes. However, current video question answer (VideoQA) datasets focus on studying causal reasoning within events, few of them genuinely incorporating human ToM. Consequently, there is a lack of development in ToM reasoning tasks within the area of VideoQA. This paper presents BDIQA, the first benchmark to explore the cognitive reasoning capabilities of VideoQA models in the context of ToM. BDIQA is inspired by the cognitive development of children's ToM and addresses the current deficiencies in machine ToM within datasets and tasks. Specifically, it offers tasks at two difficulty levels, assessing Belief, Desire and Intention (BDI) reasoning in both simple and complex scenarios. We conduct evaluations on several mainstream methods of VideoQA and diagnose their capabilities with zero-shot, few-shot and supervised learning. We find that the performance of pre-trained models on cognitive reasoning tasks remains unsatisfactory. To counter this challenge, we undertake thorough analysis and experimentation, ultimately presenting two guidelines to enhance cognitive reasoning derived from ablation analysis.
Xin Lin 0001, Liang He 0001
AAAI2
2024 MNER-MI: A Multi-image Dataset for Multimodal Named Entity Recognition in Social Media
abstract
Recently, multimodal named entity recognition (MNER) has emerged as a vital research area within named entity recognition. However, current MNER datasets and methods are predominantly based on text and a single accompanying image, leaving a significant research gap in MNER scenarios involving multiple images. To address the critical research gap and enhance the scope of MNER for real-world applications, we propose a novel human-annotated MNER dataset with multiple images called MNER-MI. Additionally, we construct a dataset named MNER-MI-Plus, derived from MNER-MI, to ensure its generality and applicability. Based on these datasets, we establish a comprehensive set of strong and representative baselines and we further propose a simple temporal prompt model with multiple images to address the new challenges in multi-image scenarios. We have conducted extensive experiments to demonstrate that considering multiple images provides a significant improvement over a single image and can offer substantial benefits for MNER. Furthermore, our proposed method achieves state-of-the-art results on both MNER-MI and MNER-MI-Plus, demonstrating its effectiveness. The datasets and source code can be found at https://github.com/JinFish/MNER-MI.
Shizhou Huang, Bo Xu 0023, Changqun Li, Jiabo Ye, Xin Lin 0001
LREC/COLING5
2024 VG-Annotator: Vision-Language Models as Query Annotators for Unsupervised Visual Grounding
abstract
Visual grounding focuses on localizing objects referred to by natural language queries. Existing fully and weakly supervised methods rely on a mass of language queries for training. However, collecting natural language queries corresponding to specific objects by annotators is expensive. To reduce the reliance on human-written queries, we propose a novel unsupervised visual grounding framework named VG-Annotator. Different from the existing unsupervised methods that rely on manually designed rules to link objects and language queries. The key idea of VG-Annotator lies in that vision-language pre-trained (VLP) generation models can be language query annotators. Thanks to the powerful multi-modal understanding ability implicitly learned from large-scale pre-training, we consider stimulating models to explicitly generate appropriate descriptions for specific objects in natural language. To this end, we explore a series of multi-modal instructions to indicate which object should be described. We also introduce a supervised fine-tuning process to teach the vision-language models to follow the instructions. Extensive experiments show that the proposed method obtains high-quality language queries. The visual grounding model trained with the generated queries outperforms state-of-the-art unsupervised methods on five widely used datasets.
Jiabo Ye, Xiaoshan Yang, Zhenru Zhang, Anwen Hu, Ming Yan 0008, Ji Zhang 0011, Liang He 0001, Xin Lin 0001
ICME9
2024 A Sentimental Prompt Framework with Visual Text Encoder for Multimodal Sentiment Analysis
abstract
Recently, multimodal sentiment analysis from social media posts has received increasing attention, as it can effectively improve single-modality-based sentiment analysis by leveraging the complementary information between text and images. Despite their success, current methods still suffer from two weaknesses: (1) the current methods for obtaining image representations do not obtain sentiment information, which leads to a significant gap between image representations and results; (2) the current methods ignore the sentiments expressed by the symbols (emoticons, emojis) in the text, but these symbols can effectively reflect the user's sentiments. To address these issues, we propose a sentimental prompt framework with visual text encoder (SPFVTE). Specifically, for the first problem, instead of using the image representation directly, we project the image representation as a prompt and utilize the prompt learning to capture sentimental information in images by learning a sentiment-specific prompt. For the second problem, considering that people get the meanings of emojis and emoticons from their graphics, we propose to render the text as an image and use a visual text encoder to capture the sentiments contained in emojis and emoticons. We have conducted experiments on three public multimodal sentiment datasets, and the experimental results show that our method can significantly and consistently outperform the state-of-the-art methods. The datasets and source code can be found at https://github.com/JinFish/SPFVTE.
Shizhou Huang, Bo Xu 0023, Changqun Li, Jiabo Ye, Xin Lin 0001
ICMR5
2024 A Review on Machine Theory of Mind
abstract
Theory of Mind (ToM) is the ability to attribute mental states to others, an important component of human cognition. At present, there has been growing interest in the artificial intelligence (AI) with cognitive abilities, for example in healthcare and the motoring industry. Research indicates that infants exhibit early signs in cognitive and social understanding, including some basic abilities related to beliefs, desires, and intentions (BDIs). Thus, the ability to attribute BDIs to others is also crucial for the development of machine ToM. In this article, we review recent progress in machine ToM on BDIs. And we shall introduce the experiments, datasets, and methods of machine ToM on these three aspects, summarize the development of different tasks and datasets in recent years, and compare well-behaved models in aspects of advantages, limitations, and applicable conditions, hoping that this study can guide researchers to quickly keep up with latest trend in this field. Unlike other domains with a specific task and resolution framework, machine ToM lacks a unified instruction and a series of standard evaluation tasks, which make it difficult to formally compare the proposed models. And the existing models still cannot exhibit the same ToM reasoning ability as real humans, lack of transferability, interpretability, few-shot learning, etc. We argue that, one method to address this difficulty is now to present a standard assessment criteria and dataset, better a large-scale dataset covered multiple aspects of ToM. Besides, for developing an AI of ToM, it requires the cooperation of experts from various domains.
Xin Lin 0001, Liang He 0001
IEEE Trans. Comput. Soc. Syst.4
2024 UniQRNet: Unifying Referring Expression Grounding and Segmentation with QRNet
abstract
Referring expression comprehension aims to align natural language queries with visual scenes, which requires establishing fine-grained correspondence between vision and language. This has important applications in multi-modal reasoning systems. Existing methods typically use text-agnostic visual backbones to extract features independently without considering the specific text input. However, we argue that the extracted visual features can be inconsistent with the referring expression, which hurts multi-modal understanding. To address this, we first propose Query-modulated Refinement Network (QRNet) that leverages language guidance to guide visual feature extraction. However, it only focuses on the grounding task that can only provide coarse-grained annotations in the form of bounding box coordinates. The guidance for the visual backbone is indirect, and the inconsistent issue still exists. To this end, we further propose UniQRNet, a multi-task framework over the QRNet to learn referring expression grounding and segmentation jointly. The framework introduces a multi-task head that leverages fine-grained pixel-level supervision from the segmentation task to directly guide the intermediate layers of QRNet to learn text-consistent visual features. Besides, UniQRNet also includes a loss balance strategy that allows two types of supervision signals to cooperate and optimize the model together. We conduct the most comprehensive comparison experiment covering four major datasets, ten evaluation set and three evaluation metrics used in previous work. UniQRNet outperforms previous state-of-the-art methods by a large margin on both referring comprehensive grounding (1.8%~5.09%) and segmentation tasks (0.57%~5.56%). Ablation and analysis reveal that UniQRNet can improve the consistency of visual features with text input and can bring significant performance improvement.
Jiabo Ye, Ming Yan 0008, Haiyang Xu 0001, Qinghao Ye, Yaya Shi, Xiaoshan Yang, Xuwu Wang, Ji Zhang 0011, Liang He 0001, Xin Lin 0001
ACM Trans. Multim. Comput. Commun. Appl.11
2023 A Disentangled-Attention Based Framework with Persona-Aware Prompt Learning for Dialogue Generation
abstract
Endowing dialogue agents with personas is the key to delivering more human-like conversations. However, existing persona-grounded dialogue systems still lack informative details of human conversations and tend to reply with inconsistent and generic responses. One of the main underlying causes is that pre-defined persona sentences are generally short and merely superficial descriptions of personal attributes, making appropriate persona selection and understanding non-trivial. Another challenge is that it is crucial to consider the context and the conversation flow to dynamically determine when to invoke different types of persona signals. To address these problems, we propose a disentangled-attention based pre-training architecture, which incorporates persona-aware prompt learning to bridge the connection between the selected persona and response generation. Our model first exploits the conversation flow to select context-relevant personas, and subsequently enriches the superficial persona descriptions with extra personality traits through persona-aware prompting. Finally, the decoder leverages a disentangled-attention mechanism to flexibly control the reliance on personas and dialogue contexts, and incorporates A*-like keyword-based heuristic estimates for controllable generation. Extensive experiments show that our approach can outperform strong baselines and deliver more consistent and engaging responses on the PERSONA-CHAT dataset.
Pingsheng Liu, Zhengjie Huang, Xiechi Zhang, Gerard de Melo, Xin Lin 0001, Liang Pang 0001, Liang He 0001
AAAI6
2023 A Unified Visual Prompt Tuning Framework with Mixture-of-Experts for Multimodal Information Extraction
Bo Xu 0023, Shizhou Huang, Ming Du 0002, Hongya Wang, Yanghua Xiao, Xin Lin 0001
DASFAA (3)7
2023 Pseudo-Query Generation For Semi-Supervised Visual Grounding With Knowledge Distillation
abstract
Visual grounding is a crucial multi-modal job for locating the objects that the referring queries refer to in images. In recent years, both fully-supervised and weakly-supervised algorithms rely on a large number of query annotations. However, collecting queries in natural language is labor-intensive, which limits the application scenarios of these methods. To overcome this weakness, we propose a novel semi-supervised visual grounding framework. The framework consists of two effective techniques: a prompt enhanced pseudo-query generator utilizing objects without query annotation to produce high-quality pseudo-queries; a knowledge distillation mechanism using a teacher network to stabilize the training process of the student network. Experiment results show that our proposed framework dramatically outperforms the existing methods under the three different label proportions on the three commonly used visual grounding datasets.
Jianglin Jin, Jiabo Ye, Xin Lin 0001, Liang He 0001
ICASSP3
2023 HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization
Zefeng Cai, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Xin Lin 0001, Liang He 0001, Daxin Jiang
ICLR6
2023 Image Alone Are Not Enough: A General Semantic-Augmented Transformer-Based Framework for Image Captioning
abstract
Image captioning has long been widely regarded as a modal transformation task from visual to linguistic modality. Most current research focuses on the information transformation between single modalities dominated by visual features, while less attention is paid to the interaction between visual features and linguistic features. This rigid single-modal conversion method is prone to information confusion and loss during the conversion process, making it difficult for the model to generate accurate and detailed captions. In this paper, we propose a general Semantic-Augmented Transformer-Based (SAT) framework to facilitate smoother transformation between the two modalities. In the encoding stage, we use the fine-grained description of each region to fuse with the corresponding image features to make the image feature representation closer to the text feature representation. In the decoding stage, the caption's part-of-speech information is used as prior knowledge to constrain the model to pay more attention to the details in the image rather than only to the prominent entities for fine-grained captions. We extensively evaluate our framework on various state-of-the-art transformer-based models. Experiments demonstrate that these models have superior performance on the MS-COCO dataset under our framework.
Xin Lin 0001, Liang He 0001
IJCNN2
2022 Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual Grounding
abstract
Visual grounding focuses on establishing fine-grained alignment between vision and natural language, which has essential applications in multimodal reasoning systems. Existing methods use pre-trained query-agnostic visual backbones to extract visual feature maps independently without considering the query information. We argue that the visual features extracted from the visual backbones and the features really needed for multimodal reasoning are inconsistent. One reason is that there are differences between pre-training tasks and visual grounding. Moreover, since the backbones are query-agnostic, it is difficult to completely avoid the inconsistency issue by training the visual backbone end-to-end in the visual grounding framework. In this paper, we propose a Query-modulated Refinement Network (QRNet) to address the inconsistent issue by adjusting intermediate features in the visual backbone with a novel Query-aware Dynamic Attention (QD-ATT) mechanism and query-aware multiscale fusion. The QD-ATT can dynamically compute query-dependent visual attention at the spatial and channel levels of the feature maps produced by the visual backbone. We apply the QRNet to an end-to-end visual grounding framework. Extensive experiments show that the proposed method outperforms state-of-the-art methods on five widely used datasets. Our code is available at https://github.com/LukeForeverYoung/QRNet.
Jiabo Ye, Ming Yan 0008, Xiaoshan Yang, Xuwu Wang, Ji Zhang 0011, Liang He 0001, Xin Lin 0001
CVPR8
2022 Curriculum Prompt Learning with Self-Training for Abstractive Dialogue Summarization
abstract
Succinctly summarizing dialogue is a task of growing interest, but inherent challenges, such as insufficient training data and low information density impede our ability to train abstractive models.In this work, we propose a novel curriculum-based prompt learning method with self-training to address these problems.Specifically, prompts are learned using a curriculum learning strategy that gradually increases the degree of prompt perturbation, thereby improving the dialogue understanding and modeling capabilities of our model.Unlabeled dialogue is incorporated by means of self-training so as to reduce the dependency on labeled data.We further investigate topic-aware prompts to better plan for the generation of summaries.Experiments confirm that our model substantially outperforms strong baselines and achieves new state-of-the-art results on the AMI and ICSI datasets.Human evaluations also show the superiority of our model with regard to the summary generation quality.
Changqun Li, Xin Lin 0001, Gerard de Melo, Liang He 0001
EMNLP3
2022 Inferring substitutable and complementary products with Knowledge-Aware Path Reasoning based on dynamic policy network
Zijing Yang, Jiabo Ye, Xin Lin 0001, Liang He 0001
Knowl. Based Syst.4
2021 Looking Wider for Better Adaptive Representation in Few-Shot Learning
abstract
Building a good feature space is essential for the metric-based few-shot algorithms to recognize a novel class with only a few samples. The feature space is often built by Convolutional Neural Networks (CNNs). However, CNNs primarily focus on local information with the limited receptive field, and the global information generated by distant pixels is not well used. Meanwhile, having a global understanding of the current task and focusing on distinct regions of the same sample for different queries are important for the few-shot classification. To tackle these problems, we propose the Cross Non-Local Neural Network (CNL) for capturing the long-range dependency of the samples and the current task. CNL extracts the task-specific and context-aware features dynamically by strengthening the features of the sample at a position via aggregating information from all positions of itself and the current task. To reduce losing important information, we maximize the mutual information between the original and refined features as a constraint. Moreover, we add a task-specific scaling to deal with multi-scale and task-specific features extracted by CNL. We conduct extensive experiments for validating our proposed algorithm, which achieves new state-of-the-art performances on two public benchmarks.
Jiabao Zhao, Yifan Yang 0001, Xin Lin 0001, Jing Yang 0023, Liang He 0001
AAAI3
2021 Cross-Modal Knowledge Distillation For Fine-Grained One-Shot Classification
abstract
Few-shot learning can recognize a novel category based on only a few samples because it learns to learn from a lot of labeled samples during the training process. When data is insufficient, the performance is affected. And it is expensive to obtain a large-scale finegrained dataset with annotation. In this paper, we adopt domain- specific knowledge to fill the gap of insufficient annotated data. We propose a cross-modal knowledge distillation (CMKD) framework to do fine-grained one-shot classification and propose the Spatial Relation Loss (SRL) to transfer cross-modal information, which can tackle the semantic gap between multimodal features. The teacher network distills the spatial relationship of the samples as a soft target for training a unimodal student network. Notably, the student network makes predictions only based on a few samples without any external knowledge in the application. This model-agnostic framework will be well adapted to other few-shot models. Extensive experimental results on benchmarks demonstrate that CMKD can make full use of cross-modal knowledge in image and text few-shot classification. CKMD improves the performances of the student networks significantly, even if it is a state-of-the-art student network.
Jiabao Zhao, Xin Lin 0001, Yifan Yang 0001, Jing Yang 0023, Liang He 0001
ICASSP2
2021 Employer-Employee Network for Conversational Recommendation
abstract
Traditional recommendation systems model user preferences based on past historical behaviors, thus unable to obtain dynamic user preferences. The conversational recommendation system (CRS) combines the conversational module with the recommendation module and overcomes the limitations by directly asking the user's preference for attributes. However, the existing CRS methods lack effective information propagation among various modules, making the model lack the basis for making correct decisions. In this paper, we propose an Employer-Employee network, which decomposes the actions into two stages, which are completed by two networks respectively. The Employer Network is responsible for analyzing information and making decisions (query or recommend), and the Employee Network is responsible for collecting information and performing tasks. Our contributions can be highlighted in three aspects: We first emphasize the importance of information propagation among multiple modules in the conversational recommendation system. Secondly, we propose an Employer-Employee (EE) network, which transforms each turn of action into a two-stage decision-making task handed over to two networks to complete. Thirdly, we conduct experiments on multiple datasets, and the experimental results show that our model achieves competitive performance compared with state-of-the-art baselines.
Zijing Yang, Xin Lin 0001, Liang He 0001, Yixin Chen 0004
IJCNN2
2021 One-Stage Visual Grounding via Semantic-Aware Feature Filter
abstract
Visual grounding has attracted much attention with the popularity of vision language. Existing one-stage methods are far ahead of two-stage methods in speed. However, these methods fuse the textual feature and visual feature map by simply concatenation, which ignores the textual semantics and limits these models' ability in cross-modal understanding. To overcome this weakness, we propose a semantic-aware framework that utilizes both queries' structured knowledge and context-sensitive representations to filter the visual feature maps to localize the referents more accurately. Our framework contains an entity filter, an attribute filter, and a location filter. These three filters filter the input visual feature map step by step according to each query's aspects respectively. A grounding module further regresses the bounding boxes to localize the referential object. Experiments on various commonly used datasets show that our framework achieves a real-time inference speed and outperforms all state-of-the-art methods.
Jiabo Ye, Xin Lin 0001, Liang He 0001, Dingbang Li, Qin Chen 0001
ACM Multimedia2
2021 VColor*: a practical approach for coloring large graphs
Yun Peng 0002, Xin Lin 0001, Byron Choi, Bingsheng He
Frontiers Comput. Sci.2
2021 Entity-level sentiment prediction in Danmaku video interaction
Qingchun Bai, Jie Zhou 0015, Yuanbin Wu, Xin Lin 0001, Liang He 0001
J. Supercomput.6
2020 Optimizing Knowledge Graphs through Voting-based User Feedback
abstract
Knowledge graphs have been used in a wide range of applications to support search, recommendation, and question answering (Q&A). For example, in Q&A systems, given a new question, we may use a knowledge graph to automatically identify the most suitable answers based on similarity evaluation. However, such systems may suffer from two major limitations. First, the knowledge graph constructed based on source data may contain errors. Second, the knowledge graph may become out of date and cannot quickly adapt to new knowledge. To address these issues, in this paper, we propose an interactive framework that refines and optimizes knowledge graphs through user votes. We develop an efficient similarity evaluation notion, called extended inverse P-distance, based on which the graph optimization problem can be formulated as a signomial geometric programming problem. We then propose a basic single-vote solution and a more advanced multi-vote solution for graph optimization. We also propose a split-and-merge optimization strategy to scale up the multi-vote solution. Extensive experiments based on real-life and synthetic graphs demonstrate the effectiveness and efficiency of our proposed framework.
Ruida Yang, Xin Lin 0001, Jianliang Xu, Yan Yang 0008, Liang He 0001
ICDE2
2020 Knowledge-Based Fine-Grained Classification For Few-Shot Learning
abstract
The small inter-class variance and the large intra-class variance make the few-shot and fine-grained image classification more difficult because the machine cannot obtain enough information from only a few images. The external knowledge contains more semantics and can support the model to extract important features, while most of existing few-shot learning algorithms only focus on leveraging the visual features from images, little attention has been paid to the cross-modal external knowledge. In this paper, we propose a knowledge-based fine-grained classification mechanism for few-shot learning, which can overcome the difficulty of only obtaining limited and discriminative features from unimodal samples. We extract the visual features and the knowledge features from textual descriptions and a domain-specific knowledge graph at global and local levels to build the semantic space. To tackle the gap between multimodal features, we propose a mirror framework, named Mirror Mapping Network (MMN), to map the multimodal features into the same semantic space with two directions. Extensive experimental results show that our method outperforms the state-of-the-art.
Jiabao Zhao, Xin Lin 0001, Jie Zhou 0015, Jing Yang 0023, Liang He 0001
ICME2
2020 Answering range-based reverse kNN and why-not reverse kNN queries
Zhefan Zhong, Xin Lin 0001, Liang He 0001
Frontiers Comput. Sci.2
2019 Sentiment Commonsense Induced Sequential Neural Networks for Sentiment Classification
abstract
Although neural networks achieve promising performance in sentence level sentiment classification, most of them are not aware of sentiment commonsense, such as sentiment polarity tags (Positive or Negative) for words, which explicitly determine the sentiment of the sentence in most cases. In this paper, we propose an auxiliary tagging task to integrate sentiment commonsense into sequential neural networks (such as LSTM). We employ the advantage of multitask learning to achieve two goals simultaneously: 1) the sequential learning task accounts for incorporating the semantic information of the surrounding words; 2) the word tagging task ensures the sequential representation still retains the corresponding word tagging information. Besides, considering the most direct way to introduce sentiment information into models as additional knowledge, we further incorporate the additional knowledge enhancing tagging task model to strengthen the effect of sentiment commonsense. We prove the effectiveness of the sentiment commonsense by extensive experiments. The results show that our models exhibit consistent superiority over competitors on three real-word datasets. Specifically, we obtain an accuracy of 55.2%, which is a new state-of-the-art for SST-fine dataset.
Xin Lin 0001, Yanghua Xiao, Liang He 0001
CIKM2
2019 Answering why-not questions on KNN queries
Zhefan Zhong, Xin Lin 0001, Liang He 0001, Jing Yang 0023
Frontiers Comput. Sci.2
2019 Cleaning uncertain graphs via noisy crowdsourcing
Yongcheng Wu, Xin Lin 0001, Yan Yang 0008, Liang He 0001
World Wide Web2
2018 Enabling Uneven Task Difficulty in Micro-Task Crowdsourcing
abstract
In micro-task crowdsourcing markets such as Amazon's Mechanical Turk, how to obtain high quality result without exceeding the limited budgets is one main challenge. The existing theory and practice of crowdsourcing suggests that uneven task difficulty plays a crucial role to task quality. Yet, it lacks a clear identifying method to task difficulty, which hinders effective and efficient execution of micro-task crowdsourcing. This paper explores the notion of task difficulty and its influence to crowdsourcing, and presents a difficulty-based crowdsourcing method to optimize the crowdsourcing process. We firstly identify task difficulty feature based on a local estimation method in the real crowdsourcing context, followed by proposing an optimization method to improve the accuracy of results, while reducing the overall cost. We conduct a series of experimental studies to evaluate our method, which show that our difficulty-based crowdsourcing method can accurately identify the task difficulty feature, improve the quality of task performance and reduce the cost significantly, and thus demonstrate the effectiveness of task difficulty as task modeling property.
Yuling Sun, Jing Yang 0023, Xin Lin 0001, Liang He 0001
GROUP4
2018 An Effective Method for Identifying Unknown Unknowns with Noisy Oracle
Xin Lin 0001, Yanghua Xiao, Jing Yang 0023, Liang He 0001
ICCBR2
2018 Human-Powered Data Cleaning for Probabilistic Reachability Queries on Uncertain Graphs
abstract
In this paper, we consider probabilistic reachability queries on uncertain graphs. To make the results more informative, we adopt a crowdsourcing-based approach to clean the uncertain edges. One important problem is how to efficiently select a limited set of edges for cleaning that maximizes the quality improvement. We prove that the edge selection problem is #P-hard. In light of the hardness of the problem, we propose a series of edge selection algorithms, followed by a number of optimization techniques and pruning heuristics for minimizing the computation time. Our experimental results demonstrate that our proposed techniques outperform a random selection by up to 27 times in terms of the result quality improvement and the brute-force solution by up to 60 times in terms of the elapsed time.
Xin Lin 0001, Yun Peng 0002, Jianliang Xu, Byron Choi
ICDE1
2018 Reducing Uncertainty of Probabilistic Top-k Ranking via Pairwise Crowdsourcing
abstract
In this paper, we propose a novel pairwise crowd-sourcing model to reduce the uncertainty of top-k ranking using a crowd of domain experts. Given a crowdsourcing task of limited budget, we propose efficient algorithms to select the best object pairs for crowdsourcing that will bring in the highest quality improvement. Extensive experiments show that our proposed solutions outperform a random selection method by up to 30 times in terms of quality improvement of probabilistic top-kranking queries. In terms of efficiency, our proposed solutions can reduce the elapsed time of a brute-force algorithm from several days to one minute.
Xin Lin 0001, Jianliang Xu, Haibo Hu 0001, Zhe Fan
ICDE1
2018 iZone: Efficient Influence Zone Evaluation over Geo-Textual Data
abstract
Owing to the widespread use of location-aware devices and the increased popularity of micro-blogging applications, we are witnessing a rapid proliferation of geo-textual data. In this demonstration, we present iZone, an efficient system for determining influence zones over geo-textual data. Specifically, iZone allows users to browse geo-textual objects, evaluate the influence zones of specified geo-textual objects, and obtain explanations of the evaluation results. The iZone system adopts a browser-server model. The server side integrates two types of spatial keyword search, namely top-k spatial keyword query and reverse top-k keyword-based location query, to support the functionality of the system. A variety of spatial indexes are employed to enhance the efficiency of the system. The browser side provides a map-based GUI interface, which enables convenient and user-friendly interaction with the system. Using a real hotel dataset from Hong Kong, iZone offers hands-on experience with influence zone evaluation in real-life applications.
Qing Liu 0008, Zijin Feng, Xike Xie, Jianliang Xu, Xin Lin 0001, Christian S. Jensen
ICDE5
2017 Reducing Unknown Unknowns with Guidance in Image Caption
Mengjun Ni, Jing Yang 0023, Xin Lin 0001, Liang He 0001
ICANN (2)3
2017 Reverse Keyword-Based Location Search
abstract
The proliferation of geo-textual data gives prominence to spatial keyword search. The basic top-k spatial keyword query, returns k geo-textual objects that rank the highest according to their textual relevance and spatial proximity to query keywords and a query location. We define, study, and provide means of computing the reverse top-k keyword-based location query. This new type of query takes a set of keywords, a query object q, and a number k as arguments, and it returns a spatial region such that any top-k spatial keyword query with the query keywords and a location in this region would contain object q in its result. This query targets applications in market analysis, geographical planning, and location optimization, and it may support applications related to safe zones and influence zones that are used widely in location-based services. We show that computing an exact query result requires evaluating and merging a set of weighted Voronoi cells, which is expensive. We therefore devise effective algorithms that approximate result regions with quality guarantees. We develop novel pruning techniques on top of an index, and we offer a series of optimization techniques that aim to further accelerate query processing. Empirical studies suggest that the proposed query processing is efficient and scalable.
Xike Xie, Xin Lin 0001, Jianliang Xu, Christian S. Jensen
ICDE2
2017 Human-Powered Data Cleaning for Probabilistic Reachability Queries on Uncertain Graphs
abstract
Uncertain graph models are widely used in real-world applications such as knowledge graphs and social networks. To capture the uncertainty, each edge in an uncertain graph is associated with an existential probability that signifies the likelihood of the existence of the edge. One notable issue of querying uncertain graphs is that the results are sometimes uninformative because of the edge uncertainty. In this paper, we consider probabilistic reachability queries, which are one of the fundamental classes of graph queries. To make the results more informative, we adopt a crowdsourcing-based approach to clean the uncertain edges. However, considering the time and monetary cost of crowdsourcing, it is a problem to efficiently select a limited set of edges for cleaning that maximizes the quality improvement. We prove that the edge selection problem is #P-hard. In light of the hardness of the problem, we propose a series of edge selection algorithms, followed by a number of optimization techniques and pruning heuristics for reducing the computation time. Our experimental results demonstrate that our proposed techniques outperform a random selection by up to 27 times in terms of the result quality improvement and the brute-force solution by up to 60 times in terms of the elapsed time.
Xin Lin 0001, Yun Peng 0002, Byron Choi, Jianliang Xu
IEEE Trans. Knowl. Data Eng.1
2017 Reducing Uncertainty of Probabilistic Top-k Ranking via Pairwise Crowdsourcing
abstract
Probabilistic top-k ranking is an important and well-studied query operator in uncertain databases. However, the quality of top-k results might be heavily affected by the ambiguity and uncertainty of the underlying data. Uncertainty reduction techniques have been proposed to improve the quality of top-k results by cleaning the original data. Unfortunately, most data cleaning models aim to probe the exact values of the objects individually and therefore do not work well for subjective data types, such as user ratings, which are inherently probabilistic. In this paper, we propose a novel pairwise crowdsourcing model to reduce the uncertainty of top-k ranking using a crowd of domain experts. Given a crowdsourcing task of limited budget, we propose efficient algorithms to select the best object pairs for crowdsourcing that will bring in the highest quality improvement. Extensive experiments show that our proposed solutions outperform a random selection method by up to 30 times in terms of quality improvement of probabilistic top-k ranking queries. In terms of efficiency, our proposed solutions can reduce the elapsed time of a brute-force algorithm from several days to one minute.
Xin Lin 0001, Jianliang Xu, Haibo Hu 0001, Zhe Fan
IEEE Trans. Knowl. Data Eng.1
2016 Answering why-not spatial keyword top-k queries via keyword adaption
abstract
Web objects, often associated with descriptive text documents, are increasingly being geo-tagged. A spatial keyword top-k query retrieves the best k such objects according to a scoring function that considers both spatial distance and textual similarity. However, it is in some cases difficult for users to identify the exact keywords that describe their query intent. After a user issues an initial query and gets back the result, the user may find that some expected objects are missing and may wonder why. Answering the resulting why-not questions can aid users in retrieving better results. However, no existing techniques are able to answer why-not questions by adapting the query keywords. We propose techniques capable of adapting an initial set of query keywords so that expected, but missing, objects enter the result along with other relevant objects. We develop a basic algorithm with a set of optimizations that sequentially examines a sequence of candidate keyword sets. In addition, we present an index-based bound-and-prune algorithm that is able to determine the best sample out of a set of candidates in just one pass of index traversal, thus speeding up the query processing. We also extend the proposed algorithms to handle multiple missing objects. Extensive experimental results offer insight into the efficiency of the proposed techniques in terms of running time and I/O cost.
Lei Chen 0031, Jianliang Xu, Xin Lin 0001, Christian S. Jensen, Haibo Hu 0001
ICDE3
2016 Reverse keyword search for spatio-textual top-k queries in location-based services
abstract
This paper proposes a novel query paradigm, namely reverse keyword search for spatio-textual top-k queries (RST Q). It returns the keywords under which a target object will be a spatio-textual top-k result. To efficiently process the new query, we devise a novel hybrid index KcR-tree to store and summarize the spatial and textual information of objects. To further improve the performance, we propose three query optimization techniques, i.e., KcR*-tree, lazy upper-bound updating, and keyword set filtering. We also extend RST Q to allow the input location to be a spatial region instead of a point. Experimental results demonstrate the efficiency of our proposed query techniques in terms of both the computational cost and I/O cost.
Xin Lin 0001, Jianliang Xu, Haibo Hu 0001
ICDE1
2015 Answering why-not questions on spatial keyword top-k queries
abstract
Large volumes of geo-tagged text objects are available on the web. Spatial keyword top-k queries retrieve k such objects with the best score according to a ranking function that takes into account a query location and query keywords. In this setting, users may wonder why some known object is unexpectedly missing from a result; and understanding why may aid users in retrieving better results. While spatial keyword querying has been studied intensively, no proposals exist for how to offer users explanations of why such expected objects are missing from results. We provide techniques that allow the revision of spatial keyword queries such that their results include one or more desired, but missing objects. In doing so, we adopt a query refinement approach to provide a basic algorithm that reduces the problem to a two-dimensional geometrical problem. To improve performance, we propose an index-based ranking estimation algorithm that prunes candidate results early. Extensive experimental results offer insight into design properties of the proposed techniques and suggest that they are efficient in terms of both running time and I/O cost.
Lei Chen 0031, Xin Lin 0001, Haibo Hu 0001, Christian S. Jensen, Jianliang Xu
ICDE2
2015 Reverse Keyword Search for Spatio-Textual Top-$k$ Queries in Location-Based Services
abstract
Spatio-textual queries retrieve the most similar objects with respect to a given location and a keyword set. Existing studies mainly focus on how to efficiently find the top-k result set given a spatio-textual query. Nevertheless, in many application scenarios, users cannot precisely formulate their keywords and instead prefer to choose them from some candidate keyword sets. Moreover, in information browsing applications, it is useful to highlight the objects with the tags (keywords) under which the objects have high rankings. Driven by these applications, we propose a novel query paradigm, namely reverse keyword search for spatio-textual top-k queries (RSTQ). It returns the keywords under which a target object will be a spatio-textual top-k result. To efficiently process the new query, we devise a novel hybrid index KcR-tree to store and summarize the spatial and textual information of objects. By accessing the high-level nodes of KcR-tree, we can estimate the rankings of the target object without accessing the actual objects. To further improve the performance, we propose three query optimization techniques, i.e., KcR*-tree, lazy upper-bound updating, and keyword set filtering. We also extend RSTQ to allow the input location to be a spatial region instead of a point. Extensive experimental evaluation demonstrates the efficiency of our proposed query techniques in terms of both the computational cost and I/O cost.
Xin Lin 0001, Jianliang Xu, Haibo Hu 0001
IEEE Trans. Knowl. Data Eng.1
2014 Authenticating Location-Based Skyline Queries in Arbitrary Subspaces
abstract
With the ever-increasing use of smartphones and tablet devices, location-based services (LBSs) have experienced explosive growth in the past few years. To scale up services, there has been a rising trend of outsourcing data management to Cloud service providers, which provide query services to clients on behalf of data owners. However, in this data-outsourcing model, the service provider can be untrustworthy or compromised, thereby returning incorrect or incomplete query results to clients, intentionally or not. Therefore, empowering clients to authenticate query results is imperative for outsourced databases. In this paper, we study the authentication problem for location-based arbitrary-subspace skyline queries (LASQs), which represent an important class of LBS applications. We propose a basic Merkle Skyline R-tree method and a novel Partial S4-tree method to authenticate one-shot LASQs. For the authentication of continuous LASQs, we develop a prefetching-based approach that enables clients to compute new LASQ results locally during movement, without frequently contacting the server for query re-evaluation. Experimental results demonstrate the efficiency of our proposed methods and algorithms under various system settings.
Xin Lin 0001, Jianliang Xu, Haibo Hu 0001, Wang-Chien Lee
IEEE Trans. Knowl. Data Eng.1
2013 An Improved Discriminative Category Matching in Relation Identification
Yongliang Sun, Jing Yang 0023, Xin Lin 0001
NLDB3
2013 Range-Based Skyline Queries in Mobile Environments
abstract
Skyline query processing for location-based services, which considers both spatial and nonspatial attributes of the objects being queried, has recently received increasing attention. Existing solutions focus on solving point- or line-based skyline queries, in which the query location is an exact location point or a line segment. However, due to privacy concerns and limited precision of localization devices, the input of a user location is often a spatial range. This paper studies a new problem of how to process such range-based skyline queries. Two novel algorithms are proposed: one is index-based (I-SKY) and the other is not based on any index (N-SKY). To handle frequent movements of the objects being queried, we also propose incremental versions of I-SKY and N-SKY, which avoid recomputing the query index and results from scratch. Additionally, we develop efficient solutions for probabilistic and continuous range-based skyline queries. Experimental results show that our proposed algorithms well outperform the baseline algorithm that adopts the existing line-based skyline solution. Moreover, the incremental versions of I-SKY and N-SKY save substantial computation cost, especially when the objects move frequently.
Xin Lin 0001, Jianliang Xu, Haibo Hu 0001
IEEE Trans. Knowl. Data Eng.1
2012 Continuous Skyline Queries with Integrity Assurance in Outsourced Spatial Databases
Xin Lin 0001, Jianliang Xu, Junzhong Gu
WAIM1
2011 Authentication of location-based skyline queries
abstract
In outsourced spatial databases, the location-based service (LBS) provides query services to the clients on behalf of the data owner. However, if the LBS is not trustworthy, it may return incorrect or incomplete query results. Thus, authentication is needed to verify the soundness and completeness of query results. In this paper, we study the authentication problem for location-based skyline queries, which have recently been receiving increasing attention in LBS applications. We propose two authentication methods: one based on the traditional MR-tree index and the other based on a newly developed MR-Sky-tree. Experimental results demonstrate the efficiency of our proposed methods in terms of the authentication cost.
Xin Lin 0001, Jianliang Xu, Haibo Hu 0001
CIKM1
2005 A RDF-Based Context Filtering System in Pervasive Environment
Xin Lin 0001, Shanping Li, Jian Xu 0001
ICIC (2)1
2005 An Efficient Context Modeling and Reasoning System in Pervasive Environment: Using Absolute and Relative Context Filtering Technology
Xin Lin 0001, Shanping Li, Jian Xu 0001
WAIM1