Lingyun Song

dblp:152/7478 · DBLP profile ↗
← Back
33ranked-venue papers
16as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 11 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction Tuning
abstract
Multimodal Large Language Models (MLLMs) employing the Mixture-of-Experts (MoE) structure exhibit encouraging results in visual language tasks. However, they struggle with catastrophic forgetting due to a lack of effective collaboration among experts and negative transfer across tasks. This happens because the router typically employed in MoE for managing expert assignments is inadequate when there are significant shifts in data distribution across various tasks. A drop in the effectiveness of earlier tasks is caused by negative transfer, which occurs due to conflicts in shared knowledge between tasks, disturbing the knowledge already acquired. To address these issues, we propose the Knowledge Space Synergy Framework in Mixture of Experts (KSS-MoE) for Continual Visual Instruction Tuning (CVIT). It dynamically combines the knowledge subspaces of experts to improve the integration of fine-grained complementary knowledge and collaborative abilities of experts, thus addressing the limitations of the basic router. Furthermore, we introduce a general expert that maintains orthogonal subspaces for shared knowledge, enabling effective cross-task knowledge utilization while reducing negative transfer. Extensive experiments conducted on eight CVIT tasks confirm the excellence of KSS-MoE, showcasing its top-tier performance.
Lingyun Song, Ziyao Chen, Kang Pan, Xiaolin Han 0002, Xinbiao Gan, Yudai Pan, Xiaofan Sun, Xuequn Shang 0001
AAAI1
2026 Robust Spatial-Temporal Similar Trajectory Search via Structure-Enhanced Domain-Invariant Learning
Xiaolin Han 0002, Yonghao Zhou, Chenhao Ma 0001, Lingyun Song, Xinbiao Gan, Xuequn Shang 0001
ICDE4
2025 Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language Models
abstract
Large Language Models (LLMs) may suffer from hallucinations in real-world applications due to the lack of relevant knowledge. In contrast, knowledge graphs encompass extensive, multi-relational structures that store a vast array of symbolic facts. Consequently, integrating LLMs with knowledge graphs has been extensively explored, with Knowledge Graph Question Answering (KGQA) serving as a critical touchstone for the integration. This task requires LLMs to answer natural language questions by retrieving relevant triples from knowledge graphs. However, existing methods face two significant challenges: *excessively long reasoning paths distracting from the answer generation*, and *false-positive relations hindering the path refinement*. In this paper, we propose an iterative interactive KGQA framework that leverages the interactive learning capabilities of LLMs to perform reasoning and Debating over Graphs (DoG). Specifically, DoG employs a subgraph-focusing mechanism, allowing LLMs to perform answer trying after each reasoning step, thereby mitigating the impact of lengthy reasoning paths. On the other hand, DoG utilizes a multi-role debate team to gradually simplify complex questions, reducing the influence of false-positive relations. This debate mechanism ensures the reliability of the reasoning process. Experimental results on five public datasets demonstrate the effectiveness and superiority of our architecture. Notably, DoG outperforms the state-of-the-art method ToG by 23.7% and 9.1% in accuracy on WebQuestions and GrailQA, respectively. Furthermore, the integration experiments with various LLMs on the mentioned datasets highlight the flexibility of DoG.
Jie Ma 0001, Zhitao Gao 0003, Qi Chai, Wangchun Sun, Pinghui Wang, Hongbin Pei, Lingyun Song, Jun Liu 0002
AAAI8
2025 PLMTox: Low-Rank Adaptation on Protein Language Models for Peptide Toxicity Prediction
abstract
Peptides, as biologically active substances between small molecules and proteins, play a crucial role in drug development due to their high specificity. However, peptides may trigger toxic reactions when exerting a therapeutic effect. Recently, deep learning has emerged as a leading paradigm for peptide toxicity prediction. With the development of protein language models (PLMs), their advantages in protein sequence analysis provide a new technological path for peptide toxicity prediction, but their practical application is limited by the high computational resource requirements. To address the above problems, this paper proposes a peptide toxicity prediction model, PLMTox, that combines PLMs and low-rank adaptation (LoRA) technology. We use protein language models to learn evolutionary information and deep semantic features behind peptide sequences. Simultaneously, PLMTox employs low-rank matrix decomposition to retain the understanding ability of the pre-trained model while realizing the efficient task adaptation of parameters, thus reducing the demand for computational resources. For different datasets, PLMTox (AUC$=0.963$,$\text{SP}=0.960$) exhibits good performance and robustness. These experiment results suggest that PLMTox is expected to promote the safety assessment process of peptide drug development.
Lingyun Song, Zilin Ren, Xuequn Shang 0001
BIBM4
2025 RADIO: Effective and Efficient Anomalous Subgraph Discovery in Financial Networks
Xiaolin Han 0002, Chenhao Ma 0001, Lingyun Song, Xuequn Shang 0001
DASFAA (2)4
2025 STREAM: Hierarchical Dynamic Traffic Pattern Inference for Sparse Trajectory Recovery
abstract
Trajectory data are crucial in intelligent transportation management, road network optimization, and urban mobility analysis. Many downstream applications, such as trajectory prediction and travel time estimation, rely on high-resolution trajectory data. However, real-world trajectories are often sparse due to GPS signal loss and power constraints. Existing trajectory recovery methods often struggle to utilize the latent hierarchical traffic conditions, and they often overlook complex movement semantics. To address these limitations, we propose sparse trajectory recovery with hierarchical dynamic traffic pattern inference (STREAM), a unified framework that collectively infers latent global and local traffic conditions from observed trajectories. By modeling these multi-scale dependencies in its encoder, STREAM enables the decoder to accurately reconstruct missing trajectory points. Additionally, our model effectively captures multi-step movement patterns to enhance the accuracy of next-location inference. Extensive experiments on real-world datasets demonstrate that our model outperforms nine existing competitors with an average improvement of 42.52% in trajectory recovery.
Xiaolin Han 0002, Tianwen Zhang, Gaukhar Issayeva, Chenhao Ma 0001, Lingyun Song, Xuequn Shang 0001
ICDM6
2025 AACoT: Chain-of-Thought Fine-Tuning via Associative Memory and Adaptive Error Correction
abstract
Traditional Chain-of-Thought (CoT) approaches in large language models (LLMs) often miss long-range semantic dependencies. As a result, early reasoning errors may cause subsequent cascading failures. To address these issues, we introduce AACoT, a fine-tuning framework that integrates associative memory and adaptive error correction within the CoT reasoning process. The AACoT memory functions in a dualmode capacity that differentiates entity-level knowledge from relation-level knowledge, facilitating dynamic knowledge interaction and efficient retrieval for reasoning. The adaptive error correction mechanism monitors the reasoning process, backtracks upon error detection, and regenerates the corrected reasoning paths. To improve robustness, a prompt refinement module adjusts short-term memory by collecting frequent error patterns to direct future reasoning, and a memory warm-up strategy loads crucial knowledge in advance of inference to minimize dependency on additional training. In the inference process, AACoT produces several reasoning paths and employs weighted voting to determine the final result. Results from experiments conducted on mathematical reasoning benchmarks reveal significant improvements in accuracy, validating that AACoT provides a clear and effective method for enhancing complex reasoning in foundational LLMs.
Ruiyue Wang, Lingyun Song, Xinbiao Gan, Yudai Pan, Xuequn Shang 0001
ICPADS2
2025 Multi-Scale Temporal Neural Network for Stock Trend Prediction Enhanced by Temporal Hyepredge Learning
abstract
Existing research in Stock Trend Prediction (STP) focuses on temporal features extracted from a temporal sequence of stock data with a look-back window, which frequently leads to the omission of important periodic patterns, such as weekly and monthly variations in stock prices. Furthermore, these methods examine stocks individually, ignoring the temporal variation patterns among stocks that share higher-order relationships, like those within the same industry. These relationships typically provide contextual insights into market investments influencing stock price fluctuations. To tackle these issues, we propose a Multi-Scale Temporal Neural Network (MSTNN) framework tailored for STP. This architecture explores the periodic fluctuation behaviors of individual stocks through an innovative 3D convolutional neural network, alongside examining temporal variation patterns of stocks linked to specific industries via a temporal hypergraph attention mechanism. Empirical results from two real-world benchmark datasets show that MSTNN significantly outperforms prior state-of-the-art STP methods. The code of our MSTNN is available at https://github.com/sunlitsong/MSTNN.
Lingyun Song, Siyu Chen 0024, Xinbiao Gan, Binze Shi, Jie Ma 0001, Yudai Pan, Xuequn Shang 0001
IJCAI1
2025 Metapath and Hypergraph Structure-based Multi-Channel Graph Contrastive Learning for Student Performance Prediction
abstract
Considerable attention has been paid to predicting student performance on exercises. The performance of prior studies is determined by the quality of the trait features of students and exercises. Nevertheless, most of the prior study primarily examines simple pairwise interactions in learning trait features, like those between students and exercises or exercises and concepts, while disregarding the complex higher-order interactions that typically exist among these components, which in turn hinders the prediction results. In this paper, we using an innovative Multi-Channel Graph Contrastive Learning (MCGCL) framework that integrates various high-order interactions for predicting student performance. MCGCL characterizes graph structures reflecting various high-order relationships among students, exercises, and concepts through multiple channels, thereby enhancing the trait features of both students and exercises. Moreover, graph contrastive learning is employed to enhance the representation of trait features acquired from high-order graph structures in diverse views. Extensive experiments on real-world datasets show that MCGCL achieves state-of-the-art results on the task of predicting student performance. The code is available at https://github.com/sunlitsong/MCGCL.
Lingyun Song, Xiaofan Sun, Xinbiao Gan, Yudai Pan, Xiaolin Han 0002, Jie Ma 0001, Jun Liu 0002, Xuequn Shang 0001
IJCAI1
2025 TempASD: Temporal Anomalous Subgraph Discovery in Large-Scale Dynamic Financial Networks
abstract
In this paper, we investigate the discovery of temporal anomalous subgraphs in large-scale financial networks, aiming to identify abnormal transaction behaviors among users over time. This task is crucial for the real-time detection of transaction anomalies in financial networks, such as money laundering and trading fraud. However, it poses significant challenges due to the diverse distribution of transactions, the dynamic nature of temporal networks, and the absence of theoretical foundation. To tackle these challenges, we introduce a novel Temporal Anomalous Subgraph Discovery (TempASD) algorithm with theoretical analysis. First, we propose a temporal candidate detection module that quickly pinpoints abnormal candidates by detecting anomalies in both the temporal structure and transaction distribution. Then, we introduce a carefully crafted reinforcement-learning-based refiner to optimize these candidates toward the most abnormal directions. We conducted extensive evaluations against thirteen advanced competitors. TempASD achieves an average improvement of 7x in abnormal degree compared to the state-of-the-art and is efficient in large-scale dynamic financial networks.
Xiaolin Han 0002, Chenhao Ma 0001, Lingyun Song, Reynold Cheng, Xuequn Shang 0001
KDD (2)4
2025 Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs
abstract
Knowledge graph-based retrieval-augmented generation seeks to mitigate hallucinations in Large Language Models (LLMs) caused by insufficient or outdated knowledge. However, existing methods often fail to fully exploit the prior knowledge embedded in knowledge graphs (KGs), particularly their structural information and explicit or implicit constraints. The former can enhance the faithfulness of LLMs' reasoning, while the latter can improve the reliability of response generations. Motivated by these, we propose a trustworthy reasoning framework, termed Deliberation over Priors (\texttt{DP}), which sufficiently utilizes the priors contained in KGs. Specifically, \texttt{DP} adopts a progressive knowledge distillation strategy that integrates structural priors into LLMs through a combination of supervised fine-tuning and Kahneman-Tversky Optimization, thereby improving the faithfulness of relation path generation. Furthermore, our framework employs a reasoning-introspection strategy, which guides LLMs to perform refined reasoning verification based on extracted constraint priors, ensuring the reliability of response generation. Extensive experiments on three benchmark datasets demonstrate that \texttt{DP} achieves new state-of-the-art performance, especially a H@1 improvement of 13% on the ComplexWebQuestions dataset, and generates highly trustworthy responses. We also conduct various analyses to verify its flexibility and practicality. Code is available at [https://github.com/mira-ai-lab/Deliberation-on-Priors](https://github.com/mira-ai-lab/Deliberation-on-Priors).
Jie Ma 0001, Ning Qu, Zhitao Gao 0003, Jun Liu 0002, Hongbin Pei, Jiang Xie 0002, Lingyun Song, Pinghui Wang
NeurIPS8
2025 GraphCom: Communication Hierarchy-aware Graph Engine for Distributed Model Training
abstract
Efficient processing of large-scale graphs with billions to trillions of edges is essential for training graph-based large language models (LLMs) in web-scale systems. The increasing complexity and size of these models create significant communication challenges due to the extensive message exchanges required across distributed nodes. Current graph engines struggle to effectively scale across hundreds of computing nodes because they often overlook variations in communication costs within the interconnection hierarchy. This paper presents GraphCom, a communication-efficient message graph engine for graph processing on supercomputers. Our key idea is to leverage the network topology information to perform communication hierarchy-aware message aggregation, where messages are (i) gathered to the responsible nodes (referred to as monitors) in the source domains, (ii) transferred between monitors, and (iii) scattered to the target nodes in the target domains. GraphCom's aggregation is more aggressive in that each source domain (instead of the source node). We have implemented GraphCom on top of MPI. We demonstrate GraphCom's effectiveness with synthetic benchmarks and real-world graphs, utilizing up to 79,024 nodes and over 1.2 million processor cores, demonstrating that GraphCom surpasses leading graph- parallel systems and state-of-the-art counterparts in both throughput and scalability. Moreover, we have deployed GraphCom on a production supercomputer, where it consistently outperforms the top solutions on the Graph500 list. These results highlight the potential GraphCom has to significantly improve the efficiency of distributed large-scale graph-based LLM training by optimizing communication between distributed systems, making it an invaluable graph engine for distributed training tasks on web-scale graphs.
Xinbiao Gan, Qiang Zhang 0053, Lingyun Song, Bo Yang 0023, Jie Liu 0002, Kai Lu 0001
WWW5
2025 Logic-Aware Knowledge Graph Reasoning for Structural Sparsity under Large Language Model Supervision
abstract
Knowledge Graph (KG) reasoning aims to predict missing entities in incomplete triples, which requires adequate structural information to derive accurate embeddings. However, KGs in the real world are not as dense as the idealized benchmarks, where sparse graph structures restrict the comprehensive structural information for superior performance. Although the logical semantics in KGs shows its potential in alleviating the impact of structural sparsity, there still exist some challenges. The deficient supervision and the semantic gap of logic make it difficult to introduce logical semantics in sparse KG reasoning. To this end, we propose a novel KG reasoning approach LoLLM injecting logic with the supervised information supplied by the Large Language Model (LLM), which is proved to be effective in evaluating and scoring. Firstly, LoLLM derives structural embeddings employing a graph convolutional network (GCN) with relation-aware and triple-aware attention. LoLLM secondly constructs reasoning paths instantiated from the first-order logic rules extracted from sparse KGs, and injects the logical semantics by a designed LLM-enhanced tuning strategy. We propose a textual loss (TL) and a logical loss (LL) in the optimization and obtain logical tuning embeddings of KG in this process. Finally, LoLLM fuses structural embeddings from the GCN and logical tuning embeddings from the LLM-enhanced tuning for scoring and incomplete triple prediction. Extensive experiments on two sparse KGs and a benchmark show that LoLLM outperforms state-of-the-art structure-based and Language Model (LM)-augmented baselines. Moreover, the logic rules with corresponding confidences provide explicit explanations as an interpretable paradigm.
Yudai Pan, Jiajie Hong, Tianzhe Zhao, Lingyun Song, Jun Liu 0002, Xuequn Shang 0001
WWW4
2024 Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
abstract
Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are prone to overlearning dataset biases, resulting in poor robustness. Furthermore, current datasets may not provide a precise diagnostic for these methods. To tackle these challenges, firstly, we propose a novel dataset, *MUSIC-AVQA-R*, crafted in two steps: rephrasing questions within the test split of a public dataset (*MUSIC-AVQA*) and subsequently introducing distribution shifts to split questions. The former leads to a large, diverse test space, while the latter results in a comprehensive robustness evaluation on rare, frequent, and overall questions. Secondly, we propose a robust architecture that utilizes a multifaceted cycle collaborative debiasing strategy to overcome bias learning. Experimental results show that this architecture achieves state-of-the-art performance on MUSIC-AVQA-R, notably obtaining a significant improvement of 9.32\%. Extensive ablation experiments are conducted on the two datasets mentioned to analyze the component effectiveness within the debiasing strategy. Additionally, we highlight the limited robustness of existing multi-modal QA methods through the evaluation on our dataset. We also conduct experiments combining various baselines with our proposed strategy on two datasets to verify its plug-and-play capability. Our dataset and code are available at <https://github.com/reml-group/MUSIC-AVQA-R>.
Jie Ma 0001, Pinghui Wang, Wangchun Sun, Lingyun Song, Hongbin Pei, Jun Liu 0002, Youtian Du
NeurIPS5
2024 Enhancing Enterprise Credit Risk Assessment with Cascaded Multi-level Graph Representation Learning
Lingyun Song, Yacong Tan, Zhanhuai Li, Xuequn Shang 0001
Neural Networks1
2024 A Multi-Group Multi-Stream attribute Attention network for fine-grained zero-shot learning
Lingyun Song, Xuequn Shang 0001, Ruizhi Zhou, Jun Liu 0002, Jie Ma 0001, Zhanhuai Li, Mingxuan Sun 0001
Neural Networks1
2024 FMSA-SC: A Fine-Grained Multimodal Sentiment Analysis Dataset Based on Stock Comment Videos
abstract
Previous Sentiment Analysis (SA) studies have demonstrated that exploring sentiment cues from multiple synchronized modalities can effectively improve the SA results. Unfortunately, until now there is no publicly available dataset for multimodal SA of the stock market. Existing datasets for stock market SA only provide textual stock comments, which usually contain words with ambiguous sentiments or even sarcasm words expressing opposite sentiments of literal meaning. To address this issue, we introduce a Fine-grained Multimodal Sentiment Analysis dataset built upon 1,247 Stock Comment videos, called FMSA-SC. It provides both multimodal sentiment annotations for the videos and unimodal sentiment annotations for the textual, visual, and acoustic modalities of the videos. In addition, FMSA-SC also provides fine-grained annotations that align text at the phrase level with visual and acoustic modalities. Furthermore, we present a new fine-grained multimodal multi-task framework as the baseline for multimodal SA on the FMSA-SC. Data and codes are available athttps://github.com/sunlitsong/FMSA-SC-dataset.git.
Lingyun Song, Siyu Chen 0024, Ziyang Meng 0002, Mingxuan Sun 0001, Xuequn Shang 0001
IEEE Trans. Multim.1
2023 A deep cross-modal neural cognitive diagnosis framework for modeling student performance
Lingyun Song, Xuequn Shang 0001, Jun Liu 0002, Mengzhen Yu, Yu Lu 0003
Expert Syst. Appl.1
2023 Meta-HGT: Metapath-aware HyperGraph Transformer for heterogeneous information network embedding
Lingyun Song, Guangtao Wang, Xuequn Shang 0001
Neural Networks2
2023 Answering knowledge-based visual questions via the exploration of Question Purpose
Lingyun Song, Jianao Li, Jun Liu 0002, Xuequn Shang 0001, Mingxuan Sun 0001
Pattern Recognit.1
2023 Attribute-Guided Multiple Instance Hashing Network for Cross-Modal Zero-Shot Hashing
abstract
Cross-Modal Zero-Shot Hashing (CMZSH) is an important image retrieval technique, e.g., Text Based Image Retrieval. Most of existing CMZSH methods mainly use semantic attributes as guidance to generate hash codes for both the images and texts of seen and unseen categories. However, existing CMZSH methods only focus on learning global attribute vectors and hash codes for images, which mixes up information of complex semantics and background clutters, and thus impedes the retrieval performance. To solve this issue, we propose an Attribute-Guided Multiple Instance Hashing (AG-MIH) network for CMZSH, where each instance represents one image region. Instead of generating global image hash codes, AG-MIH can effectively learn instance-level hash codes based on instance attributes. To improve the attribute learning for instances, AG-MIH can exploia novel 2-D Category-Attribute Relation (CAR) layer, which uses different matching templates to model the relationships between each instance and the attributes for different categories. Under the guidance of semantic attributes, AG-MIH can effectively learn hash codes for each visual instance and texts by a Multi-stream Instance Hashing Refinement (MIHR) procedure. In the MIHR, the pseudo supervisions for the instance-level attributes and hash codes in each stream are from its proceeding stream. Empirical studies on benchmark datasets show that AG-MIH achieves state-of-the-art performance on both cross-modal and single-modal zero-shot image retrieval tasks.
Lingyun Song, Xuequn Shang 0001, Mingxuan Sun 0001
IEEE Trans. Multim.1
2022 MMAN: Metapath Based Multi-Level Graph Attention Networks for Heterogeneous Network Embedding (Student Abstract)
abstract
Current Heterogeneous Network Embedding (HNE) models can be roughly divided into two types, i.e., relation-aware and metapath-aware models. However, they either fail to represent the non-pairwise relations in heterogeneous graph, or only capable of capturing local information around target node. In this paper, we propose a metapath based multilevel graph attention networks (MMAN) to jointly learn node embeddings on two substructures, i.e., metapath based graphs and hypergraphs extracted from original heterogeneous graph. Extensive experiments on three benchmark datasets for node classification and node clustering demonstrate the superiority of MMAN over the state-of-the-art works.
Lingyun Song, Xuequn Shang 0001
AAAI2
2022 Topology Imbalance and Relation Inauthenticity Aware Hierarchical Graph Attention Networks for Fake News Detection
abstract
Fake news detection is a challenging problem due to its tremendous real-world political and social impacts. Recent fake news detection works focus on learning news features from News Propagation Graph (NPG). However, little attention is paid to the issues of both authenticity of the relationships and topology imbalance in the structure of NPG, which trick existing methods and thus lead to incorrect prediction results. To tackle these issues, in this paper, we propose a novel Topology imbalance and Relation inauthenticity aware Hierarchical Graph Attention Networks (TR-HGAN) to identify fake news on social media. Specifically, we design a new topology imbalance smoothing strategy to measure the topology weight of each node. Besides, we adopt a hierarchical-level attention mechanism for graph convolutional learning, which can adaptively identify the authenticity of relationships by assigning appropriate weights to each of them. Experiments on real-world datasets demonstrate that TR-HGAN significantly outperforms state-of-the-art methods.
Lingyun Song, Xuequn Shang 0001
COLING2
2022 A deep grouping fusion neural network for multimedia content understanding
abstract
Abstract How Deep Neural Networks (DNNs) best cope with the understanding of multimedia contents still remains an open problem, mainly due to two factors. First, conventional DNNs cannot effectively learn the representations of the images with sparse visual information. For example, the images describing knowledge concepts in textbooks. Second, existing DNNs cannot effectively capture the fine‐grained interactions between the images and text descriptions. To address these issues, we propose a deep Cross‐Media Grouping Fusion Network (CMGFN), which mainly has two distinctive properties: 1) CMGFN can effectively learn visual features from the images with sparse visual information. This is achieved by first progressively adjusting the attention of convolution filters to valuable visual regions, and then enhancing the use of key visual information in feature construction. 2) By a cross‐media grouping co‐attention mechanism, CMGFN can effectively use the interactions between visual features of different semantics and textual descriptions, to learn cross‐media features representing different fine‐grained semantics in different groups. Empirical studies demonstrate that CMGFN not only achieves state‐of‐the‐art performance on the multimedia documents containing sparse visual information, but also shows superior general applicability on other multimedia data, e.g., the multimedia fake news.
Lingyun Song, Mengzhen Yu, Xuequn Shang 0001, Yu Lu 0003, Jun Liu 0002, Zhanhuai Li
IET Image Process.1
2022 Consensus cubature filtering based on Gaussian process for distributed sensor network with model uncertainty
Lingyun Song, Zhongliang Jing, Peng Dong 0001, Ruping Zou, Shichao Chen
Signal Process.1
2021 Weakly Supervised Group Mask Network for Object Detection
Lingyun Song, Jun Liu 0002, Mingxuan Sun 0001, Xuequn Shang 0001
Int. J. Comput. Vis.1
2019 Connecting Language to Images: A Progressive Attention-Guided Network for Simultaneous Image Captioning and Language Grounding
abstract
Image captioning and visual language grounding are two important tasks for image understanding, but are seldom considered together. In this paper, we propose a Progressive Attention-Guided Network (PAGNet), which simultaneously generates image captions and predicts bounding boxes for caption words. PAGNet mainly has two distinctive properties: i) It can progressively refine the predictive results of image captioning, by updating the attention map with the predicted bounding boxes. ii) It learns bounding boxes of the words using a weakly supervised strategy, which combines the frameworks of Multiple Instance Learning (MIL) and Markov Decision Process (MDP). By using the attention map generated in the captioning process, PAGNet significantly reduces the search space of the MDP. We conduct experiments on benchmark datasets to demonstrate the effectiveness of PAGNet and results show that PAGNet achieves the best performance.
Lingyun Song, Jun Liu 0002, Buyue Qian, Yihe Chen
AAAI1
2019 An Interpretable Fast Model for Predicting The Risk of Heart Failure
abstract
Lately, thanks to the huge amount of Electronic Health Records (EHR) data, deep learning models have been successfully applied to a variety of clinical prediction problems. Existing state-of-the-art clinical predicting models are usually built with recurrent neural network (RNN) and attention mechanism. However, such RNN based approaches mainly suffer from three limitations on clinical predictions, which if addressed would significantly widen their applicability. (i) Accuracy: The performance of RNN based models drops fast when the length of EHR sequences increases. (ii) Interpretability: The prediction results of RNN based models are hard to interpret due to the nature of deep models. (iii) Efficiency: The sequential property of RNN based models makes the parallelization of computation impossible, and accordingly hurts the efficiency of such models in practice. In this paper, we propose an efficient attention-based model to address the above three challenges simultaneously. In the context of heart failure prediction task, we demonstrate the interpretation capability of our model by visualizing relative connections between events and prediction result, and the high computational efficiency comparing to other baseline methods. Meanwhile, we show that the accuracy of our prediction model is comparable or better than those of other state-of-the-art prediction models in healthcare applications.
Xianli Zhang, Buyue Qian, Xiaoyu Li 0007, Jishang Wei, Yingqian Zheng, Lingyun Song
SDM6
2018 A Deep Multi-Modal CNN for Multi-Instance Multi-Label Image Classification
abstract
Deep convolutional neural networks (CNNs) have shown superior performance on the task of single-label image classification. However, the applicability of CNNs to multi-label images still remains an open problem, mainly because of two reasons. First, each image is usually treated as an inseparable entity and represented as one instance, which mixes the visual information corresponding to different labels. Second, the correlations amongst labels are often overlooked. To address these limitations, we propose a deep multi-modal CNN for multi-instance multi-label image classification, called MMCNN-MIML. By combining CNNs with multi-instance multi-label (MIML) learning, our model represents each image as a bag of instances for image classification and inherits the merits of both CNNs and MIML. In particular, MMCNN-MIML has three main appealing properties: 1) it can automatically generate instance representations for MIML by exploiting the architecture of CNNs; 2) it takes advantage of the label correlations by grouping labels in its later layers; and 3) it incorporates the textual context of label groups to generate multi-modal instances, which are effective in discriminating visually similar objects belonging to different groups. Empirical studies on several benchmark multi-label image data sets show that MMCNN-MIML significantly outperforms the state-of-the-art baselines on multi-label image classification tasks.
Lingyun Song, Jun Liu 0002, Buyue Qian, Mingxuan Sun 0001, Samar Abbas
IEEE Trans. Image Process.1
2017 Sparse Relational Topical Coding on multi-modal data
Lingyun Song, Jun Liu 0002, Minnan Luo, Buyue Qian
Pattern Recognit.1
2016 Sparse Multi-Modal Topical Coding for Image Annotation
Lingyun Song, Minnan Luo, Jun Liu 0002, Lingling Zhang 0005, Buyue Qian, Haifei Li 0002
Neurocomputing1
2016 Size-invariant extended visual cryptography with embedded watermark based on error diffusion
Bin Yan 0001, Ya-Fei Wang, Lingyun Song
Multim. Tools Appl.3
2014 Faceted Exploring for Domain Knowledge over Linked Open Data
abstract
The rapidly increasing RDF data in the Linked Open Data (LOD) community project is a valuable resource for obtaining domain knowledge. However, RDF data of specific topics also shows a trend of being more decentralized and fragmented, which makes it difficult and inefficient for the users to get an overview of a specific topic and retrieve the desired information. In this paper, we demonstrate a novel system called KFM, which can aggregate the distributed RDF data of a topic according to the facets of this topic. KFM provides a new way for users to obtain and explore domain knowledge in the LOD cloud.
Meng Wang 0009, Jun Liu 0002, Wei Zhang 0053, Lingyun Song, Siyu Yao
CIKM6