Yin Luo

dblp:83/1956 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 6 since 2021Security and privacy · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A wavelet-enhanced dual-branch LipFormer with Hybrid-GRPO for robust algorithmic trading
Yin Luo, Yuling Huang
Expert Syst. Appl.1
2026 mChartQA and mChartQABench: A multimodal-only solution for complex chart question-answering
Jingxuan Wei, Nan Xu 0004, Guiyong Chang, Yin Luo, Bihui Yu, Ruifeng Guo
Pattern Recognit.4
2025 Seeking multi-view commonality and peculiarity: A novel decoupling method for lung cancer subtype classification
Ziyu Gao, Yin Luo, Chi Cao, Houzhou Jiang, Ao Li 0001
Expert Syst. Appl.2
2024 Dual Complex Number Knowledge Graph Embeddings
abstract
Knowledge graph embedding, which aims to learn representations of entities and relations in large scale knowledge graphs, plays a crucial part in various downstream applications. The performance of knowledge graph embedding models mainly depends on the ability of modeling relation patterns, such as symmetry/antisymmetry, inversion and composition (commutative composition and non-commutative composition). Most existing methods fail in modeling the non-commutative composition patterns. Several methods support this kind of pattern by modeling in quaternion space or dihedral group. However, extending to such sophisticated spaces leads to a substantial increase in the amount of parameters, which greatly reduces the parameter efficiency. In this paper, we propose a new knowledge graph embedding method called dual complex number knowledge graph embeddings (DCNE), which maps entities to the dual complex number space, and represents relations as rotations in 2D space via dual complex number multiplication. The non-commutativity of the dual complex number multiplication empowers DCNE to model the non-commutative composition patterns. In the meantime, modeling relations as rotations in 2D space can effectively improve the parameter efficiency. Extensive experiments on multiple benchmark knowledge graphs empirically show that DCNE achieves significant performance in link prediction and path query answering.
Yao Dong 0003, Qingchao Kong, Lei Wang 0135, Yin Luo
LREC/COLING4
2024 PromISe: Releasing the Capabilities of LLMs with Prompt Introspective Search
abstract
The development of large language models (LLMs) raises the importance of assessing the fairness and completeness of various evaluation benchmarks. Regrettably, these benchmarks predominantly utilize uniform manual prompts, which may not fully capture the expansive capabilities of LLMs—potentially leading to an underestimation of their performance. To unlock the potential of LLMs, researchers pay attention to automated prompt search methods, which employ LLMs as optimizers to discover optimal prompts. However, previous methods generate the solutions implicitly, which overlook the underlying thought process and lack explicit feedback. In this paper, we propose a novel prompt introspective search framework, namely PromISe, to better release the capabilities of LLMs. It converts the process of optimizing prompts into an explicit chain of thought, through a step-by-step procedure that integrates self-introspect and self-refine. Extensive experiments, conducted over 73 tasks on two major benchmarks, demonstrate that our proposed PromISe significantly boosts the performance of 12 well-known LLMs compared to the baseline approach. Moreover, our study offers enhanced insights into the interaction between humans and LLMs, potentially serving as a foundation for future designs and implementations. Keywords: large language models, prompt search, self-introspect, self-refine
Minzheng Wang 0001, Nan Xu 0004, Jiahao Zhao 0001, Yin Luo, Wenji Mao
LREC/COLING4
2024 HNCGAT: a method for predicting plant metabolite-protein interaction using heterogeneous neighbor contrastive graph attention network
abstract
The prediction of metabolite-protein interactions (MPIs) plays an important role in plant basic life functions. Compared with the traditional experimental methods and the high-throughput genomics methods using statistical correlation, applying heterogeneous graph neural networks to the prediction of MPIs in plants can reduce the cost of manpower, resources, and time. However, to the best of our knowledge, applying heterogeneous graph neural networks to the prediction of MPIs in plants still remains under-explored. In this work, we propose a novel model named heterogeneous neighbor contrastive graph attention network (HNCGAT), for the prediction of MPIs in Arabidopsis. The HNCGAT employs the type-specific attention-based neighborhood aggregation mechanism to learn node embeddings of proteins, metabolites, and functional-annotations, and designs a novel heterogeneous neighbor contrastive learning framework to preserve heterogeneous network topological structures. Extensive experimental results and ablation study demonstrate the effectiveness of the HNCGAT model for MPI prediction. In addition, a case study on our MPI prediction results supports that the HNCGAT model can effectively predict the potential MPIs in plant.
Yin Luo
Briefings Bioinform.3
2024 A signal-diffusion-based unsupervised contrastive representation learning for spatial transcriptomics analysis
abstract
MOTIVATION: Spatial transcriptomics allows for the measurement of high-throughput gene expression data while preserving the spatial structure of tissues and histological images. Integrating gene expression, spatial information, and image data to learn discriminative low-dimensional representations is critical for dissecting tissue heterogeneity and analyzing biological functions. However, most existing methods have limitations in effectively utilizing spatial information and high-resolution histological images. We propose a signal-diffusion-based unsupervised contrast learning method (SDUCL) for learning low-dimensional latent embeddings of cells/spots. RESULTS: SDUCL integrates image features, spatial relationships, and gene expression information. We designed a signal diffusion microenvironment discovery algorithm, which effectively captures and integrates interaction information within the cellular microenvironment by simulating the biological signal diffusion process. By maximizing the mutual information between the local representation and the microenvironment representation of cells/spots, SDUCL learns more discriminative representations. SDUCL was employed to analyze spatial transcriptomics datasets from multiple species, encompassing both normal and tumor tissues. SDUCL performed well in downstream tasks such as clustering, visualization, trajectory inference, and differential gene analysis, thereby enhancing our understanding of tissue structure and tumor microenvironments. AVAILABILITY AND IMPLEMENTATION: https://github.com/WeiMin-Li-visual/SDUCL.
Weimin Li 0001, Fangfang Liu 0008, Yin Luo, Zhongkun Zuo
Bioinform.5
2023 Crude oil price forecasting base on an EEMD and multi-scale time series analysis combination methods
abstract
In this paper, a nonlinear time series analysis method is used to study and analyses crude oil price forecasting, from the perspective of multi-scale analysis, the index is used to construct proxy variables of interest to investors, and the Ensemble Empirical Mode Decomposition (EEMD) method is applied to decompose the crude oil price and Baidu index time series into several independent and different scales of the superposition of modular functions and residual terms, respectively, to extract the volatility characteristics of the series at different time scales. The underlying modal components are divided into high-frequency components, low-frequency components, and residual terms according to the frequency, representing short-term fluctuations, medium-term trends, and long-term trends, respectively. Then, a multi-scale Granger causality test model based on EEMD and a multi-scale sign transfer first method are constructed. Finally, the multi-scale causal relationship between crude oil prices and investors’ attentions in different market capitalizations is systematically investigated by dividing crude oil by market capitalization. Firstly, EEMD decomposes the original time series of oil price and uses the wavelet threshold denoising method to obtain the effective information of high-frequency modal components; secondly, the decomposed modal components are reconstructed using the fine-to-coarse method to obtain the reconstructed components from high to low; then the reconstructed components are predicted using the neural network; finally, the results are obtained by simple summation of the reconstructed series. Compared with the LSTM integrated prediction model with EEMD decomposition, the proposed model improves the oil price prediction accuracy. The two decomposition-integrated deep learning oil price forecasting models proposed in this paper use degree crude oil prices for empirical analysis, and both improve the accuracy of oil price forecasting to some extent. The empirical results show the effectiveness of these two forecasting methods in nonlinear irregular time series forecasting.
Yin Luo, LiFei Ke, Jianwen Qin, Pingping Liu
IEEE Big Data2
2023 Neuro-Logic Learning for Relation Reasoning over Event Knowledge Graph
abstract
Security-related events are continuously emerging around the world, such as social unrest, armed conflicts and natural disasters. Although event knowledge graphs provide a concise and intuitive way to organize event information, they often suffer from the incompleteness problem. To solve this problem, the knowledge graph reasoning technique is usually adopted to infer new facts based on existing relations of entities. However, existing reasoning methods often concentrate on the entity relations, without considering the attributes of nodes for event knowledge graphs, such as event descriptions and event time, which contain rich information of events and are vital for event relation reasoning. Existing methods are also deficient in interpretability, which significantly benefits security-related decision-making scenarios. In this paper, we propose a Neuro-Logic Learning (NLL) method for relation reasoning over event knowledge graph. Our method first adopts different event attributes to learn the semantic representations of events, and ranks candidate event relation triples according to their similarity scores. The proposed method then develops the neural logic rule leaner to obtain interpretable reasoning rules. We construct a new large-scale event relation dataset EventKG-R based on EventKG to evaluate the effectiveness of our proposed method. Experimental results show that our proposed method achieves superior performances compared to the state-of-the-art baselines.
Qingchao Kong, Yin Luo, Wenji Mao
ISI3
2023 Boosting Domain-Specific Question Answering Through Weakly Supervised Self-Training
abstract
Question answering systems have emerged as an important research area in the field of natural language processing, enabling the provision of accurate answers to user queries in a more efficient and user-friendly manner. These systems have much practical significance especially in security-related applications, where intelligence analysts can easily access pertinent information from diverse sources to accelerate the decision-making process. In the context of security-related scenarios, users tend to interact with a question answering system with specific domain-oriented purposes. However, most existing question answering systems focus on open-domain situations with abundant labeled data as well as structured data, and domain-specific question answering methods are in urgent need. Domain-specific question answering faces the challenge of the low-resource issue caused by data scarcity. To address this challenge, in this paper, we propose a weakly supervised self-training method for domain-specific question answering based on the Retriever-Reader framework. For the retriever module, during the self-training process, we develop two strategies for generating pseudo-labels to augment the labeled dataset, including high confidence sampling and random negative sampling. For the reader module, we adopt the pre-trained language model and fine-tune the generative reader using limited labeled datasets. To evaluate our proposed method, we construct the first Chinese financial question answering dataset of textual document. Experimental results demonstrate that our proposed method can significantly improve the performances of the baseline method through the self-training process.
Minzheng Wang 0001, Jia Cao, Qingchao Kong, Yin Luo
ISI4
2023 A Cognitive Knowledge Enriched Joint Framework for Social Emotion and Cause Mining
Xinglin Xiao, Yin Luo, Wenji Mao
KSEM (2)3
2023 CARL: Cross-Aligned Representation Learning for Multi-view Lung Cancer Histology Classification
Yin Luo, Wei Liu 0235, Qilong Song, Xuhong Min, Ao Li 0001
MICCAI (5)1
2023 A cross-lingual transfer learning method for online COVID-19-related hate speech detection
Alexis Pengfei Zhao, Daniel Dajun Zeng, Paul Jen-Hwa Hu, Qingpeng Zhang, Yin Luo, Zhidong Cao
Expert Syst. Appl.7
2023 A Deep Learning Approach for Semantic Analysis of COVID-19-Related Stigma on Social Media
abstract
The rapid spread of the pandemic of coronavirus disease of 2019 (COVID-19) has created an unprecedented, global health disaster. During the outburst period, the paucity of knowledge and research aggravated devastating panic and fears that lead to social stigma and created serious obstacles to contain the disastrous epidemic. We propose a deep learning-based method to detect stigmatized contents on online social network (OSN) platforms in the early stage of COVID-19. Our method performs a semantic-based quantitative analysis to unveil essential spatial-temporal characteristics of COVID-19 stigmatization for timely alerts and risk mitigation. Empirical evaluations are carried out to examine our method’s predictive utilities. The visualization results of the co-occurrence network using Gephi indicate two distinct groups of stigmatized words that pertain to people in Wuhan and their dietary behaviors, respectively. Netizens’ participations and stigmatizations in the Hubei region, where the COVID-19 broke out, are twice ($p < 0.05$) and four ($p < 0.01$) times more frequent and intense than those in other parts of China, respectively. Also, the number of COVID-19 patients is correlated with COVID-19-related stigma significantly (correlation coefficient = 0.838,$p < 0.01$). The responses to individual users’ posts have the power law distribution, while posts by official media appear to attract more responses (e.g., likes, replies, and forward). Our method can help platforms and government agencies manage public health disasters through effective identification and detailed analyses of social stigma on social media.
Zhidong Cao, Alexis Pengfei Zhao, Paul Jen-Hwa Hu, Daniel Dajun Zeng, Yin Luo
IEEE Trans. Comput. Soc. Syst.6
2022 Manmade-Target Three-Dimensional Reconstruction Using Multi-View Radar Images
abstract
Manmade-target three-dimensional (3D) reconstruction is an attractive topic and also a challenge in radar imaging field. The factorization based 3D reconstruction algorithm using multi-view radar images provides a way without extra configuration requirements. The precondition of it is that the reconstructed points' positions are retrievable in all images, but it is hard to be satisfied in practice due to the points loss problem caused by occlusion and scattering variation during view change. The points loss problem reduces the reconstructed points and weaken the details. We propose a modi-fied manmade-target 3D reconstruction algorithm. Firstly, we calculate homography matrix after feature matching of each adjacent image pair and improve its accuracy through non-linear optimization. Then, we put forward a track generation algorithm under the guidance of the homography matrix to estimate strong points' positions in each image and a track filter to refine the estimation. Finally, we achieve the 3D coordinates through the orthographic factorization method. The measured data processing results demonstrate the validity and the reconstruction quality improvement of the proposed method.
Yin Luo, Si-Wei Chen 0001, Xuesong Wang 0003
IGARSS1
2021 A Central Opinion Extraction Framework for Boosting Performance on Sentiment Analysis
abstract
With the rapid development of the Internet, mining opinions and emotions from the explosive growth of user-generated content is a key field of social media analysis. However, the expression forms of the central opinion which strongly expresses the essential points and converges the main sentiments of the overall document are diverse in practice, such as sequential sentences, a sentence fragment, or an individual sentence. Previous research studies on sentiment analysis based on document level and sentence level fail to deal with this actual situation uniformly. To address this issue, we propose a Central Opinion Extraction (COE) framework to boost performance on sentiment analysis with social media texts. Our framework first extracts a span-level central opinion text, which expresses the essential opinion related to sentiment representation among the whole text, and then uses extracted textual span to boost the performance of sentiment classifiers. The experimental results on a public dataset show the effectiveness of our framework for boosting the performance on document-level sentiment analysis task.
Nan Xu 0004, Wenji Mao, Yin Luo
ISI4
2021 Extracting Impacts of Non-pharmacological Interventions for COVID-19 From Modelling Study
abstract
COVID-19 pandemic continues to rampage in the world. Before the achievement of global herd immunity, non-pharmacological interventions(NPIs) are crucial to mitigate the pandemic. Although various NPIs have been put into practice, there are many concerns about the impacts and effectiveness of these NPIs. COVID-19 modelling study (CMS) in epidemiology can provide evidence to solve the aforementioned concerns. It is time-consuming to collect evidence manually when dealing with the vast amount of CMS papers. Accordingly, we seek to accelerate evidence collection by developing an information extraction model to automatically identify evidence from CMS papers. This work presents a novel COVID-19 Non-pharmacological Interventions Evidence (CNPIE) Corpus, which contains 597 abstracts of COVID-19 modelling study with richly annotated entities and relations of the impacts of NPIs. We design a semi-supervised document-level information extraction model (SS-DYGIE++) which can jointly extract entities and relations. Our model outperforms previous baselines in both entity recognition and relation extraction tasks by a large margin. The proposed work can be applied towards automatic evidence extraction in the public health domain for assisting the public health decision-making of the government.
Yunrong Yang, Zhidong Cao, Alexis Pengfei Zhao, Daniel Dajun Zeng, Qingpeng Zhang, Yin Luo
ISI6
2019 A novel oversampling method based on SeqGAN for imbalanced text classification
abstract
Nowadays, investors prefer to learn the latest trends of stock market through the Internet, grasp the best time to buy and sell stocks, and simultaneously release their comments during the trading process through various online media. Therefore, many researchers utilize NLP approaches to extract sentiment tendency of investors from online comments. However, since there exists noise in corpora related to stock, when extracting emotional tendency from these corpora by using text classification, it is inevitable to face the imbalanced dataset problem, which is harmful to traditional machine learning classifiers. Our research innovatively proposes a preprocessing method based on SeqGAN algorithm, in order to solve the imbalanced classification problem. We carry out experiments on different corpora, and results show that compared with other four oversampling methods, the oversampling method based on SeqGAN can improve the performance of text classifier at best both in binary class and multiclass text classification, which proves to be an effective method for preprocessing text data.
Yin Luo, Haishan Feng, Xuanlong Weng, Huang Zheng
IEEE BigData1
2019 A Framework for Policy Information Popularity Prediction in New Media
abstract
With the rapid development and wide application of new media, predicting the popularity of policy information on new media is of great significance for understanding and managing public opinion. However, the complexity of the diffusion patterns of policy information has brought great challenges for predicting the popularity of such information. Inspired by the methods of popularity prediction for short text information from social networks, we propose a framework for the popularity prediction of policy information. In our framework, first, the features of policy information are extracted from three dimensions: contextual information, social information and textual information. Then, effective features, such as the topic distribution, popularity competition intensity and hot information relevance, are identified by empirical analysis. Finally, the effective features are input into the prediction model to predict the popularity of policy information. We evaluate the performance of our proposed framework using a real-world dataset and the experimental results show that the framework can efficiently predict the popularity of policy information and that the features that we used are effective in improving the accuracy of policy information popularity prediction. The accurate prediction result could benefit policy makers, allowing them to make better decisions, understand and manage public opinion.
Yin Luo, Lei Wang 0062, Yanni Hao, Daniel Dajun Zeng
ISI1
2019 Identifying Risks of the Internet Finance Platforms Using Multi-Source Text Data
abstract
With the explosion of the Internet Finance Platforms, identifying the risks of these platforms is of growing significance, which can help discover problematic platforms in time and ensure the healthy development of the Internet finance industry. In this paper, we design a risk index system to measure the quantitative risk of the Internet finance platforms, and propose a deep neural network based model, CBiGRU-RI, to identify the risks of the platforms using multi-source text data. We conducted comparative experiments with various baseline models on real-world data. The experimental results show that our proposed model can identify the risks of platforms more effectively than the baseline methods.
Donglei Zhang, Yin Luo
ISI5
2019 A Modified Cartesian Factorized Back-Projection Algorithm for Highly Squint Spotlight Synthetic Aperture Radar Imaging
abstract
Highly squint synthetic aperture radar (SAR) increases the flexibility and aspect-information of observation but poses several challenges to frequency-domain algorithms. The back-projection (BP) algorithm is recognized to produce the best results for highly squint mode but at high computational expense. Several fast BP algorithms have been developed to enhance the efficiency of the BP integral. The Cartesian factorized BP (CFBP) algorithm has been proposed to improve the performance. However, CFBP is ineffective in the highly squint case because its range model (RM) is imprecise. In this letter, we propose a modified CFBP (MCFBP) algorithm for highly squint spotlight SAR imaging. We employ a transformed Cartesian coordinate system to treat the transceiver mode as approximately side-looking mode. According to the transformed coordinate system, we propose a modified RM (MRM) to reduce the range error. Based on the MRM, we derive modified image spectrum compression steps and the Nyquist sampling rate of the subaperture image. MCFBP inherits CFBP's accuracy and efficiency advantages. Results of simulations and experiments performed by the X-band airborne SAR system with a maximum bandwidth of 1.2 GHz validate the improved performance of the proposed algorithm relative to BP and fast factorized BP.
Yin Luo, Fengjun Zhao, Ning Li 0002, Heng Zhang 0007
IEEE Geosci. Remote. Sens. Lett.1
2019 Local spatial correlation-based stripe non-uniformity correction algorithm for single infrared images
Bo Zhou 0006, Yin Luo, Baoguo Chen, Mingchang Wang, Kun Liang 0003
Signal Process. Image Commun.2
2018 An Autofocus Cartesian Factorized Backprojection Algorithm for Spotlight Synthetic Aperture Radar Imaging
abstract
A backprojection (BP) algorithm is recognized as an ideal method for high-resolution synthetic aperture radar (SAR) imaging. Several fast BP algorithms have been developed to enhance the efficiency of the BP integral. The Cartesian factorized BP (CFBP) algorithm is proposed recently to avoid massive interpolations and improve the performance. However, integrating autofocus techniques with the CFBP has not been discussed. In this letter, an autofocus CFBP algorithm is proposed to compatibly combine the autofocus processing within the CFBP. After modifying the spectrum compression step in the CFBP, the approximate Fourier transformation (FT) relationship between the modified compensated subaperture images and the corresponding range-compressed phase history data in the Cartesian coordinate is revealed. The phase error is obtained by the multiple aperture map drift method, and the singular value decomposition total least square method is combined to improve the estimate robustness. Employing the range blocking method, the range variance of the phase error is compensated. The proposed algorithm inherits the advantages of the CFBP. Experiments performed by the X-band airborne SAR system with a maximum bandwidth of 1.2 GHz validate the proposed approaches.
Yin Luo, Fengjun Zhao, Ning Li 0002, Heng Zhang 0007
IEEE Geosci. Remote. Sens. Lett.1
2016 Meme extraction and tracing in crisis events
abstract
The proliferation of social media has increased the competition among different memes, which can be free texts, trending catchphrases, or micro media. As human attention is limited, these memes compete with each other, and go in and out of popularity at a rapid pace, sometimes even faster than we can recognize. Popular memes often shape the mindsets of online communities, and also shed light on their future tendencies. Considering the huge volume of memes generated and their continuous mutations, extracting and tracing online memes automatically is rather challenging. In this paper, we propose an automatic meme extraction algorithm. The proposed algorithm extracts massive memes based on phrases independency, and clusters phrase variants of a single meme efficiently. Evaluation on measles outbreak in the USA in 2015 indicates that the proposed algorithm could extract typical memes reflecting the fierce campaign between the pro-vaccination community and the anti-vaccination community. In both communities, memes are power-law distributed, and popular ones have many variants that appear more frequently. By tracing the evolution of online memes, we uncover that popular memes converge and generate peaks at times. Though the pro-vaccination community and the anti-vaccination community may focus on similar memes, they comprehend memes from totally different perspectives and deliver opposing opinions of measles vaccination.
Saike He, Xiaolong Zheng 0001, Zhijun Chang, Yin Luo, Daniel Dajun Zeng
ISI5
2006 An Approach to SOA-Based Bioinformatics Grid
abstract
Bioinformatics grid has emerged to meet the sharp requirement for computation and storage resource in the bioinformatics research field. In this paper, a design of SOA-based bioinformatics grid system is put forward under the WSRF specification. It aggregates existing tools such as BLAST, wrapped as Web services, through WS-ServiceGroup, and supports the development, deployment, execution and monitoring of complex bioinformatics grid application. Based on the supporting environment, applications can access bioinformatics resource in a transparent manner. This paper also presents a use case to show the process of workflow building, execution and monitoring
Guoshi Xu, Yin Luo, Huashan Yu, Zhuoqun Xu
APSCC2