EDBT 2026 Demo / reviewers in the wild / expert
Xiaosu Wang
dblp:305/0146
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0002-8180-8604ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | How Does Textual Information Selection Influence Time Series Forecasting? A Cross-modal Perspective on Financial Volatility PredictionabstractReal-world time series are often accompanied by textual descriptions as collateral information, especially in finance. Financial volatility prediction based on earnings call transcripts is such a typical scenario. Nevertheless, current studies rarely focus on the cross-modal impact of textual information on time series forecasting, if we regard time series as another modality. Consequently, we propose a Denoised Financial Volatility Prediction (DeFVP) model to investigate the impact of varying textual selection on time series forecasting in the volatility prediction scenario. To this end, we jointly train two neural network modules: a Differentiable Binary Selector (DBS) and a Volatility Predictor (VP). Firstly, the DBS module assigns learnable Bernoulli variables to sentences for identifying helpful sentences. Next, the VP module makes predictions by imposing selected textual influence into time series modeling. We conduct extensive experiments on three real-world datasets, demonstrating the significant and synergistic impact of textual selection on time series forecasting.1 Yun Xiong, Xiaosu Wang, Yao Zhang 0009 |
ICME | 3 |
| 2023 | Automatic ICD Coding Based on Segmented ClinicalBERT with Hierarchical Tree Structure Learning
Beichen Kang, Xiaosu Wang, Yun Xiong, Yao Zhang 0009, Chaofan Zhou, Yangyong Zhu, Jiawei Zhang 0001, Chunlei Tang |
DASFAA (4) | 2 |
| 2023 | Reducing Negative Effects of the Biases of Language Models in Zero-Shot SettingabstractPre-trained language models (PLMs) such as GPTs have been revealed to be biased towards certain target classes because of the prompt and the model's intrinsic biases. In contrast to the fully supervised scenario where there are a large number of costly labeled samples that can be used to fine-tune model parameters to correct for biases, there are no labeled samples available for the zero-shot setting. We argue that a key to calibrating the biases of a PLM on a target task in zero-shot setting lies in detecting and estimating the biases, which remains a challenge. In this paper, we first construct probing samples with the randomly generated token sequences, which are simple but effective in detecting inputs for stimulating GPTs to show the biases; and we pursue an in-depth research on the plausibility of utilizing class scores for the probing samples to reflect and estimate the biases of GPTs on a downstream target task. Furtherly, in order to effectively utilize the probing samples and thus reduce negative effects of the biases of GPTs, we propose a lightweight model Calibration Adapter (CA) along with a self-guided training strategy that carries out distribution-level optimization, which enables us to take advantage of the probing samples to fine-tune and select only the proposed CA, respectively, while keeping the PLM encoder frozen. To demonstrate the effectiveness of our study, we have conducted extensive experiments, where the results indicate that the calibration ability acquired by CA on the probing samples can be successfully transferred to reduce negative effects of the biases of GPTs on a downstream target task, and our approach can yield better performance than state-of-the-art (SOTA) models in zero-shot settings. Xiaosu Wang, Yun Xiong, Beichen Kang, Yao Zhang 0009, Philip S. Yu, Yangyong Zhu |
WSDM | 1 |
| 2022 | Composition-based Heterogeneous Graph Multi-channel Attention Network for Multi-aspect Multi-sentiment ClassificationabstractAspect-based sentiment analysis (ABSA) has drawn more and more attention because of its extensive applications. However, towards the sentence carried with more than one aspect, most existing works generate an aspect-specific sentence representation for each aspect term to predict sentiment polarity, which neglects the sentiment relationship among aspect terms. Besides, most current ABSA methods focus on sentences containing only one aspect term or multiple aspect terms with the same sentiment polarity, which makes ABSA degenerate into sentence-level sentiment analysis. In this paper, to deal with this problem, we construct a heterogeneous graph to model inter-aspect relationships and aspect-context relationships simultaneously and propose a novel Composition-based Heterogeneous Graph Multi-channel Attention Network (CHGMAN) to encode the constructed heterogeneous graph. Meanwhile, we conduct extensive experiments on three datasets: MAMSATSA, Rest14, and Laptop14, experimental results show the effectiveness of our method. Yun Xiong, Zhongchen Miao, Xiaosu Wang, Hongrun Ren, Yao Zhang 0009, Yangyong Zhu |
COLING | 5 |
| 2021 | C2BERT: Cross-contrast BERT for Chinese Biomedical Sentence RepresentationabstractPre-trained language models (PLMs), such as BERT, have achieved great success on various natural language processing (NLP) tasks. Nevertheless, we observe that PLM-derived native Chinese biomedical sentence representations are somehow collapsed, which means PLMs induce a non-smooth anisotropic semantic space of Chinese biomedical sentences and most sentences are mapped into a small area and therefore produce high similarity. Such PLM-derived native sentence representations poorly capture semantic meaning of Chinese biomedical sentences.To alleviate the aforementioned collapse issue, we then propose a novel contrastive learning framework, named Cross-contrast BERT (C2 BERT), that advances the state-of-the-art Chinese biomedical sentence embeddings. C2 BERT proposes to derive positive/negative samples from two transformer-based different PLMs; this design decision reflects our philosophy that our goal is to conflate the knowledge stored in different PLMs to produce Chinese biomedical sentence embeddings, rather than introducing new noise. Moreover, without costly further pretraining, C2 BERT exploits contrastive learning as an auxiliary training objective during fine-tuning with supervision from biomedical sentence-related tasks. We demonstrate with extensive experiments that our C2 BERT model is more effective than competitive baselines on diverse Chinese biomedical sentence-related tasks. Xiaosu Wang, Yun Xiong, Yao Zhang 0009, Yangyong Zhu |
BIBM | 1 |
| 2021 | Improving Chinese Character Representation with Formation Graph Attention NetworkabstractChinese characters are often composed of subcharacter components which are also semantically informative, and the component-level internal semantic features of a Chinese character inherently bring with additional information that benefits the semantic representation of the character. Therefore, there have been several studies that utilized subcharacter component information (e.g. radical, fine-grained components and stroke n-grams) to improve Chinese character representation. Xiaosu Wang, Yun Xiong, Jingwen Yue, Yangyong Zhu, Philip S. Yu |
CIKM | 1 |
| 2021 | BioHanBERT: A Hanzi-aware Pre-trained Language Model for Chinese Biomedical Text MiningabstractUnsupervised pre-trained language models (PLMs) have boosted the development of effective biomedical text mining models. But the biomedical texts contain a huge number of long-tail concepts and terminologies, which makes further pre-training on biomedical corpora relatively expensive (more biomedical corpora and more pre-training steps are needed). Nonetheless, this problem receives less attention in recent studies. In Chinese biomedical text, concepts and terminologies consist of Chinese characters, and Chinese characters are often composed of sub-character components which are also semantically informative; thus in order to enhance the semantics of biomedical concepts and terminologies, the use of a Chinese character’s component-level internal semantic information also appears to be reasonable.In this paper, we propose a novel hanzi-aware pre-trained language model for Chinese biomedical text mining, referred to as BioHanBERT (hanzi-aware BERT for Chinese biomedical text mining), utilizing the component-level internal semantic information of Chinese characters to enhance the semantics of Chinese biomedical concepts and terminologies, and thereby to reduce further pre-training costs. BioHanBERT first employs a Chinese character encoder to extract the component-level internal semantic feature of each Chinese character, and then fuse the character’s internal semantic feature and its contextual embedding extracted by BERT to enrich the representations of the concepts or terminologies containing the character. The results of extensive experiments show that our model is able to consistently outperform current state-of-the-art (SOTA) models in a wide range of Chinese biomedical natural language processing (NLP) tasks. Xiaosu Wang, Yun Xiong, Jingwen Yue, Yangyong Zhu, Philip S. Yu |
ICDM | 1 |