Shasha Mo

dblp:166/7129 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MSRNet: Multi-scale Spatiotemporal Retention Network for Motor Imagery Recognition
Shasha Mo, Guanglei Song, Shuo Tan
KSEM (3)2
2025 Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
abstract
Byte Pair Encoding (BPE) serves as a foundation method for text tokenization in the Natural Language Processing (NLP) field. Despite its wide adoption, the original BPE algorithm harbors an inherent flaw: it inadvertently introduces a frequency imbalance for tokens in the text corpus. Since BPE iteratively merges the most frequent token pair in the text corpus to generate a new token and keeps all generated tokens in the vocabulary, it unavoidably holds tokens that primarily act as components of a longer token and appear infrequently on their own. We term such tokens as Scaffold Tokens. Due to their infrequent occurrences in the text corpus, Scaffold Tokens pose a learning imbalance issue. To address that issue, we propose Scaffold-BPE, which incorporates a dynamic scaffold token removal mechanism by parameter-free, computation-light, and easy-to-implement modifications to the original BPE method. This novel approach ensures the exclusion of low-frequency Scaffold Tokens from the token representations for given texts, thereby mitigating the issue of frequency imbalance and facilitating model training. On extensive experiments across language modeling and even machine translation, Scaffold-BPE consistently outperforms the original BPE, well demonstrating its effectiveness.
Haoran Lian, Yizhe Xiong, Jianwei Niu 0002, Shasha Mo, Zhenpeng Su, Zijia Lin, Hui Chen 0013, Jungong Han, Guiguang Ding
AAAI4
2025 LBPE: Long-token-first Tokenization to Improve Large Language Models
abstract
The prevalent use of Byte Pair Encoding (BPE) in Large Language Models (LLMs) facilitates robust handling of subword units and avoids issues of out-of-vocabulary words. Despite its success, a critical challenge persists: long tokens, rich in semantic information, have fewer occurrences in tokenized datasets compared to short tokens, which can result in imbalanced learning issue across different tokens. To address that, we propose LBPE, which prioritizes long tokens during the encoding process. LBPE generates tokens according to their descending order of token length rather than their ranks in the vocabulary, granting longer tokens higher priority during the encoding process. Consequently, LBPE smooths the frequency differences between short and long tokens, and thus mitigates the learning imbalance. Extensive experiments across diverse language modeling tasks demonstrate that LBPE consistently outperforms the original BPE, well demonstrating its effectiveness.
Haoran Lian, Yizhe Xiong, Zijia Lin, Jianwei Niu 0002, Shasha Mo, Hui Chen 0013, Guiguang Ding
ICASSP5
2024 LogicST: A Logical Self-Training Framework for Document-Level Relation Extraction with Incomplete Annotations
abstract
Document-level relation extraction (DocRE)aims to identify relationships between entities within a document.Due to the vast number of entity pairs, fully annotating all fact triplets is challenging, resulting in datasets with numerous false negative samples.Recently, selftraining-based methods have been introduced to address this issue.However, these methods are purely black-box and sub-symbolic, making them difficult to interpret and prone to overlooking symbolic interdependencies between relations.To remedy this deficiency, our insight is that symbolic knowledge, such as logical rules, can be used as diagnostic tools to identify conflicts between pseudo-labels.By resolving these conflicts through logical diagnoses, we can correct erroneous pseudolabels, thus enhancing the training of neural models.To achieve this, we propose Log-icST, a neural-logic self-training framework that iteratively resolves conflicts and constructs the minimal diagnostic set for updating models.Extensive experiments demonstrate that LogicST significantly improves performance and outperforms previous state-of-the-art methods.For instance, LogicST achieves an increase of 7.94% in F1 score compared to CAST (Tan et al., 2023a) on the DocRED benchmark (Yao et al., 2019).Additionally, LogicST is more time-efficient than its self-training counterparts, requiring only 10% of the training time of CAST.Code is available at https: //github.com/XingYing-stack/LogicST.
Shengda Fan, Shasha Mo, Jianwei Niu 0002
EMNLP3
2023 HSRG-WSD: A Novel Unsupervised Chinese Word Sense Disambiguation Method Based on Heterogeneous Sememe-Relation Graph
Meng Lyu, Shasha Mo
ICIC (4)2
2023 GenBoost: Generative Modeling and Boosted Learning for Multi-hop Question Answering over Incomplete Knowledge Graphs
abstract
Multi-hop question answering over incomplete knowledge graphs involves iteratively reasoning on the provided question and graph to find answers, while also tackling the inherent sparsity problem in the graph. Walk-based methods transform the reasoning process into a graph traversal task; however, they encounter challenges in convergence and stability due to the extensive action space and sensitivity to missing triples. On the other hand, embedding-based methods address the problem of missing triples but compromise interpretability in answer selection because of their black-box nature. We present GenBoost, a Generate-then-Boost framework for generative reasoning. By transforming the question-answering task into an inference path generation task, GenBoost effectively addresses existing limitations and offers a more efficient and interpretable approach for answer selection. GenBoost possesses two key features: (1) The reasoning procedure does not explicitly rely on the existing triples in the knowledge graph. By combining graph traversal and link prediction, our approach mitigates the impact of knowledge graph incompleteness. (2) Each entity in the reasoning path is generated autoregressively, providing insights into the decision-making process during multi-hop reasoning and enhancing interpretability. Extensive experiments conducted on incomplete knowledge graphs have demonstrated the effectiveness of our approach.
Jianwei Niu 0002, Shasha Mo
ICPADS3
2023 Topic-Aware Modeling for Unsupervised Extractive Summarization
abstract
The recent success of extractive summarization depends on the availability of large-scale annotated datasets. Existing unsupervised approaches are mostly directed graph based by combining location information with centrality computing. These methods tend to generate summaries with two problems, one is low topic coverage of the source document called the facet bias problem, and the other is continuous position distribution of extracted sentences called the position bias problem. To solve these problems, we propose the topic-aware centrality-based sum-marization method (TACSUM). Specifically, we employ clustering techniques to explicitly model the topics of the document and define the metrics for topic consistency and topic coverage to improve the performance of summarization. The metric topic consistency is used to guide the calculation of centrality, which solves the position bias problem and achieves a more general effect in different scenarios. We combine the metric topic coverage with the centrality to enhance the topic awareness of the model, which ensures the selected sentences are important and diverse. Numerical experimental results on four datasets show that our method outperforms previous unsupervised methods, especially in long document domains. Extensive analyses confirm that our method can generate high-quality summaries by eliminating position bias and facet bias problems.
Zhihao Fan, Huiyong Li 0005, Shasha Mo, Jianwei Niu 0002
IJCNN3
2022 Key Mention Pairs Guided Document-Level Relation Extraction
abstract
Document-level Relation Extraction (DocRE) aims at extracting relations between entities in a given document. Since different mention pairs may express different relations or even no relation, it is crucial to identify key mention pairs responsible for the entity-level relation labels. However, most recent studies treat different mentions equally while predicting the relations between entities, leading to sub-optimal performance. To this end, we propose a novel DocRE model called Key Mention pairs Guided Relation Extractor (KMGRE) to directly model mention-level relations, containing two modules: a mention-level relation extractor and a key instance classifier. These two modules could be iteratively optimized with an EM-based algorithm to enhance each other. We also propose a new method to solve the multi-label problem in optimizing the mention-level relation extractor. Experimental results on two public DocRE datasets demonstrate that the proposed model is effective and outperforms previous state-of-the-art models.
Jianwei Niu 0002, Shasha Mo, Shengda Fan
COLING3
2022 CETA: A Consensus Enhanced Training Approach for Denoising in Distantly Supervised Relation Extraction
abstract
Distantly supervised relation extraction aims to extract relational facts from texts but suffers from noisy instances. Existing methods usually select reliable sentences that rely on potential noisy labels, resulting in wrongly selecting many noisy training instances or underutilizing a large amount of valuable training data. This paper proposes a sentence-level DSRE method beyond typical instance selection approaches by preventing samples from falling into the wrong classification space on the feature space. Specifically, a theorem for denoising and the corresponding implementation, named Consensus Enhanced Training Approach (CETA), are proposed in this paper. By training the model with CETA, samples of different classes are separated, and samples of the same class are closely clustered in the feature space. Thus the model can easily establish the robust classification boundary to prevent noisy labels from biasing wrongly labeled samples into the wrong classification space. This process is achieved by enhancing the classification consensus between two discrepant classifiers and does not depend on any potential noisy labels, thus avoiding the above two limitations. Extensive experiments on widely-used benchmarks have demonstrated that CETA significantly outperforms the previous methods and achieves new state-of-the-art results.
Ruri Liu, Shasha Mo, Jianwei Niu 0002, Shengda Fan
COLING2
2022 Boosting Document-Level Relation Extraction by Mining and Injecting Logical Rules
abstract
Document-level relation extraction (DocRE)aims at extracting relations of all entity pairs in a document.A key challenge to DocRE lies in the complex interdependency between the relations of entity pairs.Unlike most prior efforts focusing on implicitly powerful representations, the recently proposed LogiRE (Ru et al., 2021) explicitly captures the interdependency by learning logical rules.However, Lo-giRE requires extra parameterized modules to reason merely after training backbones, and this disjointed optimization of backbones and extra modules may lead to sub-optimal results.In this paper, we propose MILR, a logic enhanced framework that boosts DocRE by Mining and Injecting Logical Rules.MILR first mines logical rules from annotations based on frequencies.Then in training, consistency regularization is leveraged as an auxiliary loss to penalize instances that violate mined rules.Finally, MILR infers from a global perspective based on integer programming.Compared with LogiRE, MILR does not introduce extra parameters and injects logical rules during both training and inference.Extensive experiments on two benchmarks demonstrate that MILR not only improves the relation extraction performance (1.1%-3.8%F1) but also makes predictions more logically consistent (over 4.5% Logic).More importantly, MILR also consistently outperforms LogiRE on both counts.Code is available at https:// github.com/XingYing-stack/MILR.
Shengda Fan, Shasha Mo, Jianwei Niu 0002
EMNLP2
2021 DAT: Training Deep Networks Robust To Label-Noise by Matching the Feature Distributions
abstract
In real application scenarios, the performance of deep networks may be degraded when the dataset contains noisy labels. Existing methods for learning with noisy labels are limited by two aspects. Firstly, methods based on the noise probability modeling can only be applied to class-level noisy labels. Secondly, others based on the memorization effect outperform in synthetic noise but get weak promotion in real-world noisy datasets. To solve these problems, this paper proposes a novel label-noise robust method named Discrepant Adversarial Training (DAT). The DAT method has ability of enforcing prominent feature extraction by matching feature distribution between clean and noisy data. Therefore, under the noise-free feature representation, the deep network can simply output the correct result. To better capture the divergence between the noisy and clean distribution, a new metric is designed to change the distribution divergence into computable. By minimizing the proposed metric with a min-max training of discrepancy on classifiers and generators, DAT can match noisy data to clean data in the feature space. To the best of our knowledge, DAT is the first to address the noisy label problem from the perspective of the feature distribution. Experiments on synthetic and real-world noisy datasets demonstrate that DAT can consistently outperform other state-of-the-art methods. Codes are available at https://github.com/Tyqnn0323/DAT.
Yuntao Qu, Shasha Mo, Jianwei Niu 0002
CVPR2
2021 Explore Better Relative Position Embeddings from Encoding Perspective for Transformer Models
abstract
Relative position embedding (RPE) is a successful method to explicitly and efficaciously encode position information into Transformer models.In this paper, we investigate the potential problems in Shaw-RPE and XL-RPE, which are the most representative and prevalent RPEs, and propose two novel RPEs called Low-level Fine-grained High-level Coarse-grained (LFHC) RPE and Gaussian Cumulative Distribution Function (GCDF) RPE.LFHC-RPE is an improvement of Shaw-RPE, which enhances the perception ability at medium and long relative positions.GCDF-RPE utilizes the excellent properties of the Gaussian function to amend the prior encoding mechanism in XL-RPE.Experimental results on nine authoritative datasets demonstrate the effectiveness of our methods empirically.Furthermore, GCDF-RPE achieves the best overall performance among five different RPEs.
Anlin Qu, Jianwei Niu 0002, Shasha Mo
EMNLP (1)3
2021 Enhancing Transformer with Horizontal and Vertical Guiding Mechanisms for Neural Language Modeling
abstract
Language modeling is an important problem in Natural Language Processing (NLP), and the multi-layer Transformer network is currently the most advanced and effective model for this task. However, there exist two inherent defects in its multi-head self-attention structure: (1) attention information loss: the lower-level attention weights cannot be explicitly passed through upper layers, which may lead the network lose some pivotal attention information captured by lower-level layers; (2) multi-head bottleneck: the dimension of each head in vanilla Transformer is relatively small and the process of each head is independent, which introduces an expressive bottleneck and makes subspace learning inadequate constitutionally. To overcome these two weaknesses, a novel neural architecture named Guide-Transformer is proposed in this paper. The Guide-Transformer utilizes horizontal and vertical attention information to guide the original process of the multi-head self-attention sublayer without introducing excessive complexity. The experimental results on three authoritative language modeling benchmarks demonstrate the effectiveness of Guide-Transformer. For the popular perplexity (ppl) and bits-per-character (bpc) evaluation metrics, Guide-Transformer achieves moderate improvements over the powerful baseline model.
Anlin Qu, Jianwei Niu 0002, Shasha Mo
ICC3
2019 A Novel Method Based on OMPGW Method for Feature Extraction in Automatic Music Mood Classification
abstract
Music mood is useful for music-related applications such as music retrieval or recommendation, which represents the inherent emotional expression of music signals. In this paper, a novel technique is proposed for music signal analysis in the view of emotions, which is based on the orthogonal matching pursuit, Gabor functions, and the Wigner distribution function. The technique, called the OMPGW method, consists of three-level schemes: the low-level, the middle-level and the high-level schemes. For the low-level schemes, the orthogonal matching pursuit combined with Gabor functions is proposed to provide an adaptive time-frequency decomposition of music signals. Compared with other algorithms for signal analysis, the proposed algorithm can achieve a higher spatial and temporal resolution and give a better interpret of the music signal structures. In the middle-level schemes, the Wigner distribution function is applied to obtain the time-frequency energy distribution of the results from the low-level schemes. High-level schemes are used to describe the modeling of audio features, the procedure of music mood classification. A classifier based on support vector machines is utilized to model the extracted features with the proposed technique regarding the emotion models. Several experiments are conducted with four datasets, and better results are achieved with the proposed method. In music mood classification experiments, music clips are classified into different kinds of mood clusters, and mean accuracy of 69.53 percent on our dataset can be achieved using the OMPGW method.
Shasha Mo, Jianwei Niu 0002
IEEE Trans. Affect. Comput.1
2018 A Novel Method of Articles Rating Based on Concerns Tracking and Matching for Public Opinion Recommendation
abstract
Public opinion events on the Internet are gaining more and more attention from the supervisory institutions for the possibility of malicious guide. Since the number of the events on the Internet is quite enormous, the process of supervision often costs a lot of manpower, which is contrary to the purposes and objectives of Sustainable Computing. However, most traditional methods for news recommendation are designed for netizens who do not have specific responsibilities like supervisory institutions. It is also difficult for supervisory institutions themselves to rate the public opinion articles, which is indispensable for recommendation. In this paper, a novel articles rating method based on tracking and matching (ARTM), is proposed for public opinion recommendation. The ARTM method can mine institution concerns from the browsing history and keep them updating automatically with the changing of institution attention. The processing flow of ARTM is as follows. Firstly, a set of institution concerns are established in terms of three aspects: fixed concerns, potential concerns and reading preferences. Then ratings of public opinion articles are computed by measuring the similarities between the vector of article keywords and the vector of institution concerns. Finally, articles are sorted by ratings and high- ranking articles are added to the recommendation list. In addition, the proposed rating algorithm and tracking algorithm can also be used as standalone modules for other services. In the end, comprehensive evaluation of the proposed method based on real data (78 supervisory institutions browsing history in one month) is made. Experimental results show that the proposed ARTM method can significantly improve recommendation efficiency.
Jianwei Niu 0002, Yanyan Guo, Shasha Mo
ICC3
2018 Affective Analysis for Video Frames Using ConvLSTM Network
abstract
With the rapid development of various online video sharing platforms, large numbers of videos are produced every day. Video affective content analysis has become an active research area in recent years, since emotion plays an important role in the classification and retrieval of videos. In this work, we explore to train very deep convolutional networks using ConvLSTM layers to add more expressive power for video affective content analysis models. Network-in-network principles, batch normalization, and convolution auto-encoder are applied to ensure the effectiveness of the model. Then an extended emotional representation model is used as an emotional annotation. In addition, we set up a database containing two thousand fragments to validate the effectiveness of the proposed model. Experimental results on the proposed data set show that deep learning approach based on ConvLSTM outperforms the traditional baseline and reaches the state-of-the-art system.
Jianwei Niu 0002, Shasha Mo, Yanyan Guo, Lei Wang 0037
ICC3
2018 A novel feature set for video emotion recognition
Shasha Mo, Jianwei Niu 0002, Yiming Su, Sajal K. Das 0001
Neurocomputing1
2017 LWTP: An Improved Automatic Image Annotation Method Based on Image Segmentation
Jianwei Niu 0002, Shasha Mo
CollaborateCom3
2017 Sentiment Analysis of Chinese Words Using Word Embedding and Sentiment Morpheme Matching
Jianwei Niu 0002, Mingsheng Sun, Shasha Mo
CollaborateCom3
2017 A Novel Affective Visualization System for Videos Based on Acoustic and Visual Features
Jianwei Niu 0002, Yiming Su, Shasha Mo
MMM (2)3
2015 An Estimation Algorithm for Phase Errors in Synthetic Aperture Radar Imagery
abstract
This letter proposes a novel algorithm, which is based on the generalized method of moments (GMM), for the estimation and correction of phase errors induced in synthetic aperture radar (SAR) imagery. The GMM algorithm is used to replace the original phase-estimation kernel in the basic structure of the phase-gradient-autofocus algorithm. Since this novel algorithm does not require the observed signal to be a certain distribution model, it is able to estimate arbitrary phase errors. The GMM algorithm has the ability of estimating range-dependent phase errors, which makes it an efficient estimator. As a result, higher accuracy of the estimated phase errors and a better focused image can be achieved. Excellent results have been obtained in autofocusing and imaging experiments on real SAR data.
Shasha Mo, Chang Liu 0041
IEEE Geosci. Remote. Sens. Lett.1