Jia Xu 0004

dblp:95/3616-4 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-4077-230XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 ResFormer: All-Time Reservoir Memory for Long Sequence Classification
abstract
Sequence classification is essential in NLP for understanding and categorizing language patterns in tasks like sentiment analysis, intent detection, and topic classification.Transformerbased models, despite achieving state-of-the-art performance, have inherent limitations due to quadratic time and memory complexity, restricting their input length.Although extensive efforts have aimed at reducing computational demands, processing extensive contexts remains challenging.To overcome these limitations, we propose Res-Former, a novel neural network architecture designed to model varying context lengths efficiently through a cascaded methodology.Res-Former integrates an reservoir computing network featuring a nonlinear readout to effectively capture long-term contextual dependencies in linear time.Concurrently, short-term dependencies within sentences are modeled using a conventional Transformer architecture with fixed-length inputs.Experiments demonstrate that ResFormer significantly outperforms baseline models of DeepSeek-Qwen and ModernBERT, delivering an accuracy improvement of up to +22.3% on the EmoryNLP dataset and consistent gains on MultiWOZ, MELD, and IEMOCAP.In addition, ResFormer exhibits reduced memory consumption, underscoring its effectiveness and efficiency in modeling extensive contextual information.
Jia Xu 0004
EMNLP2
2023 ConceptX: A Framework for Latent Concept Analysis
abstract
The opacity of deep neural networks remains a challenge in deploying solutions where explanation is as important as precision. We present ConceptX, a human-in-the-loop framework for interpreting and annotating latent representational space in pre-trained Language Models (pLMs). We use an unsupervised method to discover concepts learned in these models and enable a graphical interface for humans to generate explanations for the concepts. To facilitate the process, we provide auto-annotations of the concepts (based on traditional linguistic ontologies). Such annotations enable development of a linguistic resource that directly represents latent concepts learned within deep NLP models. These include not just traditional linguistic concepts, but also task-specific or sensitive concepts (words grouped based on gender or religious connotation) that helps the annotators to mark bias in the model. The framework consists of two parts (i) concept discovery and (ii) annotation platform.
Firoj Alam, Fahim Dalvi, Nadir Durrani, Hassan Sajjad 0001, Abdul Rafae Khan, Jia Xu 0004
AAAI6
2023 Probabilistic Robustness for Data Filtering
abstract
We introduce our probabilistic robustness rewarded data optimization (PRoDO) approach as a framework to enhance the model's generalization power by selecting training data that optimizes our probabilistic robustness metrics.We use proximal policy optimization (PPO) reinforcement learning to approximately solve the computationally intractable training subset selection problem.The PPO's reward is defined as our (α, ϵ, γ)-Robustness that measures performance consistency over multiple domains by simulating unknown test sets in real-world scenarios using a leaving-one-out strategy.We demonstrate that our PRoDO effectively filters data that lead to significantly higher prediction accuracy and robustness on unknown-domain test sets.Our experiments achieve up to +17.2% increase of accuracy (+25.5% relatively) in sentiment analysis, and -28.05 decrease of perplexity (-32.1% relatively) in language modeling.In addition, our probabilistic (α, ϵ, γ)-Robustness definition serves as an evaluation metric with higher levels of agreement with human annotations than typical performance-based metrics.
Yu Yu 0004, Abdul Rafae Khan, Shahram Khadivi, Jia Xu 0004
EACL4
2023 Learning Uncertainty for Unknown Domains with Zero-Target-Assumption
Yu Yu 0004, Hassan Sajjad 0001, Jia Xu 0004
ICLR3
2022 Measuring Robustness for NLP
abstract
The quality of Natural Language Processing (NLP) models is typically measured by the accuracy or error rate of a predefined test set. Because the evaluation and optimization of these measures are narrowed down to a specific domain like news and cannot be generalized to other domains like Twitter, we often observe that a system reported with human parity results generates surprising errors in real-life use scenarios. We address this weakness with a new approach that uses an NLP quality measure based on robustness. Unlike previous work that has defined robustness using Minimax to bound worst cases, we measure robustness based on the consistency of cross-domain accuracy and introduce the coefficient of variation and (epsilon, gamma)-Robustness. Our measures demonstrate higher agreements with human evaluation than accuracy scores like BLEU on ranking Machine Translation (MT) systems. Our experiments of sentiment analysis and MT tasks show that incorporating our robustness measures into learning objectives significantly enhances the final NLP prediction accuracy over various domains, such as biomedical and social media.
Yu Yu 0004, Abdul Rafae Khan, Jia Xu 0004
COLING3
2022 Can Data Diversity Enhance Learning Generalization?
abstract
This paper introduces our Diversity Advanced Actor-Critic reinforcement learning (A2C) framework (DAAC) to improve the generalization and accuracy of Natural Language Processing (NLP). We show that the diversification of training samples alleviates overfitting and improves model generalization and accuracy. We quantify diversity on a set of samples using the max dispersion, convex hull volume, and graph entropy based on sentence embeddings in high-dimensional metric space. We also introduce A2C to select such a diversified training subset efficiently. Our experiments achieve up to +23.8 accuracy increase (38.0% relatively) in sentiment analysis, -44.7 perplexity decrease (37.9% relatively) in language modeling, and consistent improvements in named entity recognition over various domains. In particular, our method outperforms both domain adaptation and generalization baselines without using any target domain knowledge.
Yu Yu 0004, Shahram Khadivi, Jia Xu 0004
COLING3
2022 Byte-based Multilingual NMT for Endangered Languages
abstract
Multilingual neural machine translation (MNMT) jointly trains a shared model for translation with multiple language pairs. However, traditional subword-based MNMT approaches suffer from out-of-vocabulary (OOV) issues and representation bottleneck, which often degrades translation performance on certain language pairs. While byte tokenization is used to tackle the OOV problems in neural machine translation (NMT), until now its capability has not been validated in MNMT. Additionally, existing work has not studied how byte encoding can benefit endangered language translation to our knowledge. We propose a byte-based multilingual neural machine translation system (BMNMT) to alleviate the representation bottleneck and improve translation performance in endangered languages. Furthermore, we design a random byte mapping method with an ensemble prediction to enhance our model robustness. Experimental results show that our BMNMT consistently and significantly outperforms subword/word-based baselines on twelve language pairs up to +18.5 BLEU points, an 840% relative improvement.
Jia Xu 0004
COLING2
2022 Discovering Latent Concepts Learned in BERT
Fahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu 0004, Hassan Sajjad 0001
ICLR5
2022 Learning by Interpreting
abstract
This paper introduces a novel way of enhancing NLP prediction accuracy by incorporating model interpretation insights. Conventional efforts often focus on balancing the trade-offs between accuracy and interpretability, for instance, sacrificing model performance to increase the explainability. Here, we take a unique approach and show that model interpretation can ultimately help improve NLP quality. Specifically, we employ our learned interpretability results using attention mechanisms, LIME, and SHAP to train our model. We demonstrate a significant increase in accuracy of up to +3.4 BLEU points on NMT and up to +4.8 points on GLUE tasks, verifying our hypothesis that it is possible to achieve better model learning by incorporating model interpretation knowledge.
Xuting Tang, Abdul Rafae Khan, Shusen Wang, Jia Xu 0004
IJCAI4
2022 Analyzing Encoded Concepts in Transformer Language Models
abstract
Hassan Sajjad, Nadir Durrani, Fahim Dalvi, Firoj Alam, Abdul Khan, Jia Xu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, Firoj Alam, Abdul Rafae Khan, Jia Xu 0004
NAACL-HLT6
2022 A clustering framework for lexical normalization of Roman Urdu
abstract
Abstract Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges during automatic language processing. In this article, we present a feature-based clustering framework for the lexical normalization of Roman Urdu corpora, which includes a phonetic algorithm UrduPhone, a string matching component, a feature-based similarity function, and a clustering algorithm Lex-Var. UrduPhone encodes Roman Urdu strings to their pronunciation-based representations. The string matching component handles character-level variations that occur when writing Urdu using Roman script. The similarity function incorporates various phonetic-based, string-based, and contextual features of words. The Lex-Var algorithm is a variant of the k-medoids clustering algorithm that groups lexical variations of words. It contains a similarity threshold to balance the number of clusters and their maximum similarity. The framework allows feature learning and optimization in addition to the use of predefined features and weights. We evaluate our framework extensively on four real-world datasets and show an F-measure gain of up to 15% from baseline methods. We also demonstrate the superiority of UrduPhone and Lex-Var in comparison to respective alternate algorithms in our clustering framework for the lexical normalization of Roman Urdu.
Abdul Rafae Khan, Asim Karim, Hassan Sajjad 0001, Faisal Kamiran, Jia Xu 0004
Nat. Lang. Eng.5
2021 Grouping Words with Semantic Diversity
abstract
Karine Chubarian, Abdul Rafae Khan, Anastasios Sidiropoulos, Jia Xu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Karine Chubarian, Abdul Rafae Khan, Anastasios Sidiropoulos, Jia Xu 0004
NAACL-HLT4
2020 Coding Textual Inputs Boosts the Accuracy of Neural Networks
abstract
Natural Language Processing (NLP) tasks are usually performed word by word on textual inputs.We can use arbitrary symbols to represent the linguistic meaning of a word and use these symbols as inputs.As "alternatives" to a text representation, we introduce Soundex, MetaPhone, NYSIIS, logogram to NLP, and develop fixed-output-length coding and its extension using Huffman coding.Each of those codings combines different character/digital sequences and constructs a new vocabulary based on codewords.We find that the integration of those codewords with text provides more reliable inputs to Neural-Networkbased NLP systems through redundancy than text-alone inputs.Experiments demonstrate that our approach outperforms the state-ofthe-art models on the application of machine translation, language modeling, and part-ofspeech tagging.The source code is available at https://github.com/abdulrafae/coding nmt.
Abdul Rafae Khan, Jia Xu 0004, Weiwei Sun 0007
EMNLP (1)2
2018 Assessing Quality Estimation Models for Sentence-Level Prediction
abstract
This paper provides an evaluation of a wide range of advanced sentence-level Quality Estimation models, including Support Vector Regression, Ride Regression, Neural Networks, Gaussian Processes, Bayesian Neural Networks, Deep Kernel Learning and Deep Gaussian Processes. Beside the accurateness, our main concerns are also the robustness of Quality Estimation models. Our work raises the difficulty in building strong models. Specifically, we show that Quality Estimation models often behave differently in Quality Estimation feature space, depending on whether the scale of feature space is small, medium or large. We also show that Quality Estimation models often behave differently in evaluation settings, depending on whether test data come from the same domain as the training data or not. Our work suggests several strong candidates to use in different circumstances.
Hoang Cuong, Jia Xu 0004
COLING2
2018 WCS: Weighted Component Stitching for Sparse Network Localization
Tianyuan Sun, Yongcai Wang, Deying Li 0001, Zhaoquan Gu, Jia Xu 0004
IEEE/ACM Trans. Netw.5
2017 Efficient Online Model Adaptation by Incremental Simplex Tableau
abstract
Online multi-kernel learning is promising in the era of mobile computing, in which a combined classifier with multiple kernels are offline trained, and online adapts to personalized features for serving the end user precisely and smartly. The online adaptation is mainly carried out at the end-devices, which requires the adaptation algorithms to be light, efficient and accurate. Previous results focused mainly on efficiency. This paper proposes an novel online model adaptation framework for not only efficiency but also optimal online adaptation. At first, an online optimal incremental simplex tableau (IST)algorithm is proposed, which approaches the model adaption by linear programming and produces the optimized model update in each step when a personalized training data is collected.But keeping online optimal in each step is expensive and may cause over-fitting especially when the online data is noisy. A Fast-IST approach is therefore proposed, which measures the deviation between the training data and the current model. It schedules updating only when enough deviation is detected. The efficiency of each update is further enhanced by running IST only limited iterations, which bounds the computation complexity. Theoretical analysis and extensive evaluations show that Fast-IST saves computation cost greatly, while achieving speedy and accurate model adaptation.It provides better model adaptation speed and accuracy while using even lower computing cost than the state-of-the art.
Zhixian Lei, Xuehan Ye, Yongcai Wang, Deying Li 0001, Jia Xu 0004
AAAI5
2016 On the Power and Limits of Distance-Based Learning
abstract
We initiate the study of low-distortion finite metric embeddings in multi-class (and multi-label) classification where (i) both the space of input instances and the space of output classes have combinatorial metric structure and (ii) the concepts we wish to learn are low-distortion embeddings. We develop new geometric techniques and prove strong learning lower bounds. These provable limits hold even when we allow learners and classifiers to get advice by one or more experts. Our study overwhelmingly indicates that post-geometry assumptions are necessary in multi-class classification, as in natural language processing (NLP). Technically, the mathematical tools we developed in this work could be of independent interest to NLP. To the best of our knowledge, this is the first work which formally studies classification problems in combinatorial spaces. and where the concepts are low-distortion embeddings.
Periklis A. Papakonstantinou, Jia Xu 0004, Guang Yang 0020
ICML2
2014 Bagging by Design (on the Suboptimality of Bagging)
abstract
Bagging (Breiman 1996) and its variants is one of the most popular methods in aggregating classifiers and regressors. Originally, its analysis assumed that the bootstraps are built from an unlimited, independent source of samples, therefore we call this form of bagging ideal-bagging. However in the real world, base predictors are trained on data subsampled from a limited number of training samples and thus they behave very differently. We analyze the effect of intersections between bootstraps, obtained by subsampling, to train different base predictors. Most importantly, we provide an alternative subsampling method called design-bagging based on a new construction of combinatorial designs, and prove it universally better than bagging. Methodologically, we succeed at this level of generality because we compare the prediction accuracy of bagging and design-bagging relative to the accuracy ideal-bagging. This finds potential applications in more involved bagging-based methods. Our analytical results are backed up by experiments on classification and regression settings.
Periklis A. Papakonstantinou, Jia Xu 0004, Zhu Cao
AAAI2
2014 An ant colony optimization method to detect communities in social networks
abstract
Community detection is an important task in social network analysis. It aims to partition the network into clusters so that interactions among members within a cluster are considerably more frequent than that across clusters. A typical instantiation is to maximize the modularity of clusters which is a NP-hard problem, and thus, heuristic and meta-heuristic algorithms are employed as approximation. We present a novel divisive algorithm based on ant colony optimization to detect hierarchical community structure by maximizing the modularity. Our algorithm splits the network into two local communities iteratively and incorporates both heuristic information and pheromone trails. Experimental results on a set of synthetic benchmarks and real-world networks verified that our algorithm is highly effective for hierarchical community structure detection.
Saeed Haji Seyed Javadi, Shahram Khadivi, Mohammad Ebrahim Shiri, Jia Xu 0004
ASONAM4
2014 Query Lattice for Translation Retrieval
Meiping Dong, Yong Cheng 0003, Yang Liu 0005, Jia Xu 0004, Maosong Sun 0001, Tatsuya Izuha
COLING4
2013 Salient object detection in image sequences via spatial-temporal cue
abstract
Contemporary video search and categorization are non-trivial tasks due to the massively increasing amount and content variety of videos. We put forward the study of visual saliency models in video. Such a model is employed to identify salient objects from the image background. Starting from the observation that motion information in video often attracts more human attention compared to static images, we devise a region contrast based saliency detection model using spatial-temporal cues (RCST). We introduce and study four saliency principles to realize the RCST. This generalizes the previous static image for saliency computational model to video. We conduct experiments on a publicly available video segmentation database where our method significantly outperforms seven state-of-the-art methods with respect to PR curve, ROC curve and visual comparison.
Chuang Gan 0001, Zengchang Qin, Jia Xu 0004, Tao Wan 0001
VCIP3
2011 Enhancing Chinese Word Segmentation Using Unlabeled Data
Weiwei Sun 0007, Jia Xu 0004
EMNLP2
2011 Generating Virtual Parallel Corpus: A Compatibility Centric Method
Jia Xu 0004, Weiwei Sun 0007
MTSummit1
2011 Parallel Corpus Refinement as an Outlier Detection Algorithm
Kaveh Taghipour, Shahram Khadivi, Jia Xu 0004
MTSummit3
2008 Phrase Table Training for Precision and Recall: What Makes a Good Phrase and a Good Phrase Pair?
Yonggang Deng, Jia Xu 0004
ACL2
2008 Bayesian Semi-Supervised Chinese Word Segmentation for Statistical Machine Translation
Jia Xu 0004, Jianfeng Gao 0001, Kristina Toutanova, Hermann Ney
COLING1
2007 Domain dependent statistical machine translation
Jia Xu 0004, Yonggang Deng, Hermann Ney
MTSummit1
2006 Error Analysis of Statistical Machine Translation Output
David Vilar, Jia Xu 0004, Luis Fernando D'Haro, Hermann Ney
LREC2
2005 Sentence segmentation using IBM word alignment model 1
Jia Xu 0004, Richard Zens, Hermann Ney
EAMT1