EDBT 2026 Demo / reviewers in the wild / expert
Yuejun Guo 0001
dblp:151/4762
· DBLP profile ↗
39ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0002-5535-2420ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorSecurity and privacy · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CRX-ray: Large-Scale Detection of API Key Leakage in Browser ExtensionsabstractThe rapid proliferation of AI-enabled browser extensions has introduced significant security vulnerabilities. These client-side applications, distributed with exposed source files, frequently embed API keys from AI platforms - credentials designed to track usage for billing and prevent misuse. The exposure of these API keys poses substantial financial and operational risks to extension developers. This study presents the first comprehensive security analysis of API key leakage in browser extensions. We systematically analyzed 163,924 extensions across Chrome, Firefox, and Edge stores, uncovering 3,677 unique leaked API keys across 4,145 extensions. Most critically we identified 300 exposed AI platform keys across 309 extensions that collectively serve 1,045,339 users. Furthermore, our analysis reveals prominent reuse of API keys across different extensions, along with instances of multiple keys being used by single extensions. In this paper, we introduce the CRX-ray detection framework to identify API key leakage in browser extensions. By open-sourcing CRX-ray, we aim to empower developers to identify and mitigate API key leakage, fostering the development of more secure browser extensions that protect both developers and users. Zhi Wang 0014, Valerio Bucci, Yuejun Guo 0001, Wanpeng Li |
AsiaCCS | 7 |
| 2026 | Developing Intelligent Chatbots for Telecom Security Support: A Comparative Study of Large Language Model Utilization Strategies
Yuejun Guo 0001, Qiang Tang 0001, Duy Cu Nguyen |
ICAART (3) | 1 |
| 2026 | Variable Renaming-Based Adversarial Test Generation for Code Model: Benchmark and EnhancementabstractRobustness testing is essential for evaluating deep learning models, particularly under unforeseen circumstances. Adversarial test generation, a fundamental approach in robustness testing, is prevalent in computer vision and natural language processing, and it has gained considerable attention in code tasks recently. The Variable Renaming-Based Adversarial Test Generation (VRTG), which deceives models by altering variable names, is a key focus. VRTG involves substitution construction and variable name searching, but its systematic design remains a challenge due to the empirical nature of these components. This article introduces the first benchmark to examine the impact of various substitutions and search algorithms on VRTG effectiveness, exploring improvements for existing VRTGs. Our benchmark includes three substitution construction types, six substitution position rank ways and seven search algorithms. Analysis of four code understanding tasks and three pre-trained code models using our benchmark reveals that combining RNNS and Genetic Algorithm with code-based substitution is more effective for VRTG construction. Notably, this method outperforms the advanced black-box variable renaming test generation technique, ALERT, by up to 22.57%. Yuejun Guo 0001, Maxime Cordy, Yves Le Traon |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2026 | CodeS+: Towards Assessing the Generalization Ability of Code Models Under Distribution Shift
Ziyue Shi, Junjie Wang 0007, Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Sen Chen 0001, Mike Papadakis, Yves Le Traon, Yongqiang Lyu 0001 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Automatic Assessment of Corporate Sustainability Due Diligence Process Capability and CS3D Compliance
Stéphane Cortina, Ruben J. Marques, Yuejun Guo 0001, Joscha Krause |
EuroSPI (1) | 3 |
| 2025 | Boosting source code learning with text-oriented data augmentation: an empirical study
Zeming Dong, Yuejun Guo 0001, Zhenya Zhang 0001, Maxime Cordy, Mike Papadakis, Yves Le Traon, Jianjun Zhao 0001 |
Empir. Softw. Eng. | 3 |
| 2025 | Assessing the Robustness of Test Selection Methods for Deep Neural NetworksabstractRegularly testing deep learning-powered systems on newly collected data is critical to ensure their reliability, robustness, and efficacy in real-world applications. This process is demanding due to the significant time and human effort required for labeling new data. While test selection methods alleviate manual labor by labeling and evaluating only a subset of data while meeting testing criteria, we observe that such methods with reported promising results are simply evaluated, e.g., testing on original test data. The question arises: are they always reliable? In this paper, we explore when and to what extent test selection methods fail. First, we identify potential pitfalls of 11 selection methods based on their construction. Second, we conduct a study to empirically confirm the existence of these pitfalls. Furthermore, we demonstrate how pitfalls can break the reliability of these methods. Concretely, methods for fault detection suffer from data that are: 1) correctly classified but uncertain, or 2) misclassified but confident. Remarkably, the test relative coverage achieved by such methods drops by up to 86.85%. Besides, methods for performance estimation are sensitive to the choice of intermediate-layer output. The effectiveness of such methods can be even worse than random selection when using an inappropriate layer. Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Wei Ma 0014, Mike Papadakis, Lei Ma 0003, Yves Le Traon |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | Poster: Automated Dependency Mapping for Web API Security Testing Using Large Language ModelsabstractDependency extraction is crucial in web API security testing, as it helps identify the required API sequences to exploit a vulnerability. Traditional methods are generally rule-based and require extensive manual analysis of API specification documents by domain experts to formulate appropriate rules. This manual process is not only time-consuming and labor-intensive but also prone to missing dependencies and inaccuracies, which can compromise the effectiveness of security testing. In this paper, we explore the potential of large language models (LLMs) to automate dependency mapping in web APIs. By leveraging the capabilities of advanced LLMs such as GPT-3.5, Mistral-7B-Instruct, and Llama-3-8B-Instruct, which include understanding and generating natural language, we aim to streamline the dependency mapping process, reducing the need for manual analysis and enhancing accuracy. Our preliminary experiments demonstrate that this approach can effectively build dependency mappings, offering a a promising alternative to traditional rule-based approaches. Wanpeng Li, Yuejun Guo 0001 |
CCS | 2 |
| 2024 | Outside the Comfort Zone: Analysing LLM Capabilities in Software Vulnerability Detection
Yuejun Guo 0001, Constantinos Patsakis, Qiang Tang 0001, Fran Casino |
ESORICS (1) | 1 |
| 2024 | Towards Exploring the Limitations of Test Selection Techniques on Graph Neural Networks: An Empirical StudyabstractAbstract Graph Neural Networks (GNNs) have gained prominence in various domains, such as social network analysis, recommendation systems, and drug discovery, due to their ability to model complex relationships in graph-structured data. GNNs can exhibit incorrect behavior, resulting in severe consequences. Therefore, testing is necessary and pivotal. However, labeling all test inputs for GNNs can be prohibitively costly and time-consuming, especially when dealing with large and complex graphs. In response to these challenges, test selection has emerged as a strategic approach to alleviate labeling expenses. The objective of test selection is to select a subset of tests from the complete test set. While various test selection techniques have been proposed for traditional deep neural networks (DNNs), their adaptation to GNNs presents unique challenges due to the distinctions between DNN and GNN test data. Specifically, DNN test inputs are independent of each other, whereas GNN test inputs (nodes) exhibit intricate interdependencies. Therefore, it remains unclear whether DNN test selection approaches can perform effectively on GNNs. To fill the gap, we conduct an empirical study that systematically evaluates the effectiveness of various test selection methods in the context of GNNs, focusing on three critical aspects: 1) Misclassification detection : selecting test inputs that are more likely to be misclassified; 2) Accuracy estimation : selecting a small set of tests to precisely estimate the accuracy of the whole testing set; 3) Performance enhancement : selecting retraining inputs to improve the GNN accuracy. Our empirical study encompasses 7 graph datasets and 8 GNN models, evaluating 22 test selection approaches. Our study includes not only node classification datasets but also graph classification datasets. Our findings reveal that: 1) In GNN misclassification detection, confidence-based test selection methods, which perform well in DNNs, do not demonstrate the same level of effectiveness; 2) In terms of GNN accuracy estimation, clustering-based methods, while consistently performing better than random selection, provide only slight improvements; 3) Regarding selecting inputs for GNN performance improvement, test selection methods, such as confidence-based and clustering-based test selection methods, demonstrate only slight effectiveness; 4) Concerning performance enhancement, node importance-based test selection methods are not suitable, and in many cases, they even perform worse than random selection. Xueqi Dang, Wei Ma 0014, Yuejun Guo 0001, Mike Papadakis, Maxime Cordy, Yves Le Traon |
Empir. Softw. Eng. | 4 |
| 2024 | On the effectiveness of hybrid pooling in mixup-based graph learning for language processing
Zeming Dong, Zhenya Zhang 0001, Yuejun Guo 0001, Maxime Cordy, Mike Papadakis, Yves Le Traon, Jianjun Zhao 0001 |
J. Syst. Softw. | 4 |
| 2024 | KAPE: kNN-based Performance Testing for Deep Code SearchabstractCode search is a common yet important activity of software developers. An efficient code search model can largely facilitate the development process and improve the programming quality. Given the superb performance of learning the contextual representations, deep learning models, especially pre-trained language models, have been widely explored for the code search task. However, studies mainly focus on proposing new architectures for ever-better performance on designed test sets but ignore the performance on unseen test data where only natural language queries are available. The same problem in other domains, e.g., CV and NLP, is usually solved by test input selection that uses a subset of the unseen set to reduce the labeling effort. However, approaches from other domains are not directly applicable and still require labeling effort. In this article, we propose the k NN-b a sed p erformance t e sting ( KAPE ) to efficiently solve the problem without manually matching code snippets to test queries. The main idea is to use semantically similar training data to perform the evaluation. Extensive experiments on six programming language datasets, three state-of-the-art pre-trained models, and seven baseline methods demonstrate that KAPE can effectively assess the model performance (e.g., CodeBERT achieves MRR 0.5795 on JavaScript) with a slight difference (e.g., 0.0261). Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Yves Le Traon |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | Test Optimization in DNN Testing: A SurveyabstractThis article presents a comprehensive survey on test optimization in deep neural network (DNN) testing. Here, test optimization refers to testing with low data labeling effort. We analyzed 90 papers, including 43 from the software engineering (SE) community, 32 from the machine learning (ML) community, and 15 from other communities. Our study: (i) unifies the problems as well as terminologies associated with low-labeling cost testing, (ii) compares the distinct focal points of SE and ML communities, and (iii) reveals the pitfalls in existing literature. Furthermore, we highlight the research opportunities in this domain. Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Lei Ma 0003, Mike Papadakis, Yves Le Traon |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | LaF: Labeling-free Model Selection for Automated Deep Neural Network ReusingabstractApplying deep learning (DL) to science is a new trend in recent years, which leads DL engineering to become an important problem. Although training data preparation, model architecture design, and model training are the normal processes to build DL models, all of them are complex and costly. Therefore, reusing the open-sourced pre-trained model is a practical way to bypass this hurdle for developers. Given a specific task, developers can collect massive pre-trained deep neural networks from public sources for reusing. However, testing the performance (e.g., accuracy and robustness) of multiple deep neural networks (DNNs) and recommending which model should be used is challenging regarding the scarcity of labeled data and the demand for domain expertise. In this article, we propose a labeling-free (LaF) model selection approach to overcome the limitations of labeling efforts for automated model reusing. The main idea is to statistically learn a Bayesian model to infer the models’ specialty only based on predicted labels. We evaluate LaF using nine benchmark datasets, including image, text, and source code, and 165 DNNs, considering both the accuracy and robustness of models. The experimental results demonstrate that LaF outperforms the baseline methods by up to 0.74 and 0.53 on Spearman’s correlation and Kendall’s τ, respectively. Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Yves Le Traon |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | Active Code Learning: Benchmarking Sample-Efficient Training of Code ModelsabstractThe costly human effort required to prepare the training data of machine learning (ML) models hinders their practical development and usage in software engineering (ML4Code), especially for those with limited budgets. Therefore, efficiently training models of code with less human effort has become an emergent problem. Active learning is such a technique to address this issue that allows developers to train a model with reduced data while producing models with desired performance, which has been well studied in computer vision and natural language processing domains. Unfortunately, there is no such work that explores the effectiveness of active learning for code models. In this paper, we bridge this gap by building the first benchmark to study this critical problem - active code learning. Specifically, we collect 11 acquisition functions (which are used for data selection in active learning) from existing works and adapt them for code-related tasks. Then, we conduct an empirical study to check whether these acquisition functions maintain performance for code data. The results demonstrate that feature selection highly affects active learning and using output vectors to select data is the best choice. For the code summarization task, active code learning is ineffective which produces models with over a 29.64% gap compared to the expected performance. Furthermore, we explore future directions of active code learning with an exploratory study. We propose to replace distance calculation methods with evaluation metrics and find a correlation between these evaluation-based distance methods and the performance of code models. Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Lei Ma 0003, Mike Papadakis, Yves Le Traon |
IEEE Trans. Software Eng. | 2 |
| 2023 | Towards Understanding Model Quantization for Reliable Deep Neural Network DeploymentabstractDeep Neural Networks (DNNs) have gained considerable attention in the past decades due to their astounding performance in different applications, such as natural language modeling, self-driving assistance, and source code understanding. With rapid exploration, more and more complex DNN architectures have been proposed along with huge pre-trained model parameters. A common way to use such DNN models in user-friendly devices (e.g., mobile phones) is to perform model compression before deployment. However, recent research has demonstrated that model compression, e.g., model quantization, yields accuracy degradation as well as output disagreements when tested on unseen data. Since the unseen data always include distribution shifts and often appear in the wild, the quality and reliability of models after quantization are not ensured. In this paper, we conduct a comprehensive study to characterize and help users understand the behaviors of quantization models. Our study considers four datasets spanning from image to text, eight DNN architectures including both feed-forward neural networks and recurrent neural networks, and 42 shifted sets with both synthetic and natural distribution shifts. The results reveal that 1) data with distribution shifts lead to more disagreements than without. 2) Quantization-aware training can produce more stable models than standard, adversarial, and Mixup training. 3) Disagreements often have closer top-1 and top-2 output probabilities, and Margin is a better indicator than other uncertainty metrics to distinguish disagreements. 4) Retraining the model with disagreements has limited efficiency in removing disagreements. We release our code and models as a new benchmark for further study of model quantization. Yuejun Guo 0001, Maxime Cordy, Xiaofei Xie, Wei Ma 0014, Mike Papadakis, Yves Le Traon |
CAIN | 2 |
| 2023 | An Empirical Study of the Imbalance Issue in Software Vulnerability Detection
Yuejun Guo 0001, Qiang Tang 0001, Yves Le Traon |
ESORICS (4) | 1 |
| 2023 | Aries: Efficient Testing of Deep Neural Networks via Labeling-Free Accuracy EstimationabstractDeep learning (DL) plays a more and more important role in our daily life due to its competitive performance in industrial application domains. As the core of DL-enabled systems, deep neural networks (DNNs) need to be carefully evaluated to ensure the produced models match the expected requirements. In practice, the de facto standard to assess the quality of DNNs in the industry is to check their performance (accuracy) on a collected set of labeled test data. However, preparing such labeled data is often not easy partly because of the huge labeling effort, i.e., data labeling is labor-intensive, especially with the massive new incoming unlabeled data every day. Recent studies show that test selection for DNN is a promising direction that tackles this issue by selecting minimal representative data to label and using these data to assess the model. However, it still requires human effort and cannot be automatic. In this paper, we propose a novel technique, named Aries, that can estimate the performance of DNNs on new unlabeled data using only the information obtained from the original test data. The key insight behind our technique is that the model should have similar prediction accuracy on the data which have similar distances to the decision boundary. We performed a large-scale evaluation of our technique on two famous datasets, CIFAR-10 and Tiny-ImageNet, four widely studied DNN models including ResNetl0l and DenseNetl21, and 13 types of data transformation methods. Results show that the estimated accuracy by Aries is only 0.03% - 2.60% off the true accuracy. Besides, Aries also outperforms the state-of-the-art labeling-free methods in 50 out of 52 cases and selection-labeling-based methods in 96 out of 128 cases. Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Lei Ma 0003, Yves Le Traon |
ICSE | 2 |
| 2023 | MUTEN: Mutant-Based Ensembles for Boosting Gradient-Based Adversarial AttackabstractMutation testing (MT) for deep learning (DL) has gained huge attention in the past few years. However, how MT can really help DL is still unclear. In this paper, we introduce one promising direction for the usage of mutants. Specifically, since mutants can be seen as one kind of ensemble model and ensemble model can be used to boost the adversarial attack, we propose MUTEN, which applies the attack on mutants to improve the success rate of well-known attacks against gradient-masking models. Experimental results on MNIST, SVHN, and CIFAR-10 show that MUTEN can increase the success rate of four attacks by up to 45%. Furthermore, experiments on four defense approaches, bit-depth reduction, JPEG compression, Defensive distillation, and Label smoothing, demonstrate that MUTEN can break the defense models effectively by enhancing the attacks with the success rate of up to 96%. Yuejun Guo 0001, Maxime Cordy, Mike Papadakis, Yves Le Traon |
ASE | 2 |
| 2023 | MixCode: Enhancing Code Classification by Mixup-Based Data AugmentationabstractInspired by the great success of Deep Neural Networks (DNNs) in natural language processing (NLP), DNNs have been increasingly applied in source code analysis and attracted significant attention from the software engineering community. Due to its data-driven nature, a DNN model requires massive and high-quality labeled training data to achieve expert-level performance. Collecting such data is often not hard, but the labeling process is notoriously laborious. The task of DNN-based code analysis even worsens the situation because source code labeling also demands sophisticated expertise. Data augmentation has been a popular approach to supplement training data in domains such as computer vision and NLP. However, existing data augmentation approaches in code analysis adopt simple methods, such as data transformation and adversarial example generation, thus bringing limited performance superiority. In this paper, we propose a data augmentation approach MixCode that aims to effectively supplement valid training data, inspired by the recent advance named Mixup in computer vision. Specifically, we first utilize multiple code refactoring methods to generate transformed code that holds consistent labels with the original data. Then, we adapt the Mixup technique to mix the original code with the transformed code to augment the training data. We evaluate MixCode on two programming languages (Java and Python), two code tasks (problem classification and bug detection), four benchmark datasets (JAVA250, Python800, CodRep1, and Refactory), and seven model architectures (including two pretrained models CodeBERT and GraphCodeBERT). Experimental results demonstrate that MixCode outperforms the baseline data augmentation approach by up to 6.24% in accuracy and 26.06% in robustness. Zeming Dong, Yuejun Guo 0001, Maxime Cordy, Mike Papadakis, Zhenya Zhang 0001, Yves Le Traon, Jianjun Zhao 0001 |
SANER | 3 |
| 2023 | DRE: density-based data selection with entropy for adversarial-robust deep learning modelsabstractAbstract Active learning helps software developers reduce the labeling cost when building high-quality machine learning models. A core component of active learning is the acquisition function that determines which data should be selected to annotate.State-of-the-art (SOTA) acquisition functions focus on clean performance (e.g. accuracy) but disregard robustness (an important quality property), leading to fragile models with negligible robustness (less than 0.20%). In this paper, we first propose to integrate adversarial training into active learning (adversarial-robust active learning, ARAL) to produce robust models. Our empirical study on 11 acquisition functions and 15105 trained deep neural networks (DNNs) shows that ARAL can produce models with robustness ranging from 2.35% to 63.85%. Our study also reveals, however, that the acquisition functions that perform well on accuracy are worse than random sampling when it comes to robustness. Via examining the reasons behind this, we devise the density-based robust sampling with entropy (DRE) to target both clean performance and robustness. The core idea of DRE is to maintain a balance between selected data and the entire set based on the entropy density distribution. DRE outperforms SOTA functions in terms of robustness by up to 24.40%, while remaining competitive on accuracy. Additionally, the in-depth evaluation shows that DRE is applicable as a test selection metric for model retraining and stands out from all compared functions by up to 8.21% robustness. Yuejun Guo 0001, Maxime Cordy, Michail Papadakis, Yves Le Traon |
Neural Comput. Appl. | 1 |
| 2022 | Robust active learning: sample-efficient training of robust deep learning modelsabstractActive learning is an established technique to reduce the labeling cost for building high-quality machine learning models. However, state-of-the-art approaches focus on maximizing the clean performance (e.g. accuracy) but disregarding robustness. In this work, we propose Robust Active Learning, an active learning process that integrates adversarial training, the most established method to produce robust models. First, we conduct an empirical study to evaluate the effectiveness of existing approaches and uncover the characteristics of data. Then, we propose a novel approach, density-based robust sampling with entropy (DRE), to target both clean performance and robustness. Our experiments are conducted on 11 acquisition functions, 4 datasets, 6 DNN architectures, and 15105 trained DNNs. Yuejun Guo 0001, Maxime Cordy, Mike Papadakis, Yves Le Traon |
CAIN | 1 |
| 2022 | An Empirical Study on Data Distribution-Aware Test Selection for Deep Learning EnhancementabstractSimilar to traditional software that is constantly under evolution, deep neural networks need to evolve upon the rapid growth of test data for continuous enhancement (e.g., adapting to distribution shift in a new environment for deployment). However, it is labor intensive to manually label all of the collected test data. Test selection solves this problem by strategically choosing a small set to label. Via retraining with the selected set, deep neural networks will achieve competitive accuracy. Unfortunately, existing selection metrics involve three main limitations: (1) using different retraining processes, (2) ignoring data distribution shifts, and (3) being insufficiently evaluated. To fill this gap, we first conduct a systemically empirical study to reveal the impact of the retraining process and data distribution on model enhancement. Then based on our findings, we propose DAT, a novel distribution-aware test selection metric. Experimental results reveal that retraining using both the training and selected data outperforms using only the selected data. None of the selection metrics perform the best under various data distributions. By contrast, DAT effectively alleviates the impact of distribution shifts and outperforms the compared metrics by up to five times and 30.09% accuracy improvement for model enhancement on simulated and in-the-wild distribution shift scenarios, respectively. Yuejun Guo 0001, Maxime Cordy, Xiaofei Xie, Lei Ma 0003, Mike Papadakis, Yves Le Traon |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2021 | Towards Exploring the Limitations of Active Learning: An Empirical StudyabstractDeep neural networks (DNNs) are increasingly deployed as integral parts of software systems. However, due to the complex interconnections among hidden layers and massive hyperparameters, DNNs must be trained using a large number of labeled inputs, which calls for extensive human effort for collecting and labeling data. Spontaneously, to alleviate this growing demand, multiple state-of-the-art studies have developed different metrics to select a small yet informative dataset for the model training. These research works have demonstrated that DNN models can achieve competitive performance using a carefully selected small set of data. However, the literature lacks proper investigation of the limitations of data selection metrics, which is crucial to apply them in practice. In this paper, we fill this gap and conduct an extensive empirical study to explore the limits of data selection metrics. Our study involves 15 data selection metrics evaluated over 5 datasets (2 image classification tasks and 3 text classification tasks), 10 DNN architectures, and 20 labeling budgets (ratio of training data being labeled). Our findings reveal that, while data selection metrics are usually effective in producing accurate models, they may induce a loss of model robustness (against adversarial examples) and resilience to compression. Overall, we demonstrate the existence of a trade-off between labeling effort and different model qualities. This paves the way for future research in devising data selection metrics considering multiple quality criteria. Yuejun Guo 0001, Maxime Cordy, Xiaofei Xie, Wei Ma 0014, Mike Papadakis, Yves Le Traon |
ASE | 2 |
| 2021 | A scalable method to construct compact road networks from GPS trajectoriesabstractThe automatic generation of road networks from GPS tracks is a challenging problem that has been receiving considerable attention in the last years. Although dozens of methods have been proposed, current techniques suffer from two main shortcomings: the quality of the produced road networks is still far from those produced manually, and the methods are slow, making them not scalable to large inputs. In this paper, we present a fast four-step density-based approach to construct a road network from a set of trajectories. A key aspect of our method is the use of an improved version of the Slide method to adjust trajectories to build a more compact density surface. The network has comparable or better quality than that of state-of-the-art methods and is simpler (includes fewer nodes and edges). Furthermore, we also propose a split-and-merge strategy that allows splitting the data domain into smaller regions that can be processed independently, making the method scalable to large inputs. The performance of our method is evaluated with extensive experiments on urban and hiking data. Yuejun Guo 0001, Anton Bardera, Marta Fort, Rodrigo I. Silveira |
Int. J. Geogr. Inf. Sci. | 1 |
| 2020 | Trajectory Anomaly Detection Based on the Mean Distance Deviation
Xiaoyuan Hu, Qing Xu 0002, Yuejun Guo 0001 |
ICONIP (4) | 3 |
| 2019 | Eye Movement-Based Analysis on Methodologies and Efficiency in the Process of Image Noise Evaluation
Qing Xu 0002, Yuejun Guo 0001, Klaus Schöffmann |
ICANN (3) | 3 |
| 2019 | Anomaly Detection Based on the Global-Local Anomaly Score for Trajectory Data
Qing Xu 0002, Yuejun Guo 0001 |
ICONIP (5) | 4 |
| 2019 | Detail-Preserving Trajectory Summarization Based on Segmentation and Group-Based Filtering
Qing Xu 0002, Yuejun Guo 0001, Klaus Schöffmann |
MMM (2) | 4 |
| 2018 | A group-based signal filtering approach for trajectory abstraction and restoration
Yuejun Guo 0001, Qing Xu 0002, Xiaoxiao Luo, Hongjuan Bu, Mateu Sbert |
Neural Comput. Appl. | 1 |
| 2017 | The Abstraction for Trajectories with Different Numbers of Sampling Points
Peng Li 0042, Qing Xu 0002, Yuejun Guo 0001, Xiaoxiao Luo, Mateu Sbert |
ICONIP (6) | 4 |
| 2016 | Fast Agglomerative Information Bottleneck Based Trajectory Clustering
Yuejun Guo 0001, Qing Xu 0002, Sheng Liang, Mateu Sbert |
ICONIP (3) | 1 |
| 2015 | A mixed noise removal algorithm based on the maximum entropy principleabstractThis paper proposes a novel method for image denoising based on the maximum entropy principle. For an image corrupted by Gaussian and impulse noise, impulse noise with high value is detected first using rank statistics, and then removed by a local Gaussian filter. To remove the noise left, an improved self-adaptive non-local filter, with the weights obtained based on the maximum entropy principle, is performed. The experimental results demonstrate that our approach has significant ability to remove any mix of Gaussian and impulse noise in terms of quantitatively image evaluation and qualitatively visual effect. Shang Wu 0002, Qing Xu 0002, Jialang Li, Yuejun Guo 0001 |
ICME | 4 |
| 2015 | XaIBO: An Extension of aIB for Trajectory Clustering with Outlier
Yuejun Guo 0001, Qing Xu 0002, Sheng Liang, Mateu Sbert |
ICONIP (2) | 1 |
| 2015 | Trajectory Abstracting with Group-Based Signal Denoising
Xiaoxiao Luo, Qing Xu 0002, Yuejun Guo 0001, Yimin Lv |
ICONIP (3) | 3 |
| 2015 | Visualization on Agglomerative Information Bottleneck Based Trajectory ClusteringabstractUndoubtedly, visualization of the trajectory clustering outputs is very important and some researches have been done on visualization of the clustering results. Still importantly, the research on visualizing the procedure of clustering, which is also of great value, is little touched. In this paper, we propose a novel 3D visualization tool, which comprehensively illustrates the Agglomerative Information Bottleneck (AIB) based clustering scheme, to help users understand the clustering approach vividly and clearly. The point of the proposed metaphor makes use of the visualization, together with rich interactions, to demonstrate the iterative clustering procedure, the corresponding results and the clustering results. The experiment demonstrates the effectiveness of our 3D visualization tool for trajectory analysis. Qing Xu 0002, Yuejun Guo 0001, Sheng Liang |
IV | 3 |
| 2015 | Multiscale Visualization of Trajectory DataabstractThis paper proposes a novel three-dimensional (3D) visualization tool for the trajectory analysis, helping users understand the trajectory data from different perspectives. The details of a single and a set trajectories are well covered in multiscale views by the four main linked windows, namely Traj View, Color Bar, Multi Property and Track Map. We take advantage of and further improve the color bar and the parallel coordinates to effectively present the important attributes of trajectories and also the relationship between different attributes. In addition, the interactive actions, such as keyboard and mouse operations, provide a rich and wonderful user experience. Sheng Liang, Qing Xu 0002, Yuejun Guo 0001 |
IV | 3 |
| 2015 | 3D Visualization of Multiscale Video Key FramesabstractIn this paper, an innovative 3D visualization tool is proposed to facilitate the quickly browsing and understanding of the video sequence for users. Taking advantage of the major windows, our tool presents the multiscale key frames and the video content clearly and effectively. Namely, KF View provides a wonderful navigation of the video key frames with different levels of details. Frame View presents an interesting view of the whole video content. SIM View allows an expressive exploration of the similarities between key frames and also between key frames and the video frames. Importantly, together with many convenient and attractive interactions, this tool is quite efficient to help users grasp the video information soundly. Shihua Sun, Qing Xu 0002, Yuejun Guo 0001, Sheng Liang |
IV | 3 |
| 2014 | A New Scheme for Trajectory VisualizationabstractThis paper proposes a new approach for trajectory clustering and visualization. During clustering, speed and direction are both taken into account as the important properties of a trajectory. Additionally, the temporal change of the trajectories is explicitly considered in the procedure of clustering. Importantly, we introduce the uncertainty to represent the homogeneity of the updating trajectories based on their speed and direction attributes, to help observe and understand the happening of the trajectory data. Finally, based on the use of the HSV color space, we devise three patterns for graphically presenting the trajectory data in different perspectives. Namely, 'H+S' visualizes the clustered trajectories emphasizing the speed attribute of them. 'H+S+Update' depicts more information on the "latest" trajectory data, compared with the results by 'H+S'. 'H+V+Update' explicitly presents the uncertainty of the trajectory data. Our approach has shown its potential for the effective trajectory analysis. Yuejun Guo 0001, Qing Xu 0002, Xiu Li 0005, Xiaoxiao Luo, Mateu Sbert |
IV | 1 |