VLDB 2026 Research / reviewers in the wild / expert
Binbin Xu 0002
dblp:20/3602-2
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-0822-5250ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GUing: A Mobile GUI Search Engine using a Vision-Language ModelabstractGraphical User Interfaces (GUIs) are central to app development projects. App developers may use the GUIs of other apps as a means of requirements refinement and rapid prototyping or as a source of inspiration for designing and improving their own apps. Recent research has thus suggested retrieving relevant GUI designs that match a certain text query from screenshot datasets acquired through crowdsourced or automated exploration of GUIs. However, such text-to-GUI retrieval approaches only leverage the textual information of the GUI elements, neglecting visual information such as icons or background images. In addition, retrieved screenshots are not steered by app developers and lack app features that require particular input data. To overcome these limitations, this article proposes GUing, a GUI search engine based on a vision-language model called GUIClip, which we trained specifically for the problem of designing app GUIs. For this, we first collected from Google Play app introduction images which display the most representative screenshots and are often captioned (i.e., labeled) by app vendors. Then, we developed an automated pipeline to classify, crop, and extract the captions from these images. This resulted in a large dataset which we share with this article: including 303k app screenshots, out of which 135k have captions. We used this dataset to train a novel vision-language model, which is, to the best of our knowledge, the first of its kind for GUI retrieval. We evaluated our approach on various datasets from related work and in a manual experiment. The results demonstrate that our model outperforms previous approaches in text-to-GUI retrieval achieving a Recall@10 of up to 0.69 and a HIT@10 of 0.91. We also explored the performance of GUIClip for other GUI tasks, including GUI classification and sketch-to-GUI retrieval with encouraging results. Jialiang Wei, Anne-Lise Courbis, Thomas Lambolais, Binbin Xu 0002, Pierre-Louis Bernard, Gérard Dray, Walid Maalej |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | How Does a Single EEG Channel Tell Us About Brain States in Brain-Computer Interfaces?abstractOver recent decades, neuroimaging tools, partic-ularly electroencephalography (EEG), have revolutionized our understanding of the brain and its functions. EEG is extensively used in traditional brain-computer interface (BCI) systems due to its low cost, non-invasiveness, and high temporal resolution. This makes it invaluable for identifying different brain states relevant to both medical and non-medical applications. Although this practice is widely recognized, current methods are mainly confined to lab or clinical environments because they rely on data from multiple EEG electrodes covering the entire head. Nonethe-less, a significant advancement for these applications would be their adaptation for “real-world” use, using portable devices with a single-channel. In this study, we tackle this challenge through two distinct strategies: the first approach involves training models with data from multiple channels and then testing new trials on data from a single channel individually. The second method focuses on training with data from a single channel and then testing the performances of the models on data from all the other channels individually. To efficiently classify cognitive tasks from EEG data, we propose Convolutional Neural Networks (CNNs) with only a few parameters and fast learnable spectral-temporal features. We demonstrated the feasibility of these approaches on EEG data recorded during mental arithmetic and motor imagery tasks from three datasets. We achieved the highest accuracies of 100%, 91.55% and 73.45% in binary and 3-class classification on specific channels across three datasets. This study can contribute to the development of single-channel BCI and provides a robust EEG biomarker for brain states classification. Zaineb Ajra, Binbin Xu 0002, Gérard Dray, Jacky Montmain, Stéphane Perrey |
HSI | 2 |
| 2024 | Scale Factor Effects on Relative Radiometric Normalization between Multi-Resolution ImagesabstractWhen performing relative radiometric normalization (RRN) on higher spatial resolution images using a reference image with lower spatial resolution, it is common to downsample the high-resolution (HR) target image to match the spatial resolution of the low-resolution (LR) reference image for the transfer function estimation. In this study, we downsampled HR target images by pixel aggregation using different scale factors, generated transfer functions with LR reference images, and compared the similarity between the corrected images and the reference images. We examined two commonly used RRN methods that do not require pixel pairing: moment matching (MM) and histogram matching (HM). The evaluation metrics used for the comparison included the root mean square error (RMSE), the cumulative distribution function distance (CDFD) and the structural similarity index (SSIM). We found that variations in the scale factor can impact the accuracy of the multi-resolution image RRN. There exists an optimal scale factor value that results in the corrected high-resolution image being closest to the low-resolution reference image. Moreover, this optimal value changes with different evaluation metrics. Using the optimal scale factor does not increase computational costs, but does lead to improved correction results. Manchun Lei, Binbin Xu 0002 |
IGARSS | 2 |
| 2024 | A Deep Learning Method for Radiometric Harmonization of Non-Overlapping Remote Sensing ImagesabstractConventional relative radiometric normalization (RRN) methods establish mapping relationships by using overlapping areas between images to achieve radiometric alignment between them. However, these methods become inapplicable when stitching weakly overlapping or non-overlapping images. We propose a novel radiometric harmonization method that addresses radiometric alignment as a style transfer problem using the CycleGAN, a Generative Adversarial Network architecture. We use two non-overlapping image sets to train the model, and the trained model can perform style transfer between the target image set and the non-overlapping image set. The corrected image closely approximates the conventional RRN result using the real reference image (overlapping with the target image), and is significantly better than the conventional RRN results obtained using the non-overlapping pseudo-reference image. Binbin Xu 0002, Manchun Lei |
IGARSS | 1 |
| 2024 | Getting Inspiration for Feature Elicitation: App Store- vs. LLM-based ApproachabstractOver the past decade, app store (AppStore)-inspired requirements elicitation has proven to be highly beneficial. Developers often explore competitors' apps to gather inspiration for new features. With the advance of Generative AI, recent studies have demonstrated the potential of large language model (LLM)-inspired requirements elicitation. LLMs can assist in this process by providing inspiration for new feature ideas. While both approaches are gaining popularity in practice, there is a lack of insight into their differences. We report on a comparative study between AppStore- and LLM-based approaches for refining features into sub-features. By manually analyzing 1,200 sub-features recommended from both approaches, we identified their benefits, challenges, and key differences. While both approaches recommend highly relevant sub-features with clear descriptions, LLMs seem more powerful particularly concerning novel unseen app scopes. Moreover, some recommended features are imaginary with unclear feasibility, which suggests the importance of a human-analyst in the elicitation loop. Jialiang Wei, Anne-Lise Courbis, Thomas Lambolais, Binbin Xu 0002, Pierre-Louis Bernard, Gérard Dray, Walid Maalej |
ASE | 4 |
| 2023 | Zero-shot Bilingual App Reviews Mining with Large Language ModelsabstractApp reviews from app stores are crucial for improving software requirements. A large number of valuable reviews are continually being posted, describing software problems and expected features. Effectively utilizing user reviews necessitates the extraction of relevant information, as well as their subsequent summarization. Due to the substantial volume of user reviews, manual analysis is arduous. Various approaches based on natural language processing (NLP) have been proposed for automatic user review mining. However, the majority of them requires a manually crafted dataset to train their models, which limits their usage in real-world scenarios. In this work, we propose Mini-BAR, a tool that integrates large language models (LLMs) to perform zero-shot mining of user reviews in both English and French. Specifically, Mini-BAR is designed to (i) classify the user reviews, (ii) cluster similar reviews together, (iii) generate an abstractive summary for each cluster and (iv) rank the user review clusters. To evaluate the performance of Mini-BAR, we created a dataset containing 6,000 English and 6,000 French annotated user reviews and conducted extensive experiments. Preliminary results demonstrate the effectiveness and efficiency of Mini-BAR in requirement engineering by analyzing bilingual app reviews. Jialiang Wei, Anne-Lise Courbis, Thomas Lambolais, Binbin Xu 0002, Pierre-Louis Bernard, Gérard Dray |
ICTAI | 4 |
| 2023 | Boosting GUI Prototyping with Diffusion ModelsabstractGUI (graphical user interface) prototyping is a widely-used technique in requirements engineering for gathering and refining requirements, reducing development risks and increasing stakeholder engagement. However, GUI prototyping can be a time-consuming and costly process. In recent years, deep learning models such as Stable Diffusion have emerged as a powerful text-to-image tool capable of generating detailed images based on text prompts. In this paper, we propose UI-Diffuser, an approach that leverages Stable Diffusion to generate mobile UIs through simple textual descriptions and UI components. Preliminary results show that UI-Diffuser provides an efficient and cost-effective way to generate mobile GUI designs while reducing the need for extensive prototyping efforts. This approach has the potential to significantly improve the speed and efficiency of GUI prototyping in requirements engineering. Jialiang Wei, Anne-Lise Courbis, Thomas Lambolais, Binbin Xu 0002, Pierre-Louis Bernard, Gérard Dray |
RE | 4 |
| 2021 | A Multi-scale Line Feature Detection Using Second Order Semi-Gaussian Filters
Baptiste Magnier, Ghulam Sakhi Shokouh, Binbin Xu 0002, Philippe Montesinos |
CAIP (2) | 3 |