EDBT 2026 Demo / reviewers in the wild / expert
Thi-Lan Le
dblp:37/3213
· DBLP profile ↗
27ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-9541-3905ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Consistency: Explicit Boundary Learning for Semi-supervised Ovarian Tumor Segmentation
Minh-Khoa Vu, Hoang-Son Bui, Thi-Lan Le |
ICPR (2) | 3 |
| 2026 | FedSNC: Federated noise label learning with client similarity network and label correction
Thanh-Trung Giap, Thi-Lan Le, Trung Kien Tran, Thanh-Hai Tran 0001 |
Neurocomputing | 2 |
| 2025 | FedDC: Label Noise Correction With Dynamic Clients for Federated LearningabstractFederated learning (FL) is a distributed machine learning training paradigm that protects user privacy by training with user data stored on a local device called a client. In realworld FL systems, the number of clients often varies due to both external and internal factors. New clients may join the system right during the training process. Moreover, existing and new clients can have noisy labels with varying levels of noise. Although some works have been proposed to solve the noise problem in FL with the clients that are fixed and determined prior. These methods have not considered the case when new clients join. Therefore, there is a need for a framework that can handle dynamically noisy clients. In this article, we introduce a new framework namedFedDCto handle noisy labels in FL systems with dynamic clients. OurFedDCframework is built on top of the 3-stageFedCorrframework which has been designed to work with a fixed number of clients. In our framework, existing noisy clients will be identified through local intrinsic dimensionality (LID) scores. Then to identify new noisy clients, we use a loss threshold combined with the LID scoring technique in the first stage and with only the loss threshold in the second stage. Our experiments on three benchmark datasets that are CIFAR-10 and CIFAR-100 with independent and identically distributed/nonindependent and nonidentically distributed data partition and a real-world noisy dataset, Clothing1M, demonstrate thatFedDChelps mitigate the negative impact of new noisy clients and achieves outperformed accuracy compared toFedCorr. Our code is made available at:https://github.com/gttrung/FedDC Thanh-Trung Giap, Tuan-Dung Kieu, Thi-Lan Le, Thanh-Hai Tran 0001 |
IEEE Internet Things J. | 3 |
| 2025 | Text line segmentation approach combining deep learning model and traditional image processing techniques - application to transliteration of Cham manuscripts
Tien-Nam Nguyen, Jean-Christophe Burie, Thi-Lan Le, Anne-Valérie Schweyer |
Multim. Tools Appl. | 3 |
| 2025 | Towards an Online Text-Based Person Search in Vietnamese LanguageabstractIn recent years, many efforts have been dedicated to text-based person search, thanks to its potential applications in various domains. However, most of these works focus on person search via queries in English and conduct offline evaluations. Despite some promising results for text-based person search in English, several challenges still prevent its widespread use in practical situations when deployed in minor languages. This article extends person search to the Vietnamese language. In terms of linguistics, English and Vietnamese belong to two different language families. In addition to the difference in vocabulary, these two languages also have opposite word structures and syntactic structures. The contributions of the article are twofold. First, based on the network architecture of the ViTAA model [Wang et al. 2020 ], a framework for person search through Vietnamese queries has been developed. In this framework, to take into account specific characteristics of the Vietnamese language, the word-tokenizing, Parts of Speech (PoS) tagging techniques of different natural language processing tools, including Underthesea [UndertheseaNLP 2018 ], SEACoreNLP [Singapore 2021 ], and PhoNLP [Nguyen and Nguyen 2021 ], have been investigated in order to extract language elements from Vietnamese descriptions. Our investigation shows that selecting a suitable preprocessing technique can improve person search performance by 1.28% at R@1. When incorporating these preprocessing techniques with the person search model, the best accuracy was achieved with 27.08%, 51.38%, and 63.00% at rank 1, rank 5, and rank 10 respectively. Second, for the first time, an online evaluation of person search through natural language queries was conducted. A web-based application has been developed to serve online evaluation scenarios with different groups of end-users. An extensive evaluation was conducted with 30 subjects and 115 queries. Upon analyzing the experimental results, open issues and suggestions for future improvements in person search were uncovered. Thi-Hoai Phan, Hoang-Son Bui, Tri Trung Kien Le, Thi-Ngoc-Diep Do, Thuy-Binh Nguyen, Hong Quan Nguyen 0003, Thanh-Hai Tran 0001, Thi Thanh Thuy Pham, Thi-Lan Le |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 9 |
| 2024 | SovaSeg-Net: Scale Invariant Ovarian Tumors Segmentation from Ultrasound ImagesabstractOvarian tumors are becoming a significant health concern for women worldwide, requiring accurate and effective diagnostic tools for early detection and treatment. This paper presents a new method to segment ovarian tumors from ultrasound images, aiming to reduce the healthcare burden and minimize the risk of oversight by less experienced medical professionals. This method is called SovaSeg-Net, built from the encoder-decoder deep learning architecture. The encoder leverages a convolutional neural network combined with a self attention module to extract meaningful features of tumors. It was then enhanced by the SPPF to combine features from different scales, increasing the encoder’s robustness to variation in scale, shape, and deformation of tumors. We also introduce a Joint Loss function that combines conventional IoU loss with focal loss and structural similarity loss to address data imbalance issues as well as the specific properties of ovarian tumors and ultrasound images. Experiments conducted on the benchmark OTU_2D, show that the proposed method outperforms existing methods, mainly in its ability to segment small tumors. This segmentation of ovarian tumors provides essential input for subsequent analysis steps facilitating the classification of tumors according to the rules established by the IOTA group. Source code is available at https://github.com/SonBH0410/SovaSeg-Net. Huu-Phong Luong, Hoang-Son Bui, Nam-Khanh Nguyen, Thi-Loan Pham, Gia-Minh Pham, Sy-Hoang Tran, Thanh-Hai Tran 0001, Thi-Lan Le |
ICIP | 8 |
| 2023 | A lightweight graph convolutional network for skeleton-based action recognition
Dinh-Tan Pham, Quang-Tien Pham, Tien-Thanh Nguyen, Thi-Lan Le, Hai Vu |
Multim. Tools Appl. | 4 |
| 2022 | An effective method for text line segmentation in historical document imagesabstractIn this paper, we present a text-line segmentation method for historical documents. Historical documents are challenging given their characteristics of highly degradation, writing style variation and diacritics. From these observations, we proposed an effective approach for text line segmentation by analysing the properties of document layouts. We combine the idea of seam carving method with the novel cost functions to accurately split text lines. Experiments were conducted on two challenging datasets of historical documents, namely the DIVA-HisDB dataset and our ChamDoc dataset. Our methods provided good results on the DIVA-HisDB dataset with 99.36% of Line IU and 98.86% of Pixel IU. On the ChamDoc dataset, the proposed method outperformed the two baseline approaches i.e. seam carving-based and A* path planning by a large margin. Tien-Nam Nguyen, Jean-Christophe Burie, Thi-Lan Le, Anne-Valérie Schweyer |
ICPR | 3 |
| 2022 | An end-to-end framework for the detection of mathematical expressions in scientific document imagesabstractAbstract The detection of mathematical expressions is a prerequisite step for the digitisation of scientific documents. Many different multistage approaches have been proposed for the detection of expressions in document images, that is, page segmentation and expression detection. However, the detection accuracy of such methods still needs improvement owing to errors in the page segmentation of complex documents. This paper presents an end‐to‐end framework for mathematical expression detection in scientific document images without requiring optical character recognition (OCR) or document analysis techniques applied in conventional methods. The novelty of this paper is twofold. First, because document images are usually in binary form, the direct use of these images, which lack texture information as input for detection networks, may lead to an incorrect detection. Therefore, we propose the application of a distance transform to obtain a discriminating and meaningful representation of mathematical expressions in document images. Second, the transformed images are fed into the faster region with a convolutional neural network (Faster R‐CNN) optimized to improve the accuracy of the detection. The proposed framework was tested on two benchmark data sets (Marmot and GTDB). Compared with the original Faster R‐CNN, the proposed network improves the accuracies of detection of isolated and inline expressions by 5.09% and 3.40%, respectfully, on the Marmot data set, whereas those on the GTDB data set are improved by 4.04% and 4.55%. A performance comparison with conventional methods shows the effectiveness of the proposed method. Bui Hai Phong, Thang Manh Hoang, Thi-Lan Le |
Expert Syst. J. Knowl. Eng. | 3 |
| 2022 | A robust and efficient method for skeleton-based human action recognition and its application for cross-dataset evaluationabstractAbstract Skeleton‐based human action recognition has emerged recently thanks to its compactness and robustness to appearance variations. Although impressive results have been obtained in recent years, the performance of skeleton‐based action recognition methods has to be improved to be deployed in real‐time applications. Recently, a lightweight network structure named Double‐feature Double‐motion Network (DD‐Net) has been proposed for the skeleton‐based human action recognition. With high speed, the DD‐Net achieves state‐of‐the‐art performance on hand and body actions. The DD‐Net could not distinguish actions if they have a weak connection with the global trajectories. However, the DD‐Net is suitable for human action recognition where actions strongly correlate to the global trajectories. In this paper, the authors propose TD‐Net, an improved version of the DD‐Net in which a new branch is added. The new branch takes the normalised coordinates of joints (NCJ) to enrich the spatial information. On five datasets for skeleton‐based human activity recognition that are MSR‐Action3D, CMDFall, JHMDB, FPHAB, and NTU RGB + D, the TD‐Net consistently obtains superior performance compared with the baseline model DD‐Net. The proposed method outperforms different state‐of‐the‐art methods, including both hand‐designed and deep learning‐based methods on four datasets (MSR‐Action3D, CMDFall, JHMDB, and FPHAB). Furthermore, the generalisation of the proposed method is confirmed through cross‐dataset evaluation. To illustrate the potential use of the model for real‐time human action recognition, the authors have deployed an application on an edge device. The experimental result shows that the application can process up to 40 fps for pose estimation using MediaPipe. It takes only 0.04 ms to recognise an action from skeleton sequences. Tien-Thanh Nguyen, Dinh-Tan Pham, Hai Vu, Thi-Lan Le |
IET Comput. Vis. | 4 |
| 2022 | Towards a large-scale person search by vietnamese natural language: dataset and methods
Thi Thanh Thuy Pham, Hong Quan Nguyen 0003, Hoai Phan, Thi-Ngoc-Diep Do, Thuy-Binh Nguyen, Thanh-Hai Tran 0001, Thi-Lan Le |
Multim. Tools Appl. | 7 |
| 2021 | On the Use of Attention in Deep Learning Based Denoising Method for Ancient Cham Inscription Images
Tien-Nam Nguyen, Jean-Christophe Burie, Thi-Lan Le, Anne-Valérie Schweyer |
ICDAR (1) | 3 |
| 2021 | Adaptive most joint selection and covariance descriptions for a robust skeleton-based human action recognition
Van-Toi Nguyen, Tien-Nam Nguyen, Thi-Lan Le, Dinh-Tan Pham, Hai Vu |
Multim. Tools Appl. | 3 |
| 2020 | A projective chirp based stair representation and detection from monocular images and its application for the visually impaired
Hai Vu, Van-Nam Hoang, Thi-Lan Le, Thanh-Hai Tran 0001, Thi Thuy Nguyen |
Pattern Recognit. Lett. | 3 |
| 2019 | Mathematical Variable Detection in PDF Scientific Documents
Bui Hai Phong, Thang Manh Hoang, Thi-Lan Le, Akiko Aizawa |
ACIIDS (2) | 3 |
| 2019 | Effective multi-shot person re-identification through representative frames selection and temporal feature pooling
Thuy-Binh Nguyen, Thi-Lan Le, Louis Devillaine, Thuy Thi Thanh Pham, Nam Pham Ngoc 0001 |
Multim. Tools Appl. | 2 |
| 2018 | A Reliable Image-to-Video Person Re-identification Based on Feature Fusion
Thuy-Binh Nguyen, Thi-Lan Le, Dinh-Duc Nguyen, Dinh-Tan Pham |
ACIIDS (1) | 2 |
| 2018 | A multi-modal multi-view dataset for human fall analysis and preliminary investigation on modalityabstractOver the last decade, a large number of methods have been proposed for human fall detection. Most existing methods were evaluated based on trimmed datasets. More importantly, these datasets lack variety of falls, subjects, views and modalities. This paper makes two contributions in the topic of automatic human fall detection. Firstly, to address the above issues, we introduce a large continuous multimodal multivew dataset of human fall, namely CMDFALL. Our CMDFALL dataset was built by capturing activities from 50 subjects, with seven overlapped Kinect sensors and two wearable accelerometers. Each subject performs 20 activities including 8 falls of different styles and 12 daily activities. All multi-modal multi-view data (RGB, depth, skeleton, acceleration) are time-synchronized and annotated for evaluating performance of recognition algorithms of human activities or human fall in indoor environment. Secondly, based on the multimodal property of the dataset, we investigate the role of each modality to get the best results in the context of human activity recognition. To this end, we adopt existing baseline techniques which have been shown to be very efficient for each data modality such as C3D convnet on RGB; DMM-KDES on depth; Res-TCN on skeleton and 2D convnet on acceleration data. We analyze to show which modalities and their combination give the best performance. Thanh-Hai Tran 0001, Thi-Lan Le, Dinh-Tan Pham, Van-Nam Hoang, Van-Minh Khong, Quoc-Toan Tran, Thai Son Nguyen, Cuong Pham 0001 |
ICPR | 2 |
| 2018 | Acquiring qualified samples for RANSAC using geometrical constraints
Van-Hung Le 0002, Hai Vu, Thi Thuy Nguyen, Thi-Lan Le, Thanh-Hai Tran 0001 |
Pattern Recognit. Lett. | 4 |
| 2017 | Fully-automated person re-identification in multi-camera surveillance system with a robust kernel descriptor and effective shadow removal method
Thuy Thi Thanh Pham, Thi-Lan Le, Hai Vu, Trung-Kien Dao, Van-Toi Nguyen |
Image Vis. Comput. | 2 |
| 2016 | Selections of Suitable UAV Imagery's Configurations for Regions Classification
Hai Vu, Thi-Lan Le, Van Giap Nguyen, Dinh Tan Hung |
ACIIDS (1) | 2 |
| 2016 | Analytical method for multimodal localization combination using Wi-Fi and cameraabstractThis paper introduces and demonstrates a combination scheme for the combination of Wi-Fi based and camera based localization technologies to increase the applicability in indoor environments. The scheme is developed on the basis of an analytical approach by looking for maximizing the probability of appearance and the reliability of the location results provided by underlying technologies. The proposed scheme is derived in a general way such that it can be easily extended to include more localization technologies, as well as other kinds of them in the system. Nicolas des Aunais, Trung-Kien Dao, Dinh-Van Nguyen, Thanh-Thuy Pham, Eric Castelli, Thi-Lan Le, Ngoc Yen Pham |
ICARCV | 6 |
| 2016 | Indoor navigation assistance system for visually impaired people using multimodal technologiesabstractIn this paper, a complete indoor navigation assistance system for visually impaired people is introduced. Different multimedia technologies are integrated in a single system in order to provide a precise, safe and friendly navigation service. First, the environment is modeled and represented. After that, the user location is determined by combining Wi-Fi and vision information. This combination offers some benefits in comparison with single technology systems such as setup cost, computational time and accuracy. Finally, the interaction between users and the system is performed through natural Vietnamese language with the support of Vietnamese voice synthesis and recognition. The proposed the system has been successfully deployed in a school for visually impaired pupils. Evaluation with various criteria on visually impaired pupils reveals the feasibility of the solution. Trung-Kien Dao, Thanh-Hai Tran 0001, Thi-Lan Le, Hai Vu, Viet-Tung Nguyen, Dang-Khoa Mac, Ngoc-Diep Do, Thanh-Thuy Pham |
ICARCV | 3 |
| 2013 | A vision-based method for automatizing tea shoots detectionabstractCounting tender tea shoots in a sampled area is required before making a decision for plucking. However, it is a tedious task and requires a large amount of time. In this paper, we propose a vision-based method for automatically detecting and counting the number of tea shoots in an image acquired from a tea field. First, we build a parametric model of a tea-shoot's color distribution in order to roughly separate Regions-of-Interest (ROIs) of tea shoots from a complicated background. For each ROI, we then extract supportive (local) features with expectations that these features will only appear around an apical bud of tea shoots thanks to two measurements: the density of edge pixels and a statistic of gradient directions. Consequently, the extracted features are put into a mean shift cluster to locate the position of tea shoots. The proposed method is evaluated on a set of testing images with different species of tea plants and ages. The results show 86% correct tea shoots detected, whereas 25% of a false alarm rate exists. It offers an elegant way to build an assisting tool for tea harvesting. Hai Vu, Thi-Lan Le, Thanh-Hai Tran 0001, Thi Thuy Nguyen |
ICIP | 2 |
| 2009 | Surveillance Video Indexing and Retrieval Using Object Features and Semantic EventsabstractIn this paper, we propose an approach for surveillance video indexing and retrieval. The objective of this approach is to answer five main challenges we have met in this domain: (1) the lack of means for finding data from the indexed databases, (2) the lack of approaches working at different abstraction levels, (3) imprecise indexing, (4) incomplete indexing, (5) the lack of user-centered search. We propose a new data model containing two main types of extracted video contents: physical objects and events. Based on this data model, we present a new rich and flexible query language. This language works at different abstraction levels, provides both exact and approximate matching and takes into account users' interest. In order to work with the imprecise indexing, two new methods respectively for object representation and object matching are proposed. Videos from two projects which have been partially indexed are used to validate the proposed approach. We have analyzed both query language usage and retrieval results. The obtained retrieval results analyzed by the average normalized ranks are promising. The retrieval results at the object level are compared with another state of the art approach. Thi-Lan Le, Monique Thonnat, Alain Boucher, François Brémond |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2008 | A Query Language Combining Object Features and Semantic Events for Surveillance Video Retrieval
Thi-Lan Le, Monique Thonnat, Alain Boucher, François Brémond |
MMM | 1 |
| 2007 | Subtrajectory-Based Video Indexing and Retrieval
Thi-Lan Le, Alain Boucher, Monique Thonnat |
MMM (1) | 1 |