EDBT 2026 Demo / reviewers in the wild / expert
Cuong Pham 0001
dblp:20/6376-1
· DBLP profile ↗
30ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0003-0973-0889ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SiCLIP: An explainable multimodal framework for silicosis diagnosis
Tien Nguyen, Cong Tran, Cuong Pham 0001 |
Artif. Intell. Medicine | 6 |
| 2026 | NF-DCL: Enhancing video anomaly detection with synthetic normal features and Debiased Contrastive Learning
Cong Tran, Cuong Pham 0001 |
Comput. Vis. Image Underst. | 4 |
| 2026 | Unveiling open vocabulary relationships in images: A translation embedding approach guided by vision-language models
Nguyet Nguyen, Cong Tran, Anh Tuan Tran 0001, Cuong Pham 0001 |
Comput. Vis. Image Underst. | 4 |
| 2025 | REM: A Scalable Reinforced Multi-Expert Framework for Multiplex Influence MaximizationabstractIn social online platforms, identifying influential seed users to maximize influence spread is a crucial as it can greatly diminish the cost and efforts required for information dissemination. While effective, traditional methods for Multiplex Influence Maximization (MIM) have reached their performance limits, prompting the emergence of learning-based approaches. These novel methods aim for better generalization and scalability for more sizable graphs but face significant challenges, such as (1) inability to handle unknown diffusion patterns and (2) reliance on high-quality training samples. To address these issues, we propose the Reinforced Expert Maximization framework (REM). REM leverages a Propagation Mixture of Experts technique to encode dynamic propagation of large multiplex networks effectively in order to generate enhanced influence propagation. Noticeably, REM treats a generative model as a policy to autonomously generate different seed sets and learn how to improve them from a Reinforcement Learning perspective. Extensive experiments on several real-world datasets demonstrate that REM surpasses state-of-the-art methods in terms of influence spread, scalability, and inference time in influence maximization tasks. Hieu Dam, Nguyen Hoang Khoi Do, Cong Tran, Cuong Pham 0001 |
AAAI | 5 |
| 2025 | Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask TrackingabstractExisting 3D instance segmentation methods frequently encounter issues with over-segmentation, leading to redundant and inaccurate 3D proposals that complicate downstream tasks. This challenge arises from their unsupervised merging approach, where dense 2D instance masks are lifted across frames into point clouds to form 3D candidate proposals without direct supervision. These candidates are then hierarchically merged based on heuristic criteria, often resulting in numerous redundant segments that fail to combine into precise 3D proposals. To overcome these limitations, we propose a 3D-Aware 2D Mask Tracking module that uses robust 3D priors from a 2D mask segmentation and tracking foundation model (SAM-2) to ensure consistent object masks across video frames. Rather than merging all visible superpoints across views to create a 3D mask, our 3D Mask Optimization module leverages a dynamic programming algorithm to select an optimal set of views, refining the superpoints to produce a final 3D proposal for each object. Our approach achieves comprehensive object coverage within the scene, significantly reduces unnecessary proposals by half, and lowers the average runtime by a factor of 10, mitigating potential negative impacts on downstream applications. Evaluations on ScanNet200 and ScanNet++ confirm the effectiveness of our method, with improvements across Class-Agnostic, Open-Vocabulary, and Open-Ended 3D Instance Segmentation tasks. Minh Luu, Anh Tuan Tran 0001, Cuong Pham 0001, Khoi Nguyen 0001 |
CVPR | 4 |
| 2025 | SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step DiffusionabstractRecent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the speed demands required for real-world and on-device applications due to the costly multi-step inversion and sampling process involved. In response to this, we introduce SwiftEdit, a simple yet highly efficient editing tool that achieve instant text-guided image editing (in 0.23s). The advancement of SwiftEdit lies in its two novel contributions: a one-step inversion framework that enables one-step image reconstruction via inversion and a mask-guided editing technique with our proposed attention rescaling mechanism to perform localized image editing. Extensive experiments are provided to demonstrate the effectiveness and efficiency of SwiftEdit. In particular, SwiftEdit enables instant text-guided image editing, which is extremely faster than previous multi-step methods (at least 50× times faster) while maintain a competitive performance in editing results. Our project is at https://swift-edit.github.io/. Trong-Tung Nguyen, Khoi Nguyen 0001, Anh Tuan Tran 0001, Cuong Pham 0001 |
CVPR | 5 |
| 2025 | Supercharged One-Step Text-to-Image Diffusion Models with Negative Prompts
Viet Nguyen, Trung Dao, Khoi Nguyen 0001, Cuong Pham 0001, Toan Tran 0003, Anh Tuan Tran 0001 |
ICCV | 5 |
| 2025 | SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion ModelsabstractThe rapid growth of text-to-image diffusion models has raised concerns about their potential misuse in generating harmful or unauthorized contents. To address these issues, several Concept Erasure methods have been proposed. However, most of them fail to achieve both robustness, i.e., the ability to robustly remove the target concept., and effectiveness, i.e., maintaining image quality. While few recent techniques successfully achieve these goals for NSFW concepts, none could handle narrow concepts such as copyrighted characters or celebrities. Erasing these narrow concepts is critical in addressing copyright and legal concerns. However, erasing them is challenging due to their close distances to non-target neighboring concepts, requiring finer-grained manipulation. In this paper, we introduce Subspace Mapping (SuMa), a novel method specifically designed to achieve both robustness and effectiveness in easing these narrow concepts. SuMa first derives a target subspace representing the concept to be erased and then neutralizes it by mapping it to a reference subspace that minimizes the distance between the two. This mapping ensures the target concept is robustly erased while preserving image quality. We conduct extensive experiments with SuMa across four tasks: subclass erasure, celebrity erasure, artistic style erasure, and instance erasure and compare the results with current state-of-the-art methods. Our method achieves image quality comparable to approaches focused on effectiveness, while also yielding results that are on par with methods targeting completeness. Anh Tuan Tran 0001, Cuong Pham 0001 |
ICCV | 3 |
| 2025 | Class-Agnostic Repetitive Action Counting Using Wearable DevicesabstractWe present Class-agnostic Repetitive action Counting (CaRaCount), a novel approach to count repetitive human actions in the wild using wearable devices time series data. CaRaCount is the first few-shot class-agnostic method, being able to count repetitions of any action class with only a short exemplar data sequence containing a few examples from the action class of interest. To develop and evaluate this method, we collect a large-scale time series dataset of repetitive human actions in various context, containing smartwatch data from 10 subjects performing 50 different activities. Experiments on this dataset and three other activity counting datasets namely Crossfit, Recofit, and MM-Fit show that CaRaCount can count repetitive actions with low error, and it outperforms other baselines and state-of-the-art action counting methods. Finally, with a user experience study, we evaluate the usability of our real-time implementation. Our results highlight the efficiency and effectiveness of our approach when deployed outside the laboratory environments. Duc Duy Nguyen, Lam Thanh Nguyen, Cuong Pham 0001, Minh Hoai |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Count What You Want: Exemplar Identification and Few-Shot Counting of Human Actions in the WildabstractThis paper addresses the task of counting human actions of interest using sensor data from wearable devices. We propose a novel exemplar-based framework, allowing users to provide exemplars of the actions they want to count by vocalizing predefined sounds ``one'', ``two'', and ``three''. Our method first localizes temporal positions of these utterances from the audio sequence. These positions serve as the basis for identifying exemplars representing the action class of interest. A similarity map is then computed between the exemplars and the entire sensor data sequence, which is further fed into a density estimation module to generate a sequence of estimated density values. Summing these density values provides the final count. To develop and evaluate our approach, we introduce a diverse and realistic dataset consisting of real-world data from 37 subjects and 50 action categories, encompassing both sensor and audio data. The experiments on this dataset demonstrate the viability of the proposed method in counting instances of actions from new classes and subjects that were not part of the training data. On average, the discrepancy between the predicted count and the ground truth value is 7.47, significantly lower than the errors of the frequency-based and transformer-based methods. Our project, code and dataset can be found at https://github.com/cvlab-stonybrook/ExRAC. Duc Duy Nguyen, Cuong Pham 0001, Minh Hoai |
AAAI | 4 |
| 2024 | EFHQ: Multi-Purpose ExtremePose-Face-HQ DatasetabstractThe existing facial datasets, while having plentiful images at near frontal views, lack images with extreme head poses, leading to the downgraded performance of deep learning models when dealing with profile or pitched faces. This work aims to address this gap by introducing a novel dataset named Extreme Pose Face High-Quality Dataset (EFHQ), which includes a maximum of 450k high-quality images of faces at extreme poses. To produce such a massive dataset, we utilize a novel and meticulous dataset processing pipeline to curate two publicly available datasets, VFHQ and CelebV-HQ, which contain many high-resolution face videos captured in various settings. Our dataset can complement existing datasets on various facial-related tasks, such as facial synthesis with 2D/3D-aware GAN, diffusion-based text-to-image face generation, and face reenactment. Specifically, training with EFHQ helps models generalize well across diverse poses, significantly improving performance in scenarios involving extreme views, confirmed by extensive experiments. Additionally, we utilize EFHQ to define a challenging cross-view face verification benchmark, in which the performance of SOTA face recognition models drops 5-37% compared to frontal-to-frontal scenarios, aiming to stimulate studies on face recognition under severe pose conditions in the wild. Trung Tuan Dao, Duc Hong Vu, Cuong Pham 0001, Anh Tuan Tran 0001 |
CVPR | 3 |
| 2024 | Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidanceabstractWe introduce Open3DIS, a novel solution designed to tackle the problem of Open-Vocabulary Instance Segmentation within 3D scenes. Objects within 3D environments exhibit diverse shapes, scales, and colors, making precise instance-level identification a challenging task. Recent advancements in Open-Vocabulary scene understanding have made significant strides in this area by employing class-agnostic 3D instance proposal networks for object localization and learning queryable features for each 3D mask. While these methods produce high-quality instance proposals, they struggle with identifying small-scale and geometrically ambiguous objects. The key idea of our method is a new module that aggregates 2D instance masks across frames and maps them to geometrically coherent point cloud regions as high-quality object proposals addressing the above limitations. These are then combined with 3D class-agnostic instance proposals to include a wide range of objects in the real world. To validate our approach, we conducted experiments on three prominent datasets, including ScanNet200, S3DIS, and Replica, demonstrating significant performance gains in segmenting objects with diverse categories over the state-of-the-art approaches. Phuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan 0001, Anh Tuan Tran 0001, Cuong Pham 0001, Khoi Nguyen 0001 |
CVPR | 6 |
| 2024 | Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown DomainsabstractThis paper presents an innovative framework designed to train an image deblurring algorithm tailored to a specific camera device. This algorithm works by transforming a blurry input image, which is challenging to deblur, into another blurry image that is more amenable to deblurring. The transformation process, from one blurry state to another, leverages unpaired data consisting of sharp and blurry images captured by the target camera device. Learning this blur-to-blur transformation is inherently simpler than direct blur-to-sharp conversion, as it primarily involves modifying blur patterns rather than the intricate task of reconstructing fine image details. The efficacy of the proposed approach has been demonstrated through comprehensive experiments on various benchmarks, where it significantly outperforms state-of-the-art methods both quantitatively and qualitatively. Our code and data are available at https://github.com/VinAIResearch/Blur2Blur Bang-Dang Pham, Anh Tuan Tran 0001, Cuong Pham 0001, Rang Nguyen, Minh Hoai |
CVPR | 4 |
| 2024 | SwiftBrush V2: Make Your One-Step Diffusion Model Better Than Its Teacher
Trung Tuan Dao, Thuan Hoang Nguyen, Duc Hong Vu, Khoi Nguyen 0001, Cuong Pham 0001, Anh Tuan Tran 0001 |
ECCV (82) | 6 |
| 2024 | Active Learning Framework for Incomplete NetworksabstractSignificant progression has been made in active learning algorithms for graph networks in various tasks. However real-world applications frequently involve incomplete graphs with missing links, which pose the challenge that existing approaches might not adequately address. This paper presents an active learning approach tailored specifically for handling incomplete graphs, termed ALIN. Our algorithm employs graph neural networks (GNN) to generate node embeddings and calculates losses for both node classification and link prediction tasks. The losses are combined with appropriate weights and iteratively updating the GNN, ALIN efficiently queries nodes in batches, thereby achieving a balance between training feedbacks and resource utilization. Our empirical experiments have shown ALIN can surpass state-of-the-art baselines on Cora, Citeseer, Pubmed, and Coauthor-CS datasets. Tung Khong, Cong Tran, Cuong Pham 0001 |
UAI | 3 |
| 2023 | HyperCUT: Video Sequence from a Single Blurry Image using Unsupervised OrderingabstractWe consider the challenging task of training models for image-to-video deblurring, which aims to recover a sequence of sharp images corresponding to a given blurry image input. A critical issue disturbing the training of an image-to-video model is the ambiguity of the frame ordering since both the forward and backward sequences are plausible solutions. This paper proposes an effective self-supervised ordering scheme that allows training highquality image-to-video deblurring models. Unlike previous methods that rely on order-invariant losses, we assign an explicit order for each video sequence, thus avoiding the order-ambiguity issue. Specifically, we map each video sequence to a vector in a latent high-dimensional space so that there exists a hyperplane such that for every video sequence, the vectors extracted from it and its reversed sequence are on different sides of the hyperplane. The side of the vectors will be used to define the order of the corresponding sequence. Last but not least, we propose a real-image dataset for the image-to-video deblurring problem that covers a variety of popular domains, including face, hand, and street. Extensive experimental results confirm the effectiveness of our method. Code and data are available at https://github.com/VinAIResearch/HyperCUT.git Bang-Dang Pham, Anh Tuan Tran 0001, Cuong Pham 0001, Rang Nguyen, Minh Hoai |
CVPR | 4 |
| 2023 | Federated few-shot learning for cough classification with edge devices
Ngan Dao Hoang, Dat Tran-Anh, Manh Luong, Cong Tran, Cuong Pham 0001 |
Appl. Intell. | 5 |
| 2022 | QC-StyleGAN - Quality Controllable Image Generation and ManipulationabstractThe introduction of high-quality image generation models, particularly the StyleGAN family, provides a powerful tool to synthesize and manipulate images. However, existing models are built upon high-quality (HQ) data as desired outputs, making them unfit for in-the-wild low-quality (LQ) images, which are common inputs for manipulation. In this work, we bridge this gap by proposing a novel GAN structure that allows for generating images with controllable quality. The network can synthesize various image degradation and restore the sharp image via a quality control code. Our proposed QC-StyleGAN can directly edit LQ images without altering their quality by applying GAN inversion and manipulation techniques. It also provides for free an image restoration solution that can handle various degradations, including noise, blur, compression artifacts, and their mixtures. Finally, we demonstrate numerous other applications such as image degradation synthesis, transfer, and interpolation. Dat Viet Thanh Nguyen, Phong Tran The, Tan M. Dinh, Cuong Pham 0001, Anh Tuan Tran 0001 |
NeurIPS | 4 |
| 2022 | Masked face recognition with convolutional neural networks and local binary patterns
Huong Mai Nguyen, Cuong Pham 0001 |
Appl. Intell. | 3 |
| 2022 | Multi-task learning neural networks for breath sound detection and classification in pervasive healthcare
Dat Tran-Anh, Khanh Nguyen-Trong, Cuong Pham 0001 |
Pervasive Mob. Comput. | 4 |
| 2021 | Combining skeleton and accelerometer data for human fine-grained activity recognition and abnormal behaviour detection with deep temporal convolutional networks
Cuong Pham 0001, Ngon Nguyen, Van-Toi Nguyen |
Multim. Tools Appl. | 1 |
| 2018 | A multi-modal multi-view dataset for human fall analysis and preliminary investigation on modalityabstractOver the last decade, a large number of methods have been proposed for human fall detection. Most existing methods were evaluated based on trimmed datasets. More importantly, these datasets lack variety of falls, subjects, views and modalities. This paper makes two contributions in the topic of automatic human fall detection. Firstly, to address the above issues, we introduce a large continuous multimodal multivew dataset of human fall, namely CMDFALL. Our CMDFALL dataset was built by capturing activities from 50 subjects, with seven overlapped Kinect sensors and two wearable accelerometers. Each subject performs 20 activities including 8 falls of different styles and 12 daily activities. All multi-modal multi-view data (RGB, depth, skeleton, acceleration) are time-synchronized and annotated for evaluating performance of recognition algorithms of human activities or human fall in indoor environment. Secondly, based on the multimodal property of the dataset, we investigate the role of each modality to get the best results in the context of human activity recognition. To this end, we adopt existing baseline techniques which have been shown to be very efficient for each data modality such as C3D convnet on RGB; DMM-KDES on depth; Res-TCN on skeleton and 2D convnet on acceleration data. We analyze to show which modalities and their combination give the best performance. Thanh-Hai Tran 0001, Thi-Lan Le, Dinh-Tan Pham, Van-Nam Hoang, Van-Minh Khong, Quoc-Toan Tran, Thai Son Nguyen, Cuong Pham 0001 |
ICPR | 8 |
| 2018 | Delay-limited Protograph Low Density Parity Codes for Space-Time Block CodesabstractThis paper presents a framework to design protograph low density parity check (LDPC) codes for wireless communications where Alamouti space time block code (STBC) scheme is exploited to combat with the fading effect and the number of decoding iterations is limited. The decoding threshold of the resulting protograph code is about 0.7 dB better than that of the state-of-the-art code in the literature. The simulation of the proposed code is reported, and simulation results confirm the analytical advantages of the design. No error floor is found down to frame error rate (FER) of 10-6. Thuy V. Nguyen, Cuong Pham 0001, Hieu T. Nguyen 0002 |
PIMRC | 2 |
| 2016 | MobiCough: Real-Time Cough Detection and Monitoring Using Low-Cost Mobile Devices
Cuong Pham 0001 |
ACIIDS (1) | 1 |
| 2016 | Motion Primitive Forests for Human Activity Recognition Using Wearable Sensors
Nguyen Ngoc Diep, Cuong Pham 0001, Tu Minh Phuong |
PRICAI | 2 |
| 2016 | An Orientation Histogram Based Approach for Fall Detection Using Wearable Sensors
Nguyen Ngoc Diep, Cuong Pham 0001, Tu Minh Phuong |
PRICAI | 2 |
| 2013 | FoodBoard: surface contact imaging for food recognitionabstractWe describe FoodBoard, an instrumented chopping board that uses optical fibers and embedded camera imaging to identify unpackaged ingredients during food preparation on its surface. By embedding the sensing directly, and robustly, in the surface of a chopping board we also demonstrate how surface contact optical sensing can be used to realize the portability and privacy required of technology used in a setting such as a domestic kitchen. FoodBoard was subjected to a close to real-world evaluation in which 12 users prepared actual meals. FoodBoard compared favourably with existing unpackaged food recognition systems, classifying a larger number of distinct food ingredients (12 incl. meat, fruit, vegetables) with an average accuracy of 82.8%. Cuong Pham 0001, Daniel Jackson 0002, Johannes Schöning, Tom Bartindale, Thomas Plötz, Patrick Olivier |
UbiComp | 1 |
| 2013 | Real-Time Fall Detection and Activity Recognition Using Low-Cost Wearable Sensors
Cuong Pham 0001, Tu Minh Phuong |
ICCSA (1) | 1 |
| 2012 | The french kitchen: task-based learning in an instrumented kitchenabstractUbiquitous computing technologies have traditionally striven to augment objects and the environment with sensing capabilities to enable them to respond appropriately to the needs of the individuals in the environment. This paper considers how such technologies might be harnessed to support language learning, and specifically Task-Based Learning (TBL). Task-Based Learning (TBL) involves doing meaningful tasks in a foreign language, emphasising the language's use in practice. TBL is seen as a highly engaging and motivating approach to learning a language, but is difficult to do in the classroom. Here, learners typically engage in activities that only simulate 'real-world' tasks, and as such only rehearse language use, rather than applying the language in practice. In this paper, we explore how an instrumented, context-aware environment whose design is grounded in pedagogical principles can support TBL. We present the French Kitchen, an instrumented kitchen for English speakers who are learning French, and describe a 46-participant evaluation of the kitchen. Based on the evaluation, we provide a set of design recommendations for those building instrumented systems for TBL. Clare J. Hooper, Anne Preston, Madeline Balaam, Paul Seedhouse, Daniel Jackson 0002, Cuong Pham 0001, Cassim Ladha, Karim Ladha, Thomas Plötz, Patrick Olivier |
UbiComp | 6 |
| 2011 | Rapid specification and automated generation of prompting systems to assist people with dementia
Jesse Hoey, Thomas Plötz, Daniel Jackson 0002, Andrew F. Monk, Cuong Pham 0001, Patrick Olivier |
Pervasive Mob. Comput. | 5 |