Vuong Dinh An

dblp:273/8725 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2024
0009-0003-8533-9897ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot manipulation · 55% Generative modeling · 32% 3D vision · 10%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
grasping
1.522024
Grasp-Anything: Large-scale Grasp Dataset from Foundation Models · ICRA 2024
Language-driven Grasp Detection · CVPR 2024
Machine learning › Generative modeling
diffusion model
1.422024
Language-driven Grasp Detection · CVPR 2024
Language-driven Scene Synthesis using Multi-conditional Diffusion Model · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.812024
Language-driven Grasp Detection · CVPR 2024
Robotics › Robot manipulation › grasping
grasp dataset
0.812024
Grasp-Anything: Large-scale Grasp Dataset from Foundation Models · ICRA 2024
Robotics › Robot manipulation › grasping
grasp detection
0.812024
Grasp-Anything: Large-scale Grasp Dataset from Foundation Models · ICRA 2024
Robotics › Robot manipulation › grasping › grasp detection
language-driven grasp detection
0.812024
Language-driven Grasp Detection · CVPR 2024
Computer vision › 3D vision › 3d scene understanding
scene synthesis
0.712023
Language-driven Scene Synthesis using Multi-conditional Diffusion Model · NeurIPS 2023
Machine learning › Deep learning architectures and training
foundation model
0.212024
Grasp-Anything: Large-scale Grasp Dataset from Foundation Models · ICRA 2024

Methods — techniques the papers use, named apart from their topics

foundation model · 1.5diffusion model · 1.4zero-shot learning · 0.8contrastive training objective · 0.8multi-condition encoding · 0.7
YearPublicationVenuePosition
2024 Language-driven Grasp Detection
abstract
Grasp detection is a persistent and intricate challenge with various industrial applications. Recently, many meth-ods and datasets have been proposed to tackle the grasp detection problem. However, most of them do not consider using natural language as a condition to detect the grasp poses. In this paper, we introduce Grasp-Anything++, a new language-driven grasp detection dataset featuring 1M samples, over 3M objects, and upwards of 10M grasping in-structions. We utilize foundation models to create a large-scale scene corpus with corresponding images and grasp prompts. We approach the language-driven grasp detection task as a conditional generation problem. Drawing on the success of diffusion models in generative tasks and given that language plays a vital role in this task, we propose a new language-driven grasp detection method based on dif-fusion models. Our key contribution is the contrastive training objective, which explicitly contributes to the denoising process to detect the grasp pose given the language instructions. We illustrate that our approach is theoretically sup-portive. The intensive experiments show that our method outperforms state-of-the-art approaches and allows real-world robotic grasping. Finally, we demonstrate our large-scale dataset enables zero-short grasp detection and is a challenging benchmark for future work.
Vuong Dinh An, Minh Nhat Vu, Baoru Huang, Thieu Vo, Anh Nguyen 0003
CVPR1
2024 Grasp-Anything: Large-scale Grasp Dataset from Foundation Models
abstract
Foundation models such as ChatGPT have made significant strides in robotic tasks due to their universal representation of real-world domains. In this paper, we leverage foundation models to tackle grasp detection, a persistent challenge in robotics with broad industrial applications. Despite numerous grasp datasets, their object diversity remains limited compared to real-world figures. Fortunately, foundation models possess an extensive repository of real-world knowledge, including objects we encounter in our daily lives. As a consequence, a promising solution to the limited representation in previous grasp datasets is to harness the universal knowledge embedded in these foundation models. We present Grasp-Anything, a new large-scale grasp dataset synthesized from foundation models to implement this solution. Grasp-Anything excels in diversity and magnitude, boasting 1M samples with text descriptions and more than 3M objects, surpassing prior datasets. Empirically, we show that Grasp-Anything successfully facilitates zero-shot grasp detection on vision-based tasks and real-world robotic experiments. Our dataset and code are available at https://airvlab.github.io/grasp-anything/.
Vuong Dinh An, Minh Nhat Vu, Baoru Huang, Huynh Thi Thanh Binh, Thieu Vo, Andreas Kugi, Anh Nguyen 0003
ICRA1
2024 An adaptive charging scheme for large-scale wireless rechargeable sensor networks inspired by deep Q-network
Vuong Dinh An, Tran Thi Huong, Hoang Nguyen Quang Pham, Quang Minh Bui, Trang Phuong Ngo, Huynh Thi Thanh Binh
Neural Comput. Appl.1
2023 Language-driven Scene Synthesis using Multi-conditional Diffusion Model
abstract
Scene synthesis is a challenging problem with several industrial applications. Recently, substantial efforts have been directed to synthesize the scene using human motions, room layouts, or spatial graphs as the input. However, few studies have addressed this problem from multiple modalities, especially combining text prompts. In this paper, we propose a language-driven scene synthesis task, which is a new task that integrates text prompts, human motion, and existing objects for scene synthesis. Unlike other single-condition synthesis tasks, our problem involves multiple conditions and requires a strategy for processing and encoding them into a unified space. To address the challenge, we present a multi-conditional diffusion model, which differs from the implicit unification approach of other diffusion literature by explicitly predicting the guiding points for the original data distribution. We demonstrate that our approach is theoretically supportive. The intensive experiment results illustrate that our method outperforms state-of-the-art benchmarks and enables natural scene editing applications. The source code and dataset can be accessed at https://lang-scene-synth.github.io/.
Vuong Dinh An, Minh Nhat Vu, Toan Nguyen 0004, Baoru Huang, Dzung Nguyen, Thieu Vo, Anh Nguyen 0003
NeurIPS1
2020 Optimizing Charging Locations and Charging Time for Energy Depletion Avoidance in Wireless Rechargeable Sensor Networks
abstract
In recent years, Wireless Rechargeable Sensor Networks, which exploit wireless energy transfer technologies to address the energy constraint problem in traditional Wireless Sensor Networks, has emerged as a promising solution. There are two important factors that affect the performance of a charging process: charging path and charging time. In the literature, many studies have been done to propose efficient charging algorithms. However, most of the existing works focus only on optimizing the charging path. In this paper, we are the first one to jointly take into account both the charging path and charging time. Specifically, we aim at determining the optimal charging path and the charging time at each charging location to minimize the number of dead nodes. We first mathematically formulate the problem under mixed integer and linear programming. Then, we propose a periodic charging scheme, which is based on the Greedy and Genetic algorithm approaches. The experiment results show that our proposed the algorithm reduces significantly the number of dead nodes compared to a relevant benchmark.
Tran Thi Huong, Huynh Thi Thanh Binh, Phi-Le Nguyen, Doan Cao Thanh Long, Vuong Dinh An
CEC5