Duy Ta

dblp:393/1352 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Segmentation and scene understanding · 30% 3D vision · 30% Planning, search and constraint satisfaction · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene understanding
3d scene graph
0.912025
ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis · CVPR 2025
Computer vision › Segmentation and scene understanding
scene understanding
0.912025
ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis · CVPR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › hierarchical problem solving
task decomposition
0.912025
ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis · CVPR 2025
Computer vision › Vision and language › visual grounding
language grounding
0.312025
ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis · CVPR 2025

Methods — techniques the papers use, named apart from their topics

large language model · 0.9
YearPublicationVenuePosition
2025 ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis
abstract
While recent work in scene reconstruction and understanding has made strides in grounding natural language to physical 3D environments, it is still challenging to ground abstract, high-level instructions to a 3D scene. High-Level instructions might not explicitly invoke semantic elements in the scene, and even the process of breaking a high-level task into a set of more concrete subtasks —a process called hierarchical task analysis— is environment-dependent. In this work, we propose ASHiTA, the first framework that generates a task hierarchy grounded to a 3D scene graph by breaking down high-level tasks into grounded subtasks. ASHiTA alternates LLM-assisted hierarchical task analysis —to generate the task breakdown— with task-driven 3D scene graph construction to generate a suitable representation of the environment. Our experiments show that ASHiTA performs significantly better than LLM baselines in breaking down high-level tasks into environment-dependent subtasks and is additionally able to achieve grounding performance comparable to state-of-the-art methods.
Yun Chang, Leonor Fermoselle, Duy Ta, Bernadette Bucher, Luca Carlone, Jiuguang Wang
CVPR3