Sinan Tan

dblp:264/0041 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 22% Vision and language · 19% 3D vision · 19%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
embodied navigation
1.322023
Knowledge-Based Embodied Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Embodied Referring Expression for Manipulation Question Answering in Interactive Environment · ICRA 2023
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering
1.122023
Knowledge-Based Embodied Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Multi-agent Embodied Question Answering in Interactive Environments · ECCV (13) 2020
Machine learning › Generative modeling › autoregressive model
autoregressive image generation
0.912025
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation · ICLR 2025
Machine learning › Generative modeling › image generation › token-based image generation
vector-quantized image generation
0.912025
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation · ICLR 2025
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.712023
Mixed Neural Voxels for Fast Multi-view Video Synthesis · ICCV 2023
Computer vision › Vision and language › visual grounding › referring expression comprehension
embodied referring expression
0.712023
Embodied Referring Expression for Manipulation Question Answering in Interactive Environment · ICRA 2023
Computer vision › 3D vision
neural radiance field
0.712023
Mixed Neural Voxels for Fast Multi-view Video Synthesis · ICCV 2023
Computer vision › Vision and language
vision-and-language navigation
0.612022
Depth-Aware Vision-and-Language Navigation using Scene Query Attention Network · ICRA 2022
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412020
Multi-agent Embodied Question Answering in Interactive Environments · ECCV (13) 2020
Computer vision › Vision and language
vision-language generation
0.312025
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation · ICLR 2025
Machine learning › Efficient and distributed learning
efficient training
0.212023
Mixed Neural Voxels for Fast Multi-view Video Synthesis · ICCV 2023
Computer vision › 3D vision › range sensing
depth sensing
0.212022
Depth-Aware Vision-and-Language Navigation using Scene Query Attention Network · ICRA 2022

Methods — techniques the papers use, named apart from their topics

vector quantization · 0.9autoregressive modeling · 0.9variation field · 0.7neural program synthesis · 0.7modular network · 0.7mixed neural voxels · 0.7inner product time query · 0.73d semantic reconstruction · 0.73d scene graph · 0.7CNN · 0.6
YearPublicationVenuePosition
2025 A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
abstract
This work tackles the information loss bottleneck of vector-quantization (VQ) autoregressive image generation by introducing a novel model architecture called the 2-Dimensional Autoregression (DnD) Transformer. The DnD-Transformer predicts more codes for an image by introducing a new direction, **model depth**, along with the sequence length. Compared to 1D autoregression and previous work using similar 2D image decomposition such as RQ-Transformer, the DnD-Transformer is an end-to-end model that can generate higher quality images with the same backbone model size and sequence length, opening a new optimization perspective for autoregressive image generation. Furthermore, our experiments reveal that the DnD-Transformer's potential extends beyond generating natural images. It can even generate images with rich text and graphical elements in a self-supervised manner, demonstrating an understanding of these combined modalities. This has not been previously demonstrated for popular vision generative models such as diffusion models, showing a spark of vision-language intelligence when trained solely on images. Code, datasets and models are open at https://github.com/chenllliang/DnD-Transformer.
Liang Chen 0024, Sinan Tan, Zefan Cai, Weichu Xie, Haozhe Zhao, Yichi Zhang 0010, Junyang Lin, Jinze Bai, Tianyu Liu 0001, Baobao Chang
ICLR2
2025 Self-Supervised 3-D Semantic Representation Learning for Vision-and-Language Navigation
abstract
In vision-and-language navigation (VLN) tasks, most current methods primarily utilize RGB images, overlooking the rich 3-D semantic data inherent to environments. To rectify this, we introduce a novel VLN framework that integrates 3-D semantic information into the navigation process. Our approach features a self-supervised training scheme that incorporates voxel-level 3-D semantic reconstruction to create a detailed 3-D semantic representation. A key component of this framework is a pretext task focused on region queries, which determines the presence of objects in specific 3-D areas. Following this, we devise an long short-term memory (LSTM)-based navigation model that is trained using our 3-D semantic representations. To maximize the utility of these 3-D semantic representations, we implement a cross-modal distillation strategy. This strategy encourages the RGB model's outputs to emulate those from the 3-D semantic feature network, enabling the concurrent training of both branches to merge RGB and 3-D semantic data effectively. Comprehensive evaluations on both the R2R and R4R datasets reveal that our method significantly enhances performance in VLN tasks.
Sinan Tan, Kuankuan Sima, Dunzheng Wang, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Mixed Neural Voxels for Fast Multi-view Video Synthesis
abstract
Synthesizing high-fidelity videos from real-world multi-view input is challenging due to the complexities of real-world environments and high-dynamic movements. Previous works based on neural radiance fields have demonstrated high-quality reconstructions of dynamic scenes. However, training such models on real-world scenes is time-consuming, usually taking days or weeks. In this paper, we present a novel method named MixVoxels to efficiently represent dynamic scenes, enabling fast training and rendering speed. The proposed MixVoxels represents the 4D dynamic scenes as a mixture of static and dynamic voxels and processes them with different networks. In this way, the computation of the required modalities for static voxels can be processed by a lightweight model, which essentially reduces the amount of computation as many daily dynamic scenes are dominated by static backgrounds. To distinguish the two kinds of voxels, we propose a novel variation field to estimate the temporal variance of each voxel. For the dynamic representations, we design an inner product time query method to efficiently query multiple time steps, which is essential to recover the high-dynamic movements. As a result, with 15 minutes of training for dynamic scenes with inputs of 300-frame videos, MixVoxels achieves better PSNR than previous methods. For rendering, MixVoxels can render a novel view video with 1K resolution at 37 fps. Codes and trained models are available at https://github.com/fengres/mixvoxels.
Feng Wang 0034, Sinan Tan, Xinghang Li, Zeyue Tian, Huaping Liu 0001
ICCV2
2023 Embodied Referring Expression for Manipulation Question Answering in Interactive Environment
abstract
Embodied agents are expected to perform more complicated tasks in an interactive environment, with the progress of Embodied AI in recent years. Existing embodied tasks including Embodied Referring Expression (ERE) and other QA-form tasks mainly focuses on interaction in term of linguistic instruction. Therefore, enabling the agent to manipulate objects in the environment for exploration actively has become a challenging problem for the community. To solve this problem, We introduce a new embodied task: Remote Embodied Manipulation Question Answering (REMQA) to combine ERE with manipulation tasks. In REMQA task, the agent needs to navigate to a remote position and perform manipulation with the target object to answer the question. We build a benchmark dataset for the REMQA task in AI2-THOR simulator. To this end, a framework with 3D semantic reconstruction and modular network paradigms is proposed. The evaluation of the proposed framework on REMQA dataset is presented to validate its effectiveness.
Qie Sima, Sinan Tan, Huaping Liu 0001, Fuchun Sun 0001
ICRA2
2023 Knowledge-Based Embodied Question Answering
abstract
In this paper, we propose a novel Knowledge-based Embodied Question Answering (K-EQA) task, in which the agent intelligently explores the environment to answer various questions with the knowledge. Different from explicitly specifying the target object in the question as existing EQA work, the agent can resort to external knowledge to understand more complicated question such as "Please tell me what are objects used to cut food in the room?", in which the agent must know the knowledge such as "knife is used for cutting food". To address this K-EQA problem, a novel framework based on neural program synthesis reasoning is proposed, where the joint reasoning of the external knowledge and 3D scene graph is performed to realize navigation and question answering. Especially, the 3D scene graph can provide the memory to store the visual information of visited scenes, which significantly improves the efficiency for the multi-turn question answering. Experimental results have demonstrated that the proposed framework is capable of answering more complicated and realistic questions in the embodied environment. The proposed method is also applicable to multi-agent scenarios.
Sinan Tan, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Depth-Aware Vision-and-Language Navigation using Scene Query Attention Network
abstract
Vision-and-language navigation (VLN) has been an important task in the field of Robotics and Computer Vision. However, most existing vision-and-language navigation models only use features extracted from RGB observation as input, while robots can utilize depth sensors in the real world. Existing research has also shown that simply adding a depth stream to neural models could only provide a marginal improvement to the performance of the VLN task. Therefore, in our work, we develop a novel method for the VLN task using semantic map observations built from RGB-D input. We use vision-pretraining to efficiently encode the semantic map with CNN and scene query attention network by answering queries about semantic information of specific regions of a scene. The proposed method could be used with a simple model and does not require large-scale vision-language transformer pretraining, bringing a more than 10% increase in the success rate compared with a baseline model. When used together with the Speaker-Follower training technique, it achieves a success rate of 58 % on the test set for the R2R dataset in single-run setting, outperforming the previous RGB-D method and most existing RGB-only models that do not use large-scale vision-language transformers pretraining.
Sinan Tan, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001
ICRA1
2020 Multi-agent Embodied Question Answering in Interactive Environments
Sinan Tan, Weilai Xiang, Huaping Liu 0001, Di Guo 0002, Fuchun Sun 0001
ECCV (13)1