Danyang Tu

dblp:312/7130 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-9632-9131ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Non-contrast CT esophageal varices grading through clinical prior-enhanced multi-organ analysis
Xiaoming Zhang 0008, Chunli Li, Jiacheng Hao, Yuan Gao 0017, Danyang Tu, Jianyi Qiao, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002, Ke Yan 0006
Medical Image Anal.5
2025 Anatomy-Aware Low-Dose CT Denoising via Pretrained Vision Models and Semantic-Guided Contrastive Learning
Zeli Chen, Zhiyun Song, Wei Fang 0005, Jiajin Zhang, Danyang Tu, Yuxing Tang, Minfeng Xu, Xianghua Ye, Le Lu 0001, Dakai Jin
MICCAI (2)6
2024 Un-Gaze: A Unified Transformer for Joint Gaze-Location and Gaze-Object Detection
abstract
This paper proposes an efficient and effective method for joint gaze location detection (GL-D) and gaze object detection (GO-D), i.e., gaze following detection. Current approaches frame GL-D and GO-D as two separate tasks, employing a multi-stage framework where human head crops must first be detected and then be fed into a subsequent GL-D sub-network, which is further followed by an additional object detector for GO-D. In contrast, we reframe the gaze following detection task as detecting human head locations and their gaze followings simultaneously, aiming at jointly detect human gaze location and gaze object in a unified and single-stage pipeline. To this end, we propose GTR, short for Gaze following detection TRansformer, streamlining the gaze following detection pipeline by eliminating all additional components, leading to the first unified paradigm that unites GL-D and GO-D in a fully end-to-end manner. GTR enables an iterative interaction between holistic semantics and human head features through a hierarchical structure, inferring the relations of salient objects and human gaze from the global image context and resulting in an impressive accuracy. Concretely, GTR achieves a 12.1 mAP gain ($\mathbf {25.1}\%$) on GazeFollowing and a 18.2 mAP gain ($\mathbf {43.3\%}$) on VideoAttentionTarget for GL-D, as well as a 19 mAP improvement ($\mathbf {45.2\%}$) on GOO-Real for GO-D. Meanwhile, unlike existing systems detecting gaze following sequentially due to the need for a human head as input, GTR has the flexibility to comprehend any number of people’s gaze followings simultaneously, resulting in high efficiency. Specifically, GTR introduces over a$\times 9$improvement in FPS and the relative gap becomes more pronounced as the human number grows.
Danyang Tu, Wei Shen 0002, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.1
2023 MD-VQA: Multi-Dimensional Quality Assessment for UGC Live Videos
abstract
User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing of UGC live videos, effective video quality assessment (VQA) tools are needed to monitor and perceptually optimize live streaming videos in the distributing process. In this paper, we address UGC Live VQA problems by constructing a first-of-a-kind subjective UGC Live VQA database and developing an effective evaluation tool. Concretely, 418 source UGC videos are collected in real live streaming scenarios and 3,762 compressed ones at different bit rates are generated for the subsequent subjective VQA experiments. Based on the built database, we develop a Multi-12imensional VQA (MD-VQA) evaluator to measure the visual quality of UGC live videos from semantic, distortion, and motion aspects respectively. Extensive experimental results show that MD-VQA achieves state-of-the-art performance on both our UGC Live VQA database and existing compressed UGC VQA databases.
Wei Wu 0002, Wei Sun 0029, Danyang Tu, Wei Lu 0021, Xiongkuo Min, Ying Chen 0011, Guangtao Zhai
CVPR4
2023 Agglomerative Transformer for Human-Object Interaction Detection
abstract
We propose an agglomerative Transformer (AGER) that enables Transformer-based human-object interaction (HOI) detectors to flexibly exploit extra instance-level cues in a single-stage and end-to-end manner for the first time. AGER acquires instance tokens by dynamically clustering patch tokens and aligning cluster centers to instances with textual guidance, thus enjoying two benefits: 1) Integrality: each instance token is encouraged to contain all discriminative feature regions of an instance, which demonstrates a significant improvement in the extraction of different instance-level cues and subsequently leads to a new state-of-the-art performance of HOI detection with 36.75 mAP on HICO-Det. 2) Efficiency: the dynamical clustering mechanism allows AGER to generate instance tokens jointly with the feature learning of the Transformer encoder, eliminating the need of an additional object detector or instance decoder in prior methods, thus allowing the extraction of desirable extra cues for HOI detection in a single-stage and end-to-end pipeline. Concretely, AGER reduces GFLOPs by 8.5% and improves FPS by 36%, even compared to a vanilla DETR-like pipeline without extra cue extraction. The code will be available at https://github.com/six6607/AGER.git.
Danyang Tu, Wei Sun 0029, Guangtao Zhai, Wei Shen 0002
ICCV1
2023 BH-VQA: Blind High Frame Rate Video Quality Assessment
abstract
High frame rate (HFR) videos can provide consumers with a more immersive viewing experience in motion-rich scenes. However, they also pose a great challenge for video compression and transmission due to the increase in frame rates. Therefore, it is very important to choose proper frame rates and bit rates to achieve a trade-off between transmission bandwidth and visual quality. In this paper, we propose a novel Blind HFR Video Quality Assessment (BH-VQA) model by exploring the efficient and effective motion representation from the deep neural network (DNN). Concretely, we first train a baseline VQA model (i.e. a backbone network and a regressor) on a large-scale VQA database to derive a powerful quality-aware feature extractor for the spatial and motion feature extraction. Then, the HFR video is split into a sequence of video clips and the spatial features of each video clip are extracted just using the first frame of the video clip. To capture temporal distortions caused by frame rate variations and object and camera motion, we calculate deep structural similarities between continuous frames of each video clip as the motion features. Finally, the temporal quality dependencies between video clips are learned through a gated recurrent unit (GRU) network to obtain the perceptual video quality score. Experimental results show that BH-VQA achieves the best performance on two publicly available HFR VQA databases. The code of BH-VQA will be released.
Wei Lu 0021, Wei Sun 0029, Danyang Tu, Xiongkuo Min, Guangtao Zhai
ICME4
2022 End-to-End Human-Gaze-Target Detection with Transformers
abstract
In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head locations must first be detected and then be fed into the next gaze target prediction sub-network. In contrast, we redefine the HGT detection task as detecting human head locations and their gaze targets, simultaneously. By this way, our method, named Human-Gaze-Target detection TRansformer or HGTTR, streamlines the HGT detection pipeline by eliminating all other additional components. HGTTR reasons about the relations of salient objects and human gaze from the global image context. Moreover, unlike existing two-stage methods that require human head locations as input and can predict only one human's gaze target at a time, HGTTR can directly predict the locations of all people and their gaze targets at one time in an end-to-end manner. The effectiveness and robustness of our proposed method are verified with extensive experiments on the two standard benchmark datasets, GazeFollowing and VideoAttentionTarget. Without bells and whistles, HGTTR outperforms existing state-of-the-art methods by large margins (6.4 mAP gain on GazeFollowing and 10.3 mAP gain on VideoAttentionTarget) with a much simpler architecture.
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
CVPR1
2022 Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
ECCV (4)1
2022 Saliency in Augmented Reality
abstract
With the rapid development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary theory underlying AR is human visual confusion, which allows users to perceive the real-world scenes and augmented contents (virtual-world scenes) simultaneously by superimposing them together. To achieve good Quality of Experience (QoE), it is important to understand the interaction between two scenarios, and harmoniously display AR contents. However, studies on how this superimposition will influence the human visual attention are lacking. Therefore, in this paper, we mainly analyze the interaction effect between background (BG) scenes and AR contents, and study the saliency prediction problem in AR. Specifically, we first construct a Saliency in AR Dataset (SARD), which contains 450 BG images, 450 AR images, as well as 1350 superimposed images generated by superimposing BG and AR images in pair with three mixing levels. A large-scale eye-tracking experiment among 60 subjects is conducted to collect eye movement data. To better predict the saliency in AR, we propose a vector quantized saliency prediction method and generalize it for AR saliency prediction. For comparison, three benchmark methods are proposed and evaluated together with our proposed method on our SARD. Experimental results demonstrate the superiority of our proposed method on both of the common saliency prediction problem and the AR saliency prediction problem over benchmark methods. Our dataset and code are available at: https://github.com/DuanHuiyu/ARSaliency.
Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Danyang Tu, Jing Li 0026, Guangtao Zhai
ACM Multimedia4
2022 Video-based Human-Object Interaction Detection from Tubelet Tokens
abstract
We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatial-temporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related patch tokens along spatial and temporal domains, which enjoy two benefits: 1) Compactness: each token is learned by a selective attention mechanism to reduce redundant dependencies from others; 2) Expressiveness: each token is enabled to align with a semantic instance, i.e., an object or a human, thanks to agglomeration and linking. The effectiveness and efficiency of TUTOR are verified by extensive experiments. Results show our method outperforms existing works by large margins, with a relative mAP gain of $16.14\%$ on VidHOI and a 2 points gain on CAD-120 as well as a $4 \times$ speedup.
Danyang Tu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Wei Shen 0002
NeurIPS1
2021 Accurate Compensation Makes the World More Clear for the Visually Impaired
abstract
Visual impairment is one of the most serious social and public health problems in the world, therefore, it is of great theoretical and practical significance to study the image enhancement algorithms for the visually impaired, which is the basis for the development of assistive devices. In this paper, a general deep learning based image enhancement framework for the visually impaired is proposed, which can be used to enhance images to compensate for any visually impaired symptom that can be modeled. Take central vision loss as an example, we first model the central vision loss based on the contrast sensitivity function (CSF) specified by clinical indicator Pelli-Robson score and logMAR visual acuity, and then use the proposed framework to generate an image enhancement method aiming at compensating for the central vision loss. Both the simulation experiment and the patient experiment show the superiority of the proposed image enhancement method designed for the central vision loss, which also validates the effectiveness of the proposed framework.
Sijing Wu, Huiyu Duan, Xiongkuo Min, Danyang Tu, Guangtao Zhai
ICIP4