Yijin Yang

dblp:33/6721 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 "Eyes" On Me! 3- to 6-Year-Old Children Anthropomorphize Humanoid Robot with Biological and Mental Attributions
abstract
The increasing integration of humanoid robots into educational and social contexts underscores the importance of understanding how young children perceive robots. This study examines how children aged 3–6 years (N = 90, 54 females) and adults (N = 24, 11 females) attribute physical, biological, and mental properties to humanoid robots, assessing the impact of humanoid appearance on anthropomorphic perceptions. Participants engaged in a structured card-choice task featuring stimuli varying in humanoid characteristics, including an intelligent phone and robots with differing degrees of humanoid features. Results revealed that while children predominantly recognized robots as “non-living,” their attributions were significantly influenced by anthropomorphic features, particularly eyes and limbs. Younger children (3–5 years) demonstrated higher levels of anthropomorphism compared to older children and adults, frequently attributing biological and mental properties to robots with pronounced human-like features. These findings advance previous research by documenting the developmental trajectory of anthropomorphic perceptions from early preschool (ages 3–6) to adulthood. In contrast to prior literature, our study manipulates humanoid appearances along a spectrum (from non-humanoid devices to fully humanoid robots), identifying expressive eyes and articulated limbs as distinct visual features that disproportionately drive young children’s anthropomorphic attributions. This developmental perspective and systematic stimulus comparison provide novel insights with direct implications for designing educational robots tailored for young children. Future research should explore interactive stimuli to further understand the detailed effects of humanoid robotics on child development.
Keyu Mao, Zisong Li, Yijin Yang, Liqi Zhu
ACM Trans. Hum. Robot Interact.4
2025 Attention-Based Gating Network for Robust Segmentation Tracking
abstract
Visual object tracking is a challenging task that aims to accurately estimate the scale and position of a designated target. Recently, segmentation networks have proven effective in visual tracking, producing outstanding results for target scale estimation. However, segmentation-based trackers still lack robustness due to the presence of similar distractors. To mitigate this issue, we propose an Attention-based Gating Network (AGNet) that produces gating weights to diminish the impact of feature maps linked to similar distractors. Subsequently, we incorporate the AGNet into the segmentation-based tracking paradigm to achieve accurate and robust tracking. Specifically, the AGNet utilizes three cascading Multi-Head Cross-Attention (MHCA) modules to generate gating weights that govern the generation of feature maps in the baseline tracker. The proficiency of the MHCA in modeling global semantic information effectively suppresses feature maps associated with similar distractors. Additionally, we introduce a distractor-aware training strategy that leverages distractor masks to train our model. To alleviate the issue of partial occlusion, we introduce a box refinement module to enhance the accuracy of the predicted target box. Comprehensive experiments conducted on 11 challenging tracking benchmarks show that our approach significantly surpasses the baseline tracker across all metrics and achieves excellent results on multiple tracking benchmarks.
Yijin Yang, Xiaodong Gu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Conversational Drug Editing Using Retrieval and Domain Feedback
abstract
Recent advancements in conversational large language models (LLMs), such as ChatGPT, have demonstrated remarkable promise in various domains, including drug discovery. However, existing works mainly focus on investigating the capabilities of conversational LLMs on chemical reactions and retrosynthesis. While drug editing, a critical task in the drug discovery pipeline, remains largely unexplored. To bridge this gap, we propose ChatDrug, a framework to facilitate the systematic investigation of drug editing using LLMs. ChatDrug jointly leverages a prompt module, a retrieval and domain feedback module, and a conversation module to streamline effective drug editing. We empirically show that ChatDrug reaches the best performance on all 39 drug editing tasks, encompassing small molecules, peptides, and proteins. We further demonstrate, through 10 case studies, that ChatDrug can successfully identify the key substructures for manipulation, generating diverse and valid suggestions for drug editing. Promisingly, we also show that ChatDrug can offer insightful explanations from a domain-specific perspective, enhancing interpretability and enabling informed decision-making.
Shengchao Liu, Jiongxiao Wang, Yijin Yang, Ling Liu 0001, Chaowei Xiao
ICLR3
2024 ChatGPT as an Attack Tool: Stealthy Textual Backdoor Attack via Blackbox Generative Model Trigger
abstract
Jiazhao Li, Yijin Yang, Zhuofeng Wu, V.G. Vinod Vydiswaran, Chaowei Xiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiazhao Li, Yijin Yang, Zhuofeng Wu 0001, V. G. Vinod Vydiswaran, Chaowei Xiao
NAACL-HLT2
2024 Differentially Private Video Activity Recognition
abstract
In recent years, differential privacy has seen significant advancements in image classification; however, its application to video activity recognition remains under-explored. This paper addresses the challenges of applying differential privacy to video activity recognition, which primarily stem from: (1) a discrepancy between the desired privacy level for entire videos and the nature of input data processed by contemporary video architectures, which are typically short, segmented clips; and (2) the complexity and sheer size of video datasets relative to those in image classification, which render traditional differential privacy methods inadequate. To tackle these issues, we propose Multi-Clip DP-SGD, a novel framework for enforcing video-level differential privacy through clip-based classification models. This method samples multiple clips from each video, averages their gradients, and applies gradient clipping in DP-SGD without incurring additional privacy loss. Moreover, we incorporate a parameter-efficient transfer learning strategy to make the model scalable for large-scale video datasets. Through extensive evaluations on the UCF-101 and HMDB-51 datasets, our approach exhibits impressive performance, achieving 81% accuracy with a privacy budget of ϵ = 5 on UCF-101, marking a 76% improvement compared to a direct application of DP-SGD. Furthermore, we demonstrate that our transfer learning strategy is versatile and can enhance differentially private image classification across an array of datasets including CheXpert, ImageNet, CIFAR-10, and CIFAR-100.
Zelun Luo, Yuliang Zou, Yijin Yang, Zane Durante, De-An Huang, Zhiding Yu, Chaowei Xiao, Li Fei-Fei 0001, Anima Anandkumar
WACV3
2024 Learning Dynamical Position Embedding for Discriminative Segmentation Tracking
abstract
Visual tracking plays a pivotal role in intelligent transportation systems and has a wide range of practical applications such as autonomous driving and traffic counting. Recently, the attention mechanism in Transformers has been successfully applied to the field of visual tracking, leading to a significant improvement in tracking performance. However, Transformer-based trackers directly flatten two-dimensional image features into one-dimensional vectors to compute attention scores. This process unavoidably results in the omission of crucial position distribution information necessary for precise target localization. To address this issue, we propose a novel cross-attention based tracking-by-segmentation framework, called Dynamical Position Embedding based Tracking framework (DPET). DPET incorporates an additional network for modeling position information to complement the cross-attention module. To be specific, a dynamical position embedding network is introduced to adaptively encode position information. This network is then integrated into the cross-attention based feature fusion network to compensate for the loss of position distribution information. As a result, the fused feature incorporates abundant contextual semantic cues for target classification and precise position information for target localization simultaneously. To overcome the constraints imposed by bounding-boxes, a segmentation network that takes the fused feature as input is designed to achieve accurate pixel-wise tracking. Extensive experiments on eight challenging tracking benchmarks show that our DPET tracker enables real-time operations and achieves promising tracking performance on the GOT-10K benchmark. Especially, DPET tracker achieves the top accuracy scores on VOT2016, VOT2018 and VOT2019 benchmarks.
Yijin Yang, Xiaodong Gu 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Multi Feature Representation and Aggregation Network for Accurate and Robust Visual Tracking
abstract
Segmentation-based tracking paradigm has been successfully applied to tracking field and significantly improves the tracking performance. Although segmentation-based trackers are effective for target scale estimation, it makes the trackers have high requirements for the extracted target features due to the need for pixel-level segmentation. Therefore, in this article, we propose a novel multi feature representation and aggregation network and introduce it into tracking-by-segmentation framework to extract and integrate rich features for segmentation-based tracking. To be specific, the proposed approach firstly models three complementary feature representations through cross-attention, cross-correlation and dilated involution mechanisms respectively and employ a simple feature aggregation network to fuse these features. And then feeding those fusion features into a segmentation network obtains the accurate target state estimation. In addition, we introduce a bounding box refinement module to further refine the target box to alleviate the issues of partial occlusion and surrounding distractors. The extensive experimental results show that the proposed tracker achieves very promising tracking performance on seven challenging visual tracking benchmarks. Code and models are available at https://github.com/Yang428/FEAST.
Yijin Yang, Xiaodong Gu 0001
IJCNN1
2023 Learning rich feature representation and aggregation for accurate visual tracking
Yijin Yang, Xiaodong Gu 0001
Appl. Intell.1
2023 Accurate and robust visual tracking using bounding box refinement and online sample filtering
Yijin Yang, Xiaodong Gu 0001
Signal Process. Image Commun.1
2023 Joint Correlation and Attention Based Feature Fusion Network for Accurate Visual Tracking
abstract
Correlation operation and attention mechanism are two popular feature fusion approaches which play an important role in visual object tracking. However, the correlation-based tracking networks are sensitive to location information but loss some context semantics, while the attention-based tracking networks can make full use of rich semantic information but ignore the position distribution of the tracked object. Therefore, in this paper, we propose a novel tracking framework based on joint correlation and attention networks, termed as JCAT, which can effectively combine the advantages of these two complementary feature fusion approaches. Concretely, the proposed JCAT approach adopts parallel correlation and attention branches to generate position and semantic features. Then the fusion features are obtained by directly adding the location feature and semantic feature. Finally, the fused features are fed into the segmentation network to generate the pixel-wise state estimation of the object. Furthermore, we develop a segmentation memory bank and an online sample filtering mechanism for robust segmentation and tracking. The extensive experimental results on eight challenging visual tracking benchmarks show that the proposed JCAT tracker achieves very promising tracking performance and sets a new state-of-the-art on the VOT2018 benchmark.
Yijin Yang, Xiaodong Gu 0001
IEEE Trans. Image Process.1
2022 Dependent task offloading with energy-latency tradeoff in mobile edge computing
abstract
Abstract With the rapid development of Internet‐of‐Things (IoT) and mobile devices, the IoT applications become more computation‐intensive and latency‐sensitive, which bring severe challenges to the resource‐limited devices. Mobile Edge Computing has served as a key promising method to enhance the network's computing capability by enabling resource‐constrained devices to offload tasks to the edge servers. A major challenge, which has been overlooked by most existing works on task offloading, is the dependencies among tasks and subtasks. In this paper, the subtask offloading with logical dependency for IoT applications is focused on. Specifically, subtask dependent graphs are employed to explore the dependency of subtasks and consider the priority of task scheduling. Further, an offloading scheme is put forward for minimizing both task latency and energy consumption of the device with dependency guarantees for all IoT tasks in multi‐server edge networks. Finaly, the simulation results demonstrate that the overall reduction rate is around 14% and relatively stable can effectively reduce task latency in multi‐server edge networks.
Jian Chen 0002, Yuchen Zhou 0001, Long Yang 0002, Bingtao He, Yijin Yang
IET Commun.6
2021 Pixel-Attention CNN With Color Correlation Loss for Color Image Denoising
abstract
Convolutional neural networks (CNNs) have been applied to many image processing tasks and achieve great successes. In order to extract common features, every pixel in an image shares the same filters. However, pixels in different regions of an image varies dramatically and shared filters may lose some important local information. Rather than shared filters, smart filters which can be adapted to image context should be designed to better remove noise which occurs randomly in noisy image. Meanwhile, current CNN architectures compute the loss of each color channel independently, regardless of the potential color information. In this letter, we proposed a pixel-attention convolutional neural network (PACNN) with color correlation loss for the color image denoising task. The pixel-attention mechanism could generate pixel-wise attention maps which help remove random noise. The color correlation loss exploits color correlation to further improve denoising performance on color noisy images. The experimental results on several standard datasets demonstrate the state-of-the-art (SOTA) performance and the superiority of the proposed method.
Fan Jia 0007, Liyan Ma, Yijin Yang, Tieyong Zeng
IEEE Signal Process. Lett.3
2007 Efficient Adaptive Resource Allocation for Multiuser OFDM Systems with Minimum Rate Constraints
abstract
This work investigates the adaptive resource allocation scheme for the downlink of multiuser OFDM systems. We focus on the problem of maximizing the overall spectral efficiency while maintaining the QoS requirements of users, including bit error rate and individual minimum rate requirements. Under the assumption of equal power allocation, an efficient algorithm is proposed to obtain the suboptimal solution of the resource allocation. In this algorithm, we first introduce some positive multipliers, one for each user, according to their minimum rate constraints (MRC), and then a parallel subcarrier-and-bit allocation scheme is designed using these multipliers with low complexity. The numerical results show that our algorithm not only substantially reduces the computational complexity of the existing algorithm, but also provides a noticeable performance improvement.
Wei Xu 0001, Chunming Zhao 0001, Yijin Yang
ICC4
2007 MIMO Detection of 16-QAM Signaling Based on Semidefinite Relaxation
abstract
We propose a computationally efficient semidefinite relaxation (SDR) of the maximum likelihood (ML) detector for 16 quadrature amplitude modulation (16-QAM) in multiple-input multiple-output (MIMO) systems. The SDR stems from a variant of the ML problem, which utilizes set operation algorithms to formulate the alphabet constraint. Theoretical analysis and numerical simulations demonstrate that the proposed method can provide improved error performance as compared to some previous detectors.
Yijin Yang, Chunming Zhao 0001, Wei Xu 0001
IEEE Signal Process. Lett.1
2006 Performance Analysis of PCC-OFDM Systems Impaired by Carrier Frequency Offset
abstract
Orthogonal frequency division multiplexing (OFDM) is sensitive to carrier frequency offset (CFO), which destroys the orthogonality and causes inter-carrier interference (ICI). ICI self-cancellation schemes based on polynomial cancellation coding (PCC-OFDM) can evidently reduce the sensitivity of systems to CFO. In this paper, we analyze the performance of PCC-OFDM systems over additive white Gaussian noise (AWGN) channels in the presence of CFO. Two criteria are used to evaluate the effect of the CFO on performance degradation. Firstly, the closed-form expressions of the statistical average carrier-to-interference power ratio (CIR) and ICI power are presented. The closed-form expressions depend only on the normalized CFO, and are irrelevant to the number of subcarriers. Secondly, by exploiting the properties of the Beaulieu series, the effect of the CFO on symbol error rate (SER) for PCC-OFDM systems modulated with BPSK, QPSK and 16-QAM is exactly expressed as the sum of an infinite series in terms of the characteristic function (CHF) of the ICI.
Chunming Zhao 0001, Yijin Yang, Zhihua Shi
GLOBECOM3
2003 Establishment of special city GIS based on ArcObjects
abstract
Because of the complex operation and poor relevance to special use, current general GIS software cannot usually meet the need of setting up a special GIS. The ESRI product ArcGIS provides a set of COM components, called ArcObjects. Based on ArcObjects, a special GIS can be set up conveniently. To develop the management information system for Fuzhou's active faults, we use Visual Basic as our developing tool, which call ArcObjects, making good use of existing VBA codes. In this system, basic data and professional data are integrated for display, query, and analysis. With the 3D analysis components, flying animation can present city information in three-dimensional space. Besides, some professional analysis modules are implemented by advanced languages and can be easily embedded into the system. It is an excellent software environment for deep research of the active faults and urban planning. The results indicate that further development based on component object models (COM) can flexibly construct application systems and have good user interfaces.
Yijin Yang, Guihua Yu
IGARSS3