Yanting Zhang 0001

dblp:210/4905-1 · DBLP profile ↗
← Back
43ranked-venue papers
11as first author
37since 2021 · last 2026
0000-0001-6317-1956ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 16 since 2021Databases, data management, data science and information retrieval · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2Theory of computation · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Geometry-Guided Depth Correction for Metric Relative Pose Estimation
abstract
In recent years, Monocular Depth Estimation (MDE) has evolved from predicting affine-invariant relative depth to estimating metric-scale (absolute) depth. However, local geometric inconsistencies in single-view depth maps and scale inconsistencies across different views still severely hinder their practical application in 3D matching and relative pose estimation. To address these challenges, we propose a geometry-guided depth correction framework for metric-scale relative pose estimation. Our approach first leverages pre-trained foundation models to extract initial metric depth, semi-dense correspondences, and high-dimensional semantic features from dual-view images. We then introduce a local depth refinement module to correct geometric deviations. Finally, the corrected depth of stereo-matched pairs is integrated into a differentiable RANSAC framework to jointly optimize the relative pose with consistent scale. Experiments on ScanNet and 7-Scenes demonstrate that our method achieves superior performance and robustness across various challenging scenarios.
Shibin Xie, Xiaokang Fang, Yanting Zhang 0001, Shen Cai
ICMR7
2025 ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy Prediction
abstract
Inferring the 3D structure of a scene from a single image is an ill-posed and challenging problem in the field of vision-centric autonomous driving. Existing methods usually employ neural radiance fields to produce voxelized 3D occupancy, lacking instance-level semantic reasoning and temporal photometric consistency. In this paper, we propose ViPOcc, which leverages the visual priors from vision foundation models (VFMs) for fine-grained 3D occupancy prediction. Unlike previous works that solely employ volume rendering for RGB and depth image reconstruction, we introduce a metric depth estimation branch, in which an inverse depth alignment module is proposed to bridge the domain gap in depth distribution between VFM predictions and the ground truth. The recovered metric depth is then utilized in temporal photometric alignment and spatial geometric alignment to ensure accurate and consistent 3D occupancy prediction. Additionally, we also propose a semantic-guided non-overlapping Gaussian mixture sampler for efficient, instance-aware ray sampling, which addresses the redundant and imbalanced sampling issue that still exists in previous state-of-the-art methods. Extensive experiments demonstrate the superior performance of ViPOcc in both 3D occupancy prediction and depth estimation tasks on diverse public datasets.
Xijing Zhang, Tanghui Li, Yanting Zhang 0001, Rui Fan 0001
AAAI5
2025 CoCoB: Adaptive Collaborative Combinatorial Bandits for Online Recommendation
Cairong Yan, Jinyi Han, Jin Ju, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)4
2025 KG-TS: Knowledge Graph-Driven Thompson Sampling for Online Recommendation
Cairong Yan, Hualu Xu, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)3
2025 Dynamic Bidirectional Attentional Mamba Model for EEG-Based Motor Imagery Classification
Qianzi Shen, Zijian Wang 0010, Yanting Zhang 0001, Cairong Yan
ICIC (27)4
2025 MCCVM: Multi-Scale Cross-axes Conv-VMamba for Medical Image Classification
abstract
The performance of medical image classification relies on the effective capture and balance of local and global features. Although hybrid models combining Convolutional Neural Networks (CNNs) and Transformers have achieved notable success, they face two critical challenges: insufficient cross-axes information modeling, which hampers spatial dependency capture, and difficulty dynamically balancing within multi-scale features essential for complex medical images. In addition, the quadratic complexity of the self-attention mechanism in Transformers limits efficiency on high-resolution imaging data. To tackle these issues, this paper proposes the Multi-scale Cross-axes Convolutional VMamba Model (MCCVM), which combines CNNs for local feature extraction, the Mamba model for efficient global feature processing, and novel mechanisms for cross-axes information modeling and feature fusion. The MCCVM incorporates a Convolutional VMamba Fusion Block (CMF), which replaces Transformers with the Mamba framework to enhance computational efficiency while maintaining global feature extraction. A Cross-axes Attention SSM Block (CASSM) is also introduced within the Mamba structure to better model cross-axes spatial dependencies. Finally, a channel-based deep convolutional gated feature fusion network (CGFFN) is employed to dynamically balance local and global features, ensuring a comprehensive representation of medical images. Extensive experiments on medical image datasets demonstrate the superiority and effectiveness of our MCCVM.
Zijian Wang 0010, Qianzi Shen, Yanting Zhang 0001, Cairong Yan
IJCNN4
2025 MMCDSR: a Multimodal and Cross-domain Fusion Framework for Sequential Recommendation
abstract
Sequential recommendation (SR) aims to predict users’ next actions based on historical interaction sequences. Classical methods are developed based on deep learning mechanisms such as CNN, RNN, and Transformer to capture users’ behavioral sequential patterns. Despite their certain achievements, they still face challenges like data sparsity and limited understanding of item features. To address these issues, researchers have proposed the application of auxiliary information to enhance SR. In this paper, we realized that both cross-domain and multimodal information can be leveraged as auxiliary information to further improve the performance of SR, but how to make full use of them is faced with problems such as semantic inconsistency, inadequate user preferences mining, and the introduction of noise. To this end, we propose a MultiModal Cross-Domain Sequential Recommendation (MMCDSR) method, serving as a framework to jointly model modal and domain information for application in the SR. In MMCDSR, we design, (1) a semantic contrastive learning module to align modal and domain representations of items, (2) a sequential interest discovery module for capturing user preferences from different perspectives, and (3) an adaptive attention fusion module to eliminate noise features and generate the final user representation for the recommendation. Extensive experiments on six datasets from Amazon demonstrate that MMCDSR effectively leverages multimodal and cross-domain information, alleviates data sparsity issues, and significantly outperforms current baseline models in recommendation accuracy.
Yitong Xu, Guohao Sun 0001, Jinhu Lu 0002, Xiu Susie Fang, Yanting Zhang 0001
IJCNN6
2025 Precise spiking neurons for fitting any activation function in ANN-to-SNN Conversion
Qianzi Shen, Xuhang Li, Yanting Zhang 0001, Zijian Wang 0010, Cairong Yan
Appl. Intell.4
2025 Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
abstract
Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception. A key difficulty lies in the modality gap between unstructured point clouds and structured images, especially under sparse single-frame LiDAR settings. Existing methods typically extract features separately from point clouds and images, then rely on hand-crafted or learned matching strategies. This separate encoding fails to bridge the modality gap effectively, and more critically, these methods struggle with the sparsity and noise of single-frame LiDAR, often requiring point cloud accumulation or additional priors to improve reliability. Inspired by recent progress in detector-free matching paradigms, we revisit the projection-based approach and introduce the detector-free framework for direct point-pixel matching between LiDAR and camera views. To further enhance matching reliability, we introduce a repeatability scoring mechanism that acts as a soft visibility prior. This guides the network to suppress unreliable matches in regions with low intensity variation, improving robustness under sparse input. Extensive experiments on KITTI, nuScenes, and MIAS-LCEC-TF70 benchmarks demonstrate that our method achieves state-of-the-art performance, outperforming prior approaches on nuScenes (even those relying on accumulated point clouds), despite using only single-frame LiDAR.
Yanting Zhang 0001, Fangjun Ding, Shen Cai, Yanchao Dong, Rui Fan 0001
IEEE Trans Autom. Sci. Eng.3
2025 Pose-Guided Transformer for Fine-Grained Action Quality Assessment
abstract
Action Quality Assessment (AQA) is a task aimed at automatically and fairly evaluating the level of movement execution, which holds significant importance for action understanding. Previous methods, while adept at extracting video features, often neglect human regions. This leads to a limited capability to discern subtle action differences and results in a lack of interpretative depth. In this work, we propose a Pose-Guided Transformer framework, termed PGT, for assessing action quality more accurately. Essentially, this framework incorporates pose information to augment human region features during video feature extraction. The PGT framework incorporates two critical modules: a pose-guided attention layer and a global-local feature extractor. The former is designed to isolate body-specific features, effectively minimizing background noise, while the latter further delineates fine-grained features by utilizing decomposed information from various human body parts. The proposed PGT achieves significant results on various challenging AQA benchmarks. Notably, on MTL-AQA dataset, with a Spearman’s rank correlation of 0.9630. Additionally, on the AQA-7 dataset, our approach achieves an average Spearman’s rank correlation of 0.8673, further validating the effectiveness of our method. These findings demonstrate that our framework excels in the task of action quality assessment, providing a viable solution for accurate and fair evaluation of movement execution.
Yanting Zhang 0001, Wenhao Chai, Cairong Yan, Wenhai Wang, Gaoang Wang
IEEE Trans. Circuits Syst. Video Technol.1
2024 UniAP: Towards Universal Animal Perception in Vision via Few-Shot Learning
abstract
Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is challenging to design a deep learning-based perception model that can freely adapt to different animals across various perception tasks, due to the varying poses of a large diversity of animals, lacking data on rare species, and the semantic inconsistency of different tasks. We introduce UniAP, a novel Universal Animal Perception model that leverages few-shot learning to enable cross-species perception among various visual tasks. Our proposed model takes support images and labels as prompt guidance for a query image. Images and labels are processed through a Transformer-based encoder and a lightweight label encoder, respectively. Then a matching module is designed for aggregating information between prompt guidance and the query image, followed by a multi-head label decoder to generate outputs for various tasks. By capitalizing on the shared visual characteristics among different animals and tasks, UniAP enables the transfer of knowledge from well-studied species to those with limited labeled data or even unseen species. We demonstrate the effectiveness of UniAP through comprehensive experiments in pose estimation, segmentation, and classification tasks on diverse animal species, showcasing its ability to generalize and adapt to new classes with minimal labeled examples.
Meiqi Sun, Zhonghan Zhao, Wenhao Chai, Hanjun Luo, Shidong Cao, Yanting Zhang 0001, Jenq-Neng Hwang, Gaoang Wang
AAAI6
2024 Dual-Path Multimodal Optimal Transport for Composed Image Retrieval
Cairong Yan, Yanting Zhang 0001, Yongquan Wan
ACCV (6)3
2024 MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
abstract
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat.
Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo 0002, Tian Ye 0001, Yanting Zhang 0001, Yan Lu 0001, Jenq-Neng Hwang, Gaoang Wang
CVPR10
2024 Generalized Correspondence Matching via Flexible Hierarchical Refinement and Patch Descriptor Distillation
abstract
Correspondence matching plays a crucial role in numerous robotics applications. In comparison to conventional hand-crafted methods and recent data-driven approaches, there is significant interest in plug-and-play algorithms that make full use of pre-trained backbone networks for multi-scale feature extraction and leverage hierarchical refinement strategies to generate matched correspondences. The primary focus of this paper is to address the limitations of deep feature matching (DFM), a state-of-the-art (SoTA) plug-and-play correspondence matching approach. First, we eliminate the pre-defined threshold employed in the hierarchical refinement process of DFM by leveraging a more flexible nearest neighbor search strategy, thereby preventing the exclusion of repetitive yet valid matches during the early stages. Our second technical contribution is the integration of a patch descriptor, which extends the applicability of DFM to accommodate a wide range of backbone networks pre-trained across diverse computer vision tasks, including image classification, semantic segmentation, and stereo matching. Taking into account the practical applicability of our method in real-world robotics applications, we also propose a novel patch descriptor distillation strategy to further reduce the computational complexity of correspondence matching. Extensive experiments conducted on three public datasets demonstrate the superior performance of our proposed method. Specifically, it achieves an overall performance in terms of mean matching accuracy of 0.68, 0.92, and 0.95 with respect to the tolerances of 1, 3, and 5 pixels, respectively, on the HPatches dataset, outperforming all other SoTA algorithms. Our source code, demo video, and supplement are publicly available at mias.group/GCM.
Ziwei Long, Yanting Zhang 0001, Jin Wu 0002, Zhijun Fang 0001, Rui Fan 0001
ICRA3
2024 DepthCloak: Projecting Optical Camouflage Patches for Erroneous Monocular Depth Estimation of Vehicles
abstract
Adhesive adversarial patches have been common used in attacks against the computer vision task of monocular depth estimation (MDE). Compared to physical patches permanently attached to target objects, optical projection patches show great flexibility and have gained wide research attention. However, applying digital patches for direct projection may lead to partial blurring or omission of details in the captured patches, attributed to high information density, surface depth discrepancies, and non-uniform pixel distribution. To address these challenges, in this work we introduce DepthCloak, an adversarial optical patch designed to interfere with the MDE of vehicles. To this end, we first simplify the patch to a gray pattern because the projected ''black-and-white light'' has strong robustness to ambient light. We propose a generative adversarial network (GAN) based approach to simulate projections and deduce a projectable list. Then, we employ neighborhood averaging to fill sparse depth values, compress all depth values into a reduced dynamic range via nonlinear mapping, and use these values to adjust the Gaussian blur radius as weight parameters, thereby simulating depth variation effects. Finally, by integrating Moiré pattern and applying style transfer techniques, we customize adversarial patches featuring regularly arranged characteristics. We deploy DepthCloak in real driving scenarios, and extensive experiments demonstrate that DepthCloak can achieve an attack success rate of over 80% in the physical world.
Huixiang Wen, Shizong Yan, Shan Chang, Jie Xu 0061, Hongzi Zhu, Yanting Zhang 0001, Bo Li 0001
ACM Multimedia6
2024 Taming Diffusion for Fashion Clothing Generation with Versatile Condition
Yanting Zhang 0001, Jingyi Guo, Cairong Yan, Zhijun Fang 0001
PRCV (5)1
2024 DiffFashion: Reference-Based Fashion Design With Structure-Aware Transfer by Diffusion Models
abstract
Image-based fashion design with AI techniques has attracted increasing attention in recent years. We focus on the reference-based fashion design task, where we aim to combine a reference appearance image and a clothing image to generate a new fashion clothing image. Although existing diffusion-based image translation methods have enabled flexible style transfer, it is often difficult to transfer the appearance of the image realistically during reverse diffusion. When the referenced appearance domain greatly differs from the source domain, it often leads to the collapse in the translation. To tackle this issue, we present a novel diffusion model-based unsupervised structure-aware transfer method, namelyDiffFashion. Our method is free of model tuning and structure-preserving and has high flexibility in transferring from images with large domain gaps. Specifically, based on the optimal transport properties, we keep a shared latent across the clothing image and reference appearance image to bridge the gap between the two domains in the denoising process, and the latent of the reference image is gradually adapted to the clothing domain. Simultaneously, the structure is transferred from the source clothing to the output fashion image with mixed guidance, including pre-trained Vision Transformer (ViT) guidance and a foreground mask guidance, to further preserve the structure and appearance semantics from source and reference images. Our experimental results show that the proposed method outperforms state-of-the-art baseline models, generating more realistic images in the fashion design task.
Shidong Cao, Wenhao Chai, Shengyu Hao, Yanting Zhang 0001, Hangyue Chen, Gaoang Wang
IEEE Trans. Multim.4
2023 Thompson Sampling with Time-Varying Reward for Contextual Bandits
Cairong Yan, Hualu Xu, Haixia Han, Yanting Zhang 0001, Zijian Wang 0010
DASFAA (2)4
2023 TransLink: Transformer-Based Embedding for Tracklets' Global Link
abstract
Multi-object tracking (MOT) is essential to many tasks related to the smart transportation. Detecting and tracking humans on the road can give a vital feedback for either the moving vehicle or traffic control to ensure better driving safety and traffic flow. However, most trackers face a common problem of identity (ID) switch, resulting in an incomplete human trajectory prediction. In this paper, we propose a Transformer-based tracklet linking method called TransLink to mitigate the association failures. Specifically, the self-attention mechanism is well exploited to get the feature representation for tracklets, followed by a multilayer perceptron to predict the association likelihood, which can be further used in determining the tracklet association. Experiments on the MOT dataset demonstrate the effectiveness of the proposed module in lifting the tracking performances.
Yanting Zhang 0001, Shuanghong Wang, Yuxuan Fan, Gaoang Wang, Cairong Yan
ICASSP1
2023 Learning Golf Swing Key Events from Gaussian Soft Labels Using Multi-Scale Temporal MLPFormer
abstract
A complete golf swing includes several key events. The standardization of poses in each key event is directly related to the hitting effect. Thus, it is meaningful for the players to analyze their poses, especially at key frames, so as to improve swing performances. With the rapid development of deep learning techniques in computer vision, we are able to detect key frames during a golf swing. In this paper, we propose a framework to recognize key events in golf swing based on pure monocular video data. To achieve this, we have combined attention mechanism in the backbone network to extract concise features and leveraged the transformer structure to fuse multi-scale temporal information to enhance the feature representation. Besides, we also introduce Gaussian kernels into the label generation process, which can effectively solve the problem of ambiguity in detecting key events within their neighbouring similar frames. Notably, our method achieves an average recognition accuracy of 83.4% (+7.3% compared with SwingNet) for eight golf swing events on GoIfDB dataset.
Yanting Zhang 0001, Fuyu Tu, Zijian Wang 0010, Dandan Zhu 0001
IJCNN1
2023 Enhancing Multi-Behavior Recommendations Through Capturing Dynamic Preferences
abstract
Multi-behavior recommendation has gained significant attention in recent years for its ability to outperform single-behavior models. Current research related to multi-behavior models leaves room for improvement in the following two areas. First, the noise carried by individual behaviors and the additional noise generated during behavior processing is often overlooked, and these can ultimately degrade recommendation performance. Second, the specific time period of behavioral interactions and the frequency of interactions within that time period are also not taken into account. To address the above limitations, we propose a multi-behavior recommendation model integrating dynamic preferences (MB-DP) that captures dynamic interests while smoothing and denoising multi-behavior information. MB-DP extracts low and high-order semantics from various behaviors and unifies the measurements to generate interaction predictions. Additionally, it analyzes the interaction time and frequency of each behavior using gated recurrent units to capture the dynamic preferences of users and improve the prediction values. Extensive experimental results on two real-world datasets show that MB-DP significantly improves recommendation performance compared to the state-of-the-art baselines.
Cairong Yan, Xiaopeng Guan, Haixia Han, Zhaohui Zhang 0001, Yanting Zhang 0001
Int. J. Softw. Eng. Knowl. Eng.5
2023 MIN: multi-dimensional interest network for click-through rate prediction
Cairong Yan, Xiaoke Li, Yanting Zhang 0001, Zijian Wang 0010, Yongquan Wan
Knowl. Inf. Syst.3
2022 Urban Digital Twins for Intelligent Road Inspection
abstract
Urban digital twin (UDT) technologies offer new opportunities for intelligent road inspection (IRI). This paper first reviews the state-of-the-art algorithms used in the two key components of UDT-based IRI systems: (1) multi-temporal, multi-dimension, multi-score, and heterogeneous road data acquisition, and (2) road distress detection. This paper then summarizes the UDTIRI competition, organized in conjunction with IEEE Bigdata 2022. More details on our competition are available at sites.google.com/view/udtiri-workshop/bigdata-2022.
Rui Fan 0001, Yikang Zhang 0001, Sicen Guo, Jiahang Li 0001, Shuai Su, Yanting Zhang 0001, Wenshuo Wang 0001, Yu Jiang 0003, Mohammud Junaid Bocus, Xingyi Zhu
IEEE Big Data7
2022 High-fidelity 3D Model Compression based on Key Spheres
abstract
In recent years, neural signed distance function (SDF) has become one of the most effective representation methods for 3D models. By learning continuous SDFs in 3D space, neural networks can predict the distance from a given query space point to its closest object surface, whose positive and negative signs denote inside and outside of the object, respectively. Training a specific network for each 3D model, which individually embeds its shape, can realize compressed representation of objects by storing fewer network (and possibly latent) parameters. Consequently, reconstruction through network inference and surface recovery can be achieved. In this paper, we propose an SDF prediction network using explicit key spheres as input. Key spheres are extracted from the internal space of objects, whose centers either have relatively larger SDF values (sphere radii), or are located at essential positions. By inputting the spatial information of multiple spheres which imply different local shapes, the proposed method can significantly improve the reconstruction accuracy with a negligible storage cost. Compared to previous works, our method achieves the high-fidelity and high-compression 3D object coding and reconstruction. Experiments conducted on three datasets verify the superior performance of our method.
Yuanzhan Li, Yuqi Liu 0001, Shen Cai, Yanting Zhang 0001
DCC6
2022 Automatic Moving Pose Grading for Golf Swing in Sports
abstract
Swing is a critical important part in golf and it involves the whole body movement when players hit the ball. A good swing requires proper posture, which needs a lot of practice to achieve the standardized full-body coordination. Considering that the amateur players often lack necessary supervision during self-practice, we introduce an automatic moving pose grading method based on monocular swing videos. Specifically, given a swing video, the 3D human poses are firstly extracted. Then, dynamic time warping (DTW) is introduced to perform moving pose alignment between a query and a reference video. Finally, a distance-based grading strategy is proposed based on the temporally aligned swing videos. The system can effectively solve the increasing need in golf teaching, where proper posture can be greatly aided by adopting the feedback when golf players practice swing by themselves.
Yanting Zhang 0001, Qing'Ao Wang, Fuyu Tu, Zijian Wang 0010
ICIP1
2022 Attribute-Guided Fashion Image Retrieval by Iterative Similarity Learning
abstract
Image retrieval methods in the fashion field mainly take advantage of query images that reflect user needs, without considering additional keywords that users can provide to specify the attributes in their interests. To achieve the fine-grained fashion retrieval, we propose an iterative similarity learning network (ISLN) for attribute-guided image retrieval, which takes a query image and a specified attribute as input, and outputs other images with the same or similar attribute values. The core of the network is the iterative similarity learning module, which leverages the aggressive learning ability of the deep neural network (DNN) to focus on the area of interest and extract a more accurate feature embedding during the learning process of image and text semantic mapping. Extensive experiments on FashionAI and DARN (+8.33% and +10.73% in mAP) datasets show that ISLN performs better than the state-of-the-art methods in fine-grained similarity retrieval tasks.
Cairong Yan, Yanting Zhang 0001, Yongquan Wan, Dandan Zhu 0001
ICME3
2022 On-Road Pedestrian Tracking Across Multiple Moving Cameras
abstract
With the rapid development of autonomous driving, tracking on-road pedestrians raises more attention in the public. Currently, most researches focus on single camera based tracking or tracking across multiple static cameras. Tracking across multiple moving cameras has not been well studied yet. In this paper, we propose a workflow for tracking pedestrians across multiple moving cameras, leveraging the state-of-the-art single camera based tracking method of FairMOT. We consider different factors such as appearance features, motion information, and camera spatial distribution to improve the tracking performance. The experimental results carried on a multi-target multi-moving camera tracking dataset show the feasibility of the proposed scheme in solving the tracking issue in a complex environmental setting.
Yanting Zhang 0001, Shuanghong Wang, Qingxiang Wang, Qiubo Huang, Cairong Yan
ICME1
2022 JointCTR: a joint CTR prediction framework combining feature interaction and sequential behavior learning
Cairong Yan, Xiaoke Li, Yanting Zhang 0001
Appl. Intell.4
2022 Detection-by-tracking of traffic signs in videos
Yanting Zhang 0001, Zijian Wang 0010, Ruoning Song, Cairong Yan, Yonggang Qi
Appl. Intell.1
2022 Recurrent spiking neural network with dynamic presynaptic currents based on backpropagation
abstract
In recent years, spiking neural networks (SNNs), which originated from the theoretical basis of neuroscience, have attracted neuromorphic computing and brain-like computing due to their advantages, such as neural dynamics and coding mechanism, which are similar to biological neurons. SNNs have become one of the mainstream frameworks in the field of brain-like computing. However, most of the Leaky Integrate-and-Fire (LIF) neuron models currently used by SNNs based on direct training of backpropagation (BP) do not consider the changes in the recurrent connections and the dynamic strength of neuron connections over time. This study presented the LIF neuron model with recurrent connections and a method for dynamically changing the presynaptic currents. Recurrent LIF neurons have an additional cyclic connection compared with classic LIF neurons. Their postsynaptic current stimulates a change in membrane potential at the next time point. Their dynamics were more similar to the activities of biological neurons. We also proposed an efficient and flexible BP training method for recurrent LIF neurons. On the basis of the above methods, we proposed the recurrent SNN with dynamic presynaptic currents based on backpropagation (RDS-BP). We test the proposed RDS-BP on three image data sets (MNIST, Fashion-MNIST and CIFAR-10) and two text data sets (IMDB and TREC). The results showed that the performance of RDS-BP not only exceeded the naive SNN models based on BP but also exceeded the SNN methods proposed in previous studies in recent years, which had excellent performance in previous experiments. Our work provides a new LIF neuron model with a recurrent connection and dynamic presynaptic current and a BP training arrangement for the proposed neuron, which could merit developments with neuromorphic and brain-like computing.
Zijian Wang 0010, Yanting Zhang 0001, Haibo Shi, Lei Cao 0002, Cairong Yan
Int. J. Intell. Syst.2
2022 Dynamic clustering based contextual combinatorial multi-armed bandit for online recommendation
abstract
Recommender systems still face a trade-off between exploring new items to maximize user satisfaction and exploiting those already interacted with to match user interests. This problem is widely recognized as the exploration/exploitation (EE) dilemma, and the multi-armed bandit (MAB) algorithm has proven to be an effective solution. As the scale of users and items in real-world application scenarios increases, their purchase interactions become sparser. Then three issues need to be investigated when building MAB-based recommender systems. First, large-scale users and sparse interactions increase the difficulty of user preference mining. Second, traditional bandits model items as arms and cannot deal with ever-growing items effectively. Third, widely used Bernoulli-based reward mechanisms only feedback 0 or 1, ignoring rich implicit feedback such as behaviors like click and add-to-cart. To address these problems, we propose an algorithm named Dynamic Clustering based Contextual Combinatorial Multi-Armed Bandits (DC3MAB), which consists of three configurable key components. Specifically, a dynamic user clustering strategy enables different users in the same cluster to cooperate in estimating the expected rewards of arms. A dynamic item partitioning approach based on collaborative filtering significantly reduces the scale of arms and produces a recommendation list instead of one item to provide diversity. In addition, a multi-class reward mechanism based on fine-grained implicit feedback helps better capture user preferences. Extensive empirical experiments on three real-world datasets demonstrate the superiority of our proposed DC3MAB over state-of-the-art bandits (On average, +75.8% in F1 and +54.3% in cumulative reward). The source code is available at https://github.com/HaixHan/DC3MAB.
Cairong Yan, Haixia Han, Yanting Zhang 0001, Dandan Zhu 0001, Yongquan Wan
Knowl. Based Syst.3
2021 Learning Fashion Similarity Based on Hierarchical Attribute Embedding
abstract
Embedding items directly into a common feature space, and then measuring the similarity by calculating the feature distance in this space, has become the main method for similarity learning in current fashion retrieval tasks. The method is simple and efficient, but it ignores the correlation among fashion attributes and the impact of these correlations on the feature space, thereby reducing the accuracy of retrieval. Since the number of fashion attributes is large and the semantic granularity is also different, how to capture the relationship between fashion attributes and perform refined embedding to accurately represent fashion items is a challenge. In this paper, by constructing an attribute tree, we propose a hierarchical attribute embedding method for representing fashion items to enhance the relationship between attributes and use masking technology to disentangle different attributes. Based on these modules, we propose a hierarchical attribute-aware embedding network (HAEN) which takes images and attributes as input, learns multiple attribute-specific embedding spaces, and measures fine-grained similarity in the corresponding spaces. The extensive experimental result on two fashion-related public datasets FashionAI and DARN shows the superiority (+5.11% and +3.09% in MAP, respectively) of our proposed HAEN compared with state-of-the-art methods.
Cairong Yan, Anan Ding, Yanting Zhang 0001, Zijian Wang 0010
DSAA3
2021 Two-Phase Multi-armed Bandit for Online Recommendation
abstract
Personalized online recommendations strive to adapt their services to individual users by making use of both item and user information. Despite recent progress, the issue of balancing exploitation-exploration (EE) [1] remains challenging. In this paper, we model the personalized online recommendation of e-commence as a two-phase multi-armed bandit problem. This is the first time that “big arm” and “small arm” are introduced into multi-armed bandit (MAB), and a two-stage strategy is adopted to provide target users with the most suitable recommendation list. In the first phase, MAB is used to obtain an item subset that users may be interested in from a large number of items. We use item categories as arms instead of individual items in existing related models to control the arm scale and reduce computational complexity. In the second phase, we directly use the items generated in the first phase as arms of MAB and obtain rewards through fine-grained implicit feedback from users. Empirical studies on three real-world datasets show that our proposed method TPBandit performs better than state-of-the-art bandit-based recommendation methods in several evaluation metrics such as Precision, Recall, and Hit Ratio. Moreover, the two-phase method improves the recommendation performance by nearly 50% compared to the one-phase method in the best case.
Cairong Yan, Haixia Han, Zijian Wang 0010, Yanting Zhang 0001
DSAA4
2021 Vehicle 3d Localization in Road Scenes VIA a Monocular Moving Camera
abstract
Knowing the 3D locations of the surrounding vehicles is of vital importance in autonomous driving scenarios. It can be pretty challenging to make an accurate estimation from a monocular moving camera. In this paper, we present an effective vehicle 3D localization method, that utilizes 2D key-points predicted from a trained CNN to model the vehicles’ structure, from which the ground points are further inferred. An adaptive ground plane estimation method is exploited under the monocular camera for 3D geometric back-projection. Benefiting from tracking, we also take into account temporal information of the same object to ensure the trajectory consistency. Viewpoint and size knowledge are also considered for refinement. The evaluation on the KITTI benchmark for on-road vehicles shows the effectiveness of our proposed approach with promising 3D localization results.
Yanting Zhang 0001, Aotian Zheng, Yizhou Wang 0005, Jenq-Neng Hwang
ICASSP1
2021 SN-Graph: A Minimalist 3D Object Representation for Classification
abstract
Using deep learning techniques to process 3D objects has achieved many successes. However, few methods focus on the representation of 3D objects, which could be more effective for specific tasks than traditional representations, such as point clouds, voxels, and multi-view images. In this paper, we propose a Sphere Node Graph (SN-Graph) to represent 3D objects. Specifically, we extract a certain number of internal spheres (as nodes) from the signed distance field (SDF), and then establish connections (as edges) among the sphere nodes to construct a graph, which is seamlessly suitable for 3D analysis using graph neural network (GNN). Experiments conducted on the ModelNet40 dataset show that when there are fewer nodes in the graph or the tested objects are rotated arbitrarily, the classification accuracy of SN-Graph is significantly higher than the state-of-the-art methods.
Yuqi Liu 0001, Shen Cai, Yanting Zhang 0001, Yuanzhan Li, Xiaoyu Chi
ICME5
2021 A Multi-Task Learning Approach for Recommendation based on Knowledge Graph
abstract
Sparsity and cold start problem are two classic problems of collaborative filtering. To alleviate these issues, researchers usually add side information to the recommendation models to boost the performance. In this paper, we propose a multi-task learning approach for recommendation based on knowledge graph (KGeRec), which takes recommendation as the main task and the knowledge graph as an auxiliary task to provide side information for recommendation. To fully capture the correlation information between these two tasks, a feature interaction layer (FlU) based on cross networks is designed to share features between them. Besides, a side information embedding layer (SIE) is also designed in the recommendation task to exploit more feature information. We apply KGeRec to three public datasets about movie, book, and music. Experimental results show that the proposed KGeRec outperforms the state-of-the-art approaches (+2.2% in AUC, +2.6% in Accuracy, +2.5% and in F1-score, compared to the maximum value in Type I models; +1.3% in AUC, +0.8% in Accuracy, and +2% in F1-score, compared to the maximum value in Type II models) and it performs well in sparse datasets. We also validate the effectiveness of knowledge graphs in improving recommendation performance.
Cairong Yan, Yanting Zhang 0001, Zijian Wang 0010, Pengwei Wang 0001
IJCNN3
2021 Modeling Long- and Short-Term User Behaviors for Sequential Recommendation with Deep Neural Networks
abstract
In e-commerce platforms, a user's next behavior will be affected by his long-term constant interests and short-term temporal needs. Such information is usually hidden in the users' historical online behavior data, so how to capture long-term and short-term patterns becomes the key to design better recommendation models or algorithms. Current mainstream methods such as Markov chain, convolutional neural network, and recurrent neural network cannot well express the mixed dynamic characteristics. In this paper, we propose an attention-based deep neural network (ADNNet) to solve the problem. In ADNNet, a convolutional neural network is used to extract the short-term patterns in the behavior sequences, and a gated recurrent unit is used to mine the long-term patterns in the behavior sequences. The attention mechanism is adopted to help the network automatically learn the best fusion coefficient of these two patterns. Our experimental result on four real public datasets (+0.69% in Hit Ratio and +3.49% in MRR) shows the superiority of our proposed ADNNet compared with other state-of-the-art methods.
Cairong Yan, Yanting Zhang 0001, Zijian Wang 0010, Pengwei Wang 0001
IJCNN3
2020 Improved Traffic Sign Detection In Videos Through Reasoning Effective RoI Proposals
abstract
Traffic sign detection is an important task in assisted safety and autonomous driving. It is important to continuously detect the traffic signs emerged on the road. Currently, most object detection methods make independent detections based on single images. When we apply these methods directly to a video clip to detect traffic signs without taking into account temporal correlations among adjacent frames, missed detections or incorrect detections can frequently occur due to motion blur, size change, partial occlusion, and/or bad pose. In this paper, we fully exploit the temporal consistency of traffic sign detection in videos. More specifically, we incorporate information of adjacent frames with high confidence scores to enhance the discovery of potential objects in the missed or incorrect detected frames by “recovering” the missed RoI proposals or by “improving” the incorrect RoI proposals with low confidence scores. Our method can be regarded as a “detection-by-tracking” strategy, which results in a more robust detection performance in videos.
Yanting Zhang 0001, Yonggang Qi, Jie Yang 0023, Jenq-Neng Hwang
ICME1
2019 Bundle Adjustment for Monocular Visual Odometry Based on Detected Traffic Sign Features
abstract
Monocular visual odometry (VO), which is a subset of simultaneous localization and mapping (SLAM) used to determine the position and orientation of a moving object by analyzing the associated monocular camera image sequences, is a critical part in the vision system of autonomous driving. However, based on the frame-by-frame pose estimation, drift error can be incrementally accumulated. Bundle adjustment (BA) is thus introduced to deal with the error-drift problem through correlating several image frames together to optimize camera poses and extracted 3D map points simultaneously. In this paper, we propose a joint BA framework which takes into account additional constraints from the detected road traffic signs. This framework can be effectively integrated into existing VO systems, as evidenced by the improved vehicular localization accuracy in experimental performance when compared with the state-of-the-art baseline VO method.
Yanting Zhang 0001, Jie Yang 0023, Haotian Zhang 0005, Jenq-Neng Hwang
ICIP1
2018 CTSD: A Dataset for Traffic Sign Recognition in Complex Real-World Images
abstract
Traffic sign recognition (TSR) is an indispensable component for vision-based system of self-driving car. Promising results have been achieved which especially benefit from the rapid development of deep neural networks recently. However, there are few works focusing on the algorithms’ performances towards different complex conditions, such as weather and viewpoint variations. In this paper, we propose a new real-world TSR dataset, which is a dataset with several fine-grained conditions fine labeled involving weather, light condition, occlusion, distance, color fading and camera angle. Detailed and unbiased comparison results are reported about the performances of several state-of-the-arts on our proposed and five public TSR datasets. Experimental results demonstrate that current arts for TSR are still far from satisfactory especially when it comes to complex real-world cases.
Yanting Zhang 0001, Yonggang Qi, Jun Liu 0014, Jie Yang 0023
VCIP1
2018 Analyze users' online shopping behavior using interconnected online interest-product network
abstract
In recent years, with the rapid development of Internet, more and more users shop and socialize online, which produces a large amount of traffic data. Analyzing users' online shopping behavior is important for merchants to improve profit. Social networks are widely used in studying the relationship among users. In this paper, we apply the concept of social networks to online shopping behavior, and present a novel network perspective on the interconnected nature of online interest and product, allowing us to capture the attribute of online products reflecting users' sociability. Sociability in this paper is not a traditional social connection but common interests(defined by produce visiting behavior) among users. First, we build an interconnected online interest-product network which includes online shopping based social network (OSSN), online product network and intermediate layer, and we define popular online products and non-popular online products according to the number of distinct visitors of one product. Then, we analyze some indicators about OSSN, and find that OSSN has similar characteristics with traditional social networks, such as small world feature and homophily. Finally, we analyze how online products reflect users' sociability. With the number of online products users have visited increasing, the number of neighbors with similar interests first presents a positive correlation, and then disappears, which implies that the number of users' neighbors depends on the attribute of online products rather than the number of online products users have browsed. We define a formula to measure the attribute of online products reflecting users' sociability, and find that there is a distinct difference among different categories of online products, and non-popular online products often have high value of attribute to reflect users' sociability and the value is low in popular online products.
Shuangshuang Han, Yuanyuan Qiao 0002, Yanting Zhang 0001, Wenhui Lin, Jie Yang 0023
WCNC3
2018 A hybrid Markov-based model for human mobility prediction
Yuanyuan Qiao 0002, Zhongwei Si, Yanting Zhang 0001, Fehmi Ben Abdesslem, Xinyu Zhang 0017, Jie Yang 0023
Neurocomputing3
2018 A Survey on Machine Learning-Based Mobile Big Data Analysis: Challenges and Applications
abstract
This paper attempts to identify the requirement and the development of machine learning‐based mobile big data (MBD) analysis through discussing the insights of challenges in the mobile big data. Furthermore, it reviews the state‐of‐the‐art applications of data analysis in the area of MBD. Firstly, we introduce the development of MBD. Secondly, the frequently applied data analysis methods are reviewed. Three typical applications of MBD analysis, namely, wireless channel modeling, human online and offline behavior analysis, and speech recognition in the Internet of Vehicles, are introduced, respectively. Finally, we summarize the main challenges and future development directions of mobile big data analysis.
Jiyang Xie 0001, Yanting Zhang 0001, Hong Yu 0006, Jinnan Zhan, Zhanyu Ma, Yuanyuan Qiao 0002, Jianhua Zhang 0001, Jun Guo 0002
Wirel. Commun. Mob. Comput.4