Jun Bao

dblp:46/2981 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 NP-MiSR: Neural Process-based Multi-Interest Learning for Session-Based Recommendation
abstract
Session-based recommendation (SBR) aims to provide users with satisfactory suggestions via modeling preferences based on short-term, anonymous user-item interaction sequences. Traditional single interest learning methods struggle to align with the diverse nature of preferences. Recent advances resolved this bottleneck by learning multiple interest embeddings for each session. However, due to the pre-defining scheme of interest quantity (e.g. the number of interests), these approaches are deficient in adaptive ability towards distinctive preference patterns across different users. Moreover, these methods rely solely on the current session and ignore useful information from related ones. The short-term property of sessions would magnify the insufficient representation issue. To address these limitations, we propose a Neural Process-based Multi-interest learning framework for Session-based Recommendation, namely NP-MiSR. To be specific, our method enables adaptive multi-interest representation learning through two complementary mechanisms: 1) Neural Process-based Intra-session interest modeling: We employ Neural Processes to model the distribution of interests within a session, where the fixed interest configurations are no longer needed. 2) Cross-session context fusion: We extract interest distributions of similar sessions as contextual priors to refine the current session’s interest representation. Extensive experiments on three datasets demonstrate that our method consistently outperforms state-of-the-art SBR approaches with an average improvement of 38.8%. Moreover, the few-shot learning task reveals that NP-MiSR achieves a surprisingly favorable efficiency v.s. performance trade-off where utilizing only 10% of the training data attains 95% of the recommendation performance.
Jun Bao, Yiheng Jiang, Xiangfeng Liu, Yuanbo Xu
AAAI1
2025 What we Need is Explicit Controllability: Training 3D Gaze Estimator Using Only Facial Images
Tingwei Li, Jun Bao, Zhenzhong Kuang, Buyu Liu
ICCV2
2025 A Generic Framework for Evaluating Gaze Representations for Gaze Estimation
Buyu Liu, Suguo Zhu, Jun Bao
ICMR4
2024 GLOW: Global Layout Aware Attacks on Object Detection
abstract
Adversarial attacks aim to perturb images such that a predictor outputs incorrect results. Due to the limited research in structured attacks, imposing consistency checks on natural multi-object scenes is a practical defense against conventional adversarial attacks. More desired attacks should be able to fool defenses with such consistency checks. Therefore, we present the first approach GLOW that copes with various attack requests by generating global layout-aware adversar-ial attacks, in which both categorical and geometric layout constraints are explicitly established. Specifically, we focus on object detection tasks and given a victim image, GLOW first localizes victim objects according to target labels. And then it generates multiple attack plans, together with their context-consistency scores. GLOW, on the one hand, is ca-pable of handling various types of requests, including single or multiple victim objects, with or without specified victim objects. On the other hand, it produces a consistency score for each attack plan, reflecting the overall contextual consistency that both semantic category and global scene layout are considered. We conduct our experiments on MS COCO and Pascal. Extensive experimental results demonstrate that we can achieve about 30% average relative improvement compared to state-of-the-art methods in conventional single object attack request; Moreover, such superiority is also valid across more generic attack requests, under both white-box and zero-query black-box settings. Finally, we conduct comprehensive human analysis, which not only validates our claim further but also provides strong evidence that our evaluation metrics reflect human reviews well.
Jun Bao, Buyu Liu, Kui Ren 0001, Jun Yu 0002
CVPR1
2024 MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
Buyu Liu, Kai Wang 0036, Jun Bao, Tingting Han 0003, Jun Yu 0002
ACM Multimedia4
2024 Learnability Matters: Active Learning for Video Captioning
abstract
This work focuses on the active learning in video captioning. In particular, we propose to address the learnability problem in active learning, which has been brought up by collective outliers in video captioning and neglected in the literature. To start with, we conduct a comprehensive study of collective outliers, exploring their hard-to-learn property and concluding that ground truth inconsistency is one of the main causes. Motivated by this, we design a novel active learning algorithm that takes three complementary aspects, namely learnability, diversity, and uncertainty, into account. Ideally, learnability is reflected by ground truth consistency. Under the active learning scenario where ground truths are not available until human involvement, we measure the consistency on estimated ground truths, where predictions from off-the-shelf models are utilized as approximations to ground truths. These predictions are further used to estimate sample frequency and reliability, evincing the diversity and uncertainty respectively. With the help of our novel caption-wise active learning protocol, our algorithm is capable of leveraging knowledge from humans in a more effective yet intellectual manner. Results on publicly available video captioning datasets with diverse video captioning models demonstrate that our algorithm outperforms SOTA active learning methods by a large margin, e.g. we achieve about 103% of full performance on CIDEr with 25% of human annotations on MSR-VTT.
Buyu Liu, Jun Bao, Min Zhang 0005, Jun Yu 0002
NeurIPS3
2023 Follow-me: Deceiving Trackers with Fabricated Paths
abstract
Convolutional Neural Networks (CNNs) are vulnerable to adversarial attacks in which visually imperceptible perturbations can deceive CNN-based models. While current research on adversarial attacks in single object tracking exists, it overlooks a critical aspect of manipulating predicted trajectories to follow user-defined paths regardless of the actual location of the targeted object. To address this, we propose the very first white-box attack algorithm that is capable of deceiving victim trackers by compelling them to generate trajectories that adhere to predetermined counterfeit paths. Specifically, we focus on Siamese-based trackers as our victim models. Given an arbitrary counterfeit path, we first decompose it into discrete target locations in each frame, with the assumption of constant velocity. These locations are converted to heatmap anchors, which represent the offset of their location from the target object's location in the previous frame. Later on, we design a novel loss function to minimize the gap between above-mentioned anchors and our predicted ones. Finally, the gradients computed by such loss are used to update the original video, resulting in our adversarial video. To validate our ideas, we design three sets of counterfeit paths as well as novel evaluation metrics to measure the path-following properties. Experiments with two victim models on three publicly available datasets, OTB100, VOT2018, and VOT2016, demonstrate that our algorithm not only outperforms SOTA methods significantly under conventional evaluation metrics, e.g. 90% and 68.4% precision and successful rate drop on OTB100, but also follows the counterfeit paths well, which is beyond any existing attack methods. The source code is available at https://github.com/loushengtao/Follow-me.
Shengtao Lou, Buyu Liu, Jun Bao, Jiajun Ding, Jun Yu 0002
ACM Multimedia3
2023 Concept Parser With Multimodal Graph Learning for Video Captioning
abstract
Conventional video captioning methods are either stage-wise or simple end-to-end. While the former might introduce additional noise when exploiting off-the-shelf models to provide extra information, the latter suffers from lacking high-level cues. Therefore, a more desired framework should be able to capture multi-aspects of videos consistently. To this end, we present a concept-aware and task-specific model named CAT that accounts for both low-level visual and high-level concept cues, and incorporates them effectively in an end-to-end manner. Specifically, low-level visual and high-level concept features are obtained from the video transformer and concept parser of CAT. And a concept loss is further introduced to regularize the learning process of concept parser w.r.t. generated pseudo ground truth. To combine multi-level features, a caption transformer is later introduced in CAT, where visual and concept features are the inputs and caption is its output. In particular, we make critical design choices in the caption transformer to learn to exploit these cues with a multi-modal graph. This is achieved by a graph loss that enforces effective learning of intra and inter correlations between multi-level cues. Extensive experiments on three benchmark datasets demonstrate that CAT achieves 2.3 and 0.7 improvements in the CIDEr metric on MSVD and MSR-VTT compared to the state-of-the-art method SwinBERT and also achieves a competitive result on VATEX.
Bofeng Wu, Buyu Liu, Jun Bao, Peng Xi, Jun Yu 0002
IEEE Trans. Circuits Syst. Video Technol.4
2022 ESCNet: Gaze Target Detection with the Understanding of 3D Scenes
abstract
This paper aims to address the single image gaze target detection problem. Conventional methods either focus on 2D visual cues or exploit additional depth information in a very coarse manner. In this work, we propose to explicitly and effectively model 3D geometry under challenging scenario where only 2D annotations are available. We first obtain 3D point clouds of given scene with estimated depth and reference objects. Then we figure out the front-most points in all possible 3D directions of given person. These points are later leveraged in our ESCNet model. Specifically, ESCNet consists of geometry and scene parsing modules. The former produces an initial heatmap inferring the probability that each front-most point has been looking at according to estimated 3D gaze direction. And the latter further explores scene contextual cues to regulate detection results. We validate our idea on two publicly available dataset, GazeFollow and VideoAttentionTarget, and demon-strate the state-of-the-art performance. Our method also beats the human in terms of AUC on GazeFollow. Our code can be found here https://github.com/bjj9/ESCNet.
Jun Bao, Buyu Liu, Jun Yu 0002
CVPR1
2022 An Individual-Difference-Aware Model for Cross-Person Gaze Estimation
abstract
We propose a novel method on refining cross-person gaze prediction task with eye/face images only by explicitly modelling the person-specific differences. Specifically, we first assume that we can obtain some initial gaze prediction results with existing method, which we refer to as InitNet, and then introduce three modules, the Validity Module (VM), Self-Calibration (SC) and Person-specific Transform (PT) module. By predicting the reliability of current eye/face images, VM is able to identify invalid samples, e.g. eye blinking images, and reduce their effects in modelling process. SC and PT module then learn to compensate for the differences on valid samples only. The former models the translation offsets by bridging the gap between initial predictions and dataset-wise distribution. And the later learns more general person-specific transformation by incorporating the information from existing initial predictions of the same person. We validate our ideas on three publicly available datasets, EVE, XGaze, and MPIIGaze dataset. We demonstrate that our proposed method outperforms the SOTA methods significantly on all of them, e.g. respectively 21.7%, 36.0%, and 32.9% relative performance improvements. We are the winner of the GAZE 2021 EVE Challenge and our code can be found here https://github.com/bjj9/EVE_SCPT.
Jun Bao, Buyu Liu, Jun Yu 0002
IEEE Trans. Image Process.1
1996 Accurate localization of cortical convolutions in MR brain images
abstract
Analysis of brain images often requires accurate localization of cortical convolutions. Although magnetic resonance (MR) brain images offer sufficient resolution for identifying convolutions in theory, the nature of tomographic imaging prevents clear definition of convolutions in individual slices. Existing methods for solving this problem rely on heuristic adaptation of brain atlases created from a small number of individuals. These methods do not usually provide high accuracy because of large biological variations among individuals. The authors propose to localize convolutions by linking realistic visualizations of the cortical surface with the original image volume. They have developed a system so that a user can quickly localize key convolutions in several visualizations of an entire brain surface. Because of the links between the visualizations and the original volume, these convolutions are simultaneously localized in the original image slices. In the process of the authors' development, they have implemented a fast and easy method for visualizing cortical surfaces in MR images, thereby making their scheme usable in practical applications.
Yaorong Ge, J. Michael Fitzpatrick, Benoit M. Bao, Jun Bao, Robert M. Kessler, Richard A. Margolin
IEEE Trans. Medical Imaging4
1991 On the Design of a Tree Classifier and its Applicaton to speech Recognition
abstract
A new algorithm for constructing entropy reduction based decision tree classifier in two steps for large reference-class sets is proposed. The d-dimensional feature space is first mapped onto a line thus allowing a dynamic choice of features and then the resultant linear space is partitioned into two sections while minimizing the average system entropy. Classes of each section, again considered as a collection of references classes in a d-dimensional feature space, can be further split in a similar manner should the collection still be considered excessively large, thus forming a binary decision tree of nodes with overlapping members. The advantage of using such a classifier is that the need to match a test feature vector against all the references exhaustively is avoided. We demonstrate in this paper that discrete syllable recognition with dynamic programming equipped with such a classifier can reduce the recognition time by a factor of 40 to 100. The recognition speed is one third to one half of that using hidden Markov models (HMM), while the recognition rate is somewhat higher. The theory is considerably simpler than that of HMM but the decision tree can occupy a lot of memory space.
Chorkin Chan, Jun Bao
Int. J. Pattern Recognit. Artif. Intell.2
1989 A preliminary study on the static representation of short-timed speech dynamics
Chorkin Chan, Jun Bao, Jian-Xiong Wu
EUROSPEECH2