Zhibo Yang 0002

dblp:127/9460-2 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0002-7261-7275ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Systems, architecture and hardware · 2Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Unifying Top-Down and Bottom-Up Scanpath Prediction Using Transformers
abstract
Most models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT), a single model that predicts both forms of attention control. HAT uses a novel transformer-based architecture and a simplified foveated retina that collectively create a spatio-temporal awareness akin to the dynamic visual working memory of humans. HAT not only establishes a new state-of-the-art in predicting the scanpath of fixations made during target-present and target-absent visual search and "taskless" free viewing, but also makes human gaze behavior interpretable. Unlike previous methods that rely on a coarse grid of fixation cells and experience information loss due to fixation discretization, HAT features a sequential dense prediction architecture and outputs a dense heatmap for each fixation, thus avoiding discretizing fixations. HAT sets a new standard in computational attention, which emphasizes effectiveness, generality, and interpretability. HAT's demonstrated scope and applicability will likely inspire the development of new attention models that can better predict human behavior in various attention-demanding scenarios. Code is available at https://github.com/cvlab-stonybrook/HAT.
Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras
CVPR1
2024 Look Hear: Gaze Prediction for Speech-Directed Human Attention
Sounak Mondal, Seoyoung Ahn, Zhibo Yang 0002, Niranjan Balasubramanian, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai
ECCV (42)3
2023 Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention
abstract
Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goaldirected attention (search). Such models are limited in their application due to a common approach relying on trained target detectors for all possible objects, and the availability of human gaze data for their training (both not scalable). In response, we pose a new task called ZeroGaze, a new variant of zero-shot learning where gaze is predicted for never-before-searched objects, and we develop a novel model, Gazeformer; to solve the ZeroGaze problem. In contrast to existing methods using object detector modules, Gazeformer encodes the target using a natural language model, thus leveraging semantic similarities in scanpath prediction. We use a transformer-based encoder-decoder architecture because transformers are particularly useful for generating contextual representations. Gazeformer surpasses other models by a large margin (19%-70%) on the ZeroGaze setting. It also outperforms existing target-detection models on standard gaze prediction for both target-present and target-absent search tasks. In addition to its improved performance, Gazeformer is more than five times faster than the state-of-the-art target-present visual search model. Code can be found at https://github.com/cvlab-stonybrook/Gazeformer/
Sounak Mondal, Zhibo Yang 0002, Seoyoung Ahn, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai
CVPR2
2022 Target-Absent Human Attention
Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras
ECCV (4)1
2022 Hierarchical Proxy-based Loss for Deep Metric Learning
abstract
Proxy-based metric learning losses are superior to pair-based losses due to their fast convergence and low training complexity. However, existing proxy-based losses focus on learning class-discriminative features while overlooking the commonalities shared across classes which are potentially useful in describing and matching samples. Moreover, they ignore the implicit hierarchy of categories in real-world datasets, where similar subordinate classes can be grouped together. In this paper, we present a framework that leverages this implicit hierarchy by imposing a hierarchical structure on the proxies and can be used with any existing proxy-based loss. This allows our model to capture both class-discriminative features and class-shared characteristics without breaking the implicit data hierarchy. We evaluate our method on five established image retrieval datasets such as In-Shop and SOP. Results demonstrate that our hierarchical proxy-based loss framework improves the performance of existing proxy-based losses, especially on large datasets which exhibit strong hierarchical structure.
Zhibo Yang 0002, Muhammet Bastan, Xinliang Zhu, Douglas Gray 0001, Dimitris Samaras
WACV1
2022 Artificial intelligence in drug discovery: applications and techniques
abstract
Artificial intelligence (AI) has been transforming the practice of drug discovery in the past decade. Various AI techniques have been used in many drug discovery applications, such as virtual screening and drug design. In this survey, we first give an overview on drug discovery and discuss related applications, which can be reduced to two major tasks, i.e. molecular property prediction and molecule generation. We then present common data resources, molecule representations and benchmark platforms. As a major part of the survey, AI techniques are dissected into model architectures and learning paradigms. To reflect the technical development of AI in drug discovery over the years, the surveyed works are organized chronologically. We expect that this survey provides a comprehensive review on AI in drug discovery. We also provide a GitHub repository with a collection of papers (and codes, if applicable) as a learning resource, which is regularly updated.
Jianyuan Deng, Zhibo Yang 0002, Iwao Ojima, Dimitris Samaras, Fusheng Wang 0001
Briefings Bioinform.2
2021 Mosaic: Advancing User Quality of Experience in 360-Degree Video Streaming With Machine Learning
abstract
Conventional streaming solutions for streaming 360-degree panoramic videos are inefficient in that they download the entire 360-degree panoramic scene, while the user views only a small sub-part of the scene called the viewport. This can waste over 80% of the network bandwidth. We develop a comprehensive approach called Mosaic that combines a powerful neural network-based viewport prediction with a rate control mechanism that assigns rates to different tiles in the 360-degree frame such that the video quality of experience is optimized subject to a given network capacity. We model the optimization as a multi-choice knapsack problem and solve it using a greedy approach. We also develop an end-to-end testbed using standards-compliant components and provide a comprehensive performance evaluation of Mosaic along with five other streaming techniques - two for conventional adaptive video streaming and three for 360-degree tile-based video streaming. Mosaic outperforms the best of the competitions by as much as 47-191% in terms of average video quality of experience. Simulation-based evaluation as well as subjective user studies further confirm the superiority of the proposed approach.
Sohee Kim Park, Arani Bhattacharya, Zhibo Yang 0002, Samir Ranjan Das, Dimitris Samaras
IEEE Trans. Netw. Serv. Manag.3
2020 Predicting Goal-Directed Human Attention Using Inverse Reinforcement Learning
abstract
Human gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We propose the first inverse reinforcement learning (IRL) model to learn the internal reward function and policy used by humans during visual search. We modeled the viewer's internal belief states as dynamic contextual belief maps of object locations. These maps were learned and then used to predict behavioral scanpaths for multiple target categories. To train and evaluate our IRL model we created COCO-Search18, which is now the largest dataset of high-quality search fixations in existence. COCO-Search18 has 10 participants searching for each of 18 target-object categories in 6202 images, making about 300,000 goal-directed fixations. When trained and evaluated on COCO-Search18, the IRL model outperformed baseline models in predicting search fixation scanpaths, both in terms of similarity to human search behavior and search efficiency. Finally, reward maps recovered by the IRL model reveal distinctive target-dependent patterns of object prioritization, which we interpret as a learned object context.
Zhibo Yang 0002, Lihan Huang, Yupei Chen, Zijun Wei, Seoyoung Ahn, Gregory J. Zelinsky, Dimitris Samaras, Minh Hoai
CVPR1
2020 Nonverbal Behavioral Patterns Predict Social Rejection Elicited Aggression
Megan Quarmley, Zhibo Yang 0002, Shahrukh Athar, Gregory J. Zelinsky, Dimitris Samaras, Johanna M. Jarcho
FG2
2019 Advancing User Quality of Experience in 360-degree Video Streaming
abstract
Conventional streaming solutions for streaming 360-degree panoramic videos are inefficient in that they download the entire 360-degree panoramic scene, while the user views only a small sub-part of the scene called the viewport. This can waste over 80% of the network bandwidth. We develop a comprehensive approach called Mosaic that combines a powerful neural network-based viewport prediction with a rate control mechanism that assigns rates to different tiles in the 360-degree frame such that the video quality of experience is optimized subject to a given network capacity. We model the optimization as a multi-choice knapsack problem and solve it using a greedy approach. We also develop an end-to-end testbed using standards-compliant components and provide a comprehensive performance evaluation of Mosaic along with four other streaming techniques - two for conventional adaptive video streaming and two for 360-degree tile-based video streaming. Mosaic outperforms the best of the competition by as much as 50% in terms of median video quality.
Sohee Kim Park, Arani Bhattacharya, Zhibo Yang 0002, Mallesham Dasari, Samir Ranjan Das, Dimitris Samaras
Networking3
2018 Robust and Fast Decoding of High-Capacity Color QR Codes for Mobile Applications
abstract
The use of color in QR codes brings extra data capacity, but also inflicts tremendous challenges on the decoding process due to chromatic distortion-cross-channel color interference and illumination variation. Particularly, we further discover a new type of chromatic distortion in high-density color QR codes-cross-module color interference-caused by the high density which also makes the geometric distortion correction more challenging. To address these problems, we propose two approaches, LSVM-CMI and QDA-CMI, which jointly model these different types of chromatic distortion. Extended from SVM and QDA, respectively, both LSVM-CMI and QDA-CMI optimize over a particular objective function and learn a color classifier. Furthermore, a robust geometric transformation method and several pipeline refinements are proposed to boost the decoding performance for mobile applications. We put forth and implement a framework for high-capacity color QR codes equipped with our methods, called HiQ. To evaluate the performance of HiQ, we collect a challenging large-scale color QR code dataset, CUHK-CQRC, which consists of 5390 high-density color QR code samples. The comparison with the baseline method [2] on CUHK-CQRC shows that HiQ at least outperforms [2] by 188% in decoding success rate and 60% in bit error rate. Our implementation of HiQ in iOS and Android also demonstrates the effectiveness of our framework in real-world applications.
Zhibo Yang 0002, Huanle Xu, Jianyuan Deng, Chen Change Loy, Wing Cheong Lau
IEEE Trans. Image Process.1
2017 Mitigating Service Variability in MapReduce Clusters via Task Cloning: A Competitive Analysis
abstract
Measurement traces from real-world production environment show that the execution time of tasks within a MapReduce job varies widely due to the variability in machine service capacity. This variability issue makes efficient job scheduling over large-scale MapReduce clusters extremely challenging. To tackle this problem, we adopt the task cloning approach to mitigate the effect of machine variability and design corresponding scheduling algorithms so as to minimize the overall job flowtime in different scenarios. For offline scheduling where all jobs arrive at the same time, we design an$O(1)$-competitive algorithm, which gives priorities to jobs with small effective workload. We then extend this offline algorithm to yield the so-called Smallest Remaining Effective Workload based$\beta$-fraction Sharing plus Cloning algorithm (SREW+C($\beta$)) for the online case. We also show that SREW+C($\beta$) is$(1+ 2\beta + \epsilon)$-speed$O(\frac{1}{\beta \epsilon })$-competitive with respect to the sum of job flowtime within a cluster. We demonstrate via trace-driven simulations that SREW+C($\beta$) can significantly reduce the overall job flowtime by cutting down the elapsed time of small jobs substantially. In particular, SREW+C($\beta$) reduces the total job flowtime by 14, 10 and 11 percent respectively when comparing to Mantri, Dolly and Grass.
Huanle Xu, Wing Cheong Lau, Zhibo Yang 0002, Gustavo de Veciana, Hanxu Hou
IEEE Trans. Parallel Distributed Syst.3
2016 Towards robust color recovery for high-capacity color QR codes
abstract
Color brings extra data capacity for QR codes, but it also brings tremendous challenges to the decoding because of color interference and illumination variation, especially for high-density QR codes. In this paper, we put forth a framework for high-capacity QR codes, HiQ, which optimizes the decoding algorithm for high-density QR codes to achieve robust and fast decoding on mobile devices, and adopts a learning-based approach for color recovery. Moreover, we propose a robust geometric transformation algorithm to correct the geometric distortion. We also provide a challenging color QR code dataset, CUHK-CQRC, which consists of 5390 high-density color QR code samples captured by different smartphones under different lighting conditions. Experimental results show that HiQ outperforms the baseline [1] by 286% in decoding success rate and 60% in bit error rate.
Zhibo Yang 0002, Zhiyi Cheng, Chen Change Loy, Wing Cheong Lau, Chak Man Li
ICIP1
2015 Solving Large Graph Problems in MapReduce-Like Frameworks via Optimized Parameter Configuration
Huanle Xu, Ronghai Yang, Zhibo Yang 0002, Wing Cheong Lau
ICA3PP (2)3