Bingqing Zhang

dblp:202/4930 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval Via Uncertainty Minimization
abstract
Despite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and low-quality video frames. Although interactive systems have emerged to address these challenges by refining user intent through clarifying questions, current methods typically rely on heuristic or ad-hoc strategies without explicitly quantifying these uncertainties, limiting their effectiveness. Motivated by this gap, we propose UMIVR, an Uncertainty-Minimizing Interactive Text-to-Video Retrieval framework that explicitly quantifies three critical uncertainties-text ambiguity, mapping uncertainty, and frame uncertainty-via principled, training-free metrics: semantic entropy-based Text Ambiguity Score (TAS), Jensen-Shannon divergence-based Mapping Uncertainty Score (MUS), and a Temporal Quality-based Frame Sampler (TQFS). By adaptively generating targeted clarifying questions guided by these uncertainty measures, UMIVR iteratively refines user queries, significantly reducing retrieval ambiguity. Extensive experiments on multiple benchmarks validate UMIVR's effectiveness, achieving notable gains in Recall@1 (69.2\% after 10 interactive rounds) on the MSR-VTT-1k dataset, thereby establishing an uncertainty-minimizing foundation for interactive TVR.
Bingqing Zhang, Heming Du, Yang Li 0184, Xue Li 0001, Jiajun Liu 0004, Sen Wang 0001
ICCV1
2025 When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
abstract
Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes the existing datasets and methods insufficient for video temporal grounding. By revisiting the gap between current MR tasks and real-world applications, we introduce a high-quality datasets called QVHighlights Multi-Moment Dataset (QV-M$^2$), along with new evaluation metrics tailored for multi-moment retrieval (MMR). QV-M$^2$ consists of 2,212 annotations covering 6,384 video segments. Building on existing efforts in MMR, we propose a framework called FlashMMR. Specifically, we propose a Multi-moment Post-verification module to refine the moment boundaries. We introduce constrained temporal adjustment and subsequently leverage a verification module to re-evaluate the candidate segments. Through this sophisticated filtering pipeline, low-confidence proposals are pruned, and robust multi-moment alignment is achieved. We retrain and evaluate 6 existing MR methods on QV-M$^2$ and QVHighlights under both SMR and MMR settings. Results show that QV-M$^2$ serves as an effective benchmark for training and evaluating MMR models, while FlashMMR provides a strong baseline. Specifically, on QV-M$^2$, it achieves improvements over prior SOTA method by 3.00% on G-mAP, 2.70% on mAP@3+tgt, and 2.56% on mR@3. The proposed benchmark and method establish a foundation for advancing research in more realistic and challenging video temporal grounding scenarios. Code is released at https://github.com/Zhuo-Cao/QV-M2.
Heming Du, Bingqing Zhang, Xin Yu 0002, Xue Li 0001, Sen Wang 0001
NeurIPS3
2025 FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
abstract
Text-guided Video Temporal Grounding (VTG) aims to localize relevant segments in untrimmed videos based on textual descriptions, encompassing two subtasks: Moment Retrieval (MR) and Highlight Detection (HD). Although previous typical methods have achieved commendable results, it is still challenging to retrieve short video moments. This is primarily due to the reliance on sparse and limited decoder queries, which significantly constrain the accuracy of predictions. Furthermore, suboptimal outcomes often arise because previous methods rank predictions based on isolated predictions, neglecting the broader video context. To tackle these issues, we introduce FlashVTG, a framework featuring a Temporal Feature Layering (TFL) module and an Adaptive Score Refinement (ASR) module. The TFL module replaces the traditional decoder structure to capture nuanced video content variations across multiple temporal scales, while the ASR module improves prediction ranking by integrating context from adjacent moments and multi-temporal-scale features. Extensive experiments demonstrate that FlashVTG achieves state-of-the-art performance on four widely adopted datasets in both MR and HD. Specifically, on the QVHighlights dataset, it boosts mAP by 5.8% for MR and 3.3% for HD. For short-moment retrieval, FlashVTG increases mAP to 125% of previous SOTA performance. All these improvements are made without adding training burdens, underscoring its effectiveness. Our code is available at https://github.com/Zhuo-Cao/FlashVTG.
Bingqing Zhang, Heming Du, Xin Yu 0002, Xue Li 0001, Sen Wang 0001
WACV2
2025 TokenBinder: Text-Video Retrieval with One-to-Many Alignment Paradigm
abstract
Text-Video Retrieval (TVR) methods typically match query-candidate pairs by aligning text and video features in coarse-grained, fine-grained, or combined (coarse-to-fine) manners. However, these frameworks predominantly employ a one(query)-to-one(candidate) alignment paradigm, which struggles to discern nuanced differences among candidates, leading to frequent mismatches. Inspired by Comparative Judgement in human cognitive science, where decisions are made by directly comparing items rather than evaluating them independently, we propose TokenBinder. This innovative two-stage TVR framework introduces a novel one-to-many coarse-to-fine alignment paradigm, imitating the human cognitive process of identifying specific items within a large collection. Our method employs a Focused-view Fusion Network with a sophisticated cross-attention mechanism, dynamically aligning and comparing features across multiple videos to capture finer nuances and contextual variations. Extensive experiments on six benchmark datasets confirm that TokenBinder substantially outperforms existing state-of-the-art methods. These results demonstrate its robustness and the effectiveness of its fine-grained alignment in bridging intra- and inter-modality information gaps in TVR tasks. Code is avaliable at https://github.com/bingqingzhang/TokenBinder.
Bingqing Zhang, Heming Du, Xin Yu 0002, Xue Li 0001, Jiajun Liu 0004, Sen Wang 0001
WACV1
2023 Object Detection Difficulty: Suppressing Over-aggregation for Faster and Better Video Object Detection
abstract
Current video object detection (VOD) models often encounter issues with over-aggregation due to redundant aggregation strategies, which perform feature aggregation on every frame. This results in suboptimal performance and increased computational complexity. In this work, we propose an image-level Object Detection Difficulty (ODD) metric to quantify the difficulty of detecting objects in a given image. The derived ODD scores can be used in the VOD process to mitigate over-aggregation. Specifically, we train an ODD predictor as an auxiliary head of a still-image object detector to compute the ODD score for each image based on the discrepancies between detection results and ground-truth bounding boxes. The ODD score enhances the VOD system in two ways: 1) it enables the VOD system to select superior global reference frames, thereby improving overall accuracy; and 2) it serves as an indicator in the newly designed ODD Scheduler to eliminate the aggregation of frames that are easy to detect, thus accelerating the VOD process. Comprehensive experiments demonstrate that, when utilized for selecting global reference frames, ODD-VOD consistently enhances the accuracy of Global-frame-based VOD models. When employed for acceleration, ODD-VOD consistently improves the frames per second (FPS) by an average of 73.3% across 8 different VOD models without sacrificing accuracy. When combined, ODD-VOD attains state-of-the-art performance when competing with many VOD methods in both accuracy and speed. Our work represents a significant advancement towards making VOD more practical for real-world applications. The code will be released at https://github.com/bingqingzhang/odd-vod.
Bingqing Zhang, Sen Wang 0001, Yifan Liu 0001, Branislav Kusy, Xue Li 0001, Jiajun Liu 0004
ACM Multimedia1
2023 Reinforcement Learning Based User-Specific Shared Control Navigation in Crowds
abstract
Shared control is a mode where the user input is combined with a planned motion to achieve a common goal. In navigation, a shared control approach could provide a potential mobility solution for people who have a mobility impairment and find traditional powered wheelchairs unsuitable. While state-of-the-art work in shared control has demonstrated its capability in improving safety, human-machine interaction and reduce confusion, it is still challenging to use shared control in dynamic, crowded scenarios, in a way that is acceptable to users. Learning from recent advances in robot navigation, we present a reinforcement learning based framework, which allows navigation to be achieved in a user-specific shared controlled way. Our approach was trained and tested in a Unity3D based simulator. It achieved 33% fewer collisions, similar high user agreement (≤ 85%) and 27% less completion time when compared with our previous model-based method.
Bingqing Zhang, Catherine Holloway, Tom Carlson
SMC1
2022 Understanding Interactions for Smart Wheelchair Navigation in Crowds
abstract
Shared control wheelchairs can help users to navigate through crowds by enabling the person to drive the wheelchair while receiving support in avoiding pedestrians. To date, research into shared control has largely overlooked the perspectives of wheelchair users. In this paper, we present two studies that aim to address this gap. The first study involved a series of semi-structured interviews with wheelchair users which highlighted the presence of two different interaction loops, one between the user and the wheelchair and a second one between the user and the crowd. In the second study we engaged with wheelchair users and designers to co-design appropriate feedback loops for future shared control interaction interfaces. Based on the results of the co-design session, we present design implications for shared control wheelchair around the need for empathy, embodiment and social awareness; situational awareness and adaptability; and selective information management.
Bingqing Zhang, Giulia Barbareschi, Roxana Ramirez Herrera, Tom Carlson, Catherine Holloway
CHI1
2022 InvisibiliTee: Angle-Agnostic Cloaking from Person-Tracking Systems with a Tee
Yaxian Li, Bingqing Zhang, Guoping Zhao, Jiajun Liu 0004, Ziwei Wang 0003, Ji-Rong Wen
ICANN (3)2
2022 STAR-GNN: Spatial-Temporal Video Representation for Content-Based Retrieval
abstract
We propose a video feature representation learning frame-work called STAR-GNN, which applies a pluggable graph neural network component on a multi-scale lattice feature graph. The essence of STAR-GNN is to exploit both the temporal dynamics and spatial contents as well as vi-sual connections between regions at different scales in the frames. It models a video with a lattice feature graph in which the nodes represent regions of different granularity, with weighted edges that represent the spatial and temporal links. The contextual nodes are aggregated simultaneously by graph neural networks with parameters trained with re-trieval triplet loss. In the experiments, we show that STAR-GNN effectively implements a dynamic attention mechanism on video frame sequences, resulting in the emphasis for dy-namic and semantically rich content in the video, and is robust to noise and redundancies. Empirical results show that STAR-GNN achieves state-of-the-art performance for Content-Based Video Retrieval.
Guoping Zhao, Bingqing Zhang, Yaxian Li, Jiajun Liu 0004, Ji-Rong Wen
ICME2
2021 An 'Ethical Black Box', Learning From Disagreement in Shared Control Systems
abstract
Shared control, where a human user cooperates with an algorithm to operate a device, has the potential to greatly expand access to powered mobility, but also raises unique ethical challenges. A shared-control wheelchair may perform actions that do not reflect its user's intent in order to protect their safety, causing frustration or distrust in the process. Unlike physical accidents there is currently no frame-work for investigating or adjudicating these events, leading to a reduced capability to improve the shared control algorithm's user experience. In this paper we suggest a system based on the idea of an ‘ethical black box' that records the sensor context of sub-critical disagreements and collision risks in order to allow human investigators to examine them in retrospect and assess whether the algorithm has taken control from the user without justification.
Henry Eberle, Bingqing Zhang, Catalin-Stefan Teodorescu, George Walker, Tom Carlson
SMC2
2021 Pyramid regional graph representation learning for content-based video retrieval
Guoping Zhao, Yaxian Li, Jiajun Liu 0004, Bingqing Zhang, Ji-Rong Wen
Inf. Process. Manag.5
2020 OpenTraj: Assessing Prediction Complexity in Human Trajectories Datasets
Javad Amirian, Bingqing Zhang, Francisco Valente Castro, Juan José Baldelomar, Jean-Bernard Hayet, Julien Pettré
ACCV (6)2
2020 A hierarchical design for shared-control wheelchair navigation in dynamic environments
abstract
For people who have a mobility impairment and find standard wheelchairs unsuitable, a shared-controlled approach could provide a potential mobility solution. However, state-of-the-art research on shared control wheelchairs mainly focus on static environments. In this paper, we present a hierarchical design for our shared-controlled wheelchair using a velocity-based approach together with probabilistic shared control (PSC). By modifying the collision avoidance element and model the robot-pedestrian interaction based on their physical distance, we extended the implementation of PSC to dynamic environments. Our approach was tested in a Unity3D based simulator with human participants. It achieved least number of collisions while obtaining relatively low computational cost and high user agreement comparing with other state-of-the-art methods.
Bingqing Zhang, Catherine Holloway, Tom Carlson
SMC1
2019 Probabilistic Shared Control for a Smart Wheelchair: A Stochastic Model-Based Framework
abstract
This article presents progress made towards implementing a shared control framework for a smart wheelchair based upon stochastic dynamic programming (a model-based control design). First, we describe the mechanical, electrical and software design process of our instrumented wheelchair platform. Then, we detail a deterministic control-oriented model of the wheelchair motion dynamics using Euler-Lagrange equations. Finally, we discuss the development of a stochastic model of the human driver’s intention in view of using Markov chain.
Catalin-Stefan Teodorescu, Bingqing Zhang, Tom Carlson
SMC2
2018 Large-scale realistic cloud rendering based on game engine
abstract
Cloud simulation is an important research-field in computer graphics for a long time, which plays important roles in the weather forecast, flight simulation, big game scenes and so on. Large-scale cloud simulation, especially the realistic interactive cloud, is a difficult problem. This paper proposes a new method to simulate interactive large-scale cloud in games efficiently by combining 2D clouds and 3D volume clouds.
Wenfeng Hu, Yijun Luo, Bingqing Zhang
ICIS3
2017 Game special effect simulation based on particle system of Unity3D
abstract
Along With the progress of computer games, a considerable variety of game has appeared. There is an increasing demand of digital entertainment. It is common for game players to pursuit a better game interaction experience and visual experience. Game art designers pay more attention to achieve realistic effects. In these years, particle effects are widely used in game development. In this article, through the analysis of the basic principle and application method of Unity3D particle system, we design and implement common game scene effects and characters skills effects, mainly simulation and realize the effect of rain fall and design the Special “lightning” effects of characters skills.
Bingqing Zhang, Wenfeng Hu
ICIS1