VLDB 2026 Research / reviewers in the wild / expert
Yijun Qian
dblp:260/4667
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sound Event Detection With Boundary-Aware Optimization and InferenceabstractTemporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers—Recurrent Event Detection (RED) and Event Proposal Network (EPN)—which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes. Florian Schmid, Chi Ian Tang, Sanjeel Parekh, Vamsi K. Ithapu, Juan Azcarreta, Giacomo Ferroni, Yijun Qian, Arnoldas Jasonas, Cosmin Frateanu, Camilla Clark, Gerhard Widmer, Cagdas Bilen |
IEEE Signal Process. Lett. | 7 |
| 2025 | EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric PerceptionabstractModern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment, especially in resource-constrained environments. In this paper, we introduce EgoAdapt, a framework that adaptively performs cross-modal distillation and policy learning to enable efficient inference across different egocentric perception tasks, including egocentric action recognition, active speaker localization, and behavior anticipation. Our proposed policy module is adaptable to task-specific action spaces, making it broadly applicable. Experimental results on three challenging egocentric datasets EPIC-Kitchens, EasyCom, and Aria Everyday Activities demonstrate that our method significantly enhances efficiency, reducing GMACs by up to 89.09%, parameters up to 82.02%, and energy up to 9.6x, while still on-par and in many cases outperforming, the performance of corresponding state-of-the-art models. Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan, Calvin Murdock, Ishwarya Ananthabhotla, Yijun Qian, Vamsi K. Ithapu, Dinesh Manocha, Ruohan Gao |
ICCV | 7 |
| 2024 | Text Motion Translator: A Bi-directional Model for Enhanced 3D Human Motion Generation from Open-Vocabulary Descriptions
Yijun Qian, Jack Urbanek, Alex Hauptmann 0001, Jungdam Won |
ECCV (63) | 1 |
| 2023 | Integrated Aerobic Exercise into Adult Second Language Learning in Virtual Reality GameabstractThis paper introduces an integrated tool that combines aerobic exercise and second language learning (L2) for the purpose of enhancing adult cognitive health. The tool utilizes virtual reality (VR) technology to focus on auditory recognition and content comprehension aiming to improve language proficiency in potential adult L2 learners. The study investigates participants’ experiences, cognitive load demands, and usability of the integrated tool within a VR-based multitasking environment. Preliminary results indicate a positive trend in attitudes towards the tool’s usefulness, highlighting its potential for motivating exercise and learning, facilitating learning processes, and providing ease of use. We see a promising future in serious game that integrate physical exercise to motivate the uptake of health-promoting behaviors. Yijun Qian, Sarvesh Prajapati, Anna M. Schwartz, Ara Jung, Uri Seitz, Joshua Van Alfen, Lara Lewis, Miso Kim, Arthur F. Kramer, Leanne Chukoskie |
CoG | 1 |
| 2023 | Breaking The Limits of Text-conditioned 3D Motion Synthesis with Elaborative DescriptionsabstractGiven its wide applications, there is increasing focus on generating 3D human motions from textual descriptions. Differing from the majority of previous works, which regard actions as single entities and can only generate short sequences for simple motions, we propose EMS, an elaborative motion synthesis model conditioned on detailed natural language descriptions. It generates natural and smooth motion sequences for long and complicated actions by factorizing them into groups of atomic actions. Meanwhile, it understands atomic-action level attributes (e.g., motion direction, speed, and body parts) and enables users to generate sequences of unseen complex actions from unique sequences of known atomic actions with independent attribute settings and timings applied. We evaluate our method on the KIT Motion-Language and BABEL benchmarks, where it outperforms all previous state-of-the-art with noticeable margins. Yijun Qian, Jack Urbanek, Alex Hauptmann 0001, Jungdam Won |
ICCV | 1 |
| 2022 | Rethinking Zero-shot Action Recognition: Learning from Latent Atomic Actions
Yijun Qian, Lijun Yu, Wenhe Liu, Alex Hauptmann 0001 |
ECCV (4) | 1 |