EDBT 2026 Demo / reviewers in the wild / expert
Alexander Vilesov
dblp:326/3502
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-6197-734XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 37% Knowledge representation and reasoning · 29% Face, body and person analysis · 25% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 56% Digital forensics and information hiding · 44% | |
| Human-computer interaction and pervasive computing
1 paper |
Health and well-being technologies · 33% Wearable and physiological sensing · 33% Haptics and multimodal interaction · 33% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning › temporal reasoning
spatio-temporal reasoning |
0.9 | 1 | 2025 | VLM4D: Towards Spatiotemporal Awareness in Vision Language Models · ICCV 2025 |
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency |
0.9 | 1 | 2025 | VLM4D: Towards Spatiotemporal Awareness in Vision Language Models · ICCV 2025 |
Security and privacy of machine learning › adversarial attack › evasion attack
deepfake detection evasion |
0.9 | 1 | 2025 | Chimera: Creating Digitally Signed Fake Photos by Fooling Image Recapture and Deepfake Detectors · USENIX Security Symposium 2025 |
Digital forensics and information hiding › digital forensics › multimedia forensics
image forensics |
0.9 | 1 | 2025 | Chimera: Creating Digitally Signed Fake Photos by Fooling Image Recapture and Deepfake Detectors · USENIX Security Symposium 2025 |
Computer vision › Face, body and person analysis › face analysis
remote photoplethysmography |
0.8 | 1 | 2024 | Implicit Neural Models to Extract Heart Rate from Video · ECCV (83) 2024 |
Wearable and physiological sensing › vital sign monitoring
heart rate monitoring |
0.6 | 1 | 2022 | Blending camera and 77 GHz radar sensing for equitable, robust plethysmography · ACM Trans. Graph. 2022 |
Haptics and multimodal interaction
multimodal fusion |
0.6 | 1 | 2022 | Blending camera and 77 GHz radar sensing for equitable, robust plethysmography · ACM Trans. Graph. 2022 |
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
0.3 | 1 | 2025 | VLM4D: Towards Spatiotemporal Awareness in Vision Language Models · ICCV 2025 |
Security and privacy of machine learning
adversarial example |
0.3 | 1 | 2025 | Chimera: Creating Digitally Signed Fake Photos by Fooling Image Recapture and Deepfake Detectors · USENIX Security Symposium 2025 |
Wireless sensing and localization
radar sensing |
0.2 | 1 | 2022 | Blending camera and 77 GHz radar sensing for equitable, robust plethysmography · ACM Trans. Graph. 2022 |
Methods — techniques the papers use, named apart from their topics
light transport analysis · 1.1debiasing · 1.1supervised fine-tuning · 0.9adversarial perturbation · 0.94d feature field reconstruction · 0.9implicit neural representation · 0.8multimodal fusion · 0.6multi-modal fusion · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VLM4D: Towards Spatiotemporal Awareness in Vision Language ModelsabstractVision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts-abilities essential for robust dynamic real-world understanding yet notably lacking in current VLMs. In this paper, we introduce VLM4D, the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs. Our benchmark comprises diverse real-world and synthetic videos accompanied by carefully curated question-answer pairs emphasizing translational and rotational motions, perspective awareness, and motion continuity. Through comprehensive evaluations of state-of-the-art open and closed-source VLMs, we identify significant performance gaps compared to human baselines, highlighting fundamental deficiencies in existing models. Extensive analysis reveals that VLMs struggle particularly with integrating multiple visual cues and maintaining temporal coherence. We further explore promising directions, such as leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning, demonstrating their effectiveness in enhancing spatiotemporal comprehension. Our work aims to encourage deeper exploration into improving VLMs' spatial and temporal grounding, paving the way towards more capable and reliable visual intelligence for dynamic environments. Shijie Zhou 0003, Alexander Vilesov, Xuehai He, Ziyu Wan, Shuwang Zhang, Aditya Nagachandra, Di Chang, Xin Wang 0061, Achuta Kadambi |
ICCV | 2 |
| 2025 | Chimera: Creating Digitally Signed Fake Photos by Fooling Image Recapture and Deepfake Detectors
Alexander Vilesov, Jinghuai Zhang, Hossein Khalili, Achuta Kadambi, Nader Sehatbakhsh |
USENIX Security Symposium | 2 |
| 2024 | Implicit Neural Models to Extract Heart Rate from Video
Pradyumna Chari, Anirudh Bindiganavale Harish, Adnan Armouti, Alexander Vilesov, Sanjit Sarda, Laleh Jalilian, Achuta Kadambi |
ECCV (83) | 4 |
| 2022 | Blending camera and 77 GHz radar sensing for equitable, robust plethysmographyabstractWith the resurgence of non-contact vital sign sensing due to the COVID-19 pandemic, remote heart-rate monitoring has gained significant prominence. Many existing methods use cameras; however previous work shows a performance loss for darker skin tones. In this paper, we show through light transport analysis that the camera modality is fundamentally biased against darker skin tones. We propose to reduce this bias through multi-modal fusion with a complementary and fairer modality - radar. Through a novel debiasing oriented fusion framework, we achieve performance gains over all tested baselines and achieve skin tone fairness improvements over the RGB modality. That is, the associated Pareto frontier between performance and fairness is improved when compared to the RGB modality. In addition, performance improvements are obtained over the radar-based method, with small trade-offs in fairness. We also open-source the largest multi-modal remote heart-rate estimation dataset of paired camera and radar measurements with a focus on skin tone representation. Alexander Vilesov, Pradyumna Chari, Adnan Armouti, Anirudh Bindiganavale Harish, Kimaya Kulkarni, Ananya Deoghare, Laleh Jalilian, Achuta Kadambi |
ACM Trans. Graph. | 1 |