VLDB 2026 Research / reviewers in the wild / expert
Kiran K. Somasundaram
dblp:20/8395
· DBLP profile ↗
10ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-8554-9083ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Video understanding and tracking · 68% Face, body and person analysis · 15% 3D vision · 15% | |
| Human-computer interaction and pervasive computing
1 paper |
Wearable and physiological sensing · 100% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
egocentric video understanding |
3.1 | 4 | 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric Video · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives · Int. J. Comput. Vis. 2025 Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives · CVPR 2024 |
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
3d hand and body pose estimation |
1.6 | 2 | 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives · Int. J. Comput. Vis. 2025 Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives · CVPR 2024 |
Computer vision › Video understanding and tracking
activity recognition |
1.4 | 2 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 Egocentric Activity Recognition and Localization on a 3D Map · ECCV (13) 2022 |
Computer vision › Video understanding and tracking
activity understanding |
1.3 | 2 | 2024 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives · CVPR 2024 Ego4D: Around the World in 3, 000 Hours of Egocentric Video · CVPR 2022 |
Computer vision › Video understanding and tracking
activity prediction |
0.9 | 1 | 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric Video · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › 3D vision
egocentric vision |
0.9 | 1 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 |
Wearable and physiological sensing › wearable camera › egocentric vision
egocentric sensing |
0.9 | 1 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 |
Computer vision › Video understanding and tracking › egocentric video understanding
first-person activity recognition |
0.6 | 1 | 2022 | Ego4D: Around the World in 3, 000 Hours of Egocentric Video · CVPR 2022 |
Computer vision › 3D vision
3d scene reconstruction |
0.3 | 1 | 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric Video · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
0.3 | 1 | 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives · Int. J. Comput. Vis. 2025 |
Wearable and physiological sensing › wearable display
smart glasses |
0.3 | 1 | 2025 | Reading Recognition in the Wild · NeurIPS 2025 |
Multimedia analysis and retrieval › video dataset
video dataset benchmark |
0.2 | 1 | 2022 | Ego4D: Around the World in 3, 000 Hours of Egocentric Video · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.7head pose · 1.7eye gaze · 1.7multimodal dataset · 0.9benchmark tasks · 0.9egocentric vision · 0.6activity localization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reading Recognition in the WildabstractTo enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public. Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang 0002, Yuning Chai, Richard A. Newcombe, Hyo Jin Kim 0004 |
NeurIPS | 6 |
| 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions—including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. https://ego-exo4d-data.org/ Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, David Crandall, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
Int. J. Comput. Vis. | 69 |
| 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 55 |
| 2024 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multi-modal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured ego-centric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is un-precedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions-including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
CVPR | 69 |
| 2022 | Ego4D: Around the World in 3, 000 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/ Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
CVPR | 54 |
| 2022 | Egocentric Activity Recognition and Localization on a 3D Map
Miao Liu 0007, Lingni Ma, Kiran K. Somasundaram, Yin Li 0003, Kristen Grauman, James M. Rehg |
ECCV (13) | 3 |
| 2017 | An end-to-end system for crowdsourced 3D maps for autonomous vehicles: The mapping componentabstractAutonomous vehicles rely on precise high definition (HD) 3D maps for navigation. This paper presents the mapping component of an end-to-end system for crowdsourcing precise 3D maps with semantically meaningful landmarks such as traffic signs (6 dof pose, shape and size) and traffic lanes (3D splines). The system uses consumer grade parts, and in particular, relies on a single front facing camera and a consumer grade GPS. Using real-time sign and lane triangulation on-device in the vehicle, with offline sign/lane clustering across multiple journeys and offline Bundle Adjustment across multiple journeys in the backend, we construct maps with mean absolute accuracy at sign corners of less than 20 cm from 25 journeys. To the best of our knowledge, this is the first end-to-end HD mapping pipeline in global coordinates in the automotive context using cost effective sensors. Onkar Dabeer, Radhika Gowaiker, Slawomir K. Grzechnik, Mythreya J. Lakshman, Sean Lee, Gerhard Reitmayr, Arunandan Sharma, Kiran K. Somasundaram, Ravi Teja Sukhavasi, Xinzhou Wu |
IROS | 9 |
| 2013 | Proportional Fairness in LTE-Advanced Heterogeneous Networks with eICICabstractThe enhanced Inter-Cell Interference Coordination (eICIC) techniques introduced for LTE-Advanced HetNets yield significant gains by offloading macro cell users to small cells. Interference management is enabled via Almost Blank Subframes (ABS) on which the high power macro cells mute. This creates two different interference patterns corresponding to non-ABS and ABS subframes, which makes analyzing resource allocation and throughput performance non-trivial. We introduce a utility model for the eICIC framework that enables us to study resource allocation with different interference patterns. We use a semi-analytical method using this model to study resource allocation of pico cell-center users on both non-ABS and ABS subframes. We show for proportional fairness, the pico cell-center users may need to be scheduled on ABS subframes. We also demonstrate that this model can be used to determine the optimal ABS configuration for a given deployment. Kiran K. Somasundaram |
VTC Fall | 1 |
| 2009 | Performance improvements in distributed estimation and fusion induced by a trusted core
Kiran K. Somasundaram, John S. Baras |
FUSION | 1 |
| 2009 | Component Based Performance Modelling of Wireless Routing ProtocolsabstractWe propose a component based methodology for modelling and design of wireless routing protocols. Componentization is a standard methodology for analysis and synthesis of complex systems, or software. The feasibility of the component based design relies heavily on the compositionality property (i.e. system-level properties can be computed from properties of components). To provide a component based design methodology and to test compositionality for routing protocols, we have to develop a component based model of the wireless network. We present the main components of the routing protocol that should be modelled and focus on three main components: neighborhood discovery, selector of topology information to disseminate, and the path selection components. For each component, we identify the inputs, outputs, and a generic methodology for modelling. Throughout the paper, we use the Optimized Link State Routing (OLSR) protocol as a case study to demonstrate the effectiveness of our approach. Using the neighborhood discovery component, we present our design methodology and design a modified enhanced version of this component, and compare its performance to the original OLSR design. John S. Baras, Vahid Tabatabaee, Punyaslok Purkayastha, Kiran K. Somasundaram |
ICC | 4 |