VLDB 2026 Research / reviewers in the wild / expert
Shengxin Zha
dblp:126/4565
· DBLP profile ↗
13ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0002-9711-4447ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMsabstractLong-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited context windows. In this work, we introduce VideoMindPalace, a new framework inspired by the "Mind Palace", which organizes critical video moments into a topologically structured semantic graph. VideoMindPalace organizes key information through (i) hand-object tracking and interaction, (ii) clustered activity zones representing specific areas of recurring activities, and (iii) environment layout mapping, allowing natural language parsing by LLMs to provide grounded insights on spatio-temporal and 3D context. In addition, we propose the Video Mind-Palace Benchmark (VMB), to assess human-like reasoning, including spatial localization, temporal reasoning, and layout-aware sequential understanding. Evaluated on VMB and established video QA datasets, including EgoSchema, NExT-QA, IntentQA, and the Active Memories Benchmark, VideoMindPalace demonstrates notable gains in spatiotemporal coherence and human-aligned reasoning, advancing long-form video analysis capabilities in VLMs. Project website: https://oodbag.github.io/VMP/. Zeyi Huang, Yuyang Ji, Nikhil Mehta 0002, Tong Xiao 0003, Donghyun Lee 0004, Sigmund Vanvalkenburgh, Shengxin Zha, Bolin Lai, Licheng Yu, Yong Jae Lee, Miao Liu 0007 |
CVPR | 8 |
| 2024 | BEHAVIOR Vision Suite: Customizable Dataset Generation via SimulationabstractThe systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-world vision datasets rarely satisfy. While current synthetic data generators offer a promising alternative, particularly for embodied AI tasks, they often fall short for computer vision tasks due to low asset and rendering quality, limited diversity, and unrealistic physical properties. We introduce the BEHAVIOR Vision Suite (BVS), a set of tools and assets to generate fully customized synthetic data for systematic evaluation of computer vision models, based on the newly developed embodied AI benchmark, BEHAVIOR-1 K. BVS supports a large number of adjustable parameters at the scene level (e.g., lighting, object placement), the object level (e.g., joint configuration, attributes such as “filled” and “folded”), and the camera level (e.g., field of view, focal length). Researchers can arbitrarily vary these parameters during data generation to perform controlled experiments. We showcase three example application scenarios: systematically evaluating the robustness of models across different continuous axes of domain shift, evaluating scene understanding models on the same set of images, and training and evaluating simulation-to-real transfer for a novel vision task: unary and binary state prediction. Project website: https://behavior-vision-suite.github.io/ Yunhao Ge, Yihe Tang, Cem Gökmen, Chengshu Li 0001, Wensi Ai, Benjamin Jose Martinez, Arman Aydin, Mona Anvari, Ayush K. Chakravarthy, Hong-Xing Yu, Josiah Wong, Sanjana Srivastava, Sharon Lee, Shengxin Zha, Laurent Itti, Yunzhu Li, Roberto Martin Martin, Miao Liu 0007, Pengchuan Zhang, Li Fei-Fei 0001, Jiajun Wu 0001 |
CVPR | 15 |
| 2023 | EgoCom: A Multi-Person Multi-Modal Egocentric Communications DatasetabstractMulti-modal datasets in artificial intelligence (AI) often capture a third-person perspective, but our embodied human intelligence evolved with sensory input from the egocentric, first-person perspective. Towards embodied AI, we introduce the Egocentric Communications (EgoCom) dataset to advance the state-of-the-art in conversational AI, natural language, audio speech analysis, computer vision, and machine learning. EgoCom is a first-of-its-kind natural conversations dataset containing multi-modal human communication data captured simultaneously from the participants' egocentric perspectives. EgoCom includes 38.5 hours of synchronized embodied stereo audio, egocentric video with 240,000 ground-truth, time-stamped word-level transcriptions and speaker labels from 34 diverse speakers. We study baseline performance on two novel applications that benefit from embodied data: (1) predicting turn-taking in conversations and (2) multi-speaker transcription. For (1), we investigate Bayesian baselines to predict turn-taking within 5 percent of human performance. For (2), we use simultaneous egocentric capture to combine Google speech-to-text outputs, improving global transcription by 79 percent relative to a single perspective. Both applications exploit EgoCom's synchronous multi-perspective data to augment performance of embodied AI tasks. Curtis G. Northcutt, Shengxin Zha, Steven Lovegrove, Richard A. Newcombe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Pattern-Based Reconstruction of K-Level Images From CutsetsabstractWe present a pattern-based approach for reconstructing a K-level image from cutsets, dense samples taken along a family of lines or curves in two- or three-dimensional space, which break the image into blocks, each of which is typically reconstructed independently of the others. The pattern-based approach utilizes statistics of human segmentations to generate a codebook of patterns, each of which represents a pair of a block boundary specification and the corresponding pattern in the block interior. We develop the approach for rectangular cutset topologies and show that it can be extended to general periodic sampling topologies. We also show that, for bilevel cutset reconstruction, the pattern-based can be combined with the previously proposed cutset-MRF approach to substantially reduce the size of the codebook with a slight increase in reconstruction error. In addition, we present an algorithm for segmenting the cutset samples of an original grayscale or color image, followed by reconstruction of the full segmentation field via the pattern-based approach. Experimental results show that the proposed approaches outperform the cutset-MRF approaches in terms of both reconstruction error rate and perceptual quality. Moreover, this is accomplished without any side information about the structure of the block interior. Systematic comparisons of the performance of different sampling topologies are also provided. Shengxin Zha, Daizong Tian, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 1 |
| 2021 | Only Time Can Tell: Discovering Temporal Data for Temporal ModelingabstractUnderstanding temporal information and how the visual world changes over time, is a fundamental ability of intelligent systems. In video understanding, temporal information is at the core of many current challenges, including compression, efficient inference, motion estimation or summarization. However, in current video datasets it has been observed that action classes can often be recognized without any temporal information, from a single frame of video. As a result, both benchmarking and training in these datasets may give an unintentional advantage to models with strong image understanding capabilities, as opposed to those with strong temporal understanding. In other words, current datasets may not reward good temporal understanding, potentially hindering progress. In this paper we address this problem head on by identifying action classes where temporal information is actually necessary to recognize them and call these "temporal classes". Selecting temporal classes using a computational method would bias the process. Instead, we propose a methodology based on a simple and effective human annotation experiment. We remove just the temporal information, by shuffling frames in time, and measure if the action can still be recognized. Classes that cannot be recognized when frames are not in order, are included in the temporal set. We observe that this set is statistically different from other static classes, and that performance in it correlates with a network's ability to capture temporal information. Thus we use it as a benchmark on current popular networks, which reveals a series of interesting facts, like inflated convolutions bias networks towards classes where motion is not important. We also explore the effect of training on the temporal set, and observe that this leads to better generalization in unseen classes, demonstrating the need for more temporal data. We hope that the proposed dataset of temporal categories will help guide future research in temporal modeling for better video understanding. Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan 0001, Vedanuj Goswami, Matt Feiszli, Lorenzo Torresani |
WACV | 2 |
| 2021 | Hierarchical Lossy Bilevel Image Compression Based on Cutset SamplingabstractWe consider lossy compression of a broad class of bilevel images that satisfy the smoothness criterion, namely, images in which the black and white regions are separated by smooth or piecewise smooth boundaries, and especially lossy compression of complex bilevel images in this class. We propose a new hierarchical compression approach that extends the previously proposed fixed-grid lossy cutset coding (LCC) technique by adapting the grid size to local image detail. LCC was claimed to have the best rate-distortion performance of any lossy compression technique in the given image class, but cannot take advantage of detail variations across an image. The key advantages of the hierarchical LCC (HLCC) is that, by adapting to local detail, it provides constant quality controlled by a single parameter (distortion threshold), independent of image content, and better overall visual quality and rate-distortion performance, over a wider range of bitrates. We also introduce several other enhancements of LCC that improve reconstruction accuracy and perceptual quality. These include the use of multiple connection bits that provide structural information by specifying which black (or white) runs on the boundary of a block must be connected, a boundary presmoothing step, stricter connectivity constraints, and more elaborate probability estimation for arithmetic coding. We also propose a progressive variation that refines the image reconstruction as more bits are transmitted, with very small additional overhead. Experimental results with a wide variety of, and especially complex, bilevel images in the given class confirm that the proposed techniques provide substantially better visual quality and rate-distortion performance than existing lossy bilevel compression techniques, at bitrates lower than lossless compression with the JBIG or JBIG2 standards. Shengxin Zha, Thrasyvoulos N. Pappas, David L. Neuhoff |
IEEE Trans. Image Process. | 1 |
| 2020 | SF-Net: Single-Frame Supervision for Temporal Action Localization
Fan Ma, Linchao Zhu, Yi Yang 0001, Shengxin Zha, Gourab Kundu, Matt Feiszli, Zheng Shou 0001 |
ECCV (4) | 4 |
| 2016 | Generalized k-level cutset sampling and reconstructionabstractWe propose a family of cutset sampling schemes and a generalized k-level image reconstruction approach formulated under a minimum mean squared error (MMSE) framework. The k-level reconstruction approach is a direct generalization of the recently proposed pattern-based approach, and can be applied to periodic samples either on a cutset or on a grid. Our experimental results indicate that the generalization of the k-level reconstruction approach results in only a small performance loss. For rectangular cutsets, we show that the proposed approach outperforms the cutset-MRF approach as well as two inpainting approaches. Moreover, we show that combining the cutset sampling with an additional point sample inside the periodic structure outperforms k-level reconstruction from cutset sampling and point sampling under comparable sampling densities. Shengxin Zha, Thrasyvoulos N. Pappas |
ICASSP | 1 |
| 2016 | A hybrid Markov random field model for bilevel cutset reconstructionabstractWe propose a hybrid Markov random field (MRF) model and a two stage cutset-MRF approach for reconstructing bilevel images from cutsets. We show that the proposed approach leads to substantial improvements over previous cutset-MRF approaches in terms of both reconstruction error and visual quality (continuity of reconstructed segments and preservation of image structure). The proposed approach approaches the performance of pattern-based approaches without the additional memory requirements and training overhead. We also show that it outperforms inpainting approaches adapted to bilevel cutset reconstruction. Shengxin Zha, Thrasyvoulos N. Pappas |
ICIP | 1 |
| 2015 | Exploiting Image-trained CNN Architectures for Unconstrained Video ClassificationabstractWe conduct an in-depth exploration of different strategies for doing event detection in videos using convolutional neural networks (CNNs) trained for image classification. We study different ways of performing spatial and temporal pooling, feature normalization, choice of CNN layers as well as choice of classifiers. Making judicious choices along these dimensions led to a very significant increase in performance over more naive approaches that have been used till now. We evaluate our approach on the challenging TRECVID MED'14 dataset with two popular CNN architectures pretrained on ImageNet. On this MED'14 dataset, our methods, based entirely on image-trained CNN features, can outperform several state-of-the-art non-CNN models. Our proposed late fusion of CNN- and motion-based features can further increase the mean average precision (mAP) on MED'14 from 34.95% to 38.74%. The fusion approach achieves the state-of-the-art classification performance on the challenging UCF-101 dataset. Shengxin Zha, Florian Luisier, Walter Andrews, Nitish Srivastava, Ruslan Salakhutdinov |
BMVC | 1 |
| 2015 | Pattern-based k-level cutset reconstructionabstractWe propose a pattern-based approach for reconstructing k-level images from cutsets. We construct a database of k-level cutset block patterns and use it for reconstruction based on fuzzy pattern retrieval and Markov random field (MRF) energy minimization. The proposed approach outperforms previously proposed cutset-MRF approaches as well as inpainting approaches. When integrated into a lossy bilevel image coding scheme that utilizes cutsets, the proposed approach outperforms the state of the art with comparable constraints. Shengxin Zha, Thrasyvoulos N. Pappas |
ICIP | 1 |
| 2014 | Text Classification via iVector Based Feature RepresentationabstractIn this paper, we address the problem of text classification: classifying modern machine-printed text, handwritten text and historical typewritten text from degraded noisy documents. We propose a novel text classification approach based on iVector, a newly developed concept in speaker verification. To a given text line, the iVector is a fixed-length feature vector representation, transformed from a high-dimensional super vector based on means of Gaussian mixture model (GMM), where the text dependent component is separated from a universal background model (UBM) and can be represented by a low dimensional set of factors. We classify the text lines with a discriminative classifier - support vector machine (SVM) in iVector space. A baseline approach of text classification using GMM in feature space is also presented for evaluation purpose. Experimental results on an Arabic document database show accuracy of 92.04% for text line classification using the proposed method. Furthermore, the relative word error rate (WER) of 9.6% is decreased in optical character recognition (OCR) when coupled with the proposed iVector-SVM classifier. The proposed iVector-SVM approach is language independent, thus, can be applied to other scripts as well. Shengxin Zha, Xujun Peng, Huaigu Cao, Xiaodan Zhuang, Pradeep Natarajan, Premkumar Natarajan |
Document Analysis Systems | 1 |
| 2012 | Hierarchical bilevel image compression based on cutset samplingabstractWe propose a hierarchical lossy bilevel image compression method that relies on adaptive cutset sampling (along lines of a rectangular grid with variable block size) and Markov Random Field based reconstruction. It is an efficient encoding scheme that preserves image structure by using a coarser grid in smooth areas of the image and a finer grid in areas with more detail. Experimental results demonstrate that the proposed method performs as well as or better than the fixed-grid approach, and outperforms other lossy bilevel compression methods in its rate-distortion performance. Shengxin Zha, Thrasyvoulos N. Pappas, David L. Neuhoff |
ICIP | 1 |