VLDB 2026 Research / reviewers in the wild / expert
Huisheng Wang
dblp:48/6289
· DBLP profile ↗
14ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-8080-1777ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes Under Herd BehaviorabstractAligning Large Language Models (LLMs) with investor decision-making processes under herd behavior is a critical challenge in behavioral finance, which grapples with a fundamental limitation: the scarcity of real-user data needed for Supervised Fine-Tuning (SFT). While SFT can bridge the gap between LLM outputs and human behavioral patterns, its reliance on massive authentic data imposes substantial collection costs and privacy risks. We propose InvestAlign, a novel framework that constructs high-quality SFT datasets by leveraging theoretical solutions to similar and simple optimal investment problems rather than the complex scenarios. Our theoretical analysis demonstrates that training LLMs with InvestAlign-generated data achieves faster parameter convergence than using real-user data, suggesting superior learning efficiency. Furthermore, we develop InvestAgent, an LLM agent fine-tuned with InvestAlign, which shows significantly closer alignment to real-user data than pre-SFT models in both simple and complex investment problems. This highlights our proposed InvestAlign as a promising approach with the potential to address complex optimal investment problems and align LLMs with investor decision-making processes under herd behavior. Our code is publicly available at https://github.com/thu-social-network-research-group/InvestAlign. Huisheng Wang, Zhuoshi Pan, Hangjing Zhang, Mingxiao Liu, Hanqing Gao, H. Vicky Zhao |
ACL (1) | 1 |
| 2025 | SANPO: A Scene Understanding, Accessibility and Human Navigation DatasetabstractVision is essential for human navigation. The World Health Organization (WHO) estimates that 43.3 million people were blind in 2020, and this number is projected to reach 61 million by 2050. Modern scene understanding models could empower these people by assisting them with navigation, obstacle avoidance and visual recognition capabilities. The research community needs high quality datasets for both training and evaluation to build these systems. While datasets for autonomous vehicles are abundant, there is a critical gap in datasets tailored for outdoor human navigation. This gap poses a major obstacle to the development of computer vision based Assistive Technologies. To overcome this obstacle, we present SANPO, a large-scale egocentric video dataset designed for dense prediction in outdoor human navigation environments. SANPO contains 701 stereo videos of 30+ seconds captured in diverse real-world outdoor environments across four geographic locations in the USA. Every frame has a high resolution depth map and 112K frames were annotated with temporally consistent dense video panoptic segmentation labels. The dataset also includes 1961 high-quality synthetic videos with pixel accurate depth and panoptic segmentation annotations to balance the noisy real world annotations with the high precision synthetic annotations. SANPO is already publicly available and is being used by mobile applications like Project Guideline to train mobile models that help low-vision users go running outdoors independently. To preserve anonymization during peer review, we will provide a link to our dataset upon acceptance. Sagar Waghmare, Kimberly Wilber, Dave Hawkey, Matthew Wilson, Stephanie Debats, Cattalyya Nuengsigkapian, Astuti Sharma, Lars Pandikow, Huisheng Wang, Hartwig Adam, Mikhail Sirotenko |
WACV | 10 |
| 2024 | Improving Subject-Driven Image Synthesis with Subject-Agnostic GuidanceabstractIn subject-driven text-to-image synthesis, the synthesis process tends to be heavily influenced by the reference images provided by users, often overlooking crucial attributes detailed in the text prompt. In this work, we propose Subject-Agnostic Guidance (SAG), a simple yet effective solution to remedy the problem. We show that through constructing a subject-agnostic condition and applying our proposed dual classifier-free guidance, one could obtain outputs consistent with both the given subject and input text prompts. We validate the efficacy of our approach through both optimization-based and encoder-based methods. Additionally, we demonstrate its applicability in second-order customization methods, where an encoder-based model is fine-tuned with DreamBooth. Our approach is conceptually simple and requires only minimal code modifications, but leads to substantial quality improvements, as evidenced by our evaluations and user studies. Kelvin C. K. Chan, Xuhui Jia, Ming-Hsuan Yang 0001, Huisheng Wang |
CVPR | 5 |
| 2024 | VideoPrism: A Foundational Visual Encoder for Video UnderstandingabstractWe introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding by global-local distillation of semantic video embeddings and a token shuffling scheme, enabling VideoPrism to focus primarily on the video modality while leveraging the invaluable text associated with videos. We extensively test VideoPrism on four broad groups of video understanding tasks, from web video question answering to CV for science, achieving state-of-the-art performance on 31 out of 33 video understanding benchmarks. Long Zhao 0003, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou 0014, Shen Yan 0008, Jennifer J. Sun, Luke Friedman, Rui Qian 0003, Tobias Weyand, Yue Zhao 0006, Rachel Hornung, Florian Schroff, Ming-Hsuan Yang 0001, David A. Ross, Huisheng Wang, Hartwig Adam, Mikhail Sirotenko, Ting Liu 0005, Boqing Gong |
ICML | 15 |
| 2024 | VideoPoet: A Large Language Model for Zero-Shot Video GenerationabstractWe present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/ Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004 |
ICML | 28 |
| 2024 | PolyMaX: General Dense Prediction with Mask TransformerabstractDense prediction tasks, such as semantic segmentation, depth estimation, and surface normal prediction, can be easily formulated as per-pixel classification (discrete outputs) or regression (continuous outputs). This per-pixel prediction paradigm has remained popular due to the prevalence of fully convolutional networks. However, on the recent frontier of segmentation task, the community has been witnessing a shift of paradigm from per-pixel prediction to cluster-prediction with the emergence of transformer architectures, particularly the mask transformers, which directly predicts a label for a mask instead of a pixel. Despite this shift, methods based on the per-pixel prediction paradigm still dominate the benchmarks on the other dense prediction tasks that require continuous outputs, such as depth estimation and surface normal prediction. Motivated by the success of DORN and AdaBins in depth estimation, achieved by discretizing the continuous output space, we propose to generalize the cluster-prediction based method to general dense prediction tasks. This allows us to unify dense prediction tasks with the mask transformer framework. Remarkably, the resulting model PolyMaX demonstrates state-of-the-art performance on three benchmarks of NYUD-v2 dataset. We hope our simple yet effective design can inspire more research on exploiting mask transformers for more dense prediction tasks. Code and model will be made available1. Liangzhe Yuan, Kimberly Wilber, Astuti Sharma, Xiuye Gu, Siyuan Qiao, Stephanie Debats, Huisheng Wang, Hartwig Adam, Mikhail Sirotenko, Liang-Chieh Chen |
WACV | 8 |
| 2023 | Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal PerceptionabstractWe present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transformer encoder with minimal modality-specific components. IMP makes use of a novel design that combines Alternating Gradient Descent (AGD) and Mixture-of-Experts (MoE) for efficient model & task scaling. We conduct extensive empirical studies and reveal the following key insights:
1) performing gradient descent updates by alternating on diverse modalities, loss functions, and tasks, with varying input resolutions, efficiently improves the model.
2) sparsification with MoE on a single modality-agnostic encoder substantially improves the performance, outperforming dense models that use modality-specific encoders or additional fusion layers and greatly mitigating the conflicts between modalities.
IMP achieves competitive performance on a wide range of downstream tasks including video classification, image classification, image-text, and video-text retrieval. Most notably, we train a sparse IMP-MoE-L focusing on video tasks that achieves new state-of-the-art in zero-shot video classification: 77.0% on Kinetics-400, 76.8% on Kinetics-600, and 68.3% on Kinetics-700, improving the previous state-of-the-art by +5%, +6.7%, and +5.8%, respectively, while using only 15% of their total training computational cost. Hassan Akbari, Dan Kondratyuk, Yin Cui, Rachel Hornung, Huisheng Wang, Hartwig Adam |
NeurIPS | 5 |
| 2023 | MovieCLIP: Visual Scene Recognition in MoviesabstractLongform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual scenes in movies include transitions, person coverage, and a wide array of real-life and fictional scenarios. Existing visual scene datasets in movies have limited taxonomies and don’t consider the visual scene transition within movie clips. In this work, we address the problem of visual scene recognition in movies by first automatically curating a new and extensive movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets. Instead of manual annotations which can be expensive, we use CLIP to weakly label 1.12 million shots from 32K movie clips based on our proposed taxonomy. We provide baseline visual models trained on the weakly labeled dataset called MovieCLIP and evaluate them on an independent dataset verified by human raters. We show that leveraging features from models pretrained on MovieCLIP benefits downstream tasks such as multi-label scene and genre classification of web videos and movie trailers. Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Yin Cui, Kree Cole-McLaughlin, Huisheng Wang, Shri Narayanan |
WACV | 7 |
| 2021 | Spatiotemporal Contrastive Video Representation LearningabstractWe present a self-supervised Contrastive Video Representation Learning (CVRL) method to learn spatiotemporal visual representations from unlabeled videos. Our representations are learned using a contrastive loss, where two augmented clips from the same short video are pulled together in the embedding space, while clips from different videos are pushed away. We study what makes for good data augmentations for video self-supervised learning and find that both spatial and temporal information are crucial. We carefully design data augmentations involving spatial and temporal cues. Concretely, we propose a temporally consistent spatial augmentation method to impose strong spatial augmentations on each frame of the video while maintaining the temporal consistency across frames. We also propose a sampling-based temporal augmentation method to avoid overly enforcing invariance on clips that are distant in time. On Kinetics-600, a linear classifier trained on the representations learned by CVRL achieves 70.4% top-1 accuracy with a 3D-ResNet-50 (R3D-50) backbone, outperforming ImageNet supervised pre-training by 15.7% and SimCLR unsupervised pre-training by 18.8% using the same inflated R3D-50. The performance of CVRL can be further improved to 72.9% with a larger R3D-152 (2× filters) backbone, significantly closing the gap between unsupervised and supervised video representation learning. Our code and models will be available at https://github.com/tensorflow/models/tree/master/official/. Rui Qian 0003, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang 0001, Huisheng Wang, Serge J. Belongie, Yin Cui |
CVPR | 5 |
| 2009 | Rate-Distortion Optimized Scheduling for Redundant Video RepresentationsabstractThis paper extends rate-distortion optimized streaming techniques to operate on a general class of coding formats that explicitly support redundancy in their coding structure. Examples include multiple description layered coding (MDLC) and multiple independently encoded versions of a video source. Such source codecs usually produce multiple decoding paths, while previous work on video streaming has mostly focused on those encoding techniques that only generate a single decoding path. A new source model called Directed Acyclic HyperGraph is introduced to describe the dependency and redundancy relationship between different video data units with multiple decoding paths. Based on this model, we then propose two rate-distortion based packet scheduling algorithms, i.e., Lagrangian optimization and a greedy algorithm, to dynamically adjust the system's real-time redundancy to match the channel behavior. The proposed streaming system introduces two types of redundancies, namely, source redundancy and transport redundancy. This paper presents a detailed performance analysis of the individual benefits for error robustness provided by these redundancies and their interplay. Experimental results show that our proposed system with both redundancies achieves the best end-to-end performance on real-time video communication over a wide range of network scenarios. Huisheng Wang, Antonio Ortega |
IEEE Trans. Image Process. | 1 |
| 2008 | Sampling-Based Correlation Estimation for Distributed Source Coding Under Rate and Complexity ConstraintsabstractIn many practical distributed source coding (DSC) applications, correlation information has to be estimated at the encoder in order to determine the encoding rate. Coding efficiency depends strongly on the accuracy of this correlation estimation. While error in estimation is inevitable, the impact of estimation error on compression efficiency has not been sufficiently studied for the DSC problem. In this paper,we study correlation estimation subject to rate and complexity constraints, and its impact on coding efficiency in a DSC framework for practical distributed image and video applications. We focus on, in particular, applications where binary correlation models are exploited for Slepian-Wolf coding and sampling techniques are used to estimate the correlation, while extensions to other correlation models would also be briefly discussed. In the first part of this paper, we investigate the compression of binary data. We first propose a model to characterize the relationship between the number of samples used in estimation and the coding rate penalty, in the case of encoding of a single binary source. The model is then extended to scenarios where multiple binary sources are compressed, and based on the model we propose an algorithm to determine the number of samples allocated to different sources so that the overall rate penalty can be minimized, subject to a constraint on the total number of samples. The second part of this paper studies compression of continuous valued data. We propose a model-based estimation for the particular but important situations where binary bit-planes are extracted from a continuous-valued input source, and each bit-plane is compressed using DSC. The proposed model-based method first estimates the source and correlation noise models using continuous valued samples, and then uses the models to derive the bit-plane statistics analytically. We also extend the model-based estimation to the cases when bit-planes are extracted based on the significance of the data, similar to those commonly used in wavelet-based applications. Experimental results, including some based on hyperspectral image compression, demonstrate the effectiveness of the proposed algorithms. Ngai-Man Cheung, Huisheng Wang, Antonio Ortega |
IEEE Trans. Image Process. | 2 |
| 2005 | Format independent encryption of generalized scalable bit-streams enabling arbitrary secure adaptations [multimedia communication applications]abstractSecure format-independent adaptation of bit-streams during delivery is becoming increasingly important to cope with content piracy while still accommodating diverse networks, terminals and formats. Generalized scalable bit-streams are particularly advantageous in this regard, since they enable a variety of efficient and secure adaptations. Further, by associating such a bit-stream with appropriate metadata, such as those standardized in MPEG-21 Part 7, entitled digital item adaptation (DIA), the adaptation process can be fully format-independent. In this paper, to maximally secure a generalized scalable bit-stream while allowing arbitrary encrypted domain adaptations, strong progressive encryption methods are extended to multiple dimensions. Further it is shown that by appropriate modeling of such bit-streams and re-use of some DIA descriptions, the encryption and decryption engines themselves can be entirely metadata-driven and format-independent. This leads to end-to-end format-independent secure and adaptive delivery architectures for scalable bit-streams. Debargha Mukherjee, Huisheng Wang, Amir Said, Sam Liu |
ICASSP (2) | 2 |
| 2005 | Correlation estimation for distributed source coding under information exchange constraintsabstractDistributed source coding (DSC) depends strongly on accurate knowledge of correlation between sources. Previous works have reported capacity-approaching code constructions when exact knowledge of correlation is available at the encoder. However, in many applications exact correlation information may not be available, and correlation estimation is necessary. While error in estimation is inevitable, the impact of estimation error on compression efficiency has not been sufficiently studied for the DSC problem. In this paper we study correlation estimation subject to complexity constraints, and its impact on coding efficiency in a DSC framework. In particular, we consider the case where estimation entails information exchange between spatially separate sources and thus correlation estimation is subject to rate constraints. We first derive optimal strategies for information exchange that minimize the rate penalty due to inaccurate estimation, under constraints on the number of bits that can be exchanged between sources. Experimental results show that significant gain is possible by optimally exchanging information. We then derive analytical expressions to quantify the rate penalty, and analyze how rate penalty changes with a priori knowledge of correlation. In addition, we present a model-based estimation method which can achieve more accurate estimation results compared to directly inspecting the data. Ngai-Man Cheung, Huisheng Wang, Antonio Ortega |
ICIP (2) | 2 |
| 2004 | Scalable predictive coding by nested quantization with layered side information
Huisheng Wang, Antonio Ortega |
ICIP | 1 |