VLDB 2026 Research / reviewers in the wild / expert
Shihao Zou
dblp:223/4696
· DBLP profile ↗
33ranked-venue papers
11as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text RecognitionabstractMasked Image Modeling (MIM) has been widely recognized as a powerful self-supervised paradigm for learning general-purpose visual representations. However, standard MIM based on random masking tends to underperform in domain-specific tasks like Scene Text Recognition (STR), due to challenges such as information sparsity and appearance discrepancies caused by partial occlusion or distortion. To address this issue, we propose a novel pre-training framework called Appearance Discrepancy-guided Sequence Hybrid Masking (DSHM), specifically designed to learn robust representations for STR. To this end, we introduce an Appearance Discrepancy Metric that quantifies the discrepancy level of each image patch by measuring its deviation from anisotropic local discrepancy and intra-instance global style discrepancy. The resulting discrepancy scores are utilized in two key components: (1) A Sequence Hybrid Masking strategy, which prioritizes masking high-discrepancy patches in coherent block forms, thereby elevating the pretext task from simple pixel-level completion to more complex structural reasoning; (2) Discrepancy-Conditioned Tokens (DC-Tokens), which encode prior knowledge about patch difficulty into the decoder, enabling an adaptive reconstruction process and improving the model robustness under scenarios with partial occlusion or text distortion. We achieve competitive performance on multiple benchmark datasets, including common benchmarks, Union14M benchmarks, and Chinese benchmarks. Shihao Zou, Wei Wei 0002, Leyang Xu, Kaihe Xu, Wenfeng Xie |
AAAI | 1 |
| 2026 | SAM3-I: Segment Anything with InstructionsabstractJingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jincai Huang 0003, Wei Ji 0011, Qi Bi, Yongri Piao, Miao Zhang 0004, Xiaoqi Zhao 0003, Qiang Chen 0007, Shihao Zou, Huchuan Lu, Li Cheng 0001 |
ACL (1) | 11 |
| 2026 | Surgical Data Science in Time-Critical Contexts: A Roadmap Toward Brain-Inspired Computing
Yi Pan 0001, Shihao Zou, Jia-Wen Yang, Weixin Si |
J. Comput. Sci. Technol. | 2 |
| 2026 | Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding SpaceabstractMotion retrieval is crucial for motion acquisition, offering superior precision, realism, controllability, and editability compared to motion generation. Existing approaches leverage contrastive learning to construct a unified embedding space for motion retrieval from text or visual modality. However, these methods lack a more intuitive and user-friendly interaction mode and often overlook the sequential representation of most modalities for improved retrieval performance. To address these limitations, we propose a framework that aligns four modalities—text, audio, video, and motion—within a fine-grained joint embedding space, incorporating audio for the first time in motion retrieval to enhance user immersion and convenience. This fine-grained space is achieved through a sequence-level contrastive learning approach, which captures critical details across modalities for better alignment. To evaluate our framework, we augment existing text-motion datasets with synthetic but diverse audio recordings, creating two multi-modal motion retrieval datasets. Experimental results demonstrate superior performance over state-of-the-art methods across multiple sub-tasks, including an 10.16% improvement in R@10 for text-to-motion retrieval and a 25.43% improvement in R@1 for video-to-motion retrieval on the HumanML3D dataset. Furthermore, our results show that our 4-modal framework significantly outperforms its 3-modal counterpart, underscoring the potential of multi-modal motion retrieval for advancing motion acquisition. Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian 0002, Weixin Si, Shihao Zou |
IEEE Trans. Multim. | 7 |
| 2025 | TempDiffReg: Temporal Diffusion Model for Non-Rigid 2D-3D Vascular RegistrationabstractTransarterial chemoembolization (TACE) is a preferred treatment option for hepatocellular carcinoma and other liver malignancies, yet it remains a highly challenging procedure due to complex intra-operative vascular navigation and anatomical variability. Accurate and robust 2D-3D vessel registration is essential to guide microcatheter and instruments during TACE, enabling precise localization of vascular structures and optimal therapeutic targeting. To tackle this issue, we develop a coarse-to-fine registration strategy. First, we introduce a global alignment module, structure-aware perspective n-point (SA-PnP), to establish correspondence between 2D and 3D vessel structures. Second, we propose TempDiffReg, a temporal diffusion model that performs vessel deformation iteratively by leveraging temporal context to capture complex anatomical variations and local structural changes. We collected data from 23 patients and constructed 626 paired multi-frame samples for comprehensive evaluation. Experimental results demonstrate that the proposed method consistently outperforms state-of-the-art (SOTA) methods in both accuracy and anatomical plausibility. Specifically, our method achieves a mean squared error (MSE) of 0.63 mm and a mean absolute error (MAE) of 0.51 mm in registration accuracy, representing$66.7\%$lower MSE and$17.7\%$lower MAE compared to the most competitive existing approaches. It has the potential to assist less-experienced clinicians in safely and efficiently performing complex TACE procedures, ultimately enhancing both surgical outcomes and patient care. Code and data are available at: https://github.com/LZH970328/TempDiffReg.git Zehua Liu, Shihao Zou, Jincai Huang 0003, Weixin Si |
BIBM | 2 |
| 2025 | Distance-Aware and Knowledge-Driven Vision Mamba U-Net for Radiotherapy Dose PredictionabstractDose planning is essential in radiotherapy for cancer patients, yet current practice relies on iterative manual optimization, underscoring the need for automated prediction. Existing deep learning approaches remain limited because they often ignore the 3D spatial relationships between tumors and surrounding organs at risk (OARs), and clinical priors on safe dose thresholds. To overcome these limitations, we propose DKVMU-Net, a distance-aware and knowledge-driven Vision Mamba U-Net for automated dose prediction. Our framework incorporates Vision Mamba blocks to capture global, long-range dependencies from CT scans and OAR signed distance field (SDF) maps, which naturally encode spatial information. Additionally, we introduce a deformable dynamic feature enhancement module (DDFEM) for texture refinement, followed by a linear crossattention fusion module to improve cross-modality integration. A customized loss function is also designed to incorporate prior knowledge of OAR dose constraints, ensuring optimal target coverage and OAR protection. To alleviate the scarcity of doseplanning datasets, we collect an in-house radiotherapy lung cancer dataset (RLCD), consisting of CT volumes, OAR masks, and corresponding SDF maps from 116 patients. We evaluate our DKVMU-Net on both the in-house dataset and public available OpenKBP dataset. Compared with the sate-of-the-art method, our approach achieves an 11.6 % improvement in dose score (1.641 vs. 1.857) and 26.3 % in DVH score (6.481 vs. 8.799) on RLCD, and a 7.8 % improvement in dose score (2.421 vs. 2.626) and 13.9 % in DVH score (1.057 vs. 1.227) on OpenKBP. These results demonstrate the robustness and effectiveness of our approach. Yangyang Shi, Xiaoyan Kui, Yucong Zhang, Shihao Zou, Zuheng Ming, Weixin Si, Azeddine Beghdadi, Beiji Zou 0001 |
BIBM | 4 |
| 2025 | Modal Feature Optimization Network with Prompt for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis(MSA) is mostly used to understand human emotional states through multimodal. However, due to the fact that the effective information carried by multimodal is not balanced, the modality containing less effective information cannot fully play the complementary role between modalities. Therefore, the goal of this paper is to fully explore the effective information in modalities and further optimize the under-optimized modal representation.To this end, we propose a novel Modal Feature Optimization Network (MFON) with a Modal Prompt Attention (MPA) mechanism for MSA. Specifically, we first determine which modalities are under-optimized in MSA, and then use relevant prompt information to focus the model on these features. This allows the model to focus more on the features of the modalities that need optimization, improving the utilization of each modality’s feature representation and facilitating initial information aggregation across modalities. Subsequently, we design an intra-modal knowledge distillation strategy for under-optimized modalities. This approach preserves the integrity of the modal features. Furthermore, we implement inter-modal contrastive learning to better extract related features across modalities, thereby optimizing the entire network. Finally, sentiment prediction is carried out through the effective fusion of multimodal information. Extensive experimental results on public benchmark datasets demonstrate that our proposed method outperforms existing state-of-the-art models. Xiangmin Zhang, Wei Wei 0002, Shihao Zou |
COLING | 3 |
| 2025 | Medical Open Set Recognition via Intra-Class ClusteringabstractIn computational medical imaging, model's ability to identify whether a sample is from an unseen semantic category is critical in clinical deployments. However, conventional medical image recognition usually assumes a closed-set setting where all queries in testing are from pre-defined training categories and overlooks the fact that in practice it is possible to have queries from unknown categories such as unknown or unseen tissue. In this study, we particularly tackle this thorny challenge, namely Medical Open Set Recognition (MOSR), and explore it on medical image classification and diagnosis. The biggest challenge with this issue lies in deep model's overconfidence due to relatively large intra-class variance, which leads to incorrectly assigning an unknown sample to a known class with a high confidence level. To address this problem, we introduce intra-class clustering, which divides the samples assigned to each class into several low-variance sub-clusters. In addition, we propose to divide the samples uniformly to each cluster by optimal transport to achieve online clustering. Extensive experiments on 6 public medical imaging datasets demonstrate that a classification model trained with the proposed intra-class clustering dramatically alleviate the overconfidence problem with competitive accuracy and thus effective for improving MOSR performance. Our benchmarks and code will be publicly released when published. Hanqiu Deng, Shihao Zou, Xiangyun Liao, Weixin Si |
CW | 2 |
| 2025 | SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and O(T) Complexity
Shihao Zou, Yongkui Yang |
ICML | 1 |
| 2025 | Cerebrovascular Diseases Screening from Color Fundus Photography via Cross-View Fusion and Graph-Based Discrimination
Congyu Tian, Shihao Zou, Xiangyun Liao, Chubin Ou, Jianping Lv, Shanshan Wang 0002, Weixin Si |
MICCAI (12) | 2 |
| 2025 | Semantics-aware human motion generation from audio instructionsabstractRecent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions that align with the semantics of the audio. Unlike text-based interactions, audio provides a more natural and intuitive communication method. However, existing methods typically focus on matching motions with music or speech rhythms, which often results in a weak connection between the semantics of the audio and generated motions. We propose an end-to-end framework using a masked generative transformer, enhanced by a memory-retrieval attention module to handle sparse and lengthy audio inputs. Additionally, we enrich existing datasets by converting descriptions into conversational style and generating corresponding audio with varied speaker identities. Experiments demonstrate the effectiveness and efficiency of the proposed framework, demonstrating that audio instructions can convey semantics similar to text while providing more practical and user-friendly interactions. Zi-An Wang, Shihao Zou, Shiyao Yu |
Graph. Model. | 2 |
| 2025 | Intentional tendency-based dynamic heterogeneous graph network for emotion recognition in conversations
Xinyi Gan, Xianying Huang, Shihao Zou |
J. Intell. Inf. Syst. | 3 |
| 2025 | Highly Efficient 3D Human Pose Tracking From Events With Spiking Spatiotemporal TransformerabstractEvent camera, as an asynchronous vision sensor capturing scene dynamics, presents new opportunities for highly efficient 3D human pose tracking. Existing approaches typically adopt modern-day Artificial Neural Networks (ANNs), such as CNNs or Transformer, where sparse events are converted into dense images or paired with additional gray-scale images as input. Such practices, however, ignore the inherent sparsity of events, resulting in redundant computations, increased energy consumption, and potentially degraded performance. Motivated by these observations, we introduce the first sparse Spiking Neural Networks (SNNs) framework for 3D human pose tracking based solely on events. Our approach eliminates the need to convert sparse data to dense formats or incorporate additional images, thereby fully exploiting the innate sparsity of input events. Central to our framework is a novel Spiking Spatio-temporal Transformer, which enables bi-directional spatio-temporal fusion of spike pose features and provides a guaranteed similarity measurement between binary spike features in spiking attention. Moreover, we have constructed a largescale synthetic dataset, SynEventHPD, that features a broad and diverse set of 3D human motions, as well as much longer hours of event streams. Empirical experiments demonstrate the superiority of our approach over existing state-of-the-art (SOTA) ANN-based methods, requiring only 19.1% FLOPs and 3.6% energy cost. Furthermore, our approach outperforms existing SNN-based benchmarks in this task, highlighting the effectiveness of our proposed SNN framework. The dataset will be released upon acceptance, and code can be found at https://github.com/JimmyZou/HumanPoseTracking_SNN. Shihao Zou, Yuxuan Mu, Wei Ji 0011, Zi-An Wang, Xinxin Zuo, Sen Wang 0003, Weixin Si, Li Cheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Generating High-Fidelity Clothed Human Dynamics with Temporal DiffusionabstractClothed human modeling plays a crucial role in multimedia research, with applications spanning virtual reality, gaming, and fashion design. The goal is to learn clothed human dynamics from observations and then generate humans with high-fidelity clothing details for motion animation. Despite tremendous advancements in clothing shape analysis by existing approaches, the community still faces challenges in generating convincing visual effects of cloth dynamics, maintaining temporally smooth clothing details, and handling diverse clothing patterns. To address these challenges, we introduce ClothDiffuse, a temporal diffusion model that seamlessly integrates three key components into this task—temporal dynamics modeling, iterative refinement, and diversified generation. Our approach begins by using an encoder to extract high-level temporal features from input human body motions. These features are combined with a learnable pixel-aligned garment feature, serving as prior conditions for the shape decoder. The decoder then iteratively denoise Gaussian noise to produce clothing deformations over time on the input unclothed human bodies. To ensure that the results align with observations and adhere to physical plausibility for clothing shape inference, we propose two physics-inspired loss functions that preserve the intra-frame distances and inter-frame forces of clothing points. Additionally, the stochastic nature of the denoising process allows for the generation of diverse and plausible clothing shapes. Experiments show that our approach outperforms state-of-the-art methods in chamfer distance and visual effects, particularly for loose clothing such as dresses and skirts. Furthermore, our approach effectively adapts to out-of-domain clothing types and generates realistic clothes dynamics. Shihao Zou, Yuanlu Xu, Nikolaos Sarafianos, Federica Bogo, Tony Tung, Weixin Si, Li Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Tri-Modal Motion Retrieval by Learning a Joint Embedding SpaceabstractInformation retrieval is an ever-evolving and crucial re-search domain. The substantial demand for high-quality human motion data especially in online acquirement has led to a surge in human motion research works. Prior works have mainly concentrated on dual-modality learning, such as text and motion tasks, but three-modality learning has been rarely explored. Intuitively, an extra introduced modality can enrich a model's application scenario, and more importantly, an adequate choice of the extra modality can also act as an intermediary and enhance the alignment between the other two disparate modalities. In this work, we introduce LAVIMO (LAnguage-VIdeo-MOtion alignment), a novel framework for three-modality learning integrating human-centric videos as an additional modality, thereby ef-fectively bridging the gap between text and motion. More-over, our approach leverages a specially designed attention mechanism to foster enhanced alignment and synergistic effects among text, video, and motion modalities. Empirically, our results on the HumanML3D and KIT-ML datasets show that LAVIMO achieves state-of-the-art performance in various motion-related cross-modal retrieval tasks, in-cluding text-to-motion, motion-to-text, video-to-motion and motion-to-video. Our project webpage can be found in https://lavimo2023.github.io/LAVIMO/. Kangning Yin, Shihao Zou, Yuxuan Ge, Zheng Tian 0002 |
CVPR | 2 |
| 2024 | RACon: Retrieval-Augmented Simulated Character Locomotion ControlabstractIn computer animation, driving a simulated character with lifelike motion is challenging. Current generative models, though able to generalize to diverse motions, often pose challenges to the responsiveness of end-user control. To address these issues, we introduce RACon: Retrieval-Augmented Simulated Character Locomotion Control. Our end-to-end hierarchical reinforcement learning method utilizes a retriever and a motion controller. The retriever searches motion experts from a user-specified database in a task-oriented fashion, which boosts the responsiveness to the user’s control. The selected motion experts and the manipulation signal are then transferred to the controller to drive the simulated character. In addition, a retrieval-augmented discriminator is designed to stabilize the training process. Our method surpasses existing techniques in both quality and quantity in locomotion control, as demonstrated in our empirical study. Moreover, by switching extensive databases for retrieval, it can adapt to distinctive motion types at run time. We will release our code upon acceptance. Yuxuan Mu, Shihao Zou, Kangning Yin, Zheng Tian 0002, Li Cheng 0001, Weinan Zhang 0001, Jun Wang 0012 |
ICME | 2 |
| 2024 | PSAN: Prompt Semantic Augmented Network for aspect-based sentiment analysis
Xianying Huang, Shihao Zou |
Expert Syst. Appl. | 3 |
| 2024 | Multimodal Knowledge-enhanced Interactive Network with Mixed Contrastive Learning for Emotion Recognition in Conversation
Xianying Huang, Shihao Zou, Xinyi Gan |
Neurocomputing | 3 |
| 2024 | Improving conversational recommender systems via multi-preference modelling and knowledge-enhanced
Xianying Huang, Jiahao An, Shihao Zou |
Knowl. Based Syst. | 4 |
| 2023 | Multimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in ConversationabstractEmotion Recognition in Conversation (ERC) plays an important role in driving the development of human-machine interaction. Emotions can exist in multiple modalities, and multimodal ERC mainly faces two problems: (1) the noise problem in the cross-modal information fusion process, and (2) the prediction problem of less sample emotion labels that are semantically similar but different categories. To address these issues and fully utilize the features of each modality, we adopted the following strategies: first, deep emotion cues extraction was performed on modalities with strong representation ability, and feature filters were designed as multimodal prompt information for modalities with weak representation ability. Then, we designed a Multimodal Prompt Transformer (MPT) to perform cross-modal information fusion. MPT embeds multimodal fusion information into each attention layer of the Transformer, allowing prompt information to participate in encoding textual features and being fused with multi-level textual information to obtain better multimodal fusion features. Finally, we used the Hybrid Contrastive Learning (HCL) strategy to optimize the model's ability to handle labels with few samples. This strategy uses unsupervised contrastive learning to improve the representation ability of multimodal fusion and supervised contrastive learning to mine the information of labels with few samples. Experimental results show that our proposed model outperforms state-of-the-art models in ERC on two benchmark datasets. Shihao Zou, Xianying Huang |
ACM Multimedia | 1 |
| 2023 | Bi-directional Frame Interpolation for Unsupervised Video Anomaly DetectionabstractAnomaly detection in video surveillance aims to detect anomalous frames whose properties significantly differ from normal patterns. Anomalies in videos can occur in both spatial appearance and temporal motion, making unsupervised video anomaly detection challenging. To tackle this problem, we investigate forward and backward motion continuity between adjacent frames and propose a new video anomaly detection paradigm based on bi-directional frame interpolation. The proposed framework consists of an optical flow estimation network and an interpolation network jointly optimized end-to-end to synthesize a middle frame from its nearest two frames. We further introduce a novel dynamic memory mechanism to balance memory sparsity and normality representation diversity, which attenuates abnormal features in frame interpolation without affecting normal prototypes. In inference, interpolation error and dynamic memory error are fused as anomaly scores. The proposed bi-directional interpolation design improves normal frame synthesis, lowering the false alarm rate of anomaly appearance; meanwhile, the implicit "regular" motion constraint in our optical flow estimation and the novel dynamic memory mechanism play blocking roles in interpolating abnormal frames, increasing the system’s sensitivity to anomalies. Extensive experiments on public benchmarks demonstrates the superiority of the proposed framework over prior arts. Hanqiu Deng, Zhaoxiang Zhang 0003, Shihao Zou |
WACV | 3 |
| 2023 | Snipper: A Spatiotemporal Transformer for Simultaneous Multi-Person 3D Pose Estimation Tracking and Forecasting on a Video SnippetabstractMulti-person pose understanding from RGB videos involves three complex tasks: pose estimation, tracking and motion forecasting. Intuitively, accurate multi-person pose estimation facilitates robust tracking, and robust tracking builds crucial history for correct motion forecasting. Most existing works either focus on a single task or employ multi-stage approaches to solving multiple tasks separately, which tends to make sub-optimal decision at each stage and also fail to exploit correlations among the three tasks. In this paper, we propose Snipper, a unified framework to perform multi-person 3D pose estimation, tracking, and motion forecasting simultaneously in a single stage. We propose an efficient yet powerful deformable attention mechanism to aggregate spatiotemporal information from the video snippet. Building upon this deformable attention, a video transformer is learned to encode the spatiotemporal features from the multi-frame snippet and to decode informative pose features for multi-person pose queries. Finally, these pose queries are regressed to predict multi-person pose trajectories and future motions in a single shot. In the experiments, we show the effectiveness of Snipper on three challenging public datasets where our generic model rivals specialized state-of-art baselines for pose estimation, tracking, and forecasting. Code is available athttps://github.com/JimmyZou/Snipper. Shihao Zou, Yuanlu Xu, Chao Li 0021, Lingni Ma, Li Cheng 0001, Minh Vo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Human Pose and Shape Estimation From Single Polarization ImagesabstractThis paper focuses on a new problem of estimating human pose and shape from single polarization images. Polarization camera is known to be able to capture the polarization of reflected lights that preserves rich geometric cues of an object surface. Inspired by the recent applications in surface normal reconstruction from polarization images, in this paper, we attempt to estimate human pose and shape from single polarization images by leveraging the polarization-induced geometric cues. A dedicated two-stage pipeline is proposed: given a single polarization image, stage one (Polar2Normal) focuses on the fine detailed human body surface normal estimation; stage two (Polar2Shape) then reconstructs clothed human shape from the polarization image and the estimated surface normal. To empirically validate our approach, a dedicated dataset (PHSPD) is constructed, consisting of over 500 K frames with accurate pose and parametric shape annotations. Empirical evaluations on this real-world dataset as well as a synthetic dataset, SURREAL, demonstrate the effectiveness of our approach. It suggests polarization camera as a promising alternative to the more conventional RGB camera for human pose and shape estimation. Shihao Zou, Xinxin Zuo, Sen Wang 0003, Yiming Qian, Chuan Guo 0002, Li Cheng 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Generating Diverse and Natural 3D Human Motions from TextabstractAutomated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using temporal variational autoen-coder to synthesize a diverse set of human motions of the sampled lengths. Instead of directly engaging with pose sequences, we propose motion snippet code as our internal motion representation, which captures local semantic motion contexts and is empirically shown to facilitate the generation of plausible motions faithful to the input text. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions. Chuan Guo 0002, Shihao Zou, Xinxin Zuo, Sen Wang 0003, Wei Ji 0011, Li Cheng 0001 |
CVPR | 2 |
| 2022 | Action2video: Generating Videos of Human 3D Actions
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Xinshuang Liu, Shihao Zou, Minglun Gong, Li Cheng 0001 |
Int. J. Comput. Vis. | 5 |
| 2022 | Improving multimodal fusion with Main Modal Transformer for emotion recognition in conversation
Shihao Zou, Xianying Huang, Hankai Liu |
Knowl. Based Syst. | 1 |
| 2022 | Speckle Noise Spectrum at Near-Nadir Incidence Angles for a Time-Varying Sea SurfaceabstractSpeckle noise is inherent to radar measurements. For applications which need both a high temporal and high spatial resolution, a classical method for the reduction of the speckle noise by filtering the backscattered signal may not be sufficient. In particular, when radar observations are used to estimate ocean wave spectra from relative fluctuations of the radar signal within a given footprint, a method must be implemented to correct for the speckle effect in the Fourier domain (i.e., density spectrum). A theoretical background to model the speckle density spectrum for a radar with near-nadir incidences was proposed by Jackson in 1981 but it is based on a stationary sea surface assumption and ignores the variation of the main factor in the four-frequency moment near the origin. In this article, we revisit this theoretical background to extend this model to a time-varying sea surface and alleviate some assumptions on the Fresnel phase formulation. The results from the model applied in the configuration of an airborne system indicate that not only the displacement of the radar but also the dynamic properties of the sea surfaces have a significant effect on the speckle noise spectrum in certain directions of observations. The effects depend on the radar look direction in azimuth, and on sea surface conditions (wind speed, wind direction with respect to the aircraft route, surface wave spectrum). This new model is validated against observations of the airborne near-nadir incidence scatterometer—Ku-band Radar for Observation of Surfaces (KuROS). We show in particular that the errors between the experimental estimation of the omni-directional speckle noise spectrum from KuROS and the prediction by our model are below 10%. Danièle Hauser, Shihao Zou, Jianyang Si, Eva Le Merle |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | EventHPE: Event-based 3D Human Pose and Shape EstimationabstractEvent camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals are best at capturing local motions. This leads us to propose a two-stage deep learning approach, called EventHPE. The first-stage, FlowNet, is trained by unsupervised learning to infer optical flow from events. Both events and optical flow are closely related to human body dynamics, which are fed as input to the ShapeNet in the second stage, to estimate 3D human shapes. To mitigate the discrepancy between image-based flow (optical flow) and shape-based flow (vertices movement of human body shape), a novel flow coherence loss is introduced by exploiting the fact that both flows are originated from the identical human motion. An in-house event-based 3D human dataset is curated that comes with 3D pose and shape annotations, which is by far the largest one to our knowledge. Empirical evaluations on DHP19 dataset and our in-house dataset demonstrate the effectiveness of our approach. Shihao Zou, Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Pengyu Wang 0007, Xiaoqin Hu, Shoushun Chen, Minglun Gong, Li Cheng 0001 |
ICCV | 1 |
| 2020 | Learning to Communicate Implicitly by ActionsabstractIn situations where explicit communication is limited, human collaborators act by learning to: (i) infer meaning behind their partner's actions, and (ii) convey private information about the state to their partner implicitly through actions. The first component of this learning process has been well-studied in multi-agent systems, whereas the second — which is equally crucial for successful collaboration — has not. To mimic both components mentioned above, thereby completing the learning process, we introduce a novel algorithm: Policy Belief Learning (PBL). PBL uses a belief module to model the other agent's private information and a policy module to form a distribution over actions informed by the belief module. Furthermore, to encourage communication by actions, we propose a novel auxiliary reward which incentivizes one agent to help its partner to make correct inferences about its private information. The auxiliary reward for communication is integrated into the learning of the policy module. We evaluate our approach on a set of environments including a matrix game, particle environment and the non-competitive bidding problem from contract bridge. We show empirically that this auxiliary reward is effective and easy to generalize. These results demonstrate that our PBL algorithm can produce strong pairs of agents in collaborative games where explicit communication is disabled. Zheng Tian 0002, Shihao Zou, Ian Davies, Tim Warr, Lisheng Wu, Haitham Bou-Ammar, Jun Wang 0012 |
AAAI | 2 |
| 2020 | 3D Human Shape Reconstruction from a Polarization Image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang 0003, Chi Xu 0002, Minglun Gong, Li Cheng 0001 |
ECCV (14) | 1 |
| 2020 | Action2Motion: Conditioned Generation of 3D Human MotionsabstractAction recognition is a relatively established task, where given an input sequence of human motion, the goal is to predict its action category. This paper, on the other hand, considers a relatively new problem, which could be thought of as an inverse of action recognition: given a prescribed action type, we aim to generate plausible human motion sequences in 3D. Importantly, the set of generated motions are expected to maintain its diversity to be able to explore the entire action-conditioned motion space; meanwhile, each sampled sequence faithfully resembles a natural human body articulation dynamics. Motivated by these objectives, we follow the physics law of human kinematics by adopting the Lie Algebra theory to represent the natural human motions; we also propose a temporal Variational Auto-Encoder (VAE) that encourages a diverse sampling of the motion space. A new 3D human motion dataset, HumanAct12, is also constructed. Empirical experiments over three distinct human motion datasets (including ours) demonstrate the effectiveness of our approach. Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, Li Cheng 0001 |
ACM Multimedia | 4 |
| 2019 | MarlRank: Multi-agent Reinforced Learning to RankabstractWhen estimating the relevancy between a query and a document, ranking models largely neglect the mutual information among documents. A common wisdom is that if two documents are similar in terms of the same query, they are more likely to have similar relevance score. To mitigate this problem, in this paper, we propose a multi-agent reinforced ranking model, named MarlRank. In particular, by considering each document as an agent, we formulate the ranking process as a multi-agent Markov Decision Process (MDP), where the mutual interactions among documents are incorporated in the ranking process. To compute the ranking list, each document predicts its relevance to a query considering not only its own query-document features but also its similar documents' features and actions. By defining reward as a function of NDCG, we can optimize our model directly on the ranking performance measure. Our experimental results on two LETOR benchmark datasets show that our model has significant performance gains over the state-of-art baselines. We also find that the NDCG shows an overall increasing trend along with the step of interactions, which demonstrates that the mutual information among documents helps improve the ranking performance. Shihao Zou, Mohammad Akbari 0001, Jun Wang 0012, Peng Zhang 0002 |
CIKM | 1 |
| 2019 | A Regularized Opponent Model with Maximum Entropy ObjectiveabstractIn a single-agent setting, reinforcement learning (RL) tasks can be cast into an inference problem by introducing a binary random variable o, which stands for the "optimality". In this paper, we redefine the binary random variable o in multi-agent setting and formalize multi-agent reinforcement learning (MARL) as probabilistic inference. We derive a variational lower bound of the likelihood of achieving the optimality and name it as Regularized Opponent Model with Maximum Entropy Objective (ROMMEO). From ROMMEO, we present a novel perspective on opponent modeling and show how it can improve the performance of training agents theoretically and empirically in cooperative games. To optimize ROMMEO, we first introduce a tabular Q-iteration method ROMMEO-Q with proof of convergence. We extend the exact algorithm to complex environments by proposing an approximate version, ROMMEO-AC. We evaluate these two algorithms on the challenging iterated matrix game and differential game respectively and show that they can outperform strong MARL baselines. Zheng Tian 0002, Ying Wen 0001, Zhichen Gong, Faiz Punakkath, Shihao Zou, Jun Wang 0012 |
IJCAI | 5 |