EDBT 2026 Demo / reviewers in the wild / expert
Wei Liang 0008
dblp:22/849-8
· DBLP profile ↗
77ranked-venue papers
9as first author
41since 2021 · last 2026
0000-0002-7539-3107ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 3 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 4 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorSystems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RadarMP: Motion Perception for 4D mmWave Radar in Autonomous DrivingabstractAccurate 3D scene motion perception significantly enhances the safety and reliability of an autonomous driving system. Benefiting from its all-weather operational capability and unique perceptual properties, 4D mmWave radar has emerged as an essential component in advanced autonomous driving. However, sparse and noisy radar points often lead to imprecise motion perception, leaving autonomous vehicles with limited sensing capabilities when optical sensors degrade under adverse weather conditions. In this paper, we propose RadarMP, a novel method for precise 3D scene motion perception using low-level radar echo signals from two consecutive frames. Unlike existing methods that separate radar target detection and motion estimation, RadarMP jointly models both tasks in a unified architecture, enabling consistent radar point cloud generation and pointwise 3D scene flow prediction. Tailored to radar characteristics, we design specialized self-supervised loss functions guided by Doppler shifts and echo intensity, effectively supervising spatial and motion consistency without explicit annotations. Extensive experiments on the public dataset demonstrate that RadarMP achieves reliable motion perception across diverse weather and illumination conditions, outperforming radar-based decoupled motion perception pipelines and enhancing perception capabilities for full-scenario autonomous driving systems. Ruiqi Cheng, Huijun Di, Wei Liang 0008 |
AAAI | 5 |
| 2026 | Understanding Human-Centric Dynamics Through Need-Driven Interaction Modeling
Zimo Zhai, Manjie Xu, Wei Liang 0008 |
ICPR (6) | 3 |
| 2026 | DCRR++: Unsupervised Reflection Removal and Novel View Synthesis via Dual-Pixel Guided 3D Gaussian Splatting
Kailong Yu, Mina Han, Liyuan Pan, Liu Liu 0009, Miaomiao Liu 0001, Wei Liang 0008 |
Int. J. Comput. Vis. | 6 |
| 2026 | SFGFusion: Surface fitting guided 3D object detection with 4D radar and camera fusion
Xiaozhi Li, Huijun Di, Wei Liang 0008 |
Pattern Recognit. | 5 |
| 2026 | Audio-Visual LLM for Augmenting Accessibility of 360° VideoabstractCreators of 360° videos utilize affluent non-speech sounds for providing immersive experiences. The sound accessibility of such videos is essential for viewers, especially for d/Deaf and hard-of-hearing (DHH) people. In this paper, we proposeAVLLM-360, a multimodal framework using Large Language Models (LLMs) for understanding panorama video content and providing sound descriptions, which goes beyond the simple recognition of sound types.AVLLM-360integrates both visual and auditory information and bootstraps the cross-modal training from the pre-trained LLM. We also implemented a mixed-media interface that allows users to visualize the generated results hierarchically, enabling personalized customization of sound description generation when watching 360° videos. We conducted extensive experiments to evaluateAVLLM-360’s ability across a range of video understanding tasks. We also conducted qualitative studies with 12 DHH participants, evaluating the effectiveness of ourAVLLM-360using 24 360° videos (covering different genres). Qingyun Deng, Wei Liang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | What You See Is What You Wear: Crafting Garments for Diverse Avatars with Consistent Wearing EffectsabstractIn immersive VR applications, avatars shape users' identities and social presence. Personalization requires generating realistic outfits for avatars, often specified through a single reference image, while supporting seamless editing, adaptation to diverse avatars, and efficient rendering. Achieving these goals is challenging because VR avatars, even humanoid ones, exhibit substantial variations in body shapes and topologies. This inherent diversity makes it difficult to collect sufficient paired data, impeding the evolution of generalizable end-to-end image-to-garment models. To this end, we propose Tailor, a two-stage framework that dresses 3D humanoid avatars from a single reference image while preserving the wearing effects observed in the image. In the first stage, Tailor leverages a structured garment representation based on sewing patterns, enabling the network to predict garments in a low-dimensional, interpretable, and topology-independent space. In the second stage, Tailor performs instance-specific optimization to adapt the predicted sewing pattern to the avatar, ensuring consistent wearing effects across varying avatars. Furthermore, this framework also enables seamless garment editing, on-the-fly adaptation, and real-time rendering, making it particularly suitable for large-scale VR environments. Extensive experiments demonstrate that Tailor achieves results comparable to professional manual designs and produces garments that are both visually appealing and better aligned with reference styles than those generated by naive pattern-scaling baselines, as validated through human perceptual studies. Wei Liang 0008, Bing Ning |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | FloNa: Floor Plan Guided Embodied Visual NavigationabstractHumans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior knowledge, leading to limited efficiency and accuracy. To eliminate this gap, we introduce a novel navigation task: Floor Plan Visual Navigation (FloNa), the first attempt to incorporate floor plans into embodied visual navigation. While the floor plan offers significant advantages, two key challenges emerge: (1) handling the spatial inconsistency between the floor plan and the actual scene layout for collision-free navigation, and (2) aligning observed images with the floor plan sketch despite their distinct modalities. To address these challenges, we propose FloDiff, a novel diffusion policy framework incorporating a localization module to facilitate alignment between the current observation and the floor plan. We further collect 20k navigation episodes across 117 scenes in the iGibson simulator to support the training and evaluation. Extensive experiments demonstrate the effectiveness and efficiency of our framework in unfamiliar scenes using floor plan knowledge. Weiqi Huang, Wei Liang 0008, Huijun Di |
AAAI | 4 |
| 2025 | METASCENES: Towards Automated Replica Creation for Real-world 3D ScansabstractEmbodied AI (EAI) research requires high-quality, diverse 3D scenes to effectively support skill acquisition, sim-to-real transfer, and generalization. Achieving these quality standards, however, necessitates the precise replication of real-world object diversity. Existing datasets demon strate that this process heavily relies on artist-driven designs, which demand substantial human effort and present significant scalability challenges. To scalably produce realistic and interactive 3D scenes, we first present MetaScenes, a large-scale simulatable 3D scene dataset constructed from real-world scans, which includes 15366 objects spanning 831 fine-grained categories. Then, we introduce SCAN2SIM, a robust multi-modal alignment model, which enables the automated, high-quality replacement of assets, thereby eliminating the reliance on artist-driven designs for scaling 3D scenes. We further propose two benchmarks to evaluate MetaScenes: a detailed scene synthesis task focused on small item layouts for robotic manipulation and a domain transfer task in vision-and-language navigation (VLN) to validate cross-domain transfer. Results confirm MetaScenes ’s potential to enhance EAI by supporting more generalizable agent learning and sim-to-real applications, introducing new possibilities for EAI research. Huangyue Yu, Baoxiong Jia, Yixin Chen 0003, Yandan Yang, Puhao Li, Rongpeng Su, Qing Li 0003, Wei Liang 0008, Song-Chun Zhu, Tengyu Liu, Siyuan Huang 0001 |
CVPR | 9 |
| 2025 | Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied NavigationabstractEmbodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstruction, such as meshes and point clouds, but lack the ability to actively perceive and explore their environment. To address this limitation, we introduce \underline{\textbf{M}}ove \underline{\textbf{t}}o \underline{\textbf{U}}nderstand (\textbf{\model}), a unified framework that integrates active perception with \underline{\textbf{3D}} vision-language learning, enabling embodied agents to effectively explore and understand their environment. This is achieved by three key innovations: 1) Online query-based representation learning, enabling direct spatial memory construction from RGB-D frames, eliminating the need for explicit 3D reconstruction. 2) A unified objective for grounding and exploring, which represents unexplored locations as frontier queries and jointly optimizes object grounding and frontier selection. 3) End-to-end trajectory learning that combines \textbf{V}ision-\textbf{L}anguage-\textbf{E}xploration pre-training over a million diverse trajectories collected from both simulated and real-world RGB-D sequences. Extensive evaluations across various embodied navigation and question-answering benchmarks show that MTU3D outperforms state-of-the-art reinforcement learning and modular navigation approaches by 14\%, 23\%, 9\%, and 2\% in success rate on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA, respectively. \model's versatility enables navigation using diverse input modalities, including categories, language descriptions, and reference images. These findings highlight the importance of bridging visual grounding and exploration for embodied intelligence. Xilin Wang, Zhuofan Zhang, Xiaojian Ma 0001, Yixin Chen 0003, Baoxiong Jia, Wei Liang 0008, Zhidong Deng, Siyuan Huang 0001, Qing Li 0003 |
ICCV | 8 |
| 2025 | Env-Mani: Quadrupedal Robot Loco-Manipulation with Environment-in-the-LoopabstractDogs can climb onto tables using their front legs for support, enabling them to retrieve objects and significantly expand their workspace by leveraging the external environment. However, the ability of quadrupedal robots to perform similar skills remains largely unexplored. In this work, we introduce a unified, learning-based loco-manipulation framework for quadrupedal robots, allowing them to utilize the external environment as support to extend their workspace and enhance their manipulation capabilities. Specifically, our method proposes a unified policy that takes limited onboard sensors and proprioception as input, generating whole-body actions that enable the robot to manipulate objects. To guide the policy learning for environment-in-the-loop manipulation, we design a set of rewards that address challenges such as imprecise perception and center-of-mass shifts. Additionally, we employ curriculum learning to train both teacher and student policies, ensuring effective skill transfer in complex tasks. We train the policy in simulation and conduct extensive experiments, demonstrating that our approach allows robots to manipulate previously inaccessible objects, opening up new possibilities for enhancing quadrupedal robot capabilities without the need for hardware modifications or additional costs. The project page is available at https://sites.google.com/view/env-mani. Wei Liang 0008 |
IROS | 3 |
| 2025 | Enhanced Dual-Pixel Image Reflection Removal via Gaussian SplattingabstractImage de-reflection is a critical task in computer vision. Existing methods for de-reflection using monocular cameras face challenges due to the lack of depth cues to separate the transmission and reflection layers, particularly under strong illumination or multi-layer reflection scenarios. Although recent advances, such as 3D Gaussian Splatting (3DGS), utilize novel view-synthesis capabilities to separate transmitted and reflected layers, they still encounter difficulties in practice with monocular images. In this paper, we simplify the de-reflection task by combining dual-pixel (DP) technology with 3DGS, forming the first unsupervised de-reflection framework. Specifically, we propose the Dual-View Coordinated Reflection Removal (DCRR) Framework, which integrates depth cues from DP sensors with the rendering capabilities of 3DGS. The DCRR utilizes a dual-view approach that estimates the image transmission layer and opacity via differentiable rasterization with 3DGS and reconstructs the reflection layer through a lightweight multi-layer perceptron. We then present the Dual-Pixel-Driven Reflection Gaussian Pruning (DPRGP) to refine the separation process. By using the physical properties of DP sensors, DCRR achieves significant accuracy improvements in complex reflection scenarios. A real-world DP-based dataset that includes paired reflection/reflection-free images has been collected. Extensive experiments demonstrate our competitive performance compared to state-of-the-art de-reflection approaches. Kailong Yu, Liyuan Pan, Liu Liu 0009, Wei Liang 0008 |
ACM Multimedia | 4 |
| 2025 | Heterogeneous Adversarial Play in Interactive EnvironmentsabstractSelf-play constitutes a fundamental paradigm for autonomous skill acquisition, whereby agents iteratively enhance their capabilities through self-directed environmental exploration. Conventional self-play frameworks exploit agent symmetry within zero-sum competitive settings, yet this approach proves inadequate for open-ended learning scenarios characterized by inherent asymmetry. Human pedagogical systems exemplify asymmetric instructional frameworks wherein educators systematically construct challenges calibrated to individual learners' developmental trajectories. The principal challenge resides in operationalizing these asymmetric, adaptive pedagogical mechanisms within artificial systems capable of autonomously synthesizing appropriate curricula without predetermined task hierarchies. Here we present Heterogeneous Adversarial Play (HAP), an adversarial Automatic Curriculum Learning framework that formalizes teacher-student interactions as a minimax optimization wherein task-generating instructor and problem-solving learner co-evolve through adversarial dynamics. In contrast to prevailing automatic curriculum learning methodologies that employ static curricula or unidirectional task selection mechanisms, HAP establishes a bidirectional feedback system wherein instructors continuously recalibrate task complexity in response to real-time learner performance metrics. Experimental validation across multi-task learning domains demonstrates that our framework achieves performance parity with SOTA baselines while generating curricula that enhance learning efficacy in both artificial agents and human subjects. Manjie Xu, Jiayu Zhan, Wei Liang 0008, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 4 |
| 2025 | Emergency Evacuation Map Guided Navigation via Topological Alignment and VLM Reasoning
Canzhi Chen, Weiqi Huang, Huijun Di, Wei Liang 0008 |
PRCV (6) | 6 |
| 2025 | Let storytelling tell vivid stories: A multi-modal-agent-based unified storytelling framework
Jiji Tang, Chuanqi Zang, Mingtao Pei, Wei Liang 0008, Zeng Zhao |
Neurocomputing | 5 |
| 2025 | R2G: Reasoning to ground in 3D scenes
Wei Liang 0008 |
Pattern Recognit. | 3 |
| 2025 | X's Day: Personality-Driven Virtual Human Behavior GenerationabstractDeveloping convincing and realistic virtual human behavior is essential for enhancing user experiences in virtual reality (VR) and augmented reality (AR) settings. This paper introduces a novel task focused on generating long-term behaviors for virtual agents, guided by specific personality traits and contextual elements within 3D environments. We present a comprehensive framework capable of autonomously producing daily activities autoregressively. By modeling the intricate connections between personality characteristics and observable activities, we establish a hierarchical structure of Needs, Task, and Activity levels. Integrating a Behavior Planner and a World State module allows for the dynamic sampling of behaviors using large language models (LLMs), ensuring that generated activities remain relevant and responsive to environmental changes. Extensive experiments validate the effectiveness and adaptability of our approach across diverse scenarios. This research makes a significant contribution to the field by establishing a new paradigm for personalized and context-aware interactions with virtual humans, ultimately enhancing user engagement in immersive applications. Our project website is at: https://behavior.agent-x.cn/. Wei Liang 0008, Yizhuo Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | LiteAT: A Data-Lightweight and User-Adaptive VR Telepresence System for Remote EducationabstractIn educators' ongoing pursuit of enriching remote education, Virtual Reality (VR)-based telepresence has shown significant promise due to its immersive and interactive nature. Existing approaches often rely on point cloud or NeRF-based techniques to deliver realistic representations of teachers and classrooms to remote students. However, achieving low latency is non-trivial, and maintaining high-fidelity rendering under such constraints poses an even greater challenge. This paper introduces LiteAT, a data-lightweight and user-adaptive VR telepresence system, to enable real-time, immersive learning experiences. LiteAT employs a Gaussian Splatting-based reconstruction pipeline that integrates an SMPL-X-driven dynamic human model with a static classroom, supporting lightweight data transmission and high-quality rendering. To enable efficient and personalized exploration in the virtual classroom, we propose a user-adaptive viewpoint recommendation framework that dynamically suggests high-quality viewpoints tailored to user preferences. Candidate viewpoints are evaluated based on multiple visual quality factors and are continuously optimized based on recent user behavior and scene dynamics. Quantitative experiments and user studies validate the effectiveness of LiteAT across multiple evaluation metrics. LiteAT establishes a versatile and scalable foundation for immersive telepresence, potentially supporting real-time scenarios such as procedural teaching, multimodal instruction, and collaborative learning. Yuxin Shen, Wei Liang 0008, Jianzhu Ma |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceabstractDespite significant advancements in text-to-motion syn-thesis, generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models capable of jointly modeling natural language, 3D scenes, and human motion, and (ii) the generative models' in-tensive data requirements contrasted with the scarcity of comprehensive, high-quality, language-scene-motion datasets. To tackle these issues, we introduce a novel two-stage frame-work that employs scene affordance as an intermediate representation, effectively linking 3D scene grounding and conditional motion generation. Our framework comprises an Affordance Diffusion Model (ADM) for predicting ex-plicit affordance map and an Affordance-to-Motion Diffusion Model (AMDM) for generating plausible human motions. By leveraging scene affordance maps, our method overcomes the difficulty in generating human motion under multimodal condition signals, especially when training with limited data lacking extensive language-scene-motion pairs. Our exten-sive experiments demonstrate that our approach consistently outperforms all baselines on established benchmarks, in-cluding HumanML3D and HUMANISE. Additionally, we validate our model's exceptional generalization capabilities on a specially curated evaluation set featuring previously unseen descriptions and scenes. Yixin Chen 0003, Baoxiong Jia, Puhao Li, Jinlu Zhang 0001, Jingze Zhang, Tengyu Liu, Yixin Zhu 0001, Wei Liang 0008, Siyuan Huang 0001 |
CVPR | 9 |
| 2024 | Language-driven All-in-one Adverse Weather RemovalabstractAll-in-one (AiO) frameworks restore various adverse weather degradations with a single set of networks jointly. To handle various weather conditions, an AiO framework is expected to adaptively learn weather-specific knowledge for different degradations and shared knowledge for common patterns. However, existing methods: 1) rely on extra su-pervision signals, which are usually unknown in real-world applications; 2) employ fixed network structures, which re-strict the diversity of weather-specific knowledge. In this paper, we propose a Language-driven Restoration frame-work (LDR) to alleviate the aforementioned issues. First, we leverage the power of pre-trained vision-language (PVL) models to enrich the diversity of weather-specific knowl-edge by reasoning about the occurrence, type, and severity of degradation, generating description-based degradation priors. Then, with the guidance of degradation prior, we sparsely select restoration experts from a candidate list dy-namically based on a Mixture-of-Experts (MoE) structure. This enables us to adaptively learn the weather-specific and shared knowledge to handle various weather conditions (e.g., unknown or mixed weather). Experiments on exten-sive restoration scenarios show our superior performance. Hao Yang 0040, Liyuan Pan, Yan Yang 0011, Wei Liang 0008 |
CVPR | 4 |
| 2024 | Visual Loop Closure Detection with Thorough Temporal and Spatial Context ExploitationabstractDespite advancements in visual Simultaneous Localization and Mapping (SLAM), prevailing visual Loop Closure Detection (LCD) methods primarily rely on computationally intensive image similarity comparisons, neglecting temporal-spatial context during long-term exploration. To address this issue, we propose TOSA, a novel visual LCD algorithm harnessing TempOral and SpAtial context for efficient LCD. Specifically, as the agent explores through time, our approach recurrently updates a latent feature incorporating historical information via a Long Short-Term Memory (LSTM) module. Upon receiving a query frame, TOSA seamlessly fuses the latent feature with the query feature to predict the candidates’ distribution, thus averting intensive similarity computation. Additionally, TOSA integrates a temporal-spatial convolution for candidate refinement by thoroughly exploiting the temporal consistency and spatial correlation to enhance selected candidates, further boosting the performance. Extensive experiments across four standard datasets showcase the superiority of our method over existing state-of-the-art techniques, demonstrating the effectiveness of utilizing rich temporal-spatial contexts. Huijun Di, Wei Liang 0008 |
IROS | 5 |
| 2024 | Mastering Scene Rearrangement with Expert-Assisted Curriculum Learning and Adaptive Trade-Off Tree-SearchabstractScene Rearrangement Planning (SRP) has recently emerged as a crucial interior scene task; however, current approaches still face two primary issues. First, prior works define the action space of SRP using handcrafted coarse-grained actions, which are inflexible for scene arrangement transition and impractical for real-world deployment. Secondly, the scarcity of realistic indoor scene rearrangement data hinders popular data-hungry learning approaches and quantitative evaluation. To tackle these issues, we propose a fine-grained action space definition and curate a large-scale scene rearrangement dataset to facilitate the training of learning approaches and comprehensive benchmarking. Building upon this dataset, we introduce a novel framework, PLATO, designed for efficient agent training and inference. Our approach features an exPert-assisted curriculum Learning (PL) paradigm that possesses a Behavior Cloning (BC) and an offline Reinforcement Learning (RL) curriculum for agent training, along with an advanced tree-search-based planner enhanced by an Adaptive Trade-Off (ATO) strategy to improve expert agent performance further. We demonstrate the superior performance of our method over baseline agents through extensive experiments and provide a detailed analysis to elucidate its rationale. Our project website can be accessed at plato.github.io. Hanqing Wang 0001, Wei Liang 0008 |
IROS | 3 |
| 2024 | Context-Aware Head-and-Eye Motion Generation with Diffusion ModelabstractIn humanity’s ongoing quest to craft natural and realistic avatars within virtual environments, the generation of authentic eye gaze behaviors stands paramount. Eye gaze not only serves as a primary non-verbal communication cue, but it also reflects cognitive processes, intent, and attentiveness, making it a crucial element in ensuring immersive interactions. However, automatically generating these intricate gaze behaviors presents significant challenges. Traditional methods can be both time-consuming and lack the precision to align gaze behaviors with the intricate nuances of the environment in which the avatar resides. To overcome these challenges, we introduce a novel two-stage approach to generate context-aware head-and-eye motions across diverse scenes. By harnessing the capabilities of advanced diffusion models, our approach adeptly produces contextually appropriate eye gaze points, further leading to the generation of natural head-and-eye movements. Utilizing Head-Mounted Display (HMD) eye-tracking technology, we also present a comprehensive dataset, which captures human eye gaze behaviors in tandem with associated scene features. We show that our approach consistently delivers intuitive and lifelike head-and-eye motions and demonstrates superior performance in terms of motion fluidity, alignment with contextual cues, and overall user satisfaction. Yuxin Shen, Manjie Xu, Wei Liang 0008 |
VR | 3 |
| 2024 | Self-trained multi-cues model for video anomaly detection
Zhengang Nie, Wei Liang 0008, Mingtao Pei |
Multim. Tools Appl. | 3 |
| 2024 | Token labeling-guided multi-scale medical image classification
Fangyuan Yan, Wei Liang 0008, Mingtao Pei |
Pattern Recognit. Lett. | 3 |
| 2023 | Diffusion-based Generation, Optimization, and Planning in 3D ScenesabstractWe introduce the SceneDiffuser, a conditional generative model for 3D scene understanding. SceneDiffuser provides a unified model for solving scene-conditioned generation, optimization, and planning. In contrast to prior work, SceneDiffuser is intrinsically scene-aware, physics-based, and goal-oriented. With an iterative sampling strategy, SceneDiffuser jointly formulates the scene-aware generation, physics-based optimization, and goal-oriented planning via a diffusion-based denoising process in a fully differentiable fashion. Such a design alleviates the discrepancies among different modules and the posterior collapse of previous scene-conditioned generative models. We evaluate the SceneDiffuser on various 3D scene understanding tasks, including human pose and motion generation, dexterous grasp generation, path planning for 3D navigation, and motion planning for robot arms. The results show significant improvements compared with previous models, demonstrating the tremendous potential of the SceneDiffuser for the broad community of 3D scene understanding. Siyuan Huang 0001, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu 0001, Wei Liang 0008, Song-Chun Zhu |
CVPR | 7 |
| 2023 | Discovering the Real Association: Multimodal Causal Reasoning in Video Question AnsweringabstractVideo Question Answering (VideoQA) is challenging as it requires capturing accurate correlations between modalities from redundant information. Recent methods focus on the explicit challenges of the task, e.g. multimodal feature extraction, video-text alignment and fusion. Their frameworks reason the answer relying on statistical evidence causes, which ignores potential bias in the multimodal data. In our work, we investigate relational structure from a causal representation perspective on multimodal data and propose a novel inference framework. For visual data, question-irrelevant objects may establish simple matching associations with the answer. For textual data, the model prefers the local phrase semantics which may deviate from the global semantics in long sentences. Therefore, to enhance the generalization of the model, we discover the real association by explicitly capturing visual features that are causally related to the question semantics and weakening the impact of local language semantics on question answering. The experimental results on two large causal VideoQA datasets verify that our proposed framework 1) improves the accuracy of the existing VideoQA backbone, 2) demonstrates robustness on complex scenes and questions. The code will be released at https://github.com/Chuanqi-Zang/Discovering-the-Real-Association. Chuanqi Zang, Hanqing Wang 0001, Mingtao Pei, Wei Liang 0008 |
CVPR | 4 |
| 2023 | Dreamwalker: Mental Planning for Continuous Vision-Language NavigationabstractVLN-CE is a recently released embodied task, where AI agents need to navigate a freely traversable environment to reach a distant target location, given language instructions. It poses great challenges due to the huge space of possible strategies. Driven by the belief that the ability to anticipate the consequences of future actions is crucial for the emergence of intelligent and interpretable planning behavior, we propose Dreamwalker — a world model based VLN-CE agent. The world model is built to summarize the visual, topological, and dynamic properties of the complicated continuous environment into a discrete, structured, and compact representation. Dreamwalker can simulate and evaluate possible plans entirely in such internal abstract world, before executing costly actions. As opposed to existing model-free VLN-CE agents simply making greedy decisions in the real world, which easily results in shortsighted behaviors, Dreamwalker is able to make strategic planning through large amounts of "mental experiments." Moreover, the imagined future scenarios reflect our agent’s intention, making its decision-making process more transparent. Extensive experiments and ablation studies on VLN-CE dataset confirm the effectiveness of the proposed approach and outline fruitful directions for future work. Hanqing Wang 0001, Wei Liang 0008, Luc Van Gool, Wenguan Wang |
ICCV | 2 |
| 2023 | MEWL: Few-shot multimodal word learning with referential uncertaintyabstractWithout explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning. Despite recent advancements in multimodal learning, a systematic and rigorous evaluation is still missing for human-like word learning in machines. To fill in this gap, we introduce the MachinE Word Learning (MEWL) benchmark to assess how machines learn word meaning in grounded visual scenes. MEWL covers human’s core cognitive toolkits in word learning: cross-situational reasoning, bootstrapping, and pragmatic learning. Specifically, MEWL is a few-shot benchmark suite consisting of nine tasks for probing various word learning capabilities. These tasks are carefully designed to be aligned with the children’s core abilities in word learning and echo the theories in the developmental literature. By evaluating multimodal and unimodal agents’ performance with a comparative analysis of human performance, we notice a sharp divergence in human and machine word learning. We further discuss these differences between humans and machines and call for human-like few-shot word learning in machines. Guangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang 0008, Yujia Peng, Chi Zhang 0017, Yixin Zhu 0001 |
ICML | 4 |
| 2023 | Active Reasoning in an Open-World EnvironmentabstractRecent advances in vision-language learning have achieved notable success on *complete-information* question-answering datasets through the integration of extensive world knowledge. Yet, most models operate *passively*, responding to questions based on pre-stored knowledge. In stark contrast, humans possess the ability to *actively* explore, accumulate, and reason using both newfound and existing information to tackle *incomplete-information* questions. In response to this gap, we introduce **Conan**, an interactive open-world environment devised for the assessment of *active reasoning*. **Conan** facilitates active exploration and promotes multi-round abductive inference, reminiscent of rich, open-world settings like Minecraft. Diverging from previous works that lean primarily on single-round deduction via instruction following, **Conan** compels agents to actively interact with their surroundings, amalgamating new evidence with prior knowledge to elucidate events from incomplete observations. Our analysis on \bench underscores the shortcomings of contemporary state-of-the-art models in active exploration and understanding complex scenarios. Additionally, we explore *Abduction from Deduction*, where agents harness Bayesian rules to recast the challenge of abduction as a deductive process. Through **Conan**, we aim to galvanize advancements in active reasoning and set the stage for the next generation of artificial intelligence agents adept at dynamically engaging in environments. Manjie Xu, Guangyuan Jiang, Wei Liang 0008, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 3 |
| 2023 | Interactive Visual Reasoning under UncertaintyabstractOne of the fundamental cognitive abilities of humans is to quickly resolve uncertainty by generating hypotheses and testing them via active trials. Encountering a novel phenomenon accompanied by ambiguous cause-effect relationships, humans make hypotheses against data, conduct inferences from observation, test their theory via experimentation, and correct the proposition if inconsistency arises. These iterative processes persist until the underlying mechanism becomes clear. In this work, we devise the IVRE (pronounced as "ivory") environment for evaluating artificial agents' reasoning ability under uncertainty. IVRE is an interactive environment featuring rich scenarios centered around Blicket detection. Agents in IVRE are placed into environments with various ambiguous action-effect pairs and asked to determine each object's role. They are encouraged to propose effective and efficient experiments to validate their hypotheses based on observations and actively gather new information. The game ends when all uncertainties are resolved or the maximum number of trials is consumed. By evaluating modern artificial agents in IVRE, we notice a clear failure of today's learning methods compared to humans. Such inefficacy in interactive reasoning ability under uncertainty calls for future research in building human-like intelligence. Manjie Xu, Guangyuan Jiang, Wei Liang 0008, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 3 |
| 2023 | Optimizing Product Placement for Virtual StoresabstractThe recent popularity of consumer-grade virtual reality devices has enabled users to experience immersive shopping in virtual environments. As in a real-world store, the placement of products in a virtual store should appeal to shoppers, which could be time-consuming, tedious, and non-trivial to create manually. Thus, this work introduces a novel approach for automatically optimizing product placement in virtual stores. Our approach considers product exposure and spatial constraints, applying an optimizer to search for optimal product placement solutions. We conducted qualitative scene rationality and quantitative product exposure experiments to validate our approach with users. The results show that the proposed approach can synthesize reasonable product placements and increase product exposures for different virtual stores. Wei Liang 0008, Luhui Wang, Xinzhe Yu, ChangYang Li, Rawan Alghofaili, Yining Lang, Lap-Fai Yu |
VR | 1 |
| 2023 | Active Perception for Visual-Language Navigation
Hanqing Wang 0001, Wenguan Wang, Wei Liang 0008, Steven C. H. Hoi, Jianbing Shen, Luc Van Gool |
Int. J. Comput. Vis. | 3 |
| 2022 | Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language NavigationabstractSince the rise of vision-language navigation (VLN), great progress has been made in instruction following - building a follower to navigate environments under the guidance of instructions. However, far less attention has been paid to the inverse task: instruction generation - learning a speaker to generate grounded descriptions for navigation routes. Existing VLN methods train a speaker independently and often treat it as a data augmentation tool to strengthen the follower, while ignoring rich cross-task relations. Here we describe an approach that learns the two tasks simultaneously and exploits their intrinsic correlations to boost the training of each: the follower judges whether the speaker-created instruction explains the original navigation route correctly, and vice versa. Without the need of aligned instruction-path pairs, such cycle-consistent learning scheme is complementary to task-specific training targets defined on labeled data, and can also be applied over unlabeled paths (sampled without paired instructions). Another agent, called creator is added to generate counterfactual environments. It greatly changes current scenes yet leaves novel items - which are vital for the execution of original instructions - unchanged. Thus more informative training scenes are synthesized and the three agents compose a powerful VLN learning system. Extensive experiments on a standard benchmark show that our approach improves the performance of various follower models and produces accurate navigation instructions. Hanqing Wang 0001, Wei Liang 0008, Jianbing Shen, Luc Van Gool, Wenguan Wang |
CVPR | 2 |
| 2022 | Predicting Human Motion Using Key SubsequencesabstractHuman motion prediction is an important task in computer vision, and has a wide range of applications, such as autonomous driving and human-robot interaction. Usually, human motion tends to repeat itself and follows patterns that are well-represented by a few short key subsequences. Based on the above observations, we propose an attention-based feed-forward network, which is explicitly guided by the key subsequences, for human motion prediction. Specifically, we obtain the key subsequences by clustering, extract motion attention by the similarity between the observed poses and the motion context of corresponding key subsequences, and aggregate the relevant key subsequences by a graph convolutional network to predict human motion. Experimental results on public human motion datasets show that our method achieves better performance over state-of-the-art methods in motion prediction. Mingtao Pei, Wei Liang 0008 |
ICASSP | 3 |
| 2022 | HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesabstractLearning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characters of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics. To fill in the gap, we propose a large-scale and semantic-rich synthetic HSI dataset, denoted as HUMANISE, by aligning the captured human motion sequences with various 3D indoor scenes. We automatically annotate the aligned motions with language descriptions that depict the action and the individual interacting objects; e.g., sit on the armchair near the desk. HUMANIZE thus enables a new generation task, language-conditioned human motion generation in 3D scenes. The proposed task is challenging as it requires joint modeling of the 3D scene, human motion, and natural language. To tackle this task, we present a novel scene-and-language conditioned generative model that can produce 3D human motions of the desirable action interacting with the specified objects. Our experiments demonstrate that our model generates diverse and semantically consistent human motions in 3D scenes. Yixin Chen 0003, Tengyu Liu, Yixin Zhu 0001, Wei Liang 0008, Siyuan Huang 0001 |
NeurIPS | 5 |
| 2022 | Towards Versatile Embodied NavigationabstractWith the emergence of varied visual navigation tasks (e.g., image-/object-/audio-goal and vision-language navigation) that specify the target in different ways, the community has made appealing advances in training specialized agents capable of handling individual navigation tasks well. Given plenty of embodied navigation tasks and task-specific solutions, we address a more fundamental question: can we learn a single powerful agent that masters not one but multiple navigation tasks concurrently? First, we propose VXN, a large-scale 3D dataset that instantiates~four classic navigation tasks in standardized, continuous, and audiovisual-rich environments. Second, we propose Vienna, a versatile embodied navigation agent that simultaneously learns to perform the four navigation tasks with one model. Building upon a full-attentive architecture, Vienna formulates various navigation tasks as a unified, parse-and-query procedure: the target description, augmented with four task embeddings, is comprehensively interpreted into a set of diversified goal vectors, which are refined as the navigation progresses, and used as queries to retrieve supportive context from episodic history for decision making. This enables the reuse of knowledge across navigation tasks with varying input domains/modalities. We empirically demonstrate that, compared with learning each visual navigation task individually, our multitask agent achieves comparable or even better performance with reduced complexity. Hanqing Wang 0001, Wei Liang 0008, Luc Van Gool, Wenguan Wang |
NeurIPS | 2 |
| 2021 | Scene-Aware Behavior Synthesis for Virtual Pets in Mixed RealityabstractVirtual pets are an alternative to real pets, providing a substitute for people with allergies or preparing people for adopting a real pet. Recent advancements in mixed reality pave the way for virtual pets to provide a more natural and seamless experience for users. However, one key challenge is embedding environmental awareness into the virtual pet (e.g., identifying the food bowl’s location) so that they can behave naturally in the real world. Wei Liang 0008, Xinzhe Yu, Rawan Alghofaili, Yining Lang, Lap-Fai Yu |
CHI | 1 |
| 2021 | Toward Automatic Audio Description Generation for Accessible VideosabstractVideo accessibility is essential for people with visual impairments. Audio descriptions describe what is happening on-screen, e.g., physical actions, facial expressions, and scene changes. Generating high-quality audio descriptions requires a lot of manual description generation [50]. To address this accessibility obstacle, we built a system that analyzes the audiovisual contents of a video and generates the audio descriptions. The system consisted of three modules: AD insertion time prediction, AD generation, and AD optimization. We evaluated the quality of our system on five types of videos by conducting qualitative studies with 20 sighted users and 12 users who were blind or visually impaired. Our findings revealed how audio description preferences varied with user types and video types. Based on our study’s analysis, we provided recommendations for the development of future audio description generation technologies. Wei Liang 0008, Haikun Huang, Dingzeyu Li, Lap-Fai Yu |
CHI | 2 |
| 2021 | Structured Scene Memory for Vision-Language NavigationabstractRecently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following linguistic instructions. However, current VLN agents simply store their past experiences/observations as latent states in recurrent networks, failing to capture environment layouts and make long-term planning. To address these limitations, we propose a crucial architecture, called Structured Scene Memory (SSM). It is compartmentalized enough to accurately memorize the percepts during navigation. It also serves as a structured scene representation, which captures and disentangles visual and geometric cues in the environment. SSM has a collect-read controller that adaptively collects information for supporting current decision making and mimics iterative algorithms for long-range reasoning. As SSM provides a complete action space, i.e., all the navigable places on the map, a frontier-exploration based navigation decision making strategy is introduced to enable efficient and global planning. Experiment results on two VLN datasets (i.e., R2R and R4R) show that our method achieves state-of-the-art performance on several metrics. Hanqing Wang 0001, Wenguan Wang, Wei Liang 0008, Caiming Xiong, Jianbing Shen |
CVPR | 3 |
| 2021 | Climaxing VR Character with Scene-Aware Aesthetic Dress SynthesisabstractLike real humans, virtual characters also need to dress up according to different application scenarios so that the virtual character appears professionally, harmoniously, and naturally. However, manual selection is tedious, and the appearances of virtual characters usually lack variety. In this paper, we propose a new problem of synthesizing appropriate dress for a virtual character based on the scenario analysis where he/she shows up. We come up with a pipeline to tackle the scenario-aware dress synthesis problem. Firstly, given a scene, our approach predicts a dress code from the extracted high-level information in the scene, consisting of season, occasion, and scene category. Then our approach tunes the dress details to fit the aesthetic criteria and the virtual character's attributes. An optimization of a cost function implements the tuning process. We carried out experiments to validate the efficacy of the proposed approach. The perceptual study results show the good performance of our approach. Sifan Hou, Bing Ning, Wei Liang 0008 |
VR | 4 |
| 2021 | Work Surface Arrangement Optimization Driven by Human ActivityabstractIn this paper, we aim at guiding people to accomplish a personalized task, work surface organizing, in mixed reality environment, which can also be applied to intelligent robots. Through the cameras mounted in a MR device, e.g., Hololens, we firstly capture a person's daily activities in real scene when he uses the work surface. From such activities, we model the individual behavior habits and apply them to optimize the arrangement of the work surface. A cost function is defined for the optimization, considering general arrangement rules and human habitual behavior. The optimized arrangement is suggested to the user by augmenting the virtual arrangement on the real scene. To evaluate the effectiveness of our approach, we conducted experiments on a variety of scenes. Wei Liang 0008, Bing Ning, Ting Mao |
VR | 2 |
| 2020 | Active Visual Information Gathering for Vision-Language Navigation
Hanqing Wang 0001, Wenguan Wang, Tianmin Shu, Wei Liang 0008, Jianbing Shen |
ECCV (22) | 4 |
| 2020 | Photo Stand-Out: Photography with Virtual CharacterabstractIn this paper, we propose a novel optimization framework to synthesize an aesthetic pose for the virtual character with respect to the presented user's pose. Our approach applies aesthetic evaluation that exploits fully connected neural networks trained on example images. The aesthetic pose of the virtual character is obtained by optimizing a cost function that guides the rotation of each body joint angles. In our experiments, we demonstrate the proposed approach can synthesize poses for virtual characters according to user pose inputs. We also conducted objective and subjective experiments of the synthesized results to validate the efficacy of our approach. Sifan Hou, Bing Ning, Wei Liang 0008 |
ACM Multimedia | 4 |
| 2020 | Scene-Aware Background Music SynthesisabstractIn this paper, we introduce an interactive background music synthesis algorithm guided by visual content. We leverage a cascading strategy to synthesize background music in two stages: Scene Visual Analysis and Background Music Synthesis. First, seeking a deep learning-based solution, we leverage neural networks to analyze the sentiment of the input scene. Second, real-time background music is synthesized by optimizing a cost function that guides the selection and transition of music clips to maximize the emotion consistency between visual and auditory criteria, and music continuity. In our experiments, we demonstrate the proposed approach can synthesize dynamic background music for different types of scenarios. We also conducted quantitative and qualitative analysis on the synthesized results of multiple example scenes to validate the efficacy of our approach. Wei Liang 0008, Wanwan Li, Dingzeyu Li, Lap-Fai Yu |
ACM Multimedia | 2 |
| 2020 | Scene mover: automatic move planning for scene arrangement by deep reinforcement learningabstractWe propose a novel approach for automatically generating a move plan for scene arrangement. 1 Given a scene like an apartment with many furniture objects, to transform its layout into another layout, one would need to determine a collision-free move plan. It could be challenging to design this plan manually because the furniture objects may block the way of each other if not moved properly; and there is a large complex search space of move action sequences that grow exponentially with the number of objects. To tackle this challenge, we propose a learning-based approach to generate a move plan automatically. At the core of our approach is a Monte Carlo tree that encodes possible states of the layout, based on which a search is performed to move a furniture object appropriately in the current layout. We trained a policy neural network embedded with a LSTM module for estimating the best actions to take in the expansion step and simulation step of the Monte Carlo tree search process. Leveraging the power of deep reinforcement learning, the network learned how to make such estimations through millions of trials of moving objects. We demonstrated our approach for moving objects under different scenarios and constraints. We also evaluated our approach on synthetic and real-world layouts, comparing its performance with that of humans and other baseline approaches. Hanqing Wang 0001, Wei Liang 0008, Lap-Fai Yu |
ACM Trans. Graph. | 2 |
| 2019 | 3D Face Synthesis Driven by Personality ImpressionabstractSynthesizing 3D faces that give certain personality impressions is commonly needed in computer games, animations, and virtual world applications for producing realistic virtual characters. In this paper, we propose a novel approach to synthesize 3D faces based on personality impression for creating virtual characters. Our approach consists of two major steps. In the first step, we train classifiers using deep convolutional neural networks on a dataset of images with personality impression annotations, which are capable of predicting the personality impression of a face. In the second step, given a 3D face and a desired personality impression type as user inputs, our approach optimizes the facial details against the trained classifiers, so as to synthesize a face which gives the desired personality impression. We demonstrate our approach for synthesizing 3D faces giving desired personality impressions on a variety of 3D face models. Perceptual studies show that the perceived personality impressions of the synthesized faces agree with the target personality impressions specified for synthesizing the faces. Yining Lang, Wei Liang 0008, Lap-Fai Yu |
AAAI | 2 |
| 2019 | Deep Single-View 3D Object Reconstruction with Visual Hull Embeddingabstract3D object reconstruction is a fundamental task of many robotics and AI problems. With the aid of deep convolutional neural networks (CNNs), 3D object reconstruction has witnessed a significant progress in recent years. However, possibly due to the prohibitively high dimension of the 3D object space, the results from deep CNNs are often prone to missing some shape details. In this paper, we present an approach which aims to preserve more shape details and improve the reconstruction quality. The key idea of our method is to leverage object mask and pose estimation from CNNs to assist the 3D shape learning by constructing a probabilistic singleview visual hull inside of the network. Our method works by first predicting a coarse shape as well as the object pose and silhouette using CNNs, followed by a novel 3D refinement CNN which refines the coarse shapes using the constructed probabilistic visual hulls. Experiment on both synthetic data and real images show that embedding a single-view visual hull for shape refinement can significantly improve the reconstruction quality by recovering more shapes details and improving shape consistency with the input image. Hanqing Wang 0001, Jiaolong Yang, Wei Liang 0008, Xin Tong 0001 |
AAAI | 3 |
| 2019 | Virtual Agent Positioning Driven by Scene Semantics in Mixed RealityabstractWhen a user interacts with a virtual agent via a mixed reality device, such as a Hololens or a Magic Leap headset, it is important to consider the semantics of the real-world scene in positioning the virtual agent, so that it interacts with the user and the objects in the real world naturally. Mixed reality aims to blend the virtual world with the real world seamlessly. In line with this goal, in this paper, we propose a novel approach to use scene semantics to guide the positioning of a virtual agent. Such considerations can avoid unnatural interaction experiences, e.g., interacting with a virtual human floating in the air. To obtain the semantics of a scene, we first reconstruct the 3D model of the scene by using the RGB-D cameras mounted on the mixed reality device (e.g., a Hololens). Then, we employ the Mask R-CNN object detector to detect objects relevant to the interactions within the scene context. To evaluate the positions and orientations for placing a virtual agent in the scene, we define a cost function based on the scene semantics, which comprises a visibility term and a spatial term. We then apply a Markov chain Monte Carlo optimization technique to search for an optimized solution for placing the virtual agent. We carried out user study experiments to evaluate the results generated by our approach. The results show that our approach achieved a higher user evaluation score than that of the alternative approaches. Yining Lang, Wei Liang 0008, Lap-Fai Yu |
VR | 2 |
| 2019 | A deep Coarse-to-Fine network for head pose estimation from synthetic data
Wei Liang 0008, Jianbing Shen, Yunde Jia, Lap-Fai Yu |
Pattern Recognit. | 2 |
| 2019 | Comic-guided speech synthesisabstractWe introduce a novel approach for synthesizing realistic speeches for comics. Using a comic page as input, our approach synthesizes speeches for each comic character following the reading flow. It adopts a cascading strategy to synthesize speeches in two stages: Comic Visual Analysis and Comic Speech Synthesis. In the first stage, the input comic page is analyzed to identify the gender and age of the characters, as well as texts each character speaks and corresponding emotion. Guided by this analysis, in the second stage, our approach synthesizes realistic speeches for each character, which are consistent with the visual observations. Our experiments show that the proposed approach can synthesize realistic and lively speeches for different types of comics. Perceptual studies performed on the synthesis results of multiple sample comics validate the efficacy of our approach. Wenguan Wang, Wei Liang 0008, Lap-Fai Yu |
ACM Trans. Graph. | 3 |
| 2019 | Functional Workspace Optimization via Learning Personal Preferences from Virtual ExperiencesabstractThe functionality of a workspace is one of the most important considerations in both virtual world design and interior design. To offer appropriate functionality to the user, designers usually take some general rules into account, e.g., general workflow and average stature of users, which are summarized from the population statistics. Yet, such general rules cannot reflect the personal preferences of a single individual, which vary from person to person. In this paper, we intend to optimize a functional workspace according to the personal preferences of the specific individual who will use it. We come up with an approach to learn the individual's personal preferences from his activities while using a virtual version of the workspace via virtual reality devices. Then, we construct a cost function, which incorporates personal preferences, spatial constraints, pose assessments, and visual field. At last, the cost function is optimized to achieve an optimal layout. To evaluate the approach, we experimented with different settings. The results of the user study show that the workspaces updated in this way better fit the users. Wei Liang 0008, Yining Lang, Bing Ning, Lap-Fai Yu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | Tracking Occluded Objects and Recovering Incomplete Trajectories by Reasoning About Containment Relations and Human ActionsabstractThis paper studies a challenging problem of tracking severely occluded objects in long video sequences. The proposed method reasons about the containment relations and human actions, thus infers and recovers occluded objects identities while contained or blocked by others. There are two conditions that lead to incomplete trajectories: i) Contained. The occlusion is caused by a containment relation formed between two objects, e.g., an unobserved laptop inside a backpack forms containment relation between the laptop and the backpack. ii) Blocked. The occlusion is caused by other objects blocking the view from certain locations, during which the containment relation does not change. By explicitly distinguishing these two causes of occlusions, the proposed algorithm formulates tracking problem as a network flow representation encoding containment relations and their changes. By assuming all the occlusions are not spontaneously happened but only triggered by human actions, an MAP inference is applied to jointly interpret the trajectory of an object by detection in space and human actions in time. To quantitatively evaluate our algorithm, we collect a new occluded object dataset captured by Kinect sensor, including a set of RGB-D videos and human skeletons with multiple actors, various objects, and different changes of containment relations. In the experiments, we show that the proposed method demonstrates better performance on tracking occluded objects compared with baseline methods. Wei Liang 0008, Yixin Zhu 0001, Song-Chun Zhu |
AAAI | 1 |
| 2018 | Spatially Perturbed Collision Sounds Attenuate Perceived Causality in 3D Launching EventsabstractWhen a moving object collides with an object at rest, people immediately perceive a causal event: i.e., the first object has launched the second object forwards. However, when the second object's motion is delayed, or is accompanied by a collision sound, causal impressions attenuate and strengthen. Despite a rich literature on causal perception, researchers have exclusively utilized 2D visual displays to examine the launching effect. It remains unclear whether people are equally sensitive to the spatiotemporal properties of observed collisions in the real world. The present study first examined whether previous findings in causal perception with audiovisual inputs can be extended to immersive 3D virtual environments. We then investigated whether perceived causality is influenced by variations in the spatial position of an auditory collision indicator. We found that people are able to localize sound positions based on auditory inputs in VR environments, and spatial discrepancy between the estimated position of the collision sound and the visually observed impact location attenuates perceived causality. Duotun Wang, James Kubricht, Yixin Zhu 0001, Wei Liang 0008, Song-Chun Zhu, Chenfanfu Jiang, Hongjing Lu |
VR | 4 |
| 2017 | Transferring Objects: Joint Inference of Container and Human PoseabstractTransferring objects from one place to another place is a common task performed by human in daily life. During this process, it is usually intuitive for humans to choose an object as a proper container and to use an efficient pose to carry objects; yet, it is non-trivial for current computer vision and machine learning algorithms. In this paper, we propose an approach to jointly infer container and human pose for transferring objects by minimizing the costs associated both object and pose candidates. Our approach predicts which object to choose as a container while reasoning about how humans interact with physical surroundings to accomplish the task of transferring objects given visual input. In the learning phase, the presented method learns how humans make rational choices of containers and poses for transferring different objects, as well as the physical quantities required by the transfer task (e.g., compatibility between container and containee, energy cost of carrying pose) via a structured learning approach. In the inference phase, given a scanned 3D scene with different object candidates and a dictionary of human poses, our approach infers the best object as a container together with human pose for transferring a given object. Hanqing Wang 0001, Wei Liang 0008, Lap-Fai Yu |
ICCV | 2 |
| 2017 | Recognizing key segments of videos for video annotation by learning from web image sets
Hao Song 0002, Xinxiao Wu, Wei Liang 0008, Yunde Jia |
Multim. Tools Appl. | 3 |
| 2017 | Earthquake Safety Training through Virtual DrillsabstractRecent popularity of consumer-grade virtual reality devices, such as the Oculus Rift and the HTC Vive, has enabled household users to experience highly immersive virtual environments. We take advantage of the commercial availability of these devices to provide an immersive and novel virtual reality training approach, designed to teach individuals how to survive earthquakes, in common indoor environments. Our approach makes use of virtual environments realistically populated with furniture objects for training. During a training, a virtual earthquake is simulated. The user navigates in, and manipulates with, the virtual environments to avoid getting hurt, while learning the observation and self-protection skills to survive an earthquake. We demonstrated our approach for common scene types such as offices, living rooms and dining rooms. To test the effectiveness of our approach, we conducted an evaluation by asking users to train in several rooms of a given scene type and then test in a new room of the same type. Evaluation results show that our virtual reality training approach is effective, with the participants who are trained by our approach performing better, on average, than those trained by alternative approaches in terms of the capabilities to avoid physical damage and to detect potentially dangerous objects. ChangYang Li, Wei Liang 0008, Chris Quigley, Yibiao Zhao, Lap-Fai Yu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | 3D head pose estimation with convolutional neural network trained on synthetic imagesabstractIn this paper, we propose a method to estimate head pose with convolutional neural network, which is trained on synthetic head images. We formulate head pose estimation as a regression problem. A convolutional neural network is trained to learn head features and solve the regression problem. To provide annotated head poses in the training process, we generate a realistic head pose dataset by rendering techniques, in which we consider the variation of gender, age, race and expression. Our dataset includes 74000 head poses rendered from 37 head models. For each head pose, RGB image and annotated pose parameters are given. We evaluate our method on both synthetic and real data. The experiments show that our method improves the accuracy of head pose estimation. Xiabing Liu, Wei Liang 0008, Mingtao Pei |
ICIP | 2 |
| 2016 | What Is Where: Inferring Containment Relations from Videos
Wei Liang 0008, Yibiao Zhao, Yixin Zhu 0001, Song-Chun Zhu |
IJCAI | 1 |
| 2016 | Nonnegative correlation coding for image classification
Zhen Dong 0002, Wei Liang 0008, Yuwei Wu 0001, Mingtao Pei, Yunde Jia |
Sci. China Inf. Sci. | 2 |
| 2015 | Evaluating Human Cognition of Containing Relations with Physical Simulation
Wei Liang 0008, Yibiao Zhao, Yixin Zhu 0001, Song-Chun Zhu |
CogSci | 1 |
| 2015 | Human-Object Interaction Recognition by Modeling Context
Wei Liang 0008, Xiabing Liu |
ICIG (2) | 2 |
| 2015 | A Multiple Image Group Adaptation Approach for Event Recognition in Consumer Videos
Dengfeng Zhang, Wei Liang 0008, Hao Song 0002, Zhen Dong 0002, Xinxiao Wu |
ICIG (1) | 2 |
| 2015 | Heterogeneous Discriminant Analysis for Cross-View Action Recognition
Wanchen Sui, Xinxiao Wu, Wei Liang 0008, Yunde Jia |
ICONIP (4) | 4 |
| 2014 | Depth Super-resolution by Fusing Depth Imaging and Stereo Vision with Structural Determinant Information InferenceabstractIn this paper, we present a depth super-resolution framework by fusing depth imaging and stereo vision for high-resolution and high-accuracy depth maps. Depth cameras and stereo vision have their own limitations in some aspects, but their characteristics of range sensing are complementary. Thus, combining both approaches can produce more satisfactory results than either one. Unlike previous fusion methods, we initially taking the noisy depth observation from depth camera as prior information of scene structure. The prior information of scene structure is also utilized to infer structural determinant information, like depth discontinuity and occlusion, which is essential to improve the quality of depth map in the fusion process. In succession, the prior knowledge helps to overcome difficulties of intensity inconsistency in image observation from stereo vision component. Experimental results demonstrate effectiveness and accuracy of the proposed method. Yucheng Wang 0003, Huijun Di, Wei Liang 0008, Jian Zhang 0002, Yunde Jia |
ICPR | 4 |
| 2014 | Recognising human interaction from videos by a discriminative modelabstractThis study addresses the problem of recognising human interactions between two people. The main difficulties lie in the partial occlusion of body parts and the motion ambiguity in interactions. The authors observed that the interdependencies existing at both the action level and the body part level can greatly help disambiguate similar individual movements and facilitate human interaction recognition. Accordingly, they proposed a novel discriminative method, which model the action of each person by a large‐scale global feature and local body part features, to capture such interdependencies for recognising interaction of two people. A variant of multi‐class Adaboost method is proposed to automatically discover class‐specific discriminative three‐dimensional body parts. The proposed approach is tested on the authors newly introduced BIT‐interaction dataset and the UT‐interaction dataset. The results show that their proposed model is quite effective in recognising human interactions. Yu Kong 0001, Wei Liang 0008, Zhen Dong 0002, Yunde Jia |
IET Comput. Vis. | 2 |
| 2013 | Scene image retrieval via re-ranking semantic and packed dense interestpoints
Wei Liang 0008, Xinxiao Wu, Peng Teng |
Neurocomputing | 2 |
| 2012 | Face pose estimation with combined 2D and 3D HOG features
Jiaolong Yang, Wei Liang 0008, Yunde Jia |
ICPR | 2 |
| 2010 | Discriminative human action recognition in the learned hierarchical manifold space
Xinxiao Wu, Wei Liang 0008, Guangming Hou, Yunde Jia |
Image Vis. Comput. | 3 |
| 2010 | Incremental discriminant-analysis of canonical correlations for action recognition
Xinxiao Wu, Yunde Jia, Wei Liang 0008 |
Pattern Recognit. | 3 |
| 2009 | Incremental discriminative-analysis of canonical correlations for action recognitionabstractHuman action recognition is a challenging problem due to the large changes of human appearance in the cases of partial occlusions, non-rigid deformations and high irregularities. It is difficult to collect a large set of training samples with the hope of covering all possible variations of an action. In this paper, we propose an online recognition method, namely Incremental Discriminant-Analysis of Canonical Correlations (IDCC), whose discriminative model is incrementally updated to capture the changes of human appearance and thereby facilitates the recognition task in changing environments. As the training sets are acquired sequentially instead of being given completely in advance, our method is able to compute a new discriminant matrix by updating the existing one using the eigenspace merging algorithm. Experimental results on both Weizmann and KTH action data sets show that our method performs better than state-of-the-art methods on both accuracy and efficiency. Moreover, the robustness of our method is demonstrated on the irregular action recognition. Xinxiao Wu, Wei Liang 0008, Yunde Jia |
ICCV | 2 |
| 2009 | Tracking articulated objects by learning intrinsic structure of motion
Xinxiao Wu, Wei Liang 0008, Yunde Jia |
Pattern Recognit. Lett. | 2 |
| 2009 | Action recognition feedback-based framework for human pose reconstruction from monocular images
Xinxiao Wu, Wei Liang 0008, Yunde Jia |
Pattern Recognit. Lett. | 2 |
| 2008 | Human action recognition using discriminative models in the learned hierarchical manifold spaceabstractA hierarchical learning based approach for human action recognition is proposed in this paper. It consists of hierarchical nonlinear dimensionality reduction based feature extraction and cascade discriminative model based action modeling. Human actions are inferred from human body joint motions and human bodies are decomposed into several physiological body parts according to inherent hierarchy (e.g. right arm, left arm and head all belong to upper body). We explore the underlying hierarchical structures of high-dimensional human pose space using hierarchical Gaussian process latent variable model (HGPLVM) and learn a representative motion pattern set for each body part. In the hierarchical manifold space, the bottom-up cascade conditional random fields (CRFs) are used to predict the corresponding motion pattern in each manifold subspace, and then the final action label is estimated for each observation by a discriminative classifier on the current motion pattern set. Wei Liang 0008, Xinxiao Wu, Yunde Jia |
FG | 2 |
| 2007 | An airborne image stabilization Method based on the Gaussian Mixture modelabstractIn this paper, an image stabilization method based on the adaptive Gaussian mixture model (GMM) is presented in order to smooth down the airborne image vibration. Firstly, the projection algorithm is adopted for the motion estimation; Secondly, GMM parameter is obtained after analyzing characteristics of the first n images; Finally, a stable image sequence is achieved after GMM motion filter operates on the motion parameter. The experimental results show that the method has the advantage of fast speed and effectively smooth unwanted vibration of image sequences. Hongbin Deng, Yunde Jia, Yihua Xu, Wei Liang 0008 |
SMC | 4 |
| 2005 | Simulated Annealing Based Hand Tracking in a Discrete Space
Wei Liang 0008, Yunde Jia, Cheng Ge |
ACII | 1 |
| 2005 | Visual Hand Tracking Using Nonparametric Sequential Belief Propagation
Wei Liang 0008, Yunde Jia, Cheng Ge |
ICIC (1) | 1 |
| 2005 | Hand motion tracking using MDPF methodabstractHand motion tracking is a challenging problem due to the complexity of searching in a high dimensional configuration space for an optimal estimate. This paper represents the hand feasible configurations as a discrete space, which avoids learning to find parameters as general configuration space representations do, meanwhile, it arrange the discrete data on the KD-tree which supports fast nearest neighbor retrieval and it is easy to be modified when new samples are embedded. To track hand motion efficiently, this paper presents a MDPF (Multi-Directional search with Particle Filter) algorithm, in which a 'global' optimization and a 'local' optimization are combined to obtain the best matching configuration. The 'local' method, which is designed to run in multiple processors, could choose more representative samples for global efficiently, and the global method guards the tracking process towards a global minimum. The Experiment results show that this approach is robust and efficient for tracking 3D hand motion. Wei Liang 0008, Cheng Ge, Yunde Jia |
SMC | 1 |