VLDB 2026 Research / reviewers in the wild / expert
Liangyan Gui
dblp:155/5055 · also Liang-Yan Gui
· DBLP profile ↗
27ranked-venue papers
4as first author
20since 2021 · last 2025
0009-0005-8204-3577ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 3 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 13 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object InteractionsabstractAchieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate human-object coupling, variability in object geometries, and artifacts in motion capture data, such as inaccurate contacts and limited hand detail. We introduce InterMimic, a framework that enables a single policy to robustly learn from hours of imperfect MoCap data covering diverse full-body interactions with dynamic and varied objects. Our key insight is to employ a curriculum strategy - perfect first, then scale up. We first train subject-specific teacher policies to mimic, retarget, and refine motion capture data. Next, we distill these teachers into a student policy, with the teachers acting as online experts providing direct supervision, as well as high-quality references. Notably, we incorporate RL fine-tuning on the student policy to surpass mere demonstration replication and achieve higher-quality solutions. Our experiments demonstrate that InterMimic produces realistic and diverse interactions across multiple HOI datasets. The learned policy generalizes in a zero-shot manner and seamlessly integrates with kinematic generators, elevating the framework from mere imitation to generative modeling of complex human-object interactions. Sirui Xu 0002, Hung Yu Ling, Yu-Xiong Wang, Liangyan Gui |
CVPR | 4 |
| 2025 | InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction GenerationabstractWhile large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack extensive, high-quality motion and annotation and exhibit artifacts such as contact penetration, floating, and incorrect hand motions. To address these issues, we introduce InterAct, a large-scale 3D HOI benchmark featuring dataset and methodological advancements. First, we consolidate and standardize 21.81 hours of HOI data from diverse sources, enriching it with detailed textual annotations. Second, we propose a unified optimization framework to enhance data quality by reducing artifacts and correcting hand motions. Leveraging the principle of contact invariance, we maintain human-object relationships while introducing motion variations, expanding the dataset to 30.70 hours. Third, we define six benchmarking tasks and develop a unified HOI generative modeling perspective, achieving state-of-the-art performance. Extensive experiments validate the utility of our dataset as a foundational resource for advancing 3D human-object interaction generation. The dataset will be publicly accessible to support further research in the field. Sirui Xu 0002, Dongting Li 0001, Xiyan Xu, Qi Long, Ziyin Wang, Yunzhi Lu, Shuchang Dong, Hezi Jiang, Akshat Gupta, Yu-Xiong Wang, Liangyan Gui |
CVPR | 12 |
| 2025 | SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity PredictionabstractText-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion triplets, has opened new avenues for training, yielding promising results. However, existing methods struggle with precise control, often leading to misalignment between motion semantics and language instructions. In this paper, we introduce a related task — motion similarity prediction — and propose a multi-task training paradigm, where we train the model jointly on motion editing and motion similarity prediction to foster the learning of semantically meaningful representations. To complement this task, we design an advanced Diffusion-Transformer-based architecture that separately handles motion similarity prediction and motion editing. Extensive experiments demonstrate the state-of-the-art performance of our approach in both editing alignment and fidelity. Project URL: https://github.com/lzhyu/SimMotionEdit. Zhengyuan Li, Anindita Ghosh, Uttaran Bhattacharya, Liangyan Gui, Aniket Bera |
CVPR | 5 |
| 2025 | Argus: Vision-Centric Reasoning with Grounded Chain-of-ThoughtabstractRecent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often struggle with vision-centric scenarios where precise visual focus is needed for accurate reasoning. In this paper, we introduce Argus to address these limitations with a new visual attention grounding mechanism. Our approach employs object-centric grounding as visual chain-of-thought signals, enabling more effective goal-conditioned visual attention during multimodal reasoning tasks. Evaluations on diverse benchmarks demonstrate that Argus excels in both multi-modal reasoning tasks and referring object grounding tasks. Extensive analysis further validates various design choices of Argus, and reveals the effectiveness of explicit language-guided visual region-of-interest engagement in MLLMs, highlighting the importance of advancing multimodal intelligence from a visual-centric perspective. Yunze Man, De-An Huang, Guilin Liu, Shiwei Sheng, Shilong Liu 0004, Liangyan Gui, Jan Kautz, Yu-Xiong Wang, Zhiding Yu |
CVPR | 6 |
| 2025 | Floating No More: Object-Ground Reconstruction from a Single ImageabstractRecent advancements in 3D object reconstruction from single images have primarily focused on improving the accuracy of object shapes. Yet, these techniques often fail to accurately capture the inter-relation between the object, ground, and camera. As a result, the reconstructed objects often appear floating or tilted when placed on flat surfaces. This limitation significantly affects 3D-aware image editing applications, like shadow rendering and object pose manipulation. To address this issue, we introduce ORG (Object Reconstruction with Ground), a novel task aimed at reconstructing 3D object geometry in conjunction with the ground surface. Our method uses two compact pixel-level representations to depict the relationship between camera, object, and ground. Experiments show that the proposed ORG model can effectively reconstruct object-ground geometry on unseen data, significantly enhancing the quality of shadow generation and pose manipulation compared to conventional single-image 3D reconstruction techniques. The project page can be found at this website. Yunze Man, Yichen Sheng, Liangyan Gui, Yu-Xiong Wang |
CVPR | 4 |
| 2025 | Refer to Any Segmentation Mask Group with Vision-Language PromptsabstractRecent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that require user-friendly interactions driven by vision-language prompts. To bridge this gap, we introduce a novel task of omnimodal referring expression segmentation (ORES). In this task, a model produces a group of masks based on arbitrary prompts specified by text only or text plus reference visual entities. To address this new challenge, we propose a novel framework to "Refer to Any Segmentation Mask Group" (RAS), which augments segmentation models with complex multimodal interactions and comprehension via a mask-centric large multimodal model. For training and benchmarking ORES models, we create datasets MaskGroups-2M and MaskGroups-HQ to include diverse mask groups specified by text and reference entities. Through extensive evaluation, we demonstrate superior performance of RAS on our new ORES task, as well as classic referring expression segmentation (RES) and generalized referring expression segmentation (GRES) tasks. Project page: https://Ref2Any.github.io. Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu, Lingzhi Zhang, Jiuxiang Gu, Hyunjoon Jung, Liangyan Gui, Yu-Xiong Wang |
ICCV | 8 |
| 2024 | Situational Awareness Matters in 3D Vision Language ReasoningabstractBeing able to carry out complicated vision language reasoning tasks in 3D space represents a significant mile-stone in developing household robots and human-centered embodied AI. In this work, we demonstrate that a criti-cal and distinct challenge in 3D vision language reasoning is the situational awareness, which incorporates two key components: (1) The autonomous agent grounds its self-location based on a language prompt. (2) The agent answers open-ended questions from the perspective of its calculated position. To address this challenge, we introduce SIG3D, an end-to-end Situation-Grounded model for 3D vision language reasoning. We tokenize the 3D scene into sparse voxel representation, and propose a language-grounded situation estimator, followed by a situated question answering module. Experiments on the SQA3D and ScanQA datasets show that SIG3D outperforms state-of-the-art models in situational estimation and question answering by a large margin (e.g., an enhancement of over 30% on situation accuracy). Subsequent analysis corrobo-rates our architectural design choices, explores the distinct functions of visual and textual tokens, and highlights the importance of situational awareness in the domain of 3D question-answering. Project page is available at htt ps: //yunzeman.github.io/situation3d. Yunze Man, Liangyan Gui, Yu-Xiong Wang |
CVPR | 2 |
| 2024 | SOHES: Self-supervised Open-world Hierarchical Entity SegmentationabstractOpen-world entity segmentation, as an emerging computer vision task, aims at segmenting entities in images without being restricted by pre-defined classes, offering impressive generalization capabilities on unseen images and concepts. Despite its promise, existing entity segmentation methods like Segment Anything Model (SAM) rely heavily on costly expert annotators. This work presents Self-supervised Open-world Hierarchical Entity Segmentation (SOHES), a novel approach that eliminates the need for human annotations. SOHES operates in three phases: self-exploration, self-instruction, and self-correction. Given a pre-trained self-supervised representation, we produce abundant high-quality pseudo-labels through visual feature clustering. Then, we train a segmentation model on the pseudo-labels, and rectify the noises in pseudo-labels via a teacher-student mutual-learning procedure. Beyond segmenting entities, SOHES also captures their constituent parts, providing a hierarchical understanding of visual entities. Using raw images as the sole training data, our method achieves unprecedented performance in self-supervised open-world segmentation, marking a significant milestone towards high-quality open-world entity segmentation in the absence of human-annotated masks. Project page: https://SOHES.github.io. Shengcao Cao, Jiuxiang Gu, Jason Kuen, Hao Tan 0002, Ruiyi Zhang 0002, Handong Zhao, Ani Nenkova, Liangyan Gui, Tong Sun 0005, Yu-Xiong Wang |
ICLR | 8 |
| 2024 | InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionabstractText-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic human-object interaction (HOI) generation faces notable challenges, primarily due to the lack of large-scale interaction data and comprehensive descriptions that align with these interactions. This paper takes the initiative and showcases the potential of generating human-object interactions without direct training on text-interaction pair data. Our key insight in achieving this is that interaction semantics and dynamics can be decoupled. Being unable to learn interaction semantics through supervised training, we instead leverage pre-trained large models, synergizing knowledge from a large language model and a text-to-motion model. While such knowledge offers high-level control over interaction semantics, it cannot grasp the intricacies of low-level interaction dynamics. To overcome this issue, we introduce a world model designed to comprehend simple physics, modeling how human actions influence object motion. By integrating these components, our novel framework, InterDreamer, is able to generate text-aligned 3D HOI sequences without relying on paired text-interaction data. We apply InterDreamer to the BEHAVE, OMOMO, and CHAIRS datasets, and our comprehensive experimental analysis demonstrates its capability to generate realistic and coherent interaction sequences that seamlessly align with the text directives. Sirui Xu 0002, Ziyin Wang, Yu-Xiong Wang, Liangyan Gui |
NeurIPS | 4 |
| 2024 | Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingabstractComplex 3D scene understanding has gained increasing attention, with scene encoding strategies built on top of visual foundation models playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly compared to their image-based counterparts. To address this issue, we present the first comprehensive study that probes various visual encoding models for 3D scene understanding, identifying the strengths and limitations of each model across different scenarios. Our evaluation spans seven vision foundation encoders, including image, video, and 3D foundation models. We evaluate these models in four tasks: Vision-Language Scene Reasoning, Visual Grounding, Segmentation, and Registration, each focusing on different aspects of scene understanding. Our evaluation yields key intriguing findings: Unsupervised image foundation models demonstrate superior overall performance, video models excel in object-level tasks, diffusion models benefit geometric tasks, language-pretrained models show unexpected limitations in language-related tasks, and the mixture-of-vision-expert (MoVE) strategy leads to consistent performance improvement. These insights challenge some conventional understandings, provide novel perspectives on leveraging visual foundation models, and highlight the need for more flexible encoder selection in future vision-language and scene understanding tasks. Yunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert, Liangyan Gui, Yu-Xiong Wang |
NeurIPS | 5 |
| 2023 | Contrastive Mean Teacher for Domain Adaptive Object DetectorsabstractObject detectors often suffer from the domain gap between training (source domain) and real-world applications (target domain). Mean-teacher self-training is a powerful paradigm in unsupervised domain adaptation for object detection, but it struggles with low-quality pseudo-labels. In this work, we identify the intriguing alignment and synergy between mean-teacher self-training and contrastive learning. Motivated by this, we propose Contrastive Mean Teacher (CMT) - a unified, general-purpose framework with the two paradigms naturally integrated to maximize beneficial learning signals. Instead of using pseudo-labels solely for final predictions, our strategy extracts object-level features using pseudo-labels and optimizes them via contrastive learning, without requiring labels in the target domain. When combined with recent mean-teacher self-training methods, CMT leads to new state-of-the-art target-domain performance: 51.9% mAP on Foggy Cityscapes, outperforming the previously best by 2.1% mAP. Notably, CMT can stabilize performance and provide more significant gains as pseudo-label noise increases. Shengcao Cao, Dhiraj Joshi, Liangyan Gui, Yu-Xiong Wang |
CVPR | 3 |
| 2023 | SDFusion: Multimodal 3D Shape Completion, Reconstruction, and GenerationabstractIn this work, we present a novel framework built to sim-plify 3D asset generation for amateur users. To enable interactive generation, our method supports a variety of input modalities that can be easily provided by a human, in-cluding images, text, partially observed shapes and combinations of these, further allowing to adjust the strength of each input. At the core of our approach is an encoder-decoder, compressing 3D shapes into a compact latent representation, upon which a diffusion model is learned. To enable a variety of multimodal inputs, we employ task-specific encoders with dropout followed by a cross-attention mechanism. Due to its flexibility, our model naturally supports a variety of tasks, outperforming prior works on shape completion, image-based 3D reconstruction, and text-to-3D. Most interestingly, our model can combine all these tasks into one swiss-army-knife tool, enabling the user to perform shape generation using incomplete shapes, images, and textual descriptions at the same time, providing the relative weights for each input and facilitating interactivity. Despite our approach being shape-only, we further show an efficient method to texture the generated shape using large-scale text-to-image models. Yen-Chi Cheng, Hsin-Ying Lee 0001, Sergey Tulyakov, Alexander G. Schwing, Liangyan Gui |
CVPR | 5 |
| 2023 | BEV-Guided Multi-Modality Fusion for Driving PerceptionabstractIntegrating multiple sensors and addressing diverse tasks in an end-to-end algorithm are challenging yet critical topics for autonomous driving. To this end, we introduce BEVGuide, a novel Bird's Eye- View (BEV) representation learning framework, representing the first attempt to unify a wide range of sensors under direct BEV guidance in an end-to-end fashion. Our architecture accepts input from a diverse sensor pool, including but not limited to Camera, Lidar and Radar sensors, and extracts BEV feature embeddings using a versatile and general transformer backbone. We design a BEV-guided multi-sensor attention block to take queries from BEV embeddings and learn the BEV representation from sensor-specific features. BEVGuide is efficient due to its lightweight backbone design and highly flexible as it supports almost any input sensor configurations. Extensive experiments demonstrate that our framework achieves exceptional performance in BEV perception tasks with a diverse sensor set. Project page is at https://yunzeman.github.io/BEVGuide. Yunze Man, Liangyan Gui, Yu-Xiong Wang |
CVPR | 2 |
| 2023 | InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionabstractThis paper addresses a novel task of anticipating 3D human-object interactions (HOIs). Most existing research on HOI synthesis lacks comprehensive whole-body interactions with dynamic objects, e.g., often limited to manipulating small or static objects. Our task is significantly more challenging, as it requires modeling dynamic objects with various shapes, capturing whole-body motion, and ensuring physically valid interactions. To this end, we propose InterDiff, a framework comprising two key steps: (i) interaction diffusion, where we leverage a diffusion model to encode the distribution of future human-object interactions; (ii) interaction correction, where we introduce a physics-informed predictor to correct denoised HOIs in a diffusion step. Our key insight is to inject prior knowledge that the interactions under reference with respect to contact points follow a simple pattern and are easily predictable. Experiments on multiple human-object interaction datasets demonstrate the effectiveness of our method for this task, capable of producing realistic, vivid, and remarkably longterm 3D HOI predictions. Sirui Xu 0002, Zhengyuan Li, Yu-Xiong Wang, Liangyan Gui |
ICCV | 4 |
| 2023 | Stochastic Multi-Person 3D Motion Forecasting
Sirui Xu 0002, Yu-Xiong Wang, Liangyan Gui |
ICLR | 3 |
| 2023 | Learning Lightweight Object Detectors via Multi-Teacher Progressive DistillationabstractResource-constrained perception systems such as edge computing and vision-for-robotics require vision models to be both accurate and lightweight in computation and memory usage. While knowledge distillation is a proven strategy to enhance the performance of lightweight classification models, its application to structured outputs like object detection and instance segmentation remains a complicated task, due to the variability in outputs and complex internal network modules involved in the distillation process. In this paper, we propose a simple yet surprisingly effective sequential approach to knowledge distillation that progressively transfers the knowledge of a set of teacher detectors to a given lightweight student. To distill knowledge from a highly accurate but complex teacher model, we construct a sequence of teachers to help the student gradually adapt. Our progressive strategy can be easily combined with existing detection distillation mechanisms to consistently maximize student performance in various settings. To the best of our knowledge, we are the first to successfully distill knowledge from Transformer-based teacher detectors to convolution-based students, and unprecedentedly boost the performance of ResNet-50 based RetinaNet from 36.5% to 42.0% AP and Mask R-CNN from 38.2% to 42.5% AP on the MS COCO benchmark. Code available at https://github.com/Shengcao-Cao/MTPD. Shengcao Cao, James Hays, Deva Ramanan, Yu-Xiong Wang, Liangyan Gui |
ICML | 6 |
| 2023 | DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV PerceptionabstractClosing the domain gap between training and deployment and incorporating multiple sensor modalities are two challenging yet critical topics for self-driving. Existing work only focuses on single one of the above topics, overlooking the simultaneous domain and modality shift which pervasively exists in real-world scenarios. A model trained with multi-sensor data collected in Europe may need to run in Asia with a subset of input sensors available. In this work, we propose DualCross, a cross-modality cross-domain adaptation framework to facilitate the learning of a more robust monocular bird's-eye-view (BEV) perception model, which transfers the point cloud knowledge from a LiDAR sensor in one domain during the training phase to the camera-only testing scenario in a different domain. This work results in the first open analysis of cross-domain cross-sensor perception and adaptation for monocular 3D tasks in the wild. We benchmark our approach on large-scale datasets under a wide range of domain shifts and show state-of-the-art results against various baselines. Our project webpage is at https://yunzeman.github.io/DualCross. Yunze Man, Liangyan Gui, Yu-Xiong Wang |
IROS | 2 |
| 2023 | HASSOD: Hierarchical Adaptive Self-Supervised Object DetectionabstractThe human visual perception system demonstrates exceptional capabilities in learning without explicit supervision and understanding the part-to-whole composition of objects. Drawing inspiration from these two abilities, we propose Hierarchical Adaptive Self-Supervised Object Detection (HASSOD), a novel approach that learns to detect objects and understand their compositions without human supervision. HASSOD employs a hierarchical adaptive clustering strategy to group regions into object masks based on self-supervised visual representations, adaptively determining the number of objects per image. Furthermore, HASSOD identifies the hierarchical levels of objects in terms of composition, by analyzing coverage relations between masks and constructing tree structures. This additional self-supervised learning task leads to improved detection performance and enhanced interpretability. Lastly, we abandon the inefficient multi-round self-training process utilized in prior methods and instead adapt the Mean Teacher framework from semi-supervised learning, which leads to a smoother and more efficient training process. Through extensive experiments on prevalent image datasets, we demonstrate the superiority of HASSOD over existing methods, thereby advancing the state of the art in self-supervised object detection. Notably, we improve Mask AR from 20.2 to 22.5 on LVIS, and from 17.0 to 26.0 on SA-1B. Project page: https://HASSOD-NeurIPS23.github.io. Shengcao Cao, Dhiraj Joshi, Liangyan Gui, Yu-Xiong Wang |
NeurIPS | 3 |
| 2022 | Joint Forecasting of Panoptic Segmentations with Difference AttentionabstractForecasting of a representation is important for safe and effective autonomy. For this, panoptic segmentations have been studied as a compelling representation in recent work. However, recent state-of-the-art on panoptic segmentation forecasting suffers from two issues: first, individual object instances are treated independently of each other; second, individual object instance forecasts are merged in a heuristic manner. To address both issues, we study a new panoptic segmentation forecasting model that jointly forecasts all object instances in a scene using a transformer model based on ‘difference attention.’ It further refines the predictions by taking depth estimates into account. We evaluate the proposed model on the Cityscapes and AIODrive datasets. We find difference attention to be particularly suitable for forecasting because the difference of quantities like locations enables a model to explicitly reason about velocities and acceleration. Because of this, we attain state-of-the-art on panoptic segmentation forecasting metrics. Colin Graber, Cyril Jazra, Wenjie Luo 0002, Liangyan Gui, Alexander G. Schwing |
CVPR | 4 |
| 2022 | Diverse Human Motion Prediction Guided by Multi-level Spatial-Temporal Anchors
Sirui Xu 0002, Yu-Xiong Wang, Liangyan Gui |
ECCV (22) | 3 |
| 2018 | Adversarial Geometry-Aware Human Motion Prediction
Liangyan Gui, Yu-Xiong Wang, Xiaodan Liang, José M. F. Moura |
ECCV (4) | 1 |
| 2018 | Few-Shot Human Motion Prediction via Meta-learning
Liangyan Gui, Yu-Xiong Wang, Deva Ramanan, José M. F. Moura |
ECCV (8) | 1 |
| 2018 | Teaching Robots to Predict Human MotionabstractTeaching a robot to predict and mimic how a human moves or acts in the near future by observing a series of historical human movements is a crucial first step in human-robot interaction and collaboration. In this paper, we instrument a robot with such a prediction ability by leveraging recent deep learning and computer vision techniques. First, our system takes images from the robot camera as input to produce the corresponding human skeleton based on real-time human pose estimation obtained with the OpenPose library. Then, conditioning on this historical sequence, the robot forecasts plausible motion through a motion predictor, generating a corresponding demonstration. Because of a lack of high-level fidelity validation, existing forecasting algorithms suffer from error accumulation and inaccurate prediction. Inspired by generative adversarial networks (GANs), we introduce a global discriminator that examines whether the predicted sequence is smooth and realistic. Our resulting motion GAN model achieves superior prediction performance to state-of-the-art approaches when evaluated on the standard H3.6M dataset. Based on this motion GAN model, the robot demonstrates its ability to replay the predicted motion in a human-like manner when interacting with a person. Liangyan Gui, Kevin Zhang 0002, Yu-Xiong Wang, Xiaodan Liang, José M. F. Moura, Manuela M. Veloso |
IROS | 1 |
| 2018 | Factorized Convolutional Networks: Unsupervised Fine-Tuning for Image ClusteringabstractDeep convolutional neural networks (CNNs) have recognized promise as universal representations for various image recognition tasks. One of their properties is the ability to transfer knowledge from a large annotated source dataset (e.g., ImageNet) to a (typically smaller) target dataset. This is usually accomplished through supervised fine-tuning on labeled new target data. In this work, we address "unsupervised fine-tuning" that transfers a pre-trained network to target tasks with unlabeled data such as image clustering tasks. To this end, we introduce group-sparse non-negative matrix factorization (GSNMF), a variant of NMF, to identify a rich set of high-level latent variables that are informative on the target task. The resulting "factorized convolutional network" (FCN) can itself be seen as a feed-forward model that combines CNN and two-layer structured NMF. We empirically validate our approach and demonstrate state-of-the-art image clustering performance on challenging scene (MIT-67) and fine-grained (Birds-200, Flowers-102) benchmarks. We further show that, when used as unsupervised initialization, our approach improves image classification performance as well. Liangyan Gui, Liangke Gui, Yu-Xiong Wang, Louis-Philippe Morency, José M. F. Moura |
WACV | 1 |
| 2015 | Traffic flow from a low frame rate city cameraabstractTraffic flow in a city is a rich source of information about the city. Cities are being instrumented with video cameras. They can potentially generate continuously large datasets to be processed (big data). This paper reports on our current work to detect traffic flow from an on-line low quality, low frame rate city video camera. The paper details a pipeline of four main steps - background subtraction, scene geometry, car detection, and car counting, and it illustrates results obtained with processing video from a single camera. Evgeny Toropov, Liangyan Gui, Shanghang Zhang, Satwik Kottur, José M. F. Moura |
ICIP | 2 |
| 2014 | Relay Channel with Causal Channel State InformationabstractIn this paper, state dependent relay channel (SD- RC) with causal channel state information (CSI) is considered. Two different cases are investigated in which causal CSI is available: 1) only at the relay node, 2) at all the nodes. The second situation is specialized to cases where causal CSI is known: at both source and relay; at both destination and relay. We established the lower bounds for these situations. Our bound contains all the bounds in previous work, and it is strictly better under some cases. In our scheme, the CSI at relay is compressed by Wyner-Ziv compression and transmitted with the message information by exploiting Shannon strategy. The transmission of CSI has twofold effect. On the one hand, it may reduce the message rate to be relayed. On the other hand, the destination can use the received CSI to help decode message from both the source and relay. Note that if the channel between the relay and destination is good enough, the relay can transmit all the message information it got and use the extra capacity to transmit compressed CSI. In this situation the transmission of CSI doesn't reduce the message rate to be relayed, and our scheme is better than compress- forward (CF) relaying with Shannon strategy. Dajin Wang, Liangyan Gui |
VTC Fall | 2 |
| 2012 | Neighborhood Preserving Non-negative Tensor Factorization for image representationabstractNon-negative Matrix Factorization (NMF) has become a powerful tool for image representation due to its enhanced semantic interpretability under non-negativity. Unfortunately, two types of neighborhood information essential to representation are lost in NMF. For individual image, the local structure information is missing in the vectorization, which can then be avoided by Non-negative Tensor Factorization (NTF). For image data points, they often reside on a low dimensional submanifold embedded in a high dimensional ambient space. NMF and NTF are incapable of encoding the local geometrical information, which can nevertheless be resuscitated by manifold learning. To simultaneously model both of the neighborhood relationship within and among image data, this paper proposes a novel algorithm called Neighborhood Preserving Non-negative Tensor Factorization (NPNTF) by incorporating locally linear embedding regularization into tensor factorization. Experimental results on image clustering show the superior performance of NPNTF with more natural and discriminating representation ability. Yu-Xiong Wang, Liangyan Gui, Yu-Jin Zhang |
ICASSP | 2 |