EDBT 2026 Demo / reviewers in the wild / expert
Liqian Ma
dblp:145/1269
· DBLP profile ↗
30ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pairwise ratio transformation of gene expression data leads to improved checkpoint response prediction in lung cancer patientsabstractBACKGROUND: Machine learning algorithms identify patterns that would otherwise be difficult to observe in high-dimensional molecular and clinical data. For this reason, machine learning has the potential to have a profound impact on clinical decision-making and drug target discovery. However, there are technical challenges in adapting these tools for clinical use, including clinical feature engineering, model selection, and defining optimal strategies for model training. For cancer care, RNA sequencing of patient tumor biopsies has already proven to be a powerful molecular assay to characterize tumor-intrinsic and -extrinsic phenotypes influencing therapeutic response, but an optimal solution for using gene expression data to predict outcome is yet to be established. RESULTS: We developed the tauX machine learning framework to refine gene expression features and improve the predictive performance of RNA-sequencing data. The tauX framework uses aggregated ratios of positively and negatively associated predictive genes to simplify the prediction task. We showed a significant improvement in predictive performance using a large database of synthetic gene expression profiles. We also showed how the tauX framework can be used to elucidate the mechanisms of response and resistance to checkpoint blockade therapy using data from the Stand Up to Cancer (SU2C) Lung Response Cohort and The Cancer Genome Atlas (TCGA). The tauX framework achieved superior predictive performance (~ 30% improvement) compared to models built upon established feature engineering strategies or widely used cancer gene expression signatures. The tauX framework is available as a freely deployable docker container ( https://hub.docker.com/r/pfeiljx/taux ). CONCLUSION: By simultaneously modeling gene expression signatures associated with response and resistance to drug therapy, the tauX approach revealed expression patterns that can be used to improve genomic medicine strategies in several ways. Significantly, tauX allows the paired response and resistance signatures to be used to design new companion diagnostics. Application of the tauX framework can also be used to identify drug targets by indicating genes that consistently associate with resistance. Improved performance in drug response prediction using the tauX approach can support data-driven decision-making in the precision medicine space that can lead to improved clinical outcomes for patients. Jacob Pfeil, Liqian Ma, Hin Ching Lo, Tolga Turan, R. Tyler McLaughlin, Severiano Villarruel, Josue Samayoa, Kyle Halliwill |
BMC Bioinform. | 2 |
| 2026 | What you say is what you get: Text-aligned semantic reward modeling with vision-language representations for reinforcement learning
Yukun Xiao, Liqian Ma, Yisheng An, Bo Sun 0017, Yaxian Wang, Jun Liu 0002 |
Knowl. Based Syst. | 3 |
| 2026 | PRITO: Performance-Reputation Integrated Task Offloading for Reliable Vehicular Edge ComputingabstractAs vehicular edge computing (VEC) grows increasingly demanding, distributed task offloading is brought up as a potential solution. However, the high mobility and limited resources of vehicles, coupled with uncertain service reliability, make it difficult to guarantee timely and correct execution of computation-intensive tasks. Existing task offloading models overlook critical aspects such as punctuality and correctness, limiting their ability to select reliable service vehicles in dynamic and potentially adversarial environments. To address this gap, we propose a performance-reputation integrated task offloading scheme for reliable VEC. We introduce a dual-dimensional reputation model (D2Rep) that jointly evaluates vehicles based on the degree of punctuality, result correctness probability, and historical reputation value. Building on this, we develop the performance value-maximized vehicular computation offloading (PFVMax-VCO) scheme, which combines reputation value, link reliability, and projected success rate into a unified performance value to guide multi-objective optimization of delay, energy consumption, and reliability. To support the realistic evaluation, we implement a SUMO-Veins-OMNeT++ integrated simulation framework. Experimental results demonstrate that the proposed PFVMax-VCO scheme outperforms existing solutions in terms of execution cost, result correctness probability, and task success rate, while enabling robust and dynamic reputation management for VEC systems. Liqian Ma, Yisheng An, Shumei Liu, Yukun Xiao, Yonghui Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Bootstraping Clustering of Gaussians for View-consistent 3D Scene UnderstandingabstractInjecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance on 2D supervision can undermine cross-view semantic consistency and necessitate complex data preparation processes, therefore hindering view-consistent scene understanding. In this work, we present FreeGS, an unsupervised semantic-embedded 3DGS framework that achieves view-consistent 3D scene understanding without the need for 2D labels. Instead of directly learning semantic features, we introduce the IDentity-coupled Semantic Field (IDSF) into 3DGS, which captures both semantic representations and view-consistent instance indices for each Gaussian. We optimize IDSF with a two-step alternating strategy: semantics help to extract coherent instances in 3D space, while the resulting instances regularize the injection of stable semantics from 2D space. Additionally, we adopt a 2D-3D joint contrastive loss to enhance the complementarity between view-consistent 3D geometry and rich semantics during the bootstrapping process, enabling FreeGS to uniformly perform tasks such as novel-view semantic segmentation, object selection, and 3D object detection. Extensive experiments on LERF-Mask, 3D-OVS, and ScanNet datasets demonstrate that FreeGS performs comparably to state-of-the-art methods while avoiding the complex data preprocessing workload. Lu Zhang 0053, Ping Hu 0001, Liqian Ma, Yunzhi Zhuge, Huchuan Lu |
AAAI | 4 |
| 2025 | CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian SplattingabstractRecent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: cross-view semantic inconsistencies induced by occlusion, image blur, and view-dependent variations. These inconsistencies, when propagated via projection supervision, deteriorate the quality of 3D Gaussian semantic fields and introduce artifacts in the rendered outputs. To mitigate this limitation, we propose CCL-LGS, a novel framework that enforces view-consistent semantic supervision by integrating multi-view semantic cues. Specifically, our approach first employs a zero-shot tracker to align a set of SAM-generated 2D masks and reliably identify their corresponding categories. Next, we utilize CLIP to extract robust semantic encodings across views. Finally, our Contrastive Codebook Learning (CCL) module distills discriminative semantic features by enforcing intra-class compactness and inter-class distinctiveness. In contrast to previous methods that directly apply CLIP to imperfect masks, our framework explicitly resolves semantic conflicts while preserving category discriminability. Extensive experiments demonstrate that CCL-LGS outperforms previous state-of-the-art methods. Our project page is available at https://epsilontl.github.io/CCL-LGS/. Xiaomin Li 0001, Liqian Ma, Zirui Zheng, Hefei Huang, Taiqing Li, Huchuan Lu, Xu Jia 0012 |
ICCV | 3 |
| 2025 | VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical PriorabstractVideo diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos due to an inherent lack of understanding of physics, resulting in incorrect dynamics and event sequences. To address this limitation, we propose a novel two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior. In the first stage, we employ a Vision Language Model (VLM) as a coarse-grained motion planner, integrating chain-of-thought and physics-aware reasoning to predict a rough motion trajectories/changes that approximate real-world physical dynamics while ensuring the inter-frame consistency. In the second stage, we use the predicted motion trajectories/changes to guide the video generation of a VDM. As the predicted motion trajectories/changes are rough, noise is added during inference to provide freedom to the VDM in generating motion with more fine details. Extensive experimental results demonstrate that our framework can produce physically plausible motion, and comparative evaluations highlight the notable superiority of our approach over existing methods. More video results are available on our Project Page: https://madaoer.github.io/projects/physically_plausible_video_generation. Xindi Yang, Baolu Li 0001, Zhenfei Yin, Lei Bai 0001, Liqian Ma, Zhiyong Wang 0001, Jianfei Cai 0001, Tien-Tsin Wong, Huchuan Lu, Xu Jia 0012 |
ICCV | 6 |
| 2025 | Towards Volumetric Video: a Technical Overview of Immersive MediaabstractWith the continuous growth in display technologies, computer vision and graphics, Volumetric Video is emerging as a promising frontier in research. However, the increasing interest in this field has led to growing confusion regarding video representations and terminologies. To address this, we present a taxonomy of videos based on interaction degrees of freedom and stereoscopic properties. Our work provides detailed insights into the technological foundations of each immersive media format, identifies their respective strengths and weaknesses, and offers a foundational understanding of volumetric video capture, reconstruction, and display. We hope this survey will stimulate future research and applications in immersive media. Shi Pan, Hongshuai Li, Zhengxian Yang, Liqian Ma, Hua Du, Borong Lin |
MMSP | 6 |
| 2025 | Generalizable Domain Adaptation for Sim-and-Real Policy Co-TrainingabstractBehavior cloning has shown promise for robot manipulation, but real-world demonstrations are costly to acquire at scale. While simulated data offers a scalable alternative, particularly with advances in automated demonstration generation, transferring policies to the real world is hampered by various simulation and real domain gaps. In this work, we propose a unified sim-and-real co-training framework for learning generalizable manipulation policies that primarily leverages simulation and only requires a few real-world demonstrations. Central to our approach is learning a domain-invariant, task-relevant feature space. Our key insight is that aligning the joint distributions of observations and their corresponding actions across domains provides a richer signal than aligning observations (marginals) alone. We achieve this by embedding an Optimal Transport (OT)-inspired loss within the co-training framework, and extend this to an Unbalanced OT framework to handle the imbalance between abundant simulation data and limited real-world examples. We validate our method on challenging manipulation tasks, showing it can leverage abundant simulation data to achieve up to a 30\% improvement in the real-world success rate and even generalize to scenarios seen only in simulation. Liqian Ma, Zhenyang Chen, Ajay Mandlekar, Caelan Reed Garrett, Danfei Xu |
NeurIPS | 2 |
| 2025 | CharacterFactory: Sampling Consistent Characters With GANs for Diffusion ModelsabstractRecent advances in text-to-image models have opened new frontiers in human-centric generation. However, these models cannot be directly employed to generate images with consistent newly coined identities. In this work, we propose CharacterFactory, a framework that allows sampling new characters with consistent identities in the latent space of GANs for diffusion models. More specifically, we consider the word embeddings of celeb names as ground truths for the identity-consistent generation task and train a GAN model to learn the mapping from a latent space to the celeb embedding space. In addition, we design a context-consistent loss to ensure that the generated identity embeddings can produce identity-consistent images in various contexts. Remarkably, the whole model only takes 10 minutes for training, and can sample infinite characters end-to-end during inference. Extensive experiments demonstrate excellent performance of the proposed CharacterFactory on character creation in terms of identity consistency and editability. Furthermore, the generated characters can be seamlessly combined with the off-the-shelf image/video/3D diffusion models. We believe that the proposed CharacterFactory is an important step for identity-consistent character generation. Code and Gradio demo are available at: https://qinghew.github.io/CharacterFactory/. Baolu Li 0001, Xiaomin Li 0001, Bing Cao 0002, Liqian Ma, Huchuan Lu, Xu Jia 0012 |
IEEE Trans. Image Process. | 5 |
| 2025 | StableIdentity: Inserting Anybody Into Anywhere at First SightabstractRecent advances in large pretrained text-to-image generation models have shown unprecedented capabilities for high-quality human-centric generation, however, customizing face identity is still an intractable problem. Existing methods cannot ensure stable identity preservation and flexible editability, even with several images for each subject during training. In this work, we propose StableIdentity, which allows identity-consistent recontextualization with just one face image from a person seen for the first time. More specifically, we employ a face encoder with the identity prior to encode the input face, and then calibrate the face representation to align the distribution of a space with the editability prior, which is constructed from celeb names. By incorporating identity prior and editability prior, the learned identity can be injected anywhere with various contexts. In addition, we design a masked two-phase diffusion loss to boost the pixel-level perception of the input face and maintain the diversity of generation. Extensive experiments demonstrate our method outperforms previous customization methods. In addition, the learned identity can be flexibly combined with the off-theshelf modules such as ControlNet. Notably, to the best of our knowledge, we are the first to directly inject the identity learned from a single image into video/3D generation without finetuning. We believe that the proposed StableIdentity is an important step to unify image, video, and 3D customized generation models. The code is available: https://github.com/qinghew/StableIdentity. Xu Jia 0012, Xiaomin Li 0001, Taiqing Li, Liqian Ma, Yunzhi Zhuge, Huchuan Lu |
IEEE Trans. Multim. | 5 |
| 2025 | DexSim2Real$^{\mathbf{2}}$: Building Explicit World Model for Precise Articulated Object Dexterous ManipulationabstractArticulated objects are ubiquitous in daily life. In this paper, we present DexSim2Real$^{\mathbf{2}}$, a novel framework for goal-conditioned articulated object manipulation. The core of our framework is constructing an explicit world model of unseen articulated objects through active interactions, which enables sampling-based model predictive control to plan trajectories achieving different goals without requiring demonstrations or RL. It first predicts an interaction using an affordance network trained on self-supervised interaction data or videos of human manipulation. After executing the interactions on the real robot to move the object parts, we propose a novel modeling pipeline based on 3D AIGC to build a digital twin of the object in simulation from multiple frames of observations. For dexterous hands, we utilize eigengrasp to reduce the action dimension, enabling more efficient trajectory searching. Experiments validate the framework's effectiveness for precise manipulation using a suction gripper, a two-finger gripper and two dexterous hands. The generalizability of the explicit world model also enables advanced manipulation strategies like manipulating with tools. Taoran Jiang, Liqian Ma, Jing Xu 0011, Jiaojiao Meng, Weihang Chen, Zecui Zeng, Lusong Li, Dan Wu 0008, Rui Chen 0019 |
IEEE Trans. Robotics | 3 |
| 2024 | PTUS: Photo-Realistic Talking Upper-Body Synthesis via 3D-Aware Motion Decomposition WarpingabstractTalking upper-body synthesis is a promising task due to its versatile potential for video creation and consists of animating the body and face from a source image with the motion from a given driving video. However, prior synthesis approaches fall short in addressing this task and have been either limited to animating heads of a target person only, or have animated the upper body but neglected the synthesis of precise facial details. To tackle this task, we propose a Photo-realistic Talking Upper-body Synthesis method via 3D-aware motion decomposition warping, named PTUS, to both precisely synthesize the upper body as well as recover the details of the face such as blinking and lip synchronization. In particular, the motion decomposition mechanism consists of a face-body motion decomposition, which decouples the 3D motion estimation of the face and body, and a local-global motion decomposition, which decomposes the 3D face motion into global and local motions resulting in the transfer of facial expression. The 3D-aware warping module transfers the large-scale and subtle 3D motions to the extracted 3D depth-aware features in a coarse-tofine manner. Moreover, we present a new dataset, Talking-UB, which includes upper-body images with high-resolution faces, addressing the limitations of prior datasets that either consist of only facial images or upper-body images with blurry faces. Experimental results demonstrate that our proposed method can synthesize high-quality videos that preserve facial details, and achieves superior results compared to state-of-the-art cross-person motion transfer approaches. Code and collected dataset are released in https://github.com/cooluoluo/PTUS. Luoyang Lin, Zutao Jiang, Xiaodan Liang, Liqian Ma, Michael Kampffmeyer, Xiaochun Cao |
AAAI | 4 |
| 2023 | GM-NeRF: Learning Generalizable Model-Based Neural Radiance Fields from Multi-View ImagesabstractIn this work, we focus on synthesizing high-fidelity novel view images for arbitrary human performers, given a set of sparse multi-view images. It is a challenging task due to the large variation among articulated body poses and heavy self-occlusions. To alleviate this, we introduce an effective generalizable framework Generalizable Model-based Neural Radiance Fields (GM-NeRF) to synthesize free-viewpoint images. Specifically, we propose a geometry-guided attention mechanism to register the appearance code from multi-view 2D images to a geometry proxy which can alleviate the misalignment between inaccurate geometry prior and pixel space. On top of that, we further conduct neural rendering and partial gradient back propagation for efficient perceptual supervision and improvement of the perceptual quality of synthesis. To evaluate our method, we conduct experiments on synthesized datasets THuman2.0 and Multi-garment, and real-world datasets Genebody and ZJUMocap. The results demonstrate that our approach outperforms state-of-the-art methods in terms of novel view synthesis and geometric reconstruction. Jianchuan Chen, Wentao Yi, Liqian Ma, Xu Jia 0012, Huchuan Lu |
CVPR | 3 |
| 2023 | Cloth2Body: Generating 3D Human Body Mesh from 2D ClothingabstractIn this paper, we define and study a new Cloth2Body problem which has a goal of generating 3d human body meshes from a 2D clothing image. Unlike the existing human mesh recovery problem, Cloth2Body needs to address new and emerging challenges raised by the partial observation of the input and the high diversity of the output. Indeed, there are three specific challenges. First, how to locate and pose human bodies into the clothes. Second, how to effectively estimate body shapes out of various clothing types. Finally, how to generate diverse and plausible results from a 2D clothing image. To this end, we propose an end-to-end framework that can accurately estimate 3D body mesh parameterized by pose and shape from a 2D clothing image. Along this line, we first utilize Kinematics-aware Pose Estimation to estimate body pose parameters. 3D skeleton is employed as a proxy followed by an inverse kinematics module to boost the estimation accuracy. We additionally design an adaptive depth trick to align the re-projected 3D mesh better with 2D clothing image by disentangling the effects of object size and camera extrinsic. Next, we propose Physics-informed Shape Estimation to estimate body shape parameters. 3D shape parameters are predicted based on partial body measurements estimated from RGB image, which not only improves pixel-wise human-cloth alignment, but also enables flexible user editing. Finally, we design Evolution-based pose generation method, a skeleton transplanting method inspired by genetic algorithms to generate diverse reasonable poses during inference. As shown by experimental results on both synthetic and real-world data, the proposed framework achieves state-of-the-art performance and can effectively recover natural and diverse 3D body meshes from 2D images that align well with clothing. Lu Dai 0001, Liqian Ma, Shenhan Qian, Hao Liu 0026, Ziwei Liu 0002, Hui Xiong 0001 |
ICCV | 2 |
| 2023 | Sim2Real2: Actively Building Explicit Physics Model for Precise Articulated Object ManipulationabstractAccurately manipulating articulated objects is a challenging yet important task for real robot applications. In this paper, we present a novel framework called Sim2Real2to enable the robot to manipulate an unseen articulated object to the desired state precisely in the real world with no human demonstrations. We leverage recent advances in physics simulation and learning-based perception to build the interactive explicit physics model of the object and use it to plan a long-horizon manipulation trajectory to accomplish the task. However, the interactive model cannot be correctly estimated from a static observation. Therefore, we learn to predict the object affordance from a single-frame point cloud, control the robot to actively interact with the object with a one-step action, and capture another point cloud. Further, the physics model is constructed from the two point clouds. Experimental results show that our framework achieves about 70% manipulations with < 30% relative error for common articulated objects, and 30% manipulations for difficult objects. Our proposed framework also enables advanced manipulation strategies, such as manipulating with different tools. Code and videos are available on our project webpage: https://ttimelord.github.io/Sim2Real2-site/ Liqian Ma, Jiaojiao Meng, Shuntao Liu, Weihang Chen, Jing Xu 0011, Rui Chen 0019 |
ICRA | 1 |
| 2022 | Controllable Radiance Fields for Dynamic Face SynthesisabstractRecent work on 3D-aware image synthesis has achieved compelling results using advances in neural rendering. However, 3D-aware synthesis of face dynamics hasn't received much attention. Here, we study how to explicitly control generative model synthesis of face dynamics exhibiting non-rigid motion (e.g., facial expression change), while simultaneously ensuring 3D-awareness. For this we propose a Controllable Radiance Field (CoRF): 1) Motion control is achieved by embedding motion features within the layered latent motion space of a style-based generator; 2) To ensure consistency of background, motion features and subject-specific attributes such as lighting, texture, shapes, albedo, and identity, a face parsing net, a head regressor and an identity encoder are incorporated. On head image/video data we show that CoRFs are 3D-aware while enabling editing of identity, viewing directions, and motion. Peiye Zhuang, Liqian Ma, Oluwasanmi Koyejo, Alexander G. Schwing |
3DV | 2 |
| 2022 | UNIF: United Neural Implicit Functions for Clothed Human Reconstruction and Animation
Shenhan Qian, Ziwei Liu 0002, Liqian Ma, Shenghua Gao |
ECCV (3) | 4 |
| 2021 | Direct Dense Pose EstimationabstractDense human pose estimation is the problem of learning dense correspondences between RGB images and the surfaces of human bodies, which finds various applications, such as human body reconstruction, human pose transfer, and human action recognition. Prior dense pose estimation methods are all based on Mask R-CNN framework and operate in a top-down manner of first attempting to identify a bounding box for each person and matching dense correspondences in each bounding box. Consequently, these methods lack robustness due to their critical dependence on the Mask R-CNN detection, and the runtime increases drastically as the number of persons in the image increases. We therefore propose a novel alternative method for solving the dense pose estimation problem, called Direct Dense Pose (DDP). DDP first predicts the instance mask and global IUV representation separately and then combines them together. We also propose a simple yet effective 2D temporal-smoothing scheme to alleviate the temporal jitters when dealing with video data. Experiments demonstrate that DDP overcomes the limitations of previous top-down baseline methods and achieves competitive accuracy. In addition, DDP is computationally more efficient than previous dense pose estimation methods, and it reduces jitters when applied to a video sequence, which is a problem plaguing the previous methods. Liqian Ma, Lingjie Liu, Christian Theobalt, Luc Van Gool |
3DV | 1 |
| 2020 | Unselfie: Translating Selfies to Neutral-Pose Portraits in the Wild
Liqian Ma, Zhe Lin 0001, Connelly Barnes, Alexei A. Efros, Jingwan Lu |
ECCV (17) | 1 |
| 2020 | Unpaired Image-To-Image Shape Translation Across Fashion DataabstractWe address the problem of unpaired geometric image-to-image translation. Rather than transferring the style of an image as a whole, our goal is to translate the geometry of an object while preserving its appearance. Our model is trained without the need for paired images. It performs all steps of the shape transfer within a single model and without additional post-processing stages. Experiments on clothing-based datasets show the effectiveness of the proposed method. Liqian Ma, José Oramas M., Luc Van Gool, Tinne Tuytelaars |
ICIP | 2 |
| 2019 | Exemplar Guided Unsupervised Image-to-Image Translation with Semantic Consistency
Liqian Ma, Xu Jia 0012, Stamatios Georgoulis, Tinne Tuytelaars, Luc Van Gool |
ICLR (Poster) | 1 |
| 2019 | Expressive facial style transfer for personalized memes mimic
Yanlong Tang, Xiaoguang Han 0001, Yue Li 0049, Liqian Ma, Ruofeng Tong 0001 |
Vis. Comput. | 4 |
| 2018 | Customized Multi-person Tracker
Liqian Ma, Siyu Tang 0001, Michael J. Black, Luc Van Gool |
ACCV (2) | 1 |
| 2018 | Disentangled Person Image GenerationabstractGenerating novel, yet realistic, images of persons is a challenging task due to the complex interplay between the different image factors, such as the foreground, background and pose information. In this work, we aim at generating such images based on a novel, two-stage reconstruction pipeline that learns a disentangled representation of the aforementioned image factors and generates novel person images at the same time. First, a multi-branched reconstruction network is proposed to disentangle and encode the three factors into embedding features, which are then combined to re-compose the input image itself. Second, three corresponding mapping functions are learned in an adversarial manner in order to map Gaussian noise to the learned embedding feature space, for each factor, respectively. Using the proposed framework, we can manipulate the foreground, background and pose of the input image, and also sample new embedding features to generate such targeted manipulations, that provide more control over the generation process. Experiments on the Market-1501 and Deepfashion datasets show that our model does not only generate realistic person images with new foregrounds, backgrounds and poses, but also manipulates the generated factors and interpolates the in-between states. Another set of experiments on Market-1501 shows that our model can also be beneficial for the person re-identification task1. Liqian Ma, Qianru Sun, Stamatios Georgoulis, Luc Van Gool, Bernt Schiele, Mario Fritz |
CVPR | 1 |
| 2018 | Natural and Effective Obfuscation by Head InpaintingabstractAs more and more personal photos are shared online, being able to obfuscate identities in such photos is becoming a necessity for privacy protection. People have largely resorted to blacking out or blurring head regions, but they result in poor user experience while being surprisingly ineffective against state of the art person recognizers [17]. In this work, we propose a novel head inpainting obfuscation technique. Generating a realistic head inpainting in social media photos is challenging because subjects appear in diverse activities and head orientations. We thus split the task into two sub-tasks: (1) facial landmark generation from image context (e.g. body pose) for seamless hypothesis of sensible head pose, and (2) facial landmark conditioned head inpainting. We verify that our inpainting method generates realistic person images, while achieving superior obfuscation performance against automatic person recognizers. Qianru Sun, Liqian Ma, Seong Joon Oh, Luc Van Gool, Bernt Schiele, Mario Fritz |
CVPR | 2 |
| 2017 | Pose Guided Person Image GenerationabstractThis paper proposes the novel Pose Guided Person Generation Network (PG$^2$) that allows to synthesize person images in arbitrary poses, based on an image of that person and a novel pose. Our generation framework PG$^2$ utilizes the pose information explicitly and consists of two key stages: pose integration and image refinement. In the first stage the condition image and the target pose are fed into a U-Net-like network to generate an initial but coarse image of the person with the target pose. The second stage then refines the initial and blurry result by training a U-Net-like generator in an adversarial way. Extensive experimental results on both 128$\times$64 re-identification images and 256$\times$256 fashion photos show that our model generates high-quality person images with convincing details. Liqian Ma, Xu Jia 0012, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, Luc Van Gool |
NIPS | 1 |
| 2016 | A novel hierarchical Bag-of-Words model for compact action representation
Qianru Sun, Hong Liu 0008, Liqian Ma, Tianwei Zhang 0002 |
Neurocomputing | 3 |
| 2015 | Body-structure based feature representation for person re-identificationabstractPerson re-identification is valuable for intelligent video surveillance and has drawn wide attention. Although person re-identification research is making progress, it still faces some challenges such as varying poses, illumination and viewpoints. As a major aspect of person re-identification, feature representation has been widely researched. Low-level descriptors are generally used in existing works, which do not take full advantage of body structure information and result in low discrimination. In this paper, body-structure based mid-level feature representation is proposed, which introduces body structure pyramid for codebook learning and feature pooling. Additionally, low computational LLC is used to encode mid-level features. Experimental results on two challenging datasets VIPeR and CUHK01 have demonstrated that our approach outperforms the state-of-the-art methods. Hong Liu 0008, Liqian Ma, Can Wang 0006 |
ICASSP | 2 |
| 2015 | Online person orientation estimation based on classifier updateabstractPerson orientation estimation is valuable for intelligent video surveillance. Although much progress has been made in recent years, it still faces challenges such as varying poses, illuminations and viewpoints. Most existing approaches merely use appearance information or combine it with motion information. Appearance-based classifiers are trained offline without updating in real time, which can not adapt to unknown scenes. To fix it, a novel orientation estimation approach based on online appearance-based classifier update is proposed. Reliable motion direction is determined acting as pre-estimated person orientation to update the appearance-based classifier. Moreover, a novel criterion based on motion reliability is proposed to determine the motion direction. Experimental results show that the proposed approach achieves more competitive performances especially for unknown scenes. Hong Liu 0008, Liqian Ma |
ICIP | 2 |
| 2014 | Depth Motion Detection - A Novel RS-Trigger Temporal Logic based MethodabstractRecently, depth data is widely used in computer vision applications such as detection and tracking, which shows great promises in complicated environments due to its complementary natures to RGB data. However, previous works mostly use depth as an auxiliary cue of RGB data and overlook its inherent advantage on motion detection. Intrinsically different from RGB data, points in depth map essentially represents 3-D positions in the world, so depth video represents the variation of these “positions,” which is motion. Motivated by this, we proposed a novel motion detection scheme based on RS-Trigger temporal logic which best fits nature of depth data on motion detection. The proposed algorithm can fast detect motion regions in the scene without statistics of background and prior knowledge of objects to detect. In following refinement modules, a depth-invariant density-constant projection is proposed which contributes to a fast spatial clustering and accurate segmentation, for it transforms dense 3-D points cloud to depth-invariant 2-D map with density-constance, not only it overcomes depth-dependent sampling of depth sensor, but also overcomes the common ‘scale problem’ in 2-D image analysis, which makes it easy to set system parameters to de-noise and pop-out motion regions. Experimental results validate its effectiveness and efficiency. Can Wang 0006, Hong Liu 0008, Liqian Ma |
IEEE Signal Process. Lett. | 3 |