Qingmin Liao

dblp:13/322 · DBLP profile ↗
← Back
222ranked-venue papers
1as first author
105since 2021 · last 2026
0000-0002-7509-3964ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 142 · 1 first-author · 55 since 2021Artificial intelligence and machine learning · 69 · 47 since 2021Databases, data management, data science and information retrieval · 14 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AgentSwift: Efficient LLM Agent Design via Value-Guided Hierarchical Search
abstract
Large language model (LLM) agents have demonstrated strong capabilities across diverse domains, yet automated agent design remains a significant challenge. Current automated agent design approaches are often constrained by limited search spaces that primarily optimize workflows but fail to integrate crucial human-designed components like memory, planning, and tool use. Furthermore, these methods are hampered by high evaluation costs, as evaluating even a single new agent on a benchmark can require tens of dollars. The difficulty of this exploration is further exacerbated by inefficient search strategies that struggle to navigate the large design space effectively, making the discovery of novel agents a slow and resource-intensive process. To address these challenges, we propose AgentSwift, a novel framework for automated agent design. We formalize a hierarchical search space that jointly models agentic workflow and composable functional components. This structure moves beyond optimizing workflows alone by co-optimizing functional components, which enables the discovery of more complex and effective agent architectures. To make exploration within this expansive space feasible, we mitigate high evaluation costs by training a value model on a high-quality dataset, generated via a novel strategy combining combinatorial coverage and balanced Bayesian sampling for low-cost evaluation. Guiding the entire process is a hierarchical Monte Carlo Tree Search (MCTS) strategy, which is informed by uncertainty to efficiently navigate the search space. Evaluated across a comprehensive set of seven benchmarks spanning embodied, math, web, tool, and game domains, AgentSwift discovers agents that achieve an average performance gain of 8.34\% over both existing automated agent search methods and manually designed agents. Moreover, our framework exhibits steeper and more stable search trajectories. By enabling the efficient, automated composition of workflow with functional components, AgentSwift provides a scalable methodology to explore complex agent designs. Our framework serves as a launchpad for researchers to rapidly prototype and discover powerful agent architectures without the impediment of prohibitive evaluation costs.
Yu Li 0022, Lehui Li, Qingmin Liao, Jianye Hao, Kun Shao, Fengli Xu
AAAI4
2026 WeightFlow: Learning Stochastic Dynamics via Evolving Weight of Neural Network
abstract
Modeling stochastic dynamics from discrete observations is a key interdisciplinary challenge. Existing methods often fail to estimate the continuous evolution of probability densities from trajectories or face the curse of dimensionality. To address these limitations, we presents a novel paradigm: modeling dynamics directly in the weight space of a neural network by projecting the evolving probability distribution. We first theoretically establish the connection between dynamic optimal transport in measure space and an equivalent energy functional in weight space. Subsequently, we design WeightFlow, which constructs the neural network weights into a graph and learns its evolution via a graph controlled differential equation. Experiments on interdisciplinary datasets show that WeightFlow improves performance by an average of 43.02\% over state-of-the-art methods, providing an effective and scalable solution for modeling high-dimensional stochastic dynamics.
Ruikun Li 0002, Huandong Wang, Qingmin Liao, Yong Li 0008
AAAI4
2026 AXFL: Axial prior-guided cross-view fusion learning for radar semantic segmentation
Liwen Zhang 0001, Youcheng Zhang, Qingmin Liao
Expert Syst. Appl.4
2026 TOFFNet: A Texture Orientation-based Feature Fusion Network for contactless multimodal finger recognition
Zishuang Wang, Jiapeng Lin, Wenming Yang, Qingmin Liao
Pattern Recognit.5
2026 BDC-Occ: Binarized Deep Convolution Unit for Binarized Occupancy Network
abstract
Existing 3D occupancy networks demand significant hardware resources, hindering the deployment of resource-limited devices. Binarized Neural Networks (BNNs) offer a potential solution by substantially reducing computational and memory requirements. However, their performance decrease notably compared to full-precision networks. In addition, it is challenging to enhance the performance of the binarized model by increasing the number of binarized convolutional layers, which limits its practicability for 3D occupancy prediction. In this paper, we reconsider the components in binarized convolutional layers, and structures, for 3D occupancy prediction task. Two original insights into binarized convolution are presented, substantiated with theoretical proofs: (a) 1×1 binarized convolution introduces minimal binarization errors as the network deepens, and (b) binarized convolution is inferior to full-precision convolution in capturing cross-channel feature importance. Building on the above insights, we propose a novel binarized deep convolution (BDC) unit that significantly enhances performance, even when the number of binarized convolutional layers increases to meet the requirements of 3D occupancy networks. Specifically, in the BDC unit, additional binarized convolutional kernels are constrained to 1×1 to minimize the effects of binarization errors. Further, we propose a per-channel refinement branch to reweight the output via first-order approximation. Then, we partition the 3D occupancy networks into four distinct convolutional modules, employing BDC units to explore the effects of binarizing each of these modules. The proposed BDC unit minimizes binarization errors and improves perceptual capability, meeting the stringent requirements for accuracy and computational efficiency in 3D occupancy prediction. Extensive quantitative and qualitative experiments demonstrate that the proposed BDC unit achieves state-of-the-art performance in 3D occupancy prediction and 3D object detection tasks, while significantly reducing parameters and computational costs. This highlights the potential of the BDC unit as an efficient fundamental component in binarized 3D occupancy networks. Code for our paper will be released on “https://github.com/zzk785089755/BDC”.
Zongkai Zhang, Peng Ling, Zidong Xu, Wenming Yang, Qingmin Liao, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.5
2026 MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Albedo Post-Processing
abstract
Current methods for 3D generation still fall short in physically based rendering (PBR) texturing, primarily due to limited data and challenges in modeling multi-channel materials. In this work, we propose MuMA, a method for 3D PBR texturing through Multi-channel Multi-view generation and Albedo post-processing. Our approach features two key innovations: 1) we opt to model shaded and albedo appearance channels, where the shaded channels enables the integration intrinsic decomposition modules for material properties; and 2) leveraging multimodal large language models, we emulate artists' techniques for material assessment and selection. Experiments demonstrate that MuMA achieves superior results in visual quality and material fidelity compared to existing methods.
Lingting Zhu, Jingrui Ye, Zeyu Hu, Yingda Yin, Lanjiong Li, Jinnan Chen, Shengju Qian, Xin Wang 0178, Qingmin Liao, Lequan Yu
IEEE Trans. Image Process.10
2026 Efficient Feature Aggregation and Scale-Aware Regression for Monocular 3-D Object Detection
Fanqi Pu, Qingmin Liao, Wenming Yang
IEEE Trans. Intell. Transp. Syst.4
2025 DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval
abstract
Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person domain is now a emerging research topic due to the abundant knowledge of vision-language pretraining, but challenges still remain during fine-tuning: (i) Previous full-model fine-tuning in TPR is computationally expensive and prone to overfitting.(ii) Existing parameter-efficient transfer learning (PETL) for TPR lacks of fine-grained feature extraction. To address these issues, we propose Domain-Aware Mixture-of-Adapters (DM-Adapter), which unifies Mixture-of-Experts (MOE) and PETL to enhance fine-grained feature representations while maintaining efficiency. Specifically, Sparse Mixture-of-Adapters is designed in parallel to MLP layers in both vision and language branches, where different experts specialize in distinct aspects of person knowledge to handle features more finely. To promote the router to exploit domain information effectively and alleviate the routing imbalance, Domain-Aware Router is then developed by building a novel gating function and injecting learnable domain-aware prompts. Extensive experiments show that our DM-Adapter achieves state-of-the-art performance, outperforming previous methods by a significant margin.
Zimo Liu, Xiangyuan Lan, Wenming Yang, Yaowei Li 0001, Qingmin Liao
AAAI6
2025 Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN Network
abstract
Current state-of-the-art (SOTA) methods in 3D Human Pose Estimation (HPE) are primarily based on Transformers. However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work, we leverage recent advances in state space models and utilize Mamba for high-quality and efficient long-range modeling. Nonetheless, Mamba still faces challenges in precisely exploiting local dependencies between joints. To address these issues, we propose a new attention-free hybrid spatiotemporal architecture named Hybrid Mamba-GCN (Pose Magic). This architecture introduces local enhancement with GCN by capturing relationships between neighboring joints, thus producing new representations to complement Mamba's outputs. By adaptively fusing representations from Mamba and GCN, Pose Magic demonstrates superior capability in learning the underlying 3D structure. To meet the requirements of real-time inference, we also provide a fully causal version. Extensive experiments show that Pose Magic achieves new SOTA results (0.9 mm drop) while saving 74.1% FLOPs. In addition, Pose Magic exhibits optimal motion consistency and the ability to generalize to unseen sequence lengths.
Xinyi Zhang 0008, Qiqi Bao 0001, Qinpeng Cui, Wenming Yang, Qingmin Liao
AAAI5
2025 Predicting the Energy Landscape of Stochastic Dynamical System via Physics-informed Self-supervised Learning
abstract
Energy landscapes play a crucial role in shaping dynamics of many real-world complex systems. System evolution is often modeled as particles moving on a landscape under the combined effect of energy-driven drift and noise-induced diffusion, where the energy governs the long-term motion of the particles. Estimating the energy landscape of a system has been a longstanding interdisciplinary challenge, hindered by the high operational costs or the difficulty of obtaining supervisory signals. Therefore, the question of how to infer the energy landscape in the absence of true energy values is critical. In this paper, we propose a physics-informed self-supervised learning method to learn the energy landscape from the evolution trajectories of the system. It first maps the system state from the observation space to a discrete landscape space by an adaptive codebook, and then explicitly integrates energy into the graph neural Fokker-Planck equation, enabling the joint learning of energy estimation and evolution prediction. Experimental results across interdisciplinary systems demonstrate that our estimated energy has a correlation coefficient above 0.9 with the ground truth, and evolution prediction accuracy exceeds the baseline by an average of 17.65\%. The code is available at https://github.com/tsinghua-fib-lab/PESLA.
Ruikun Li 0002, Huandong Wang, Qingmin Liao, Yong Li 0008
ICLR3
2025 Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network
abstract
Reinforcement learning (RL) for continuous control often requires large amounts of online interaction data. Value-based RL methods can mitigate this burden by offering relatively high sample efficiency. Some studies further enhance sample efficiency by incorporating offline demonstration data to “kick-start” training, achieving promising results in continuous control. However, they typically compute the Q-function independently for each action dimension, neglecting interdependencies and making it harder to identify optimal actions when learning from suboptimal data, such as non-expert demonstration and online-collected data during the training process. To address these issues, we propose Auto-Regressive Soft Q-learning (ARSQ), a value-based RL algorithm that models Q-values in a coarse-to-fine, auto-regressive manner. First, ARSQ decomposes the continuous action space into discrete spaces in a coarse-to-fine hierarchy, enhancing sample efficiency for fine-grained continuous control tasks. Next, it auto-regressively predicts dimensional action advantages within each decision step, enabling more effective decision-making in continuous control tasks. We evaluate ARSQ on two continuous control benchmarks, RLBench and D4RL, integrating demonstration data into online training. On D4RL, which includes non-expert demonstrations, ARSQ achieves an average 1.62$\times$ performance improvement over SOTA value-based baseline. On RLBench, which incorporates expert demonstrations, ARSQ surpasses various baselines, demonstrating its effectiveness in learning from suboptimal online-collected data.
Jijia Liu, Qingmin Liao, Chao Yu 0005, Yu Wang 0002
ICML3
2025 EAY-Net: Edge-Aware Y-Network for Color Guided Depth Map Super-Resolution
Jiamian Bian, Xiaoyu Jin, Qingmin Liao, Wenming Yang
ICONIP (2)4
2025 SAP-SLAM: Semantic-Assisted Perception SLAM with 3D Gaussian Splatting
abstract
The integration of 3D Gaussians has introduced a novel scene representation in Simultaneous Localization and Mapping (SLAM), characterized by explicit representation and differentiable rendering capabilities that enhance scene reconstruction and understanding. However, most current SLAM systems only exploit the basic representational capacity of 3D Gaussians, neglecting their potential to offer richer information and facilitate higher-dimensional scene comprehension. Furthermore, these systems often struggle with reconstruction when encountering rapid camera movements or depth missing. Drawing inspiration from 3D language field, which explores the intrinsic relationships among scene objects, we propose SAPSLAM, a dense SLAM system that combines high-fidelity reconstruction and advanced semantic understanding. Our approach leverages pre-trained visual models to extract semantic features, which are then fused, dimensionally reduced, and encoded into the 3D Gaussian model for optimization and rendering. The integration of these features improves the systems semantic comprehension and scene representation, ultimately enabling the creation of high-precision 3D semantic maps. Additionally, we introduce a semantic-guided Gaussian densification and pruning strategy, which uses semantic consistency to prioritize attention on poorly reconstructed areas, greatly improving performance in complex scenarios. SAP-SLAM achieves competitive results on both real-world and synthetic datasets, demonstrating superior capabilities in semantic understanding and reconstruction.
Yudong Lin, Wenming Yang, Guijin Wang, Qingmin Liao
ICRA5
2025 IDEA-GP: Instruction-Driven Architecture with Efficient Online Workload Allocation for Geometric Perception
abstract
The algorithmic complexity of robotic systems presents significant challenges to achieving generalized acceleration in robot applications.On the one hand, the diversity of operators and computational flows within similar task categories prevents the reuse of specialized computational units.On the other hand, task variations and environmental dynamics can cause workload fluctuations, leading to inefficient resource utilization.This paper focuses on the geometric perception capability of robots, taking localization and mapping as the basic applications, and proposes IDEA-GP, an Instruction-Driven Architecture with Efficient online workload Allocation for Geometric Perception.Built around an array of general computational units designed for spatial positioning representations, IDEA-GP supports a wide range of robot pose-related computational tasks.IDEA-GP employs a compiler to perform online workload analysis and resource allocation.It generates instructions tailored to processing elements (PEs) to schedule computations, thereby accelerating optimization problems and enhancing geometric perception performance.Deployed on the ZCU102 evaluation board, IDEA-GP demonstrates an average speedup of 7.5× over the Intel CPU and 19.7× over the ARM CPU in Simultaneous Localization and Mapping (SLAM) tasks, and a 16.4× speedup over the Intel CPU and 41.6× over the ARM CPU in Structure from Motion (SfM) tasks.
Suquan Zhang, Yunfei Xiang, Yuanfan Xu, Qingmin Liao, Yu Wang 0002
ISCA6
2025 Predicting the Dynamics of Complex System via Multiscale Diffusion Autoencoder
abstract
Predicting the dynamics of complex systems is crucial for various scientific and engineering applications. The accuracy of predictions depends on the model's ability to capture the intrinsic dynamics. While existing methods capture key dynamics by encoding a low-dimensional latent space, they overlook the inherent multiscale structure of complex systems, making it difficult to accurately predict complex spatiotemporal evolution. Therefore, we propose a Multiscale Diffusion Prediction Network (MDPNet) that leverages the multiscale structure of complex systems to discover the latent space of intrinsic dynamics. First, we encode multiscale features through a multiscale diffusion autoencoder to guide the diffusion model for reliable reconstruction. Then, we introduce an attention-based graph neural ordinary differential equation to model the co-evolution across different scales. Extensive evaluations on representative systems demonstrate that the proposed method achieves an average prediction error reduction of 53.23% compared to baselines, while also exhibiting superior robustness and generalization.
Ruikun Li 0002, Jingwen Cheng, Huandong Wang, Qingmin Liao, Yong Li 0008
KDD (2)4
2025 OccGaussian: 3D Gaussian Splatting for Occluded Human Rendering
abstract
Rendering dynamic 3D humans from monocular videos is crucial for various applications such as virtual reality and digital entertainment. Most methods assume the human is in an unobstructed scene, while various objects may cause the occlusion of body parts in real-life scenarios. Previous method utilizing NeRF for surface rendering to recover the occluded areas, but it requiring more than one day to train and several seconds to render, failing to meet the requirements of real-time interactive applications. To address these issues, we propose OccGaussian based on 3D Gaussian Splatting, which can be trained within 6 minutes and produces high-quality human renderings up to 160 FPS with occluded input. OccGaussian initializes 3D Gaussian distributions in the canonical space, and we perform occlusion feature query at occluded regions, the aggregated pixel-align feature is extracted to compensate for the missing information. Then we use Gaussian Feature MLP to further process the aggregated feature, along with the specially designed occlusion-aware loss functions to better perceive the occluded area. Extensive experiments both in simulated and real-world occlusions, demonstrate that our method achieves superior performance compared to the state-of-the-art method. And we improving training and inference speeds by 250x and 800x, respectively. Our code will be available for research purposes.
Jingrui Ye, Qingmin Liao
ICMR3
2025 LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language Models
abstract
Policy exploration is critical in reinforcement learning (RL), where existing approaches include $\epsilon$-greedy, Gaussian process, etc. However, these approaches utilize preset stochastic processes and are indiscriminately applied in all kinds of RL tasks without considering task-specific features that influence policy exploration. Moreover, during RL training, the evolution of such stochastic processes is rigid, which typically only incorporates a decay in the variance, failing to adjust flexibly according to the agent's real-time learning status. Inspired by the analyzing and reasoning capability of large language models (LLMs), we design **LLM-Explorer** to adaptively generate task-specific exploration strategies with LLMs, enhancing the policy exploration in RL. In our design, we sample the learning trajectory of the agent during the RL training in a given task and prompt the LLM to analyze the agent's current policy learning status and then generate a probability distribution for future policy exploration. Updating the probability distribution periodically, we derive a stochastic process specialized for the particular task and dynamically adjusted to adapt to the learning process. Our design is a plug-in module compatible with various widely applied RL algorithms, including the DQN series, DDPG, TD3, and any possible variants developed based on them. Through extensive experiments on the Atari and MuJoCo benchmarks, we demonstrate LLM-Explorer's capability to enhance RL policy exploration, achieving an average performance improvement up to 37.27%. Our code is open-source at https://github.com/tsinghua-fib-lab/LLM-Explorer for reproducibility.
Qianyue Hao, Yiwen Song, Qingmin Liao, Yong Li 0008
NeurIPS3
2025 What Can RL Bring to VLA Generalization? An Empirical Study
abstract
Large Vision-Language Action (VLA) models have shown significant potential for embodied AI. However, their predominant training via supervised fine-tuning (SFT) limits generalization due to susceptibility to compounding errors under distribution shifts. Reinforcement learning (RL) offers a path to overcome these limitations by optimizing for task objectives via trial-and-error, yet a systematic understanding of its specific generalization benefits for VLAs compared to SFT is lacking. To address this, our study introduces a comprehensive benchmark for evaluating VLA generalization and systematically investigates the impact of RL fine-tuning across diverse visual, semantic, and execution dimensions. Our extensive experiments reveal that RL fine-tuning, particularly with PPO, significantly enhances generalization in semantic understanding and execution robustness over SFT, while maintaining comparable visual robustness. We identify PPO as a more effective RL algorithm for VLAs than LLM-derived methods like DPO and GRPO. We also develop a simple recipe for efficient PPO training on VLAs, and demonstrate its practical utility for improving VLA generalization. The project page is at https://rlvla.github.io
Jijia Liu, Bingwen Wei, Xinlei Chen, Qingmin Liao, Yi Wu 0013, Chao Yu 0005, Yu Wang 0002
NeurIPS5
2025 Elucidating the Solution Space of Extended Reverse-Time SDE for Diffusion Models
abstract
Sampling from Diffusion Models can alternatively be seen as solving differential equations, where there is a challenge in balancing speed and image visual quality. ODE-based samplers offer rapid sampling time but reach a performance limit, whereas SDE-based samplers achieve superior quality, albeit with longer iterations. In this work, we formulate the sampling process as an Extended Reverse-Time SDE (ER SDE), unifying prior explorations into ODEs and SDEs. Theoretically, leveraging the semi-linear structure of ER SDE solutions, we offer exact solutions and approximate solutions for VP SDE and VE SDE, respectively. Based on the approximate solution space of the ER SDE, referred to as one-step prediction errors, we yield mathematical insights elucidating the rapid sampling capability of ODE solvers and the high-quality sampling ability of SDE solvers. Additionally, we unveil that VP SDE solvers stand on par with their VE SDE counterparts. Based on these findings, leveraging the dual advantages of ODE solvers and SDE solvers, we devise efficient high-quality samplers, namely ER-SDE-Solvers. Experimental results demonstrate that ER-SDE-Solvers achieve state-of-the-art performance across all stochastic samplers while maintaining efficiency of deterministic samplers. Specifically, on the ImageNet 128 × 128 dataset, ER-SDE-Solvers obtain 8.33 FID in only 20 function evaluations. Code is available at https://github.com/QinpengCui/ER-SDE-Solver
Qinpeng Cui, Xinyi Zhang 0008, Qiqi Bao 0001, Qingmin Liao
WACV4
2025 UV Gaussians: Joint learning of mesh deformation and Gaussian textures for human avatar modeling
Yujiao Jiang, Qingmin Liao, Xiaoyu Li 0002, Qi Zhang 0029, Chaopeng Zhang, Zongqing Lu 0001, Ying Shan
Knowl. Based Syst.2
2025 UP-Person: Unified Parameter-Efficient Transfer Learning for Text-Based Person Retrieval
abstract
Text-based Person Retrieval (TPR) as a multi-modal task, which aims to retrieve the target person from a pool of candidate images given a text description, has recently garnered considerable attention due to the progress of contrastive visual-language pre-trained model. Prior works leverage pre-trained CLIP to extract person visual and textual features and fully fine-tune the entire network, which have shown notable performance improvements compared to uni-modal pre-training models. However, full-tuning a large model is prone to overfitting and hinders the generalization ability. In this paper, we propose a novelUnifiedParameter-Efficient Transfer Learning (PETL) method for Text-basedPersonRetrieval (UP-Person) to thoroughly transfer the multi-modal knowledge from CLIP. Specifically, UP-Person simultaneously integrates three lightweight PETL components including Prefix, LoRA and Adapter, where Prefix and LoRA are devised together to mine local information with task-specific information prompts, and Adapter is designed to adjust global feature representations. Additionally, two vanilla submodules are optimized to adapt to the unified architecture of TPR. For one thing, S-Prefix is proposed to boost attention of prefix and enhance the gradient propagation of prefix tokens, which improves the flexibility and performance of the vanilla prefix. For another thing, L-Adapter is designed in parallel with layer normalization to adjust the overall distribution, which can resolve conflicts caused by overlap and interaction among multiple submodules. Extensive experimental results demonstrate that our UP-Person achieves state-of-the-art results across various person retrieval datasets, including CUHK-PEDES, ICFG-PEDES and RSTPReid while merely fine-tuning 4.7% parameters. Code is available at https://github.com/Liu-Yating/UP-Person.
Yaowei Li 0001, Xiangyuan Lan, Wenming Yang, Zimo Liu, Qingmin Liao
IEEE Trans. Circuits Syst. Video Technol.6
2025 Controllable Human Trajectory Generation Using Profile-Guided Latent Diffusion
abstract
Trajectory generation is a vital element in AI applications. Firstly, it enables simulation such as traffic simulation and epidemic spreading modeling. Secondly, it can provide synthetic privacy-preserving data for training AI models. Notably, trajectory generation featuring controllable user profiles holds substantial value in generating customized mobility trajectories tailored to diverse requirements. However, relevant work is still lacking. On the one hand, traditional deep generative models fall short in guiding controllable trajectory generation due to the statistical nature of human mobility patterns and the corresponding insufficient control mechanisms. On the other hand, though the diffusion model has demonstrated strong generative capabilities in many fields, to achieve controllable generation on discrete trajectory data, we still need to redesign the structure of the continuous diffusion model. In this article, we introduce a controllable trajectory generation framework that leverages a continuous diffusion model and classifier guidance for more robust condition control. Our proposed framework comprises two modules: a latent trajectory diffusion model and a trajectory classifier for profile guidance. Experiments on two real-world mobility datasets consistently demonstrate its capability of generating trajectories matching given user profiles and conforming to human mobility patterns. Our source code and trained models are released at https://github.com/tsinghua-fib-lab/User-Profile-Guided-Latent-Diffusion .
Yiwen Song, Jingtao Ding, Qingmin Liao, Yong Li 0008
ACM Trans. Knowl. Discov. Data4
2024 UV-SAM: Adapting Segment Anything Model for Urban Village Identification
abstract
Urban villages, defined as informal residential areas in or around urban centers, are characterized by inadequate infrastructures and poor living conditions, closely related to the Sustainable Development Goals (SDGs) on poverty, adequate housing, and sustainable cities. Traditionally, governments heavily depend on field survey methods to monitor the urban villages, which however are time-consuming, labor-intensive, and possibly delayed. Thanks to widely available and timely updated satellite images, recent studies develop computer vision techniques to detect urban villages efficiently. However, existing studies either focus on simple urban village image classification or fail to provide accurate boundary information. To accurately identify urban village boundaries from satellite images, we harness the power of the vision foundation model and adapt the Segment Anything Model (SAM) to urban village segmentation, named UV-SAM. Specifically, UV-SAM first leverages a small-sized semantic segmentation model to produce mixed prompts for urban villages, including mask, bounding box, and image representations, which are then fed into SAM for fine-grained boundary identification. Extensive experimental results on two datasets in China demonstrate that UV-SAM outperforms existing baselines, and identification results over multiple years show that both the number and area of urban villages are decreasing over time, providing deeper insights into the development trends of urban villages and sheds light on the vision foundation models for sustainable cities. The dataset and codes of this study are available at https://github.com/tsinghua-fib-lab/UV-SAM.
Xin Zhang 0106, Yu Liu 0016, Yuming Lin 0003, Qingmin Liao, Yong Li 0008
AAAI4
2024 EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities
abstract
The advent of artificial intelligence has led to a growing emphasis on data-driven modeling in macroeconomics, with agent-based modeling (ABM) emerging as a prominent bottom-up simulation paradigm.In ABM, agents (e.g., households, firms) interact within a macroeconomic environment, collectively generating market dynamics.Existing agent modeling typically employs predetermined rules or learningbased neural networks for decision-making.However, customizing each agent presents significant challenges, complicating the modeling of agent heterogeneity.Additionally, the influence of multi-period market dynamics and multifaceted macroeconomic factors are often overlooked in decision-making processes.In this work, we introduce EconAgent, a large language model-empowered agent with humanlike characteristics for macroeconomic simulation.We first construct a simulation environment that incorporates various market dynamics driven by agents' decisions regarding work and consumption.Through the perception module, we create heterogeneous agents with distinct decision-making mechanisms.Furthermore, we model the impact of macroeconomic trends using a memory module, which allows agents to reflect on past individual experiences and market dynamics.Simulation experiments show that EconAgent can make realistic decisions, leading to more reasonable macroeconomic phenomena compared to existing rule-based or learning-based agents.
Nian Li 0001, Chen Gao 0001, Yong Li 0008, Qingmin Liao
ACL (1)5
2024 IRGen: Generative Modeling for Image Retrieval
Ting Zhang 0002, Dong Chen 0003, Yujing Wang 0002, Qi Chen 0009, Xing Xie 0001, Hao Sun 0015, Qi Zhang 0066, Fan Yang 0024, Mao Yang 0004, Qingmin Liao, Jingdong Wang 0001, Baining Guo
ECCV (15)12
2024 Clip-Based Synergistic Knowledge Transfer for text-based Person Retrieval
abstract
Text-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and language modalities, especially when dealing with limited large-scale datasets. In this paper, we introduce a CLIP-based Synergistic Knowledge Transfer (CSKT) approach for TPR. Specifically, to explore the CLIP’s knowledge on input side, we first propose a Bidirectional Prompts Transferring (BPT) module constructed by text-to-image and image-to-text bidirectional prompts and coupling projections. Secondly, Dual Adapters Transferring (DAT) is designed to transfer knowledge on output side of Multi-Head Self-Attention (MHA) in vision and language. This synergistic two-way collaborative mechanism promotes the early-stage feature fusion and efficiently exploits the existing knowledge of CLIP. CSKT outperforms the state-of-the-art approaches across three benchmark datasets when the training parameters merely account for 7.4% of the entire model, demonstrating its remarkable efficiency, effectiveness and generalization.
Yaowei Li 0001, Zimo Liu, Wenming Yang, Yaowei Wang 0001, Qingmin Liao
ICASSP6
2024 LR-MAE: Locate while Reconstructing with Masked Autoencoders for Point Cloud Self-supervised Learning
abstract
As an efficient self-supervised pre-training approach, Masked autoencoder (MAE) has shown promising improvement across various 3D point cloud understanding tasks. However, the pretext task of existing point-based MAE is to reconstruct the geometry of masked points only, hence it learns features at lower semantic levels which is not appropriate for high-level downstream tasks. To address this challenge, we propose a novel self-supervised approach named Locate while Reconstructing with Masked Autoencoders (LR-MAE). Specifically, a multi-head decoder is designed to simultaneously localize the global position of masked patches while reconstructing masked points, aimed at learning better semantic features that align with downstream tasks. Moreover, we design a random query patch detection strategy for 3D object detection tasks in the pre-training stage, which significantly boosts the model performance with faster convergence speed. Extensive experiments show that our LR-MAE achieves superior performance on various point cloud understanding tasks. By fine-tuning on downstream datasets, LR-MAE outperforms the Point-MAE baseline by 3.65% classification accuracy on the ScanObjectNN dataset, and significantly exceeds the 3DETR baseline by 6.1% AP50on the ScanNetV2 dataset. Code is available at https://github.com/cathy-ji/LR-MAE.
Huizhen Ji, Yaohua Zha, Qingmin Liao
ICME3
2024 SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations
abstract
Recovering photorealistic and drivable full-body avatars is crucial for numerous applications, including virtual reality, 3D games, and tele-presence. Most methods, whether reconstruction or generation, require large numbers of human motion sequences and corresponding textured meshes. To easily learn a drivable avatar, a reasonable parametric body model with unified topology is paramount. However, existing human body datasets either have images or textured models and lack parametric models which fit clothes well. We propose a new parametric model SMPLX-Lite-D, which can fit detailed geometry of the scanned mesh while maintaining stable geometry in the face, hand and foot regions. We present SMPLX-Lite dataset, the most comprehensive clothing avatar dataset with multi-view RGB sequences, keypoints annotations, textured scanned meshes, and textured SMPLX-Lite-D models. With the SMPLX-Lite dataset, we train a conditional variational autoencoder model that takes human pose and facial keypoints as input, and generates a photorealistic drivable human avatar.
Yujiao Jiang, Qingmin Liao, Xiangru Lin, Zongqing Lu 0001, Yuxi Zhao, Hanqing Wei, Jingrui Ye, Yu Zhang 0166, Zhijing Shao
ICME2
2024 Long-term Detection and Monitory of Chinese Urban Village Using Satellite Imagery
Yuming Lin 0003, Xin Zhang 0106, Yu Liu 0016, Zhenyu Han, Qingmin Liao, Yong Li 0008
IJCAI5
2024 Reschedule Diffusion-based Bokeh Rendering
Shiyue Yan, Xiaoshi Qiu, Qingmin Liao, Jing-Hao Xue
IJCAI3
2024 MetaMask: Improving Few-Shot Semantic Segmentation via Multi-Mask Calibriation
abstract
Few-shot Semantic Segmentation (FSS) aims to develop models that can segment previously unseen classes with only a few annotations. Recent approaches employ a "multi-mask" framework, which initially generates various mask proposals from query images and then matches related mask proposals to get the final output guided by support images. Despite its promise, this framework is limited by the quality of mask proposals for unseen classes and a naive mask matching process. To address such limitations, in this paper, we propose a meta-learning-based method called MetaMask. First, MetaMask builds a Support-Guided Latent Object Segmenter (SG-LOS) module, which incorporates unseen class information into mask proposal generation for query images, where episodic training is used to enhance mask generation for latent unseen classes. Second, MetaMask improves the mask-matching mechanism through our proposed Contrastive Mask Matching (CMM) module with a cross-image multi-level contrastive learning strategy, bolstering feature embedding spaces. Our method shows competitive results on two main benchmarks: 69.9% mIoU on Pascal-5ione-shot setting and 49.6% mIoU COCO-20ione-shot setting, marginally outperforming our baseline by 6.6% and 5.4%, setting a new state-of-the-art on the both Pascal-5iand COCO-20idatasets.
Li Dinghang, Zongqing Lu 0001, Weiliang Zheng, Qingmin Liao, Fan Lyu
IJCNN4
2024 Multi-Dimensional Attention on Cost Volume for Stereo Matching
abstract
Stereo matching is a fundamental research topic in computer vision tasks, and the careful processing of cost volume plays a vital role in stereo matching solutions. Previous convolutional networks have deep-layer structures but could only aggregate local regions, leading to suboptimal matching performance in areas with edges or weak textures, etc. Considering the global perception capability of the attention mechanism, we for the first time propose global attention modules directly operating on the cost volume for cost aggregation. Our proposed attention module is named Multi-Dimensional Attention (MDA) and it includes two submodules: the Cross-Disparity Attention (CDA) and the Intra-Disparity Attention (IDA). CDA accomplishes cost aggregation under different disparities, and IDA is further categorized into Channel-Wise Attention (CWA) and Disparity-Wise Attention (DWA), focusing on the similarity of structure and disparity variations within a fixed disparity. For evaluation, we conduct experiments on four publicly available datasets including KITTI 2012, KITTI 2015, Scene Flow and Middlebury, and results show that our proposed method achieves state-of-the-art (SoTA) performance in stereo matching tasks.
Zhou Jiale, Wenqin Huang, Qingmin Liao, Zongqing Lu 0001
IJCNN3
2024 Cross-Patch Relation Enhanced for Weakly Supervised Semantic Segmentation
abstract
Weakly Supervised Semantic Segmentation (WSSS) using only image-level labels relies on Class Activation Map (CAM) to produce pixel-level pseudo segmentation labels, but it struggles with limited object region activation, resulting in low-quality annotations. To address this issue, a local-to-global framework is employed to enable the model to capture details from patches randomly cropped from input images. However, the pseudo-masks generated by this approach still have an issue with object incompleteness. We notice that it is caused by the neglect of semantic relations among patches, which capture abundant contextual information. Under this observation, we present a Cross-Patch Relation Enhanced Network to improve the quality of the CAMs, leading to the generation of better pseudo segmentation labels. Specifically, a cross-patch relation attention (including the class-prototype extraction and the class-feature aggregation) is proposed to alleviate the intra-class inconsistency due to variations of contextual information across local patches. The class-prototype extraction module gathers contextual relation from all local class-region embeddings. Besides, class-feature aggregation improves class-level representations of multiple patches through feature aggregation. Extensive experimental results on two public datasets have demonstrated the effectiveness of the proposed method. Our method achieves competitive scores with state-of-the-art methods for weakly supervised semantic segmentation on both PASCAL VOC 2012 and MS-COCO 2014 benchmarks.
Huiqing Su, Wenqin Huang, Qingmin Liao, Zongqing Lu 0001
IJCNN3
2024 Region Motion-based Adaptive Composite Long-Term Reference Coding for VVC
abstract
The adoption of composite long-term reference (CLTR) in versatile video coding (VVC) has demonstrated good performance, especially for the coding of video containing a large amount of stationary background areas. However, if the long-term reference (LTR) picture is constructed by using the frames that contain lots of foreground contents, the composed LTR picture cannot offer enough background information for the coding of target frame. To effectively apply CLTR to video coding, especially to VVC, we propose a region motion-based determination method to adaptively choose long-term and short-term reference pictures for inter-picture coding. The experimental results demonstrate that our proposed method can achieve promising performance improvement compared with the results generated from VVC common test condition and the traditional LTR-based coding scheme.
Xiaozhen Zheng, Yu Liu 0091, Jianglin Wang, Zihao Ren 0002, Shuyuan Zhu, Qingmin Liao
ISCAS6
2024 Predicting Long-term Dynamics of Complex Networks via Identifying Skeleton in Hyperbolic Space
abstract
Learning complex network dynamics is fundamental for understanding, modeling, and controlling real-world complex systems. Though great efforts have been made to predict the future states of nodes on networks, the capability of capturing long-term dynamics remains largely limited. This is because they overlook the fact that long-term dynamics in complex network are predominantly governed by their inherent low-dimensional manifolds, i.e., skeletons. Therefore, we propose the Dynamics-Invariant Skeleton Neural Net}work (DiskNet), which identifies skeletons of complex networks based on the renormalization group structure in hyperbolic space to preserve both topological and dynamics properties. Specifically, we first condense complex networks with various dynamics into simple skeletons through physics-informed hyperbolic embeddings. Further, we design graph neural ordinary differential equations to capture the condensed dynamics on the skeletons. Finally, we recover the skeleton networks and dynamics to the original ones using a degree-based super-resolution module. Extensive experiments across three representative dynamics as well as five real-world and two synthetic networks demonstrate the superior performances of the proposed DiskNet, which outperforms the state-of-the-art baselines by an average of 10.18\% in terms of long-term prediction accuracy. Code for reproduction is available at: https://github.com/tsinghua-fib-lab/DiskNet.
Ruikun Li 0002, Huandong Wang, Jinghua Piao, Qingmin Liao, Yong Li 0008
KDD4
2024 Geometry-Guided Diffusion Model with Masked Transformer for Robust Multi-View 3D Human Pose Estimation
abstract
Recent research on Diffusion Models and Transformers has brought significant advancements to 3D Human Pose Estimation (HPE). Nonetheless, existing methods often fail to concurrently address the issues of accuracy and generalization. In this paper, we propose a Geometry-guided Dif fusion Model with Masked Transformer (Masked Gifformer) for robust multi-view 3D HPE. Within the framework of the diffusion model, a hierarchical multi-view trans-former-based denoiser is exploited to fit the 3D pose distribution by systematically integrating joint and view information. To address the long-standing problem of poor generalization, we introduce a fully random mask mechanism without any additional learnable modules or parameters. Furthermore, we incorporate geometric guidance into the diffusion model to enhance the accuracy of the model. This is achieved by optimizing the sampling process to minimize reprojection errors through modeling a conditional guidance distribution. Extensive experiments on two benchmarks demonstrate that Masked Gifformer effectively achieves a trade-off between accuracy and generalization. Specifically, our method outperforms other probabilistic methods by > 40% and achieves comparable results with state-of-the-art deterministic methods. In addition, our method exhibits robustness to varying camera numbers, spatial arrangements, and datasets.
Xinyi Zhang 0008, Qinpeng Cui, Qiqi Bao 0001, Wenming Yang, Qingmin Liao
ACM Multimedia5
2024 Predicting community case transfer path and processing time using decoder models
abstract
Government agencies and non-profit organizations often rely on case management systems to process the large influx of community request cases. To improve the efficiency of community case management, it's important to model how a community request case is transferred between different departments within the organization and how long it takes to resolve the case. In this paper, we propose two decoder models to predict the departmental transfer path of a given community case and estimate the total processing time based on the predicted path, trained on historical community case records. We compared our prediction results with those obtained using other common machine learning models on a dataset collected from multiple community platforms in Shenzhen, China. Experiments show that our proposed method significantly outperforms the baselines in transfer path and total processing time prediction.
Yuanbo Tang, Qingmin Liao, Yang Li 0104
MobiCom6
2024 Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs
abstract
Diffusion-based image super-resolution (SR) models have attracted substantial interest due to their powerful image restoration capabilities. However, prevailing diffusion models often struggle to strike an optimal balance between efficiency and performance. Typically, they either neglect to exploit the potential of existing extensive pretrained models, limiting their generative capacity, or they necessitate a dozens of forward passes starting from random noises, compromising inference efficiency. In this paper, we present DoSSR, a $\textbf{Do}$main $\textbf{S}$hift diffusion-based SR model that capitalizes on the generative powers of pretrained diffusion models while significantly enhancing efficiency by initiating the diffusion process with low-resolution (LR) images. At the core of our approach is a domain shift equation that integrates seamlessly with existing diffusion models. This integration not only improves the use of diffusion prior but also boosts inference efficiency. Moreover, we advance our method by transitioning the discrete shift process to a continuous formulation, termed as DoS-SDEs. This advancement leads to the fast and customized solvers that further enhance sampling efficiency. Empirical results demonstrate that our proposed method achieves state-of-the-art performance on synthetic and real-world datasets, while notably requiring $\textbf{\emph{only 5 sampling steps}}$. Compared to previous diffusion prior based methods, our approach achieves a remarkable speedup of 5-7 times, demonstrating its superior efficiency.
Qinpeng Cui, Yixuan Liu 0004, Xinyi Zhang 0008, Qiqi Bao 0001, Qingmin Liao, liwang Amd, Zicheng Liu 0001, Zhongdao Wang, Emad Barsoum
NeurIPS5
2024 AdaPKC: PeakConv with Adaptive Peak Receptive Field for Radar Semantic Segmentation
abstract
Deep learning-based radar detection technology is receiving increasing attention in areas such as autonomous driving, UAV surveillance, and marine monitoring. Among recent efforts, PeakConv (PKC) provides a solution that can retain the peak response characteristics of radar signals and play the characteristics of deep convolution, thereby improving the effect of radar semantic segmentation (RSS). However, due to the use of a pre-set fixed peak receptive field sampling rule, PKC still has limitations in dealing with problems such as inconsistency of target frequency domain response broadening, non-homogeneous and time-varying characteristic of noise/clutter distribution. Therefore, this paper proposes an idea of adaptive peak receptive field, and upgrades PKC to AdaPKC based on this idea. Beyond that, a novel fine-tuning technology to further boost the performance of AdaPKC-based RSS networks is presented. Through experimental verification using various real-measured radar data (including publicly available low-cost millimeter-wave radar dataset for autonomous driving and self-collected Ku-band surveillance radar dataset), we found that the performance of AdaPKC-based models surpasses other SoTA methods in RSS tasks. The code is available at https://github.com/lihua199710/AdaPKC.
Youcheng Zhang, ZijunHu, Pengcheng Pi, Zongqing Lu 0001, Qingmin Liao
NeurIPS7
2024 Modeling User Fatigue for Sequential Recommendation
abstract
Recommender systems filter out information that meets user interests. However, users may be tired of the recommendations that are too similar to the content they have been exposed to in a short historical period, which is the so-called user fatigue. Despite the significance for a better user experience, user fatigue is seldom explored by existing recommenders. In fact, there are three main challenges to be addressed for modeling user fatigue, including what features support it, how it influences user interests, and how its explicit signals are obtained. In this paper, we propose to model user Fatigue in interest learning for sequential Recommendations (FRec). To address the first challenge, based on a multi-interest framework, we connect the target item with historical items and construct an interest-aware similarity matrix as features to support fatigue modeling. Regarding the second challenge, built upon feature cross, we propose a fatigue-enhanced multi-interest fusion to capture long-term interest. In addition, we develop a fatigue-gated recurrent unit for short-term interest learning, with temporal fatigue representations as important inputs for constructing update and reset gates. For the last challenge, we propose a novel sequence augmentation to obtain explicit fatigue signals for contrastive learning. We conduct extensive experiments on real-world datasets, including two public datasets and one large-scale industrial dataset. Experimental results show that FRec can improve AUC and GAUC up to 0.026 and 0.019 compared with state-of-the-art models, respectively. Moreover, large-scale online experiments demonstrate the effectiveness of FRec for fatigue reduction. Our codes are released at https://github.com/tsinghua-fib-lab/SIGIR24-FRec.
Nian Li 0001, Xin Ban, Cheng Ling, Chen Gao 0001, Lantao Hu, Peng Jiang 0002, Kun Gai, Yong Li 0008, Qingmin Liao
SIGIR9
2024 Full-stage Diversified Recommendation: Large-scale Online Experiments in Short-video Platform
abstract
The recommender systems on online platforms assist users in finding personalized information, yet this also leads to the issue of limited diversity, potentially giving rise to societal issues such as filter bubbles. Despite significant progress in diversified recommendation algorithms, they have not been extensively experimented with and evaluated for effectiveness in large-scale, full-stage industrial recommender systems. Specifically, industrial recommenders usually consist of three stages of matching, ranking, and re-ranking, in which specific characteristics lead to critical challenges for promoting both recommendation diversity and user engagement. First, user interests are partially observed due to only relevance maximization. Second, item-side feature-aware bias causes imbalanced recommendations. Last, the impact of diversity perception on user engagement stresses the necessity of explicit diversity modeling. To address these challenges in industrial systems, in this work, we deploy several existing diversified algorithms in a real-world short-video platform, including exploration-exploitation, feature-aware debiasing, and diversity optimization. We conduct large-scale online A/B testing for evaluation via online metrics of user engagement and recommendation diversity. Performance improvement across full stages demonstrates the effectiveness of these simple solutions. From comparing performance across different stages and algorithms, we identify that the ranking stage is the most suitable for real-world deployment, and the combination of debiasing and diversity optimization is a promising direction in terms of diversified recommendations. This work provides experiential guidance for the large-scale deployment of diversified algorithms and the construction of a more inclusive platform on the Web.
Nian Li 0001, Yunzhu Pan, Chen Gao 0001, Depeng Jin, Qingmin Liao
WWW5
2024 VPCFormer: A transformer-based multi-view finger vein recognition model and a new benchmark
Pengyang Zhao, Yizhuo Song, Jing-Hao Xue, Shuping Zhao, Qingmin Liao, Wenming Yang
Pattern Recognit.6
2024 DSR-Diff: Depth map super-resolution with diffusion model
Huiyun Cao, Bin Xia 0014, Rui Zhu 0006, Qingmin Liao, Wenming Yang
Pattern Recognit. Lett.5
2024 Dual Correlation Network for Efficient Video Semantic Segmentation
abstract
Video data bring a big challenge to semantic segmentation due to the large volume of data and strong inter-frame redundancy. In this paper, we propose a dual local and global correlation network tailored for efficient video semantic segmentation. It consists of three modules: 1) a local attention based module, which measures correlation and achieves feature aggregation in a local region between key frame and non-key frame; 2) a consistent constraint module, which considers long-range correlation among pixels from a global view for promoting intra-frame semantic consistency of non-key frame; and 3) a key frame decision module, which selects key frames adaptively based on the ability of feature transferring. Extensive experiments on the Cityscapes and Camvid video datasets demonstrate that our proposed method could reduce inference time significantly while maintaining high accuracy. The implementation is available at https://github.com/An01168/DCNVSS.
Shumin An, Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.2
2024 DiffVein: A Unified Diffusion Network for Finger Vein Segmentation and Authentication
abstract
Finger vein authentication, recognized for its high security and specificity, has become a focal point in biometric research. Traditional methods predominantly concentrate on vein feature extraction for discriminative modeling, with a limited exploration of generative approaches. Suffering from verification failure, existing methods often fail to obtain authentic vein patterns by segmentation. To fill this gap, we introduce DiffVein, a unified diffusion model-based framework which simultaneously addresses vein segmentation and authentication tasks. DiffVein is composed of two dedicated branches: one for segmentation and the other for denoising. For better feature interaction between these two branches, we introduce two specialized modules to improve their collective performance. The first, a mask condition module, incorporates the semantic information of vein patterns from the segmentation branch into the denoising process. Additionally, we also propose a Semantic Difference Transformer (SD-Former), which employs Fourier-space self-attention and cross-attention modules to extract category embedding before feeding it to the segmentation task. In this way, our framework allows for a dynamic interplay between diffusion and segmentation embeddings, thus vein segmentation and authentication tasks can inform and enhance each other in the joint training. To further optimize our model, we introduce a Fourier-space Structural Similarity (FSSIM) loss function, which is tailored to improve the denoising network’s learning efficacy. Extensive experiments on the USM and THU-MVFV3V datasets substantiates DiffVein’s superior performance, setting new benchmarks in both vein segmentation and authentication tasks.
Wenming Yang, Qingmin Liao
IEEE Trans. Circuits Syst. Video Technol.3
2024 Study of 3D Finger Vein Biometrics on Imaging Device Design and Multi-View Verification
abstract
Finger vein recognition is an emerging biometric technology with high security and various application scenarios. Most finger vein recognition methods are based on a single view. However, the inherent problems in single-view finger vein recognition, such as limited feature, sensitivity to finger translation and rotation, and the ambiguity issue in 2D projections, hinder the improvement of the system performance. To address these problems and enhance finger vein verification performance, we employ multi-view finger vein images that are capable of providing a more comprehensive feature of 3D finger vein. Specifically, we design a novel low-cost full-view finger vein imaging device that enables full-view capture of finger veins with only a single camera and establish a multi-view finger vein dataset, named THU-MVFV. In addition, we propose a Multi-view Finger Vein Feature Encoding and Selection Network (MFV-FESNet), which is based on an improved Transformer encoder that can learn the dependencies between different views. By fusing the extracted global context feature and local dominant feature, the network can generate a feature descriptor with high discrimination. Extensive experiments are conducted on THU-MVFV and demonstrate the superior performance of the proposed model. The THU-MVFV dataset will be publicly available athttps://github.com/Finger-Vein-Dataset/THU-MVFV.
Yizhuo Song, Pengyang Zhao, Qingmin Liao, Wenming Yang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Exploit the Best of Both End-to-End and Map-Based Methods for Multi-Focus Image Fusion
abstract
Multi-focus image fusion is a technique to fuse the images focused on different depth ranges to generate an all-in-focus image. Existing deep learning approaches to multi-focus image fusion can be categorized as end-to-end methods and decision map based methods. End-to-end methods can generate natural fusion near the focus-defocus boundaries (FDB), but the output is often inconsistent with the input in the areas far from the boundaries (FFB). On the contrary, decision map based methods can preserve original images in the FFB areas, but often generate artifacts near the FDB. In this paper, we propose a dual-branch network for multi-focus image fusion (DB-MFIF) to exploit the best of both worlds, achieving better results in both FDB and FFB areas, i.e. with naturally sharper FDB areas and more consistent FFB areas with the inputs. In our DB-MFIF, an end-to-end branch and a decision map based branch are proposed to mutually assist each other. In addition, to this end, two map-based loss functions are also proposed. Experiments show that our method surpasses existing algorithms on multiple datasets, both qualitatively and quantitatively, and achieves the state-of-the-art performance. The code and model is available on GitHub:https://github.com/Zancelot/DB-MFIF.
Juncheng Zhang, Qingmin Liao, Jing-Hao Xue, Wenming Yang
IEEE Trans. Multim.2
2024 STDAN: Deformable Attention Network for Space-Time Video Super-Resolution
abstract
The target of space-time video super-resolution (STVSR) is to increase the spatial-temporal resolution of low-resolution (LR) and low-frame-rate (LFR) videos. Recent approaches based on deep learning have made significant improvements, but most of them only use two adjacent frames, that is, short-term features, to synthesize the missing frame embedding, which cannot fully explore the information flow of consecutive input LR frames. In addition, existing STVSR models hardly exploit the temporal contexts explicitly to assist high-resolution (HR) frame reconstruction. To address these issues, in this article, we propose a deformable attention network called STDAN for STVSR. First, we devise a long short-term feature interpolation (LSTFI) module that is capable of excavating abundant content from more neighboring input frames for the interpolation process through a bidirectional recurrent neural network (RNN) structure. Second, we put forward a spatial-temporal deformable feature aggregation (STDFA) module, in which spatial and temporal contexts in dynamic video frames are adaptively captured and aggregated to enhance SR reconstruction. Experimental results on several datasets demonstrate that our approach outperforms state-of-the-art STVSR methods. The code is available at https://github.com/littlewhitesea/STDAN.
Hai Wang 0020, Xiaoyu Xiang, Yapeng Tian, Wenming Yang, Qingmin Liao
IEEE Trans. Neural Networks Learn. Syst.5
2023 Dynamic Ensemble of Low-Fidelity Experts: Mitigating NAS "Cold-Start"
abstract
Predictor-based Neural Architecture Search (NAS) employs an architecture performance predictor to improve the sample efficiency. However, predictor-based NAS suffers from the severe ``cold-start'' problem, since a large amount of architecture-performance data is required to get a working predictor. In this paper, we focus on exploiting information in cheaper-to-obtain performance estimations (i.e., low-fidelity information) to mitigate the large data requirements of predictor training. Despite the intuitiveness of this idea, we observe that using inappropriate low-fidelity information even damages the prediction ability and different search spaces have different preferences for low-fidelity information types. To solve the problem and better fuse beneficial information provided by different types of low-fidelity information, we propose a novel dynamic ensemble predictor framework that comprises two steps. In the first step, we train different sub-predictors on different types of available low-fidelity information to extract beneficial knowledge as low-fidelity experts. In the second step, we learn a gating network to dynamically output a set of weighting coefficients conditioned on each input neural architecture, which will be used to combine the predictions of different low-fidelity experts in a weighted sum. The overall predictor is optimized on a small set of actual architecture-performance data to fuse the knowledge from different low-fidelity experts to make the final prediction. We conduct extensive experiments across five search spaces with different architecture encoders under various experimental settings. For example, our methods can improve the Kendall's Tau correlation coefficient between actual performance and predicted scores from 0.2549 to 0.7064 with only 25 actual architecture-performance data on NDS-ResNet. Our method can easily be incorporated into existing predictor-based NAS frameworks to discover better architectures. Our method will be implemented in Mindspore (Huawei 2020), and the example code is published at https://github.com/A-LinCui/DELE.
Junbo Zhao 0007, Xuefei Ning, Enshu Liu, Binxin Ru, Tianchen Zhao, Chen Chen 0077, Jiajin Zhang, Qingmin Liao, Yu Wang 0002
AAAI9
2023 Rethinking CNN Architectures in Transformer Detectors
Mengze Pan, Qingmin Liao
ICANN (10)3
2023 Trans-Cycle: Unpaired Image-to-Image Translation Network by Transformer
Mengze Pan, Zongqing Lu 0001, Qingmin Liao
ICANN (6)4
2023 Flowpose: Conditional Normalizing Flows for 3D Human Pose and Shape Estimation from Monocular Videos
abstract
Human motion modeling is essential for video-based 3D human pose and shape estimation. Most existing methods model human motion by learning a deterministic mapping from the input videos to the human body parameters, while the uncertainties such as occlusions and depth ambiguities are ignored. To address this problem, we propose a probabilistic model based on conditional normalizing flows called FlowPose to learn the distribution of feasible 3D human motion. This model allows access to the most likely 3D human poses given a video input, which means that more accurate and temporally coherent human poses can be obtained. Additionally, a contrastive training strategy is utilized to maximize the mutual information between video features and their 3D human poses, resulting in an improvement on feature extraction of the conditional flow model. Experimental results on two benchmarks 3DPW and Human3.6M demonstrate that our method outperforms the state-of-the-art video-based methods.
Yaoyao Du, Zixiao Zhang, Zhihao Li 0002, Qingmin Liao, Wenming Yang
ICASSP5
2023 Robust Content-Variant Reference Image Quality Assessment Via Similar Patch Matching
abstract
Although image quality assessment (IQA) methods have achieved remarkable success in the past decades, full-reference IQA is limited to reference images, while no-reference IQA has relatively poor performance. To boost the performance of IQA models in the no-reference scenario, a new class of IQA methods using content-variant high-quality images as references have emerged. However, the existing approaches do not take advantage of the content information of the content-variant reference (CVR) images, resulting in the insufficient use of high-quality reference information and the unsatisfactory robustness of the algorithm performance. To effectively utilize CVR images and make the algorithm more robust, we propose a CVR IQA scheme based on similar patch matching. For each image patch to be evaluated, the patch with the most similar content is first searched in the CVR image as the reference patch. Since the two patches are more similar, more useful reference information can be extracted. A similarity calculation module based on cross-attention is designed to find content-similar patches. Extensive experimental results show that the proposed algorithm has good performance and robustness.
Wenming Yang, Qingmin Liao
ICASSP3
2023 Spatial Correlation Fusion Network for Few-Shot Segmentation
abstract
Few-shot semantic segmentation aims to learn new knowledge rapidly with very few annotated data to segment novel classes. Recent methods follow a metric learning framework with prototypes for foreground representation [1]. However, representing support images by one or more prototypes may face problems caused by inadequate representation for segmentation, noise in complex scenes, and close semantic relation to background features. We propose a Spatial Correlation Fusion Network(SCFNet) for few-shot segmentation to address the issues. Firstly, to better capture fine-grained features, we design a Spatial Correlation Fusion module to address the loss of spatial information in support images, thus improving the performance of Few-shot segmentation. Secondly, a Prototype Contrastive Transformation(PCT) module is proposed to learn a transformation matrix for the prototype, which is capable of alleviating close semantic information and noise by adopting transformation loss. Experiments on PASCAL-5i[2] and COCO-20i[3] validate the effectiveness of our network for few-shot semantic segmentation and show our approach achieves state-of-the-art results.
Wenqi Huang 0002, Wenming Yang, Qingmin Liao
ICASSP4
2023 CDHD: Contrastive Dreamer for Hint Distillation
abstract
Replaying previous training data is the most effective approach for Class-Incremental Learning (CIL), with its performance bounded by data availability. Therefore, many recent studies consider the Data-Free Class-Incremental Learning (DFCIL) problem that requires no previous data. However, the existing methods do not consider synthesising data of heterogeneity, thus limiting models’ generalizability. Such homogenous images further hinder the knowledge distillation process when regularising only the deeper layers close to the output, resulting in catastrophic forgetting. To address these issues, we present CDHD: a contrastive dreamer for hint distillation. Our approach starts with training a generator for data synthesis. A model inversion technique is introduced to obtain a generator capable of producing heterogeneous images from the classifier by imposing the ContRastive Loss. Moreover, to better transfer the previous knowledge to the current model, we force the teacher network to provide more general knowledge to its students by enforcing the Hint Loss in shallower layers rather than only in deeper ones. We validate the performance of CDHD on CIFAR-100 for various tasks and compare it against the SOTA baseline for DFCIL, demonstrating our superiorities and thus constituting a new benchmark.
Tongyan Hua, Wenming Yang, Qingmin Liao
ICASSP5
2023 Distortion-Aware Mutual Constraint for Screen Content Image Quality Assessment
Jintong Hu, Wengming Yang, Qingmin Liao
ICIG (1)4
2023 Hard Samples Based Margin Loss for Face Verification
abstract
Although softmax loss and its variants have achieved great success in face verification, the performance is still subject to the data imbalance and early saturation problems. In this paper, we define hard samples as minority class samples and early saturation samples, in order to address both issues, we propose a new loss function termed Hard-Samples based Margin (HSM) loss. Inspired by the class-variant margin normalized softmax loss, we add larger margin on minority classes, the proposed real-class margin overcomes the negative influence from the data imbalance via making the optimization more balanced, while by expanding the margin of early saturated samples, the proposed pseudo-class margin keeps the samples away from the saturation region. Comprehensive experiments show that our HSM loss consistently surpasses the state-of-the-art loss functions on four popular face verification benchmarks.
Xiaying Bai, Wenxian Zheng, Wenming Yang, Guijin Wang, Qingmin Liao
ICIP5
2023 LLA-Flow: A Lightweight Local Aggregation on Cost Volume for Optical Flow Estimation
abstract
Lack of texture often causes ambiguity in matching, and handling this issue is an important challenge in optical flow estimation. Some methods insert stacked transformer modules that allow the network to use global information of cost volume for estimation. But the global information aggregation often incurs serious memory and time costs during training and inference, which hinders model deployment. We draw inspiration from the traditional local region constraint and design the local similarity aggregation (LSA) and the shifted local similarity aggregation (SLSA). The aggregation for cost volume is implemented with lightweight modules that act on the feature maps. Experiments on the final pass of Sintel show the lower cost required for our approach while maintaining competitive performance.
Zongqing Lu 0001, Qingmin Liao
ICIP3
2023 Boosting External-Reference Image Quality Assessment by Content-Constrain Loss and Attention-based Adaptive Feature Fusion
abstract
With the development of deep learning, image quality assessment (IQA) methods have made significant progress, but full-reference (FR) methods are limited by the reference image, and the performance of no-reference (NR) methods is relatively poor. Therefore, some researchers have introduced a new scheme, external-reference (ER) IQA, which uses an arbitrary high-quality image as the reference image (external reference image). The key to ER-IQA is how to extract useful reference information from external reference images and use it effectively. Since the content of the external reference image is independent of the distorted image, we think that the content information of the external reference image is harmful to the algorithm. Therefore, a content-constrain loss is designed for training the network to suppress the content information of external reference images. To utilize the external reference information more effectively, we design an attention-based adaptive feature fusion (AAFF) module. Experimental results demonstrate the effectiveness of the designed loss and feature fusion module.
Wenming Yang, Qingmin Liao
IJCNN3
2023 Patchmatch Stereo++: Patchmatch Binocular Stereo with Continuous Disparity Optimization
abstract
Current deep-learning-based stereo matching algorithms achieve remarkably low error rates but they suffer from the edge ambiguity effect. The primary reason is that they treat disparity estimation as a labeling problem, constructing a cost volume based on uniform discrete pixel-wise labels. It is insufficient to model the continuous disparity probability distribution (DPD), which harms the accuracy of complex regions. Moreover, current cost aggregation strategies cannot process unstructured disparity candidates very well, which is one of the bottlenecks limiting continuous modeling. We propose Patchmatch Stereo++, inspired by the traditional Patchmatch Stereo to achieve better continuous disparity optimization in deep-learning-based methods. Firstly, to model accurate continuous DPD, we introduce an adaptive dense sub-pixel sampling strategy to binocular stereo and approximate a continuous unstructured DPD for every pixel. Secondly, we design a convolution-based optimizer that can accept unstructured disparity candidates to parse the above continuous DPD in an adaptive manner and perform updates accordingly. Extensive experiments demonstrate our method has the best performance among existing stereo matching networks at the edges, both quantitatively and qualitatively. At the time of submission, compared with published works pre-trained on SceneFlow, we rank 1st in the foreground of KITTI and 2nd on SceneFlow, ETH3D under various metrics.The source code will be released.
Wenjia Ren, Qingmin Liao, Zhijing Shao, Xiangru Lin, Xin Yue, Yu Zhang 0166, Zongqing Lu 0001
ACM Multimedia2
2023 The neglected background cues can facilitate finger vein recognition
Pengyang Zhao, Shuping Zhao, Jing-Hao Xue, Wenming Yang, Qingmin Liao
Pattern Recognit.5
2023 EIFNet: An Explicit and Implicit Feature Fusion Network for Finger Vein Verification
abstract
Finger vein recognition has received more attention in recent years due to its high security and promising development potential. However, extracting complete vein patterns and obtaining features from the original images suffer from the low contrast of finger vein images, which dramatically restrains the performance of finger vein recognition algorithms. Inspired by this motivation, we propose an explicit and implicit feature fusion Network (EIFNet) for finger vein verification. It can extract more comprehensive and discriminative features by complementarily fusing the features extracted from binary vein masks and gray original images. We design a feature fusion module (FFM) acting as a bridge between mask feature extraction module (MFEM) and contextual feature extraction module (CFEM) to achieve the optimal fusion of features. To obtain more accurate vein masks, we develop a novel finger vein pattern extraction method and provide the first finger vein segmentation dataset THUFVS. We solve the difficulty of building finger vein segmentation datasets in a simple but effective way, and develop a complete process encompassing dataset creation, data augmentation refinement and network design, which refers to the Mask Generation Module (MGM), for the deep learning based finger vein pattern extraction method. Experimental results demonstrate the superior verification performance of EIFNet on three widely used datasets compared with other existing methods.
Yizhuo Song, Pengyang Zhao, Wenming Yang, Qingmin Liao, Jie Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Meta-Learning-Based Degradation Representation for Blind Super-Resolution
abstract
Blind image super-resolution (blind SR) aims to generate high-resolution (HR) images from low-resolution (LR) input images with unknown degradations. To enhance the performance of SR, the majority of blind SR methods introduce an explicit degradation estimator, which helps the SR model adjust to unknown degradation scenarios. Unfortunately, it is impractical to provide concrete labels for the multiple combinations of degradations (e.g., blurring, noise, or JPEG compression) to guide the training of the degradation estimator. Moreover, the special designs for certain degradations hinder the models from being generalized for dealing with other degradations. Thus, it is imperative to devise an implicit degradation estimator that can extract discriminative degradation representations for all types of degradations without requiring the supervision of degradation ground-truth. To this end, we propose a Meta-Learning based Region Degradation Aware SR Network (MRDA), including Meta-Learning Network (MLN), Degradation Extraction Network (DEN), and Region Degradation Aware SR Network (RDAN). To handle the lack of ground-truth degradation, we use the MLN to rapidly adapt to the specific complex degradation after several iterations and extract implicit degradation information. Subsequently, a teacher network MRDAT is designed to further utilize the degradation information extracted by MLN for SR. However, MLN requires iterating on paired LR and HR images, which is unavailable in the inference phase. Therefore, we adopt knowledge distillation (KD) to make the student network learn to directly extract the same implicit degradation representation (IDR) as the teacher from LR images. Furthermore, we introduce an RDAN module that is capable of discerning regional degradations, allowing IDR to adaptively influence various texture patterns. Extensive experiments under classic and real-world degradation settings show that MRDA achieves SOTA performance and can generalize to various degradation processes.
Bin Xia 0014, Yapeng Tian, Yulun Zhang 0001, Yucheng Hang, Wenming Yang, Qingmin Liao
IEEE Trans. Image Process.6
2023 Disentangled Modeling of Social Homophily and Influence for Social Recommendation
abstract
Social recommendation leverages social information to alleviate data sparsity and cold-start issues of collaborative filtering (CF) methods. Most existing works model user interests following the assumption ofsocial homophilybased on social-relation data. The explicit modeling ofsocial influence, which also largely affects user behaviors, has not been well explored. Considering user behaviors may be driven by social factors in today’s information services (e.g., purchasing products shared by close friends on social e-commerce applications), these methods will be suboptimal. In this work, we propose a method modeling both social homophily-aware user interests and social influence as two essential effects on user behaviors for social recommendation, named as DISGCN (short forDISentangled modeling of Social homophily and influence withGraphConvolutionalNetwork). Specifically, we devise a disentangled embedding layer to encode these two effects. Furthermore, two tailored graph convolutional layers are developed to disentangle them refinedly, leveraging the high-order embedding propagation in social-network graph from two aspects. Technically, first, the operation of attentive embedding propagation is adopted for capturing personalized social homophily-aware interests, and second, the item-gate-based embedding propagation is proposed for capturing item-specific social influence. In addition, to ensure the disentanglement of social influence, we propose a contrastive learning framework that endows corresponding embeddings with explicit semantics. Extensive experiments on two real-world datasets demonstrate the effectiveness of our proposed model. Further studies also verify the rationality and necessity of our designs. We have released the datasets and codes at this link:https://github.com/tsinghua-fib-lab/DISGCN.
Nian Li 0001, Chen Gao 0001, Depeng Jin, Qingmin Liao
IEEE Trans. Knowl. Data Eng.4
2023 Dual Polarization Modality Fusion Network for Assisting Pathological Diagnosis
abstract
Polarization imaging is sensitive to sub-wavelength microstructures of various cancer tissues, providing abundant optical characteristics and microstructure information of complex pathological specimens. However, how to reasonably utilize polarization information to strengthen pathological diagnosis ability remains a challenging issue. In order to take full advantage of pathological image information and polarization features of samples, we propose a dual polarization modality fusion network (DPMFNet), which consists of a multi-stream CNN structure and a switched attention fusion module for complementarily aggregating the features from different modality images. Our proposed switched attention mechanism could obtain the joint feature embeddings by switching the attention map of different modality images to improve their semantic relatedness. By including a dual-polarization contrastive training scheme, our method can synthesize and align the interaction and representation of two polarization features. Experimental evaluations on three cancer datasets show the superiority of our method in assisting pathological diagnosis, especially in small datasets and low imaging resolution cases. Grad-CAM visualizes the important regions of the pathological images and the polarization images, indicating that the two modalities play different roles and allow us to give insightful corresponding explanations and analysis on cancer diagnosis conducted by the DPMFNet. This technique has potential to facilitate the performance of pathological aided diagnosis and broaden the current digital pathology boundary based on pathological image features.
Lu Si, Wenming Yang, Xuewu Tian, Qingmin Liao, Hui Ma 0003
IEEE Trans. Medical Imaging8
2023 SCTANet: A Spatial Attention-Guided CNN-Transformer Aggregation Network for Deep Face Image Super-Resolution
abstract
Numerous CNN-based algorithms have been proposed to reconstruct high-quality face images. However, the inability of convolution operation to model long-distance relationships limits the performance of the CNN-based methods. Moreover, in the high-resolution (HR) image reconstruction stage, with the well decoded feature representations, more efficient architecture design can be explored to synthesize pixel-level image details. In this work, we propose a spatial attention-guided CNN-Transformer aggregation network (SCTANet) for face image super-resolution (FSR) tasks. The core component in the deep feature extraction stage is the Hybrid Attention Aggregation (HAA) block. The HAA block has two parallel paths, one for the Residual Spatial Attention (RSA) block, the other for the Multi-scale Patch embedding and Spatial-attention Masked Transformer (MPSMT) block. The HAA block combines the strengths of CNN and transformer to effectively exploit both local and global information. For the reconstruction stage, we propose to use the Sub-pixel MLP-based Upsampling (SMU) module instead of the conventional CNN architecture. The SMU module promotes the reconstruction of pixel-level image details and reduces computational complexity. Extensive experiments on both synthetic and real-world face datasets demonstrate the superiority of our proposed SCTANet over state-of-the-art methods.
Qiqi Bao 0001, Yunmeng Liu, Bowen Gang, Wenming Yang, Qingmin Liao
IEEE Trans. Multim.5
2023 APANet: Adaptive Prototypes Alignment Network for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation aims to segment novel-class objects in a given query image with only a few labeled support images. Most advanced solutions exploit a metric learning framework that performs segmentation through matching each query feature to a learned class-specific prototype. However, this framework suffers from biased classification due to incomplete feature comparisons. To address this issue, we present an adaptive prototype representation by introducing class-specific and class-agnostic prototypes and thus construct complete sample pairs for learning semantic alignment with query features. The complementary features learning manner effectively enriches feature comparison and helps yield an unbiased segmentation model in the few-shot setting. It is implemented with a two-branch end-to-end network (i.e., a class-specific branch and a class-agnostic branch), which generates prototypes and then combines query features to perform comparisons. In addition, the proposed class-agnostic branch is simple yet effective. In practice, it can adaptively generate multiple class-agnostic prototypes for query images and learn feature alignment in a self-contrastive manner. Extensive experiments on PASCAL-5$^{i}$and COCO-20$^{i}$demonstrate the superiority of our method. At no expense of inference efficiency, our model achieves state-of-the-art results in both 1-shot and 5-shot settings for semantic segmentation.
Bin-Bin Gao, Zongqing Lu 0001, Jing-Hao Xue, Chengjie Wang 0001, Qingmin Liao
IEEE Trans. Multim.6
2023 Blind JPEG Compression Artifacts Removal by Integrating Channel Regulation With Exit Strategy
abstract
Compression artifacts removal methods based on convolutional neural networks have attracted great attention. However, most existing methods require a specific trained model for a specific compression quality factor (QF), which inevitably leads to resource-consuming. Unfortunately, the QF is unknown in most practical applications, so it is intractable to choose a suitable model. In this work, we experimentally analyze the relationship between compression index estimation and compression artifacts removal. Based on the connection between them, we couple compression index estimation with compression artifacts removal into a unified network. A network named CRESNet is proposed, working for a wide range of QFs by integrating channel regulation with an exit strategy. Specifically, CRESNet adopts a multi-stage progressive structure with an exit strategy embedded to automatically select the optimal exit stage according to the estimated compression index reflecting the difficulty of the input sample. Benefiting from the exit strategy, CRESNet removes artifacts from slightly compressed images through a simple process while doing an elaborate process for severely compressed images. Furthermore, a compression-information-guided channel regulation (CICR) mechanism is developed to adaptively regulate feature maps based on the estimated compression index. CRESNet achieves a more elegant trade-off between artifacts removal and detail preservation in a resource-efficient manner. Experiments demonstrate that CRESNet achieves state-of-the-art performance.
Yunmeng Liu, Wenming Yang, Qingmin Liao
IEEE Trans. Multim.6
2023 Bi-RSTU: Bidirectional Recurrent Upsampling Network for Space-Time Video Super-Resolution
abstract
One-stage space-time video super-resolution (STVSR) aims to directly reconstruct high-resolution (HR) and high frame rate (HFR) video from its low-resolution (LR) and low frame rate (LFR) counterpart. Due to the wide application, one-stage STVSR has drawn much attention recently. However, existing one-stage methods suffer from ineffective exploration of the auxiliary information from adjacent time steps that may be useful to STVSR at the current time step. To address this issue, we propose a novel Bidirectional Recurrent Space-Time Upsampling network called Bi-RSTU for one-stage STVSR to utilize auxiliary information at various time steps. Specifically, an efficient channel attention feature interpolation (ECAFI) module is devised to synthesize the intermediate frame’s LR feature by exploiting its two neighboring LR video frame features. Subsequently, we fuse the information from the previous time step into these intermediate and neighboring features. Finally, second-order attention spindle (SOAS) blocks are stacked to form the feature reconstruction module that learns a mapping from LR fused feature space to HR feature space. Experimental results on public datasets demonstrate that our Bi-RSTU shows competitive performance compared with current two-stage and one-stage state-of-the-art STVSR methods.
Hai Wang 0020, Wenming Yang, Qingmin Liao, Jie Zhou 0001
IEEE Trans. Multim.3
2022 Efficient Non-local Contrastive Attention for Image Super-resolution
abstract
Non-Local Attention (NLA) brings significant improvement for Single Image Super-Resolution (SISR) by leveraging intrinsic feature correlation in natural images. However, NLA gives noisy information large weights and consumes quadratic computation resources with respect to the input size, limiting its performance and application. In this paper, we propose a novel Efficient Non-Local Contrastive Attention (ENLCA) to perform long-range visual modeling and leverage more relevant non-local features. Specifically, ENLCA consists of two parts, Efficient Non-Local Attention (ENLA) and Sparse Aggregation. ENLA adopts the kernel method to approximate exponential function and obtains linear computation complexity. For Sparse Aggregation, we multiply inputs by an amplification factor to focus on informative features, yet the variance of approximation increases exponentially. Therefore, contrastive learning is applied to further separate relevant and irrelevant features. To demonstrate the effectiveness of ENLCA, we build an architecture called Efficient Non-Local Contrastive Network (ENLCN) by adding a few of our modules in a simple backbone. Extensive experimental results show that ENLCN reaches superior performance over state-of-the-art approaches on both quantitative and qualitative evaluations.
Bin Xia 0014, Yucheng Hang, Yapeng Tian, Wenming Yang, Qingmin Liao, Jie Zhou 0001
AAAI5
2022 Coarse-to-Fine Embedded PatchMatch and Multi-Scale Dynamic Aggregation for Reference-Based Super-resolution
abstract
Reference-based super-resolution (RefSR) has made significant progress in producing realistic textures using an external reference (Ref) image. However, existing RefSR methods obtain high-quality correspondence matchings consuming quadratic computation resources with respect to the input size, limiting its application. Moreover, these approaches usually suffer from scale misalignments between the low-resolution (LR) image and Ref image. In this paper, we propose an Accelerated Multi-Scale Aggregation network (AMSA) for Reference-based Super-Resolution, including Coarse-to-Fine Embedded PatchMatch (CFE-PatchMatch) and Multi-Scale Dynamic Aggregation (MSDA) module. To improve matching efficiency, we design a novel Embedded PatchMacth scheme with random samples propagation, which involves end-to-end training with asymptotic linear computational cost to the input size. To further reduce computational cost and speed up convergence, we apply the coarse-to-fine strategy on Embedded PatchMacth constituting CFE-PatchMatch. To fully leverage reference information across multiple scales and enhance robustness to scale misalignment, we develop the MSDA module consisting of Dynamic Aggregation and Multi-Scale Aggregation. The Dynamic Aggregation corrects minor scale misalignment by dynamically aggregating features, and the Multi-Scale Aggregation brings robustness to large scale misalignment by fusing multi-scale information. Experimental results show that the proposed AMSA achieves superior performance over state-of-the-art approaches on both quantitative and qualitative evaluations.
Bin Xia 0014, Yapeng Tian, Yucheng Hang, Wenming Yang, Qingmin Liao, Jie Zhou 0001
AAAI5
2022 Pose-Invariant Face Recognition via Adaptive Angular Distillation
abstract
Pose-invariant face recognition is a practically useful but challenging task. This paper introduces a novel method to learn pose-invariant feature representation without normalizing profile faces to frontal ones or learning disentangled features. We first design a novel strategy to learn pose-invariant feature embeddings by distilling the angular knowledge of frontal faces extracted by teacher network to student network, which enables the handling of faces with large pose variations. In this way, the features of faces across variant poses can cluster compactly for the same person to create a pose-invariant face representation. Secondly, we propose a Pose-Adaptive Angular Distillation loss to mitigate the negative effect of uneven distribution of face poses in the training dataset to pay more attention to the samples with large pose variations. Extensive experiments on two challenging benchmarks (IJB-A and CFP-FP) show that our approach consistently outperforms the existing methods.
Zhenduo Zhang, Yongru Chen, Wenming Yang, Guijin Wang, Qingmin Liao
AAAI5
2022 An Exploratory Study of Information Cocoon on Short-form Video Platform
abstract
In recent years, short-form video platforms have emerged rapidly and attracted a large and wide variety of users, with the help of advanced recommendation algorithms. Despite the great success, the algorithms have caused some negative effects, such as information cocoon, algorithm unfairness,etc. In this work, we focus on theinformation cocoon that measures overwhelmingly homogeneity of users' video consumption. Specifically, we conduct an exploratory study of this phenomenon on a top short-form video platform, with one-year behavioral records of new users. First, we evaluate the evolution of users' information cocoons and find the limitation of the diversity of video content that users consume. In addition, we further explore user cocoons via the correlation analysis from three aspects, including user demographics, video content, and user-recommender interactions driven by algorithms and user preferences. Correspondingly, we observe that video content plays a more significant role in affecting user cocoons than demographics does. In terms of user-recommender interactions, more accurate personalization does not contribute to more severe information cocoons necessarily, while users with narrow preferences are more likely to be trapped. In summary, our study illuminates the current concern of information cocoons that may hurt user experience on short-form video platforms, and offers potential directions for mitigation implied by the correlation analysis.
Nian Li 0001, Chen Gao 0001, Jinghua Piao, Aizhen Yue, Qingmin Liao, Yong Li 0008
CIKM7
2022 SCS-Co: Self-Consistent Style Contrastive Learning for Image Harmonization
abstract
Image harmonization aims to achieve visual consistency in composite images by adapting a foreground to make it compatible with a background. However, existing methods always only use the real image as the positive sample to guide the training, and at most introduce the corresponding composite image as a single negative sample for an auxiliary constraint, which leads to limited distortion knowledge, and further causes a too large solution space, making the generated harmonized image distorted. Besides, none of them jointly constrain from the foreground selfstyle and foreground-background style consistency, which exacerbates this problem. Moreover, recent region-aware adaptive instance normalization achieves great success but only considers the global background feature distribution, making the aligned foreground feature distribution biased. To address these issues, we propose a self-consistent style contrastive learning scheme (SCS-Co). By dynamically generating multiple negative samples, our SCS-Co can learn more distortion knowledge and well regularize the generated harmonized image in the style representation space from two aspects of the foreground self-style and foreground-background style consistency, leading to a more photorealistic visual result. In addition, we propose a background-attentional adaptive instance normalization (BAIN) to achieve an attention-weighted background feature distribution according to the foreground-background feature similarity. Experiments demonstrate the superiority of our method over other state-of-the-art methods in both quantitative comparison and visual analysis.
Yucheng Hang, Bin Xia 0014, Wenming Yang, Qingmin Liao
CVPR4
2022 Sain: Similarity-Aware Video Frame Interpolation
abstract
Video frame interpolation (VFI) aims to synthesize an intermediate frame between two consecutive original frames. Most existing methods simply linearly combine the warped frames, leading to a loss of image texture. Since moving objects usually have similarities in consecutive frames, we propose a similarity-aware video frame interpolation method (SAIN) that searches patches with similar texture in the embedding space from input frames to extract features and capture image details. To gather the frame details and restore image texture, SAIN incorporates an implicit neural representation learning from similar patches to enrich image details and refine outputs in frame synthesis networks. Experiments demonstrate that SAIN preserves image texture and enhances interpolated image quality significantly.
Yue Lv, Wenming Yang, Wangmeng Zuo, Qingmin Liao, Rui Zhu 0006
ICASSP4
2022 Two-Stream Non-Uniform Concentration Reasoning Network for Single Image Air Pollution Estimation
abstract
With the increasing availability of portable cameras and smart phones, directly estimating PM2.5based on digital photography shows advantages in efficiency and economic costs. In this paper, a novel Two-stream Non-uniform Concentration Reasoning Network (TNCR-Net) is proposed for single image PM2.5concentration estimation. Motivated by locally non-uniform particle pollution concentration distribution in images, we adopt patch-based scheme and adaptive weighted average mechanism to obtain patch-wise concentration and relative weight based on spatially varying perceptual relevance of local particle pollution concentration. Then aggregate patch-wise concentrations according to relative weights. To learn more effective feature from particular pollution image, we use a two-stream network structure with the dark channel map as the input of one stream. Besides, we employ attention-based feature fusion method to flexibly aggregate the feature maps of the two streams. Experiments on real-world dataset indicate that our TNCR-Net outperforms other state-of-the-art methods with fewer parameters.
Wenming Yang, Qingmin Liao
ICIP3
2022 Quality-Oriented Feature Regression for Robust Image Similarity Metric
abstract
Full-reference image quality assessment aims to predict the perceptual quality of a distorted image based on its similarity to the pristine reference. In this paper, we propose a robust image similarity metric by fully exploring the representation power of deep learning-based features. A convolutional neu-ral network (CNN) is adopted to extract deep features from multiple scales. We show that such CNN features that con-tain multi -scale visual information are comprehensive and ro-bust enough for quality assessment. We further propose a quality-oriented feature regression (QOFR) module based on the multi-layer perceptron architecture. The QOFR module can efficiently integrate hierarchy CNN features and generate the final quality score. Extensive experiments on the bench-mark datasets demonstrate that our method achieves state-of-the-art performance with outstanding robustness and general-ization ability.
Qiqi Bao 0001, Rui Zhu 0006, Wenming Yang, Qingmin Liao
ICME5
2022 Feature Pyramid Boosting Network for Rendering Natural Bokeh
abstract
Natural bokeh is a typical characteristic of digital single-lens reflex (DSLR) cameras and high-quality lenses, which is commonly used to emphasize a subject from a distracting background. However, it is still a big challenge for mobile platforms to produce similar effects due to the small apertures of their lenses. Unlike many previous methods formulated as a two-stage task composed of depth/defocus estimation and defocus magnification, we propose a feature pyramid boosting network with novel hierarchical attention modules to render bokeh in one step. In addition, existing learning-based methods suffer from the pixel misalignment of the datasets. We present a well-aligned bokeh dataset captured by a DSLR to address this problem. Experiments show that our method can render comparable bokeh with the state-of-the-art method but requires fewer parameters.
Juncheng Zhang, Qingmin Liao
ICME3
2022 Distilling Resolution-robust Identity Knowledge for Texture-Enhanced Face Hallucination
abstract
The main focus of most existing face hallucination methods is to generate visually pleasing results. However, in many applications, the final goal is to identify the person in the low-resolution (LR) image. In this paper, we propose a texture and identity integration network (TIIN) to effectively incorporate identity information into face hallucination tasks. TIIN consists of an identity-preserving denormalization module (IDM) and an equalized texture enhance module (ETEM). The IDM exploits the identity prior and the ETEM improves image quality through histogram equalization. To extract identity information effectively, we propose a resolution-robust identity knowledge distillation network (RIKDN). RIKDN is specifically designed for LR face recognition and can be of independent interest. It employs two teacher-student streams. One stream narrows the performance gap between high-resolution (HR) and LR images. The other distills correlation information from the HR-HR teacher stream to guide learning in the LR-HR student stream. We conduct extensive experiments on multiple datasets to demonstrate the effectiveness of our methods.
Qiqi Bao 0001, Rui Zhu 0006, Bowen Gang, Pengyang Zhao, Wenming Yang, Qingmin Liao
ACM Multimedia6
2022 Unsupervised visual feature learning based on similarity guidance
Zhihao Jin, Qicong Wang, Wenming Yang, Qingmin Liao, Hongying Meng
Neurocomputing5
2022 Exploiting Multiperspective Driven Hierarchical Content-Aware Network for Finger Vein Verification
abstract
The finger vein trait has attracted widespread attention for personal authentication in recent years. However, most finger vein verification methods are performed on the single perspective, captured by a monocular near-infrared camera fixed at one side of the finger. Consequently, the contents of a single perspective have few details of the spatial network structure of the finger vein and show noticeable differences even if the posture of the same finger is slightly different. Both of them impact the verification performance. Hence, finger vein images captured from different viewpoints are considered in this work. We first design a low-cost multi-perspective based dorsal finger vein imaging device for data collection. A deep neural network named Hierarchical Content-Aware Network (HCAN) is then proposed to extract the discriminative hierarchical features of the finger vein. Specifically, HCAN is compound of a Global Stem Network (GSN) and a Local Perception Module (LPM). GSN aims to extract the latent global 3D feature from all perspectives through a recurrent neural network. It enables the model to retain the details in previous hidden states by incorporating a memory weighting strategy. LPM is designed to perceive each perspective from the aspect of image entropy. Guided by the entropy loss, LPM captures the prominent local feature and improves the discriminability and robustness of the hierarchical feature. The experimental results on the newly collected THU-MFV database demonstrate the superiority of the proposed method in comparison with other multi-perspective and single-perspective based methods.
Pengyang Zhao, Shuping Zhao, Luyang Chen, Wenming Yang, Qingmin Liao
IEEE Trans. Circuits Syst. Video Technol.5
2022 Frontal-Centers Guided Face: Boosting Face Recognition by Learning Pose-Invariant Features
abstract
In recent years, face recognition has made a remarkable breakthrough due to the emergence of deep learning. However, compared with frontal face recognition, plenty of deep face recognition models still suffer serious performance degradation when handling profile faces. To address this issue, we propose a novel Frontal-Centers Guided Loss (FCGFace) to obtain highly discriminative features for face recognition. Most existing discriminative feature learning approaches project features from the same class into a separated latent subspace. These methods only model the distribution at the identity-level but ignore the latent relationship between frontal and profile viewpoints. Different from these methods, FCGFace takes viewpoints into consideration by modeling the distribution at both the identity-level and the viewpoint-level. At the identity-level, a softmax-based loss is employed for a relatively rough classification. At the viewpoint-level, centers of frontal face features are defined to guide the optimization conducted in a more refined way. Specifically, our FCGFace is capable of adaptively adjusting the distribution of profile face features and narrowing the gap between them and frontal face features during different training stages to form compact identity clusters. Extensive experimental results on popular benchmarks, including cross-pose datasets (CFP-FP, CPLFW, VGGFace2-FP, and Multi-PIE) and non-cross-pose datasets (YTF, LFW, AgeDB-30, CALFW, IJB-B, IJB-C, and RFW), have demonstrated the superiority of our FCGFace over the SOTA competitors.
Yingfan Tao, Wenxian Zheng, Wenming Yang, Guijin Wang, Qingmin Liao
IEEE Trans. Inf. Forensics Secur.5
2022 Attention-Driven Graph Neural Network for Deep Face Super-Resolution
abstract
With the help of convolutional neural networks (CNNs), deep learning-based methods have achieved remarkable performance in face super-resolution (FSR) task. Despite their success, most of the existing methods neglect non-local correlations of face images, leaving much room for improvement. In this paper, we introduce a novel end-to-end trainable attention-driven graph neural network (AD-GNN) for more discriminative feature extraction and feature relation modeling. This is achieved by two major components. The first component is a cross-scale dynamic graph (CDG) block. The CDG block considers cross-scale relationships of patches in distant areas and employs two dynamic graphs to construct enhanced features. The second component is a series of channel attention and spatial dynamic graph (CASDG) blocks. A CASDG block has a channel-wise attention unit and a spatial-aware dynamic graph (SDG) unit. The SDG unit extracts informative features by exploring spatial non-local self-similarity information of the patches using dynamic graph convolution. Using these two components, facial details can be effectively reconstructed with the help of information supplemented by similar but spatially remote patches and structural information of faces. Extensive experiments on two public benchmarks demonstrate the superiority of AD-GNN over the state-of-the-art FSR methods.
Qiqi Bao 0001, Bowen Gang, Wenming Yang, Jie Zhou 0001, Qingmin Liao
IEEE Trans. Image Process.5
2022 Defocus Image Deblurring Network With Defocus Map Estimation as Auxiliary Task
abstract
Different from the object motion blur, the defocus blur is caused by the limitation of the cameras' depth of field. The defocus amount can be characterized by the parameter of point spread function and thus forms a defocus map. In this paper, we propose a new network architecture called Defocus Image Deblurring Auxiliary Learning Net (DID-ANet), which is specifically designed for single image defocus deblurring by using defocus map estimation as auxiliary task to improve the deblurring result. To facilitate the training of the network, we build a novel and large-scale dataset for single image defocus deblurring, which contains the defocus images, the defocus maps and the all-sharp images. To the best of our knowledge, the new dataset is the first large-scale defocus deblurring dataset for training deep networks. Moreover, the experimental results demonstrate that the proposed DID-ANet outperforms the state-of-the-art methods for both tasks of defocus image deblurring and defocus map estimation, both quantitatively and qualitatively. The dataset, code, and model is available on GitHub: https://github.com/xytmhy/DID-ANet-Defocus-Deblurring.
Qingmin Liao, Juncheng Zhang, Jing-Hao Xue
IEEE Trans. Image Process.3
2022 MDAN: Mirror Difference Aware Network for Brain Stroke Lesion Segmentation
abstract
Brain stroke lesion segmentation is of great importance for stroke rehabilitation neuroimaging analysis. Due to the large variance of stroke lesion shapes and similarities of tissue intensity distribution, it remains a challenging task. To help detect abnormalities, the anatomical symmetries of brain magnetic resonance (MR) images have been widely used as visual cues for clinical practices. However, most methods for brain images segmentation do not fully utilize structural symmetry information. This paper presents a novel mirror difference aware network (MDAN) for stroke lesion segmentation. The network uses an encoder-decoder architecture, aiming at holistically exploiting the symmetries of image features. Specifically, a differential feature augmentation (DFA) module is developed in the encoding path to highlight the semantically pathological asymmetries of features in abnormalities. In the DFA module, a Siamese contrastive supervised loss is designed to enhance discriminative features, and a mirror position-based difference augmentation (MDA) module is used to further magnify the discrepancy. Moreover, mirror feature fusion (MFF) modules are applied to efficiently fuse and transfer the information both of the original input and the horizontally flipped features to the decoding path. Extensive experiments on the Anatomical Tracings of Lesions After Stroke (ATLAS) dataset show the proposed MDAN outperforms the state-of-the-art methods.
Qiqi Bao 0001, Shiyu Mi, Bowen Gang, Wenming Yang, Jie Chen 0001, Qingmin Liao
IEEE J. Biomed. Health Informatics6
2022 Efficient Semantic Segmentation via Self-Attention and Self-Distillation
abstract
Lightweight models are pivotal in efficient semantic segmentation, but they often suffer from insufficient context information due to limited convolution and small receptive field. To address this problem, we propose a tailored approach to efficient semantic segmentation by leveraging two complementary distillation schemes for supplementing context information to small networks: 1) a self-attention distillation scheme, which transfers long-range context knowledge adaptively from large teacher networks to small student networks; and 2) a layer-wise context distillation scheme, which transfers structured context from deep layers to shallow layers within student networks for promoting semantic consistency of the shallow layers. Extensive experiments on the ADE20K, Cityscapes, and Camvid datasets well demonstrate the effectiveness of our proposal.
Shumin An, Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
IEEE Trans. Intell. Transp. Syst.2
2022 Deep Learning in Lane Marking Detection: A Survey
abstract
Lane marking detection is a fundamental but crucial step in intelligent driving systems. It can not only provide relevant road condition information to prevent lane departure but also assist vehicle positioning and forehead car detection. However, lane marking detection faces many challenges, including extreme lighting, missing lane markings, and obstacle obstructions. Recently, deep learning-based algorithms draw much attention in intelligent driving society because of their excellent performance. In this paper, we review deep learning methods for lane marking detection, focusing on their network structures and optimization objectives, the two key determinants of their success. Besides, we summarize existing lane-related datasets, evaluation criteria, and common data processing techniques. We also compare the detection performance and running time of various methods, and conclude with some current challenges and future trends for deep learning-based lane marking detection algorithm.
Youcheng Zhang, Zongqing Lu 0001, Xuechen Zhang 0003, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Intell. Transp. Syst.5
2022 Fast Extended Inductive Robust Principal Component Analysis With Optimal Mean
abstract
Inspired by the mean calculation of RPCA_OM and inductiveness of IRPCA, we first propose an inductive robust principal component analysis method with removing the optimal mean automatically, which is shorted as IRPCA_OM. Furthermore, IRPCA_OM is extended to Schatten-$p$norm and a more general framework (i.e., EIRPCA_OM) is presented. The objective function of EIRPCA_OM includes two terms, the first term is a robust reconstruction error term constrained by an$\ell _{2,1}$-norm and the second term is a regularization term constrained by a Schatten-$p$norm. The proposed EIRPCA_OM method is robust, inductive and accurate. However, on the high-dimensional data, it would spend a large computation cost in training stage. To this end, a fast version of EIRPCA_OM called as FEIRPCA_OM is proposed, and its basic idea is to eliminate the zero eigenvalues of data matrix. More importantly, an effective theoretical proof is presented to ensure that FEIRPCA_OM has faster processing speed than EIRPCA_OM when processing high-dimensional data, but without any performance loss. Based on it, we also can exchange the less performance loss for the higher computation efficiency by removing the small eigenvalues of data matrix. Experimental results on the public datasets demonstrate that FEIRPCA_OM works efficiently on the high-dimensional data.
Shuangyan Yi, Feiping Nie 0001, Yongsheng Liang 0001, Wei Liu 0065, Zhenyu He 0001, Qingmin Liao
IEEE Trans. Knowl. Data Eng.6
2022 GenDet: Meta Learning to Generate Detectors From Few Shots
abstract
Object detection has made enormous progress and has been widely used in many applications. However, it performs poorly when only limited training data is available for novel classes that the model has never seen before. Most existing approaches solve few-shot detection tasks implicitly without directly modeling the detectors for novel classes. In this article, we propose GenDet, a new meta-learning-based framework that can effectively generate object detectors for novel classes from few shots and, thus, conducts few-shot detection tasks explicitly. The detector generator is trained by numerous few-shot detection tasks sampled from base classes each with sufficient samples, and thus, it is expected to generalize well on novel classes. An adaptive pooling module is further introduced to suppress distracting samples and aggregate the detectors generated from multiple shots. Moreover, we propose to train a reference detector for each base class in the conventional way, with which to guide the training of the detector generator. The reference detectors and the detector generator can be trained simultaneously. Finally, the generated detectors of different classes are encouraged to be orthogonal to each other for better generalization. The proposed approach is extensively evaluated on the ImageNet, VOC, and COCO data sets under various few-shot detection settings, and it achieves new state-of-the-art results.
Liyang Liu, Bochao Wang, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2021 Gaussian Mixture Distribution Makes Data Uncertainty Learning Better
abstract
As a mainstream method in face recognition, extracting separable facial features by deep CNNs in the latent space has achieved remarkable success. In most existing works, people often view the embedding features as points. Dealing with entirely unconstrained face images, DUL and PFE demonstrated that point estimation shows weak robusticity on the inherent noise in the input images (data uncertainty) and introduced the distribution estimation by modeling each latent feature using a Gaussian distribution. However, these two methods only apply a unimodal Gaussian prior distribution, which is insufficient to represent wild faces with complex variations. In this paper, we propose a novel face recognition framework based on the multivariate Gaussian mixture distribution (DUL-GM). Through numerous experiments, we show that compared with the prior works, the features modeled by multivariate Gaussian mixture distribution have a better interference suppression ability and achieve state-of-the-art performance on extensive challenging benchmarks.
Hao Ai, Qingmin Liao
FG2
2021 Parallax Contextual Representations For Stereo Matching
abstract
In this work, we study the context aggregation in stereo matching from a new parallax perspective. Unlike previous works, we propose to characterize and augment a pixel with its parallax contextual representation (PCR), which has not been explored before. We also propose a new concept called disparity prototype to describe the overall representation of a disparity plane. Our proposed PCR module consists of three steps: 1) divide disparity planes for a rough estimation of disparity; 2) estimate the disparity prototypes for each disparity plane; 3) derive PCR-augmented representations with disparity prototypes. Extensive experiments on various datasets using different networks validate the effectiveness of our proposal.
Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
ICIP2
2021 A Region-Based Descriptor Network for Uniformly Sampled Keypoints
abstract
Matching keypoint pairs of different images is a basic task of computer vision. Most methods require customized extremum point schemes to obtain the coordinates of feature points with high confidence, which often need complex algorithmic design or a network with higher training difficulty and also ignore the possibility that flat regions can be used as candidate regions of matching points. In this paper, we design a region-based descriptor by combining the context features of a deep network. The new descriptor can give a robust representation of a point even in flat regions. By the new descriptor, we can obtain more high confidence matching points without extremum operation. The experimental results show that our proposed method achieves a performance comparable to state-of-the-art.
Zongqing Lu 0001, Qingmin Liao
ICIP3
2021 Towards Impartial Multi-task Learning
Liyang Liu, Yi Li 0050, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
ICLR7
2021 Disparity Estimation with Scene Depth Cues
abstract
The cost volume plays a pivotal role in stereo matching, usually working as an optimization object. However, we find it also can provide effective scene prior to guide the disparity learning, as it reflects well the depth relationship between scenario objects. Inspired by this new perspective, we propose the CSA module, which consists of a new correlation and selection (CS) layer and a new aggregation layer. The CS layer can regulate the matching costs and re-encode the feature information into the correlation volume. The aggregation layer can preserve better the depth cues of the refined cost volume, through a convolution network and a unimodalization operation. The proposed module can be trained in a supervised manner, making the extraction of scene depth cues more accurate. Extensive experiments on the Sceneflow and KITTI datasets have demonstrated that with our module embedded, SOTA networks can achieve substantially better performance.
Zongqing Lu 0001, Qingmin Liao, Jing-Hao Xue
ICME3
2021 Better Stereo Matching From Simple Yet Effective Wrangling of Deep Features
abstract
Cost volume plays a pivotal role in stereo matching. Most recent works focused on deep feature extraction and cost refinement for a more accurate cost volume. Unlike them, we probe from a different perspective: feature wrangling. We find that simple wrangling of deep features can effectively improve the construction of cost volume and thus the performance of stereo matching. Specifically, we develop two simple yet effective wrangling techniques of deep features, spatially a differentiable feature transformation and channel-wise a memory-economical feature expansion, for better cost construction. Exploiting the local ordering information provided by a differentiable rank transform, we achieve an enhancement of the search for correspondence; with the help of disparity division, our feature expansion allows for more features into the cost volume with no extra memory required. Equipped with these two feature wrangling techniques, our simple network can perform outstandingly on the widely used KITTI and Sceneflow datasets.
Zongqing Lu 0001, Qingmin Liao, Jing-Hao Xue
ICME3
2021 RGB Guided Depth Map Super-Resolution with Coupled U-Net
abstract
The depth maps captured by RGB-D cameras usually are of low resolution, entailing recent efforts to develop depth super-resolution (DSR) methods. However, several problems remain in existing DSR methods. First, conventional DSR methods often suffer from unexpected artifacts. Secondly, high-resolution (HR) RGB features and low-resolution (LR) depth features are often fused in shallow layers only. Thirdly, only the last layer of features is used for reconstruction. To address the above problems, we propose Coupled U-Net (CU-Net), a new color image guided DSR method built on two U-Net branches for HR color images and LR depth maps, respectively. The CU-Net embeds a dual skip connection structure to leverage the feature interaction of the two branches, and a multi-scale fusion to fuse the deeper and multi-scale features of two branch decoders for more effective feature reconstruction. Moreover, a channel attention module is proposed to eliminate artifacts. Extensive experiments show that the proposed CU-Net outperforms state-of-the-art methods.
Yingjie Cui, Qingmin Liao, Wenming Yang, Jing-Hao Xue
ICME2
2021 EFRNet: A Lightweight Network with Efficient Feature Fusion and Refinement for Real-Time Semantic Segmentation
abstract
To pursue high accuracy, most image semantic segmentation methods are computationally costly and thus not suitable to real-time applications. Existing lightweight methods either adopt a single branch without feature fusion, which dam-ages accuracy, or introduce extra branches for feature fusion, which harms efficiency. In this paper, we propose a lightweight network named EFRNet, with feature fusion and refinement in a single branch to achieve better balance between accuracy and efficiency in real-time semantic segmentation. Specifically, in EFRNet, we design a novel Feature Fusion Module to fuse multi-stage features in a single CNN efficiently, and we propose a lightweight Channel Attention Refinement Module to refine features with few extra parameters. Extensive experiments show that our EFRNet achieves decent accuracy with an extremely small model size and high inference speed. It achieves the best accuracy of 70.02% mIoU compared with state-of-the-art lightweight methods on CamVid with only 0.48M parameters.
Kuayue Zhang, Qingmin Liao, Juncheng Zhang, Jing-Hao Xue
ICME2
2021 Group Fisher Pruning for Practical Network Compression
abstract
Network compression has been widely studied since it is able to reduce the memory and computation cost during inference. However, previous methods seldom deal with complicated structures like residual connections, group/depth-wise convolution and feature pyramid network, where channels of multiple layers are coupled and need to be pruned simultaneously. In this paper, we present a general channel pruning approach that can be applied to various complicated structures. Particularly, we propose a layer grouping algorithm to find coupled channels automatically. Then we derive a unified metric based on Fisher information to evaluate the importance of a single channel and coupled channels. Moreover, we find that inference speedup on GPUs is more correlated with the reduction of memory rather than FLOPs, and thus we employ the memory reduction of each channel to normalize the importance. Our method can be used to prune any structures including those with coupled channels. We conduct extensive experiments on various backbones, including the classic ResNet and ResNeXt, mobile-friendly MobileNetV2, and the NAS-based RegNet, both on image classification and object detection which is under-explored. Experimental results validate that our method can effectively prune sophisticated networks, boosting inference speed without sacrificing accuracy.
Liyang Liu, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xinjiang Wang, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
ICML9
2021 Hourglass Face Detector for Hard Face
abstract
Face detection is an upstream task of facial image analysis. In many real-world scenarios, we need to detect small, occluded or dense faces that are hard to detect, but hard face detection is a challenging task in particular considering the balance between accuracy and inference speed for real-world applications. This paper proposes an Hourglass Face Detector (HFD) for hard face by developing a deep one-stage fully-convolutional hourglass network, which achieves an excellent balance between accuracy and inference speed. To this end, the HFD firstly shrinks a feature map by a series of stridden convolutional layers rather than pooling layers, so that useful subtle information is preserved better. Secondly, it exploits context information by merging fine-grained shallow feature maps with deep ones full of semantic information, making a better fusion of detailed information and semantic information to achieve a better detection of small faces. Moreover, the HFD exploits prior and multiscale information from the training data to enhance its scale-invariance and adaptability of anchor scales. Compared with the SSH and S3FD methods, the HFD can achieve a better performance in average precision on detecting hard faces as well as a quicker inference. Experiments on the WIDER FACE and FDDB datasets demonstrate the superior performance of our proposed method.
Zijun Yu, Jian Yin 0016, Wenming Yang, Jing-Hao Xue, Qingmin Liao
IJCNN6
2021 Triplet Angular Loss for Pose-Robust Face Recognition
abstract
Although face recognition has been widely applied in many areas, pose-robust face recognition is still a challenging topic due to the large pose variations in real scenes. In this paper, we propose to learn the pose-robust face representation by normalizing the profile face in feature level directly and jointly considering both intra-class compactness and inter-class separability. Our approach minimizes the angular distance between the profile face and the positive frontal anchor. And it maximizes the angular distance between the profile face and the negative frontal anchor simultaneously. Furthermore, we modify the Triplet loss and derive the Triplet Angular loss to guarantee the intra-class compactness and the inter-class separability in angular space. In this way, the faces under varying poses can cluster compactly to create a pose-robust feature representation. Extensive experiments on two challenging benchmarks (CFP-FP and IJB-A) illustrate that our approach achieves a competitive performance in the field of pose-robust face recognition.
Zhenduo Zhang, Yongru Chen, Wenming Yang, Guijin Wang, Qingmin Liao
IJCNN5
2021 TimNet: A text-image matching network integrating multi-stage feature extraction with multi-scale metrics
Xiaoqi Zheng, Yingfan Tao, Ruikai Zhang, Wenming Yang, Qingmin Liao
Neurocomputing5
2021 Ripple-GAN: Lane Line Detection With Ripple Lane Line Detection Network and Wasserstein GAN
abstract
With artificial intelligence technology being advanced by leaps and bounds, intelligent driving has attracted a huge amount of attention recently in research and development. In intelligent driving, lane line detection is a fundamental but challenging task particularly under complex road conditions. In this paper, we propose a simple yet appealing network called Ripple Lane Line Detection Network (RiLLD-Net), to exploit quick connections and gradient maps for effective learning of lane line features. RiLLD-Net can handle most common scenes of lane line detection. Then, in order to address challenging scenarios such as occluded or complex lane lines, we propose a more powerful network called Ripple-GAN, by integrating RiLLD-Net, confrontation training of Wasserstein generative adversarial networks, and multi-target semantic segmentation. Experiments show that, especially for complex or obscured lane lines, Ripple-GAN can produce a superior detection performance to other state-of-the-art methods.
Youcheng Zhang, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Intell. Transp. Syst.5
2021 Understanding Urban Dynamics via State-Sharing Hidden Markov Model
abstract
With the ever-increasing urbanization process, systematically modeling people's activities in the urban space is being recognized as a crucial socioeconomic task. It is extremely challenging due to the lack of reliable data and suitable methods, yet the emergence of population-scale urban mobility data sheds new light on it. However, recent works on discovering activity patterns from urban mobility data are still limited in terms of concisely and specifically modeling the temporal dynamics of people's urban activities. To bridge the gap, we present a State-sharing Hidden Markov Model (SSHMM), a novel time-series modeling method that uncovers urban dynamics with massive urban mobility data. SSHMM models the urban dynamics from two aspects. First, it extracts the urban states from the whole city, which captures the volume of population flows as well as the frequency of each type of Point of Interests (PoIs) visited. Second, it characterizes the urban dynamics of each urban region as the state transition on the shared-states, which reveals distinct daily rhythms of urban activities. We evaluate our method via large-scale real-life mobility dataset. The results demonstrate that SSHMM learns semantics-rich urban dynamics, which are highly correlated with the functions of the region. Besides, it recovers the urban dynamics in different time slots with RMSE of 0.0793 when only learn limited states for the whole city, which outperforms the general HMM by 54.2 percent.
Tong Xia, Yong Li 0008, Fengli Xu, Qingmin Liao, Depeng Jin
IEEE Trans. Knowl. Data Eng.5
2021 Class-Variant Margin Normalized Softmax Loss for Deep Face Recognition
abstract
In deep face recognition, the commonly used softmax loss and its newly proposed variations are not yet sufficiently effective to handle the class imbalance and softmax saturation issues during the training process while extracting discriminative features. In this brief, to address both issues, we propose a class-variant margin (CVM) normalized softmax loss, by introducing a true-class margin and a false-class margin into the cosine space of the angle between the feature vector and the class-weight vector. The true-class margin alleviates the class imbalance problem, and the false-class margin postpones the early individual saturation of softmax. With negligible computational complexity increment during training, the new loss function is easy to implement in the common deep learning frameworks. Comprehensive experiments on the LFW, YTF, and MegaFace protocols demonstrate the effectiveness of the proposed CVM loss function.
Wanping Zhang, Yongru Chen, Wenming Yang, Guijin Wang, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Neural Networks Learn. Syst.6
2021 Clustering Through Probability Distribution Analysis Along Eigenpaths
abstract
Data clustering is one of the most fundamental techniques in exploratory data analysis. It is widely used for determining the underlying data structure, classifying natural data and compressing data in engineering, business management, social statistics, computer science, and medicine. Under the assumption that clusters are high density regions in the feature space separated by relatively low density neighbors, a novel approach is proposed for modeling any high dimensional clustering problem as a one-dimensional analysis of the probability distribution. First, a special path between two vertexes, namely eigenpath, is defined in this paper to represent their close connection. Second, we propose the connectedness index based on the eigenpath for quantitatively describing the connection between two vertexes. Third, the connectedness index is applied to the candidates of cluster centers and measures the connection between different candidates. Then an indicative curve can be drawn with the knowledge of connectedness index. This approach not only provides effective indicative curve for unknown data sets but also facilitates eliminating the curse of dimensionality partly as well as correctly recognizes arbitrary cluster forms and automatically excludes outliers. Extensive experiments showed the effectiveness and efficiency of the proposed approach.
Wenming Yang, Changqing Hui, Daren Sun, Xiang Sun 0003, Qingmin Liao
IEEE Trans. Syst. Man Cybern. Syst.5
2020 Noise-Sampling Cross Entropy Loss: Improving Disparity Regression Via Cost Volume Aware Regularizer
abstract
Recent end-to-end deep neural networks for disparity regression have achieved the state-of-the-art performance. However, many well-acknowledged specific properties of disparity estimation are omitted in these deep learning algorithms. Especially, matching cost volume, one of the most important procedure, is treated as a normal intermediate feature for the following softargmin regression, lacking explicit constraints compared with those traditional algorithms. In this paper, inspired by previous canonical definition of cost volume, we propose the noise-sampling cross entropy loss function to regularize the cost volume produced by deep neural networks to be unimodal and coherent. Extensive experiments validate that the proposed noise-sampling cross entropy loss can not only help neural networks learn more informative cost volume, but also lead to better stereo matching performance compared with several representative algorithms.
Zongqing Lu 0001, Xuechen Zhang 0003, Qingmin Liao
ICIP5
2020 Lightweight Single Image Super-Resolution Through Efficient Second-Order Attention Spindle Network
abstract
Recent years have witnessed great success of applying deep convolutional neural networks (CNNs) to single image super-resolution (SISR). However, most of these algorithms focus on increasing modeling capability through developing deeper and wider networks, improving the performance but at a cost of huge computation. Targeting at a better trade-off between efficiency and effectiveness, we propose ESASN, an efficient second-order attention spindle network for lightweight SISR. ESASN is built upon efficient second-order attention spindle (ESAS) blocks, each of which contains two well-designed new modules, efficient multi-scale (EMS) module and second-order attention (SOA) module. EMS reduces a considerable number of parameters while retaining the multi-scale structure to explore rich features. SOA further rescales the multi-scale feature maps, capturing the inter-dependencies among channels pixel-wisely with little additional cost. Both qualitative and quantitative experimental results demonstrate that the combination of EMS and SOA works out favorably for SISR, lifting the performance with fewer parameters. Code is available at https://github.com/yiyunchen/ESASN.
Jing-Hao Xue, Wenming Yang, Qingmin Liao
ICME5
2020 Two-Stage Adaptive Object Scene Flow Using Hybrid CNN-CRF Model
abstract
Scene flow estimation based on stereo sequences is a comprehensive task relevant to disparity and optical flow. Some existing methods are time-consuming and often fail in the presence of reflective surfaces. In this paper, we propose a two-stage adaptive object scene flow estimation method using a hybrid CNN-CRF model (ACOSF), which benefits from high-quality features and the structured modelling capability. Meanwhile, in order to balance the computational efficiency and accuracy, we employ adaptive iteration for energy function optimization, which is flexible and efficient for various scenes. Besides, we utilize high-quality pixel selection to reduce the computation time with only a slight decrease in accuracy. Our method achieves competitive results with the state-of-the-art, which ranks second on the challenging KITTI 2015 scene flow benchmark.
Qingmin Liao
ICPR3
2020 Attention Cube Network for Image Restoration
abstract
Recently, deep convolutional neural network (CNN) have been widely used in image restoration and obtained great success. However, most of existing methods are limited to local receptive field and equal treatment of different types of information. Besides, existing methods always use a multi-supervised method to aggregate different feature maps, which can not effectively aggregate hierarchical feature information. To address these issues, we propose an attention cube network (A-CubeNet) for image restoration for more powerful feature expression and feature correlation learning. Specifically, we design a novel attention mechanism from three dimensions, namely spatial dimension, channel-wise dimension and hierarchical dimension. The adaptive spatial attention branch (ASAB) and the adaptive channel attention branch (ACAB) constitute the adaptive dual attention module (ADAM), which can capture the long-range spatial and channel-wise contextual information to expand the receptive field and distinguish different types of information for more effective feature representations. Furthermore, the adaptive hierarchical attention module (AHAM) can capture the long-range hierarchical contextual information to flexibly aggregate different feature maps by weights depending on the global context. The ADAM and AHAM cooperate to form an 'attention in attention' structure, which means AHAM's inputs are enhanced by ASAB and ACAB. Experiments demonstrate the superiority of our method over state-of-the-art image restoration methods in both quantitative comparison and visual analysis.
Yucheng Hang, Qingmin Liao, Wenming Yang, Jie Zhou 0001
ACM Multimedia2
2020 Emotion Recognition with Facial Landmark Heatmaps
Siyi Mo, Wenming Yang, Guijin Wang, Qingmin Liao
MMM (1)4
2020 Defocus map estimation from a single image using improved likelihood feature and edge-based basis
Qingmin Liao, Jing-Hao Xue, Fei Zhou 0001
Pattern Recognit.2
2020 Real-MFF: A large realistic multi-focus image dataset with ground truth
Juncheng Zhang, Qingmin Liao, Wenming Yang, Jing-Hao Xue
Pattern Recognit. Lett.2
2020 Classifier shared deep network with multi-hierarchy loss for low resolution face recognition
Jingna Sun, Yehu Shen, Wenming Yang, Qingmin Liao
Signal Process. Image Commun.4
2020 Inter-class angular margin loss for face recognition
Jingna Sun, Wenming Yang, Riqiang Gao, Jing-Hao Xue, Qingmin Liao
Signal Process. Image Commun.5
2020 An α-Matte Boundary Defocus Model-Based Cascaded Network for Multi-Focus Image Fusion
abstract
Capturing an all-in-focus image with a single camera is difficult since the depth of field of the camera is usually limited. An alternative method to obtain the all-in-focus image is to fuse several images that are focused at different depths. However, existing multi-focus image fusion methods cannot obtain clear results for areas near the focused/defocused boundary (FDB). In this paper, a novel α-matte boundary defocus model is proposed to generate realistic training data with the defocus spread effect precisely modeled, especially for areas near the FDB. Based on this α-matte defocus model and the generated data, a cascaded boundary-aware convolutional network termed MMF-Net is proposed and trained, aiming to achieve clearer fusion results around the FDB. Specifically, the MMF-Net consists of two cascaded subnets for initial fusion and boundary fusion. These two subnets are designed to first obtain a guidance map of FDB and then refine the fusion near the FDB. Experiments demonstrate that with the help of the new α-matte boundary defocus model, the proposed MMF-Net outperforms the state-of-the-art methods both qualitatively and quantitatively.
Qingmin Liao, Juncheng Zhang, Jing-Hao Xue
IEEE Trans. Image Process.2
2020 LCSCNet: Linear Compressing-Based Skip-Connecting Network for Image Super-Resolution
abstract
In this paper, we develop a concise but efficient network architecture called linear compressing based skipconnecting network (LCSCNet) for image super-resolution. Compared with two representative network architectures with skip connections, ResNet and DenseNet, a linear compressing layer is designed in LCSCNet for skip connection, which connects former feature maps and distinguishes them from newly-explored feature maps. In this way, the proposed LCSCNet enjoys the merits of the distinguish feature treatment of DenseNet and the parametereconomic form of ResNet. Moreover, to better exploit hierarchical information from both low and high levels of various receptive fields in deep models, inspired by gate units in LSTM, we also propose an adaptive element-wise fusion strategy with multisupervised training. Experimental results in comparison with state-of-the-art algorithms validate the effectiveness of LCSCNet.
Wenming Yang, Xuechen Zhang 0003, Yapeng Tian, Wei Wang 0194, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Image Process.6
2020 DeepApp: Predicting Personalized Smartphone App Usage via Context-Aware Multi-Task Learning
abstract
Smartphone mobile application (App) usage prediction, i.e., which Apps will be used next, is beneficial for user experience improvement. Through an in-depth analysis on a real-world dataset, we find that App usage is highly spatio-temporally correlated and personalized. Given the ability to model complex spatio-temporal contexts, we aim to apply deep learning to achieve high prediction accuracy. However, the personalization yields a problem: training one network for each individual suffers from data scarcity, yet training one deep neural network for all users often fails to uncover user preference. In this article, we propose a novel App usage prediction framework, named DeepApp , to achieve context-aware prediction via multi-task learning. To tackle the challenge of data scarcity, we train one general network for multiple users to share common patterns. To better utilize the spatio-temporal contexts, we supplement a location prediction task in the multi-task learning framework to learn spatio-temporal relations. As for the personalization, we add a user identification task to capture user preference. We evaluate DeepApp on the large-scale dataset by extensive experiments. Results demonstrate that DeepApp outperforms the start-of-the-art baseline by 6.44%.
Tong Xia, Yong Li 0008, Jie Feng 0002, Depeng Jin, Hengliang Luo, Qingmin Liao
ACM Trans. Intell. Syst. Technol.7
2020 An Equalized Margin Loss for Face Recognition
abstract
In this paper, we propose a new loss function, termed the equalized margin (EqM) loss, which is designed to make both intra-class scopes and inter-class margins similar over all classes, such that all the classes can be evenly distributed on the hypersphere of the feature space. The EqM loss controls both the lower limit of intra-class similarity by exploiting hard-sample mining and the upper limit of inter-class similarity by assuring equalized margins. Therefore, using the EqM loss, we can not only obtain more discriminative features, but also overcome the negative impacts from the data imbalance on the inter-class margins. We also observe that the EqM loss is stable with the variation of the scale in normalized Softmax. Furthermore, by conducting extensive experiments on LFW, YTF, CFP, MegaFace and IJB-B, we are able to verify the effectiveness and superiority of the EqM loss, compared with other state-of-the-art loss functions for face recognition.
Jingna Sun, Wenming Yang, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Multim.4
2019 A Universal Fusion Strategy for Image Super-Resolution Jointly from External and Internal Examples
Wei Wang 0194, Xuesen Shang, Wenming Yang, Canrong Zhang, Qingmin Liao
ICIG (1)5
2019 Temporal Feature Enhancing Network for Human Pose Estimation in Videos
abstract
Although state-of-the-art methods for human pose estimation have achieved superior results on the single image, their performance on videos usually deteriorates dramatically due to motion blur and occlusion. Since there is close temporal correlation among video frames, exploiting the contextual information properly can be helpful to tackle the problem. In this paper, we present a Temporal Feature Enhancing Network (TFEN) for video human pose estimation. It boosts the per-frame features by utilizing motion information in terms of optical flow and conducting temporal feature encoding by the convolution gated recurrent units (convGRU). It is an end-to-end learning framework and can extend any image based algorithm to video pose estimation. The experimental results validate the effectiveness of the proposed approach on two large-scale video pose estimation benchmarks.
Haihan Li, Wenming Yang, Qingmin Liao
ICIP3
2019 Exploring Discriminative Features in Mueller Matrix Images for Electrospinning Classification
abstract
Polarization images, which are captured in lights with different polarization angles, can extract more detail information about samples. Generally, there are two ways to make use of polarization images: direct processing of original polarization images and processing of Mueller matrix (MM) images. Since MM has clear physical meaning and each element in it represents a specific characteristic about samples, studying the relationship between the elements in MM and samples is a meaningful topic. In this paper, an importance sorting algorithm is proposed to explore discriminative elements in MM. Firstly, a linear weighted feature fusion method is proposed and three distances are defined to form the target function. Then, a convex quadratic programming model is built, with an algorithm to search the optimal solution. Finally, discriminative elements are choosed for classification according to the optimal weight vector. Experiments conducted on an electrospinning dataset show that the proposed method not only provides a consistent importance order of elements in MM, but also helps to find discriminative feature combinations for classification, which is useful for explaining of the polarization characteristics of samples. The source code is available at: https://github.com/madd2014/ImportanceSort.
Zongqing Lu 0001, Youcheng Zhang, Qingmin Liao
ICIP4
2019 Estimating Human Shape Under Clothing from Single Frontal View Point Cloud of a Dressed Human
abstract
Estimating human shape under clothing is a challenging task. We propose the first method to estimate accurate shape parameters from single-frame frontal view point cloud. To account for casual clothing, we personalize the original SMPL model to describe clothing as deviation from naked human parametric model, define a novel method to search for corresponding vertex pairs, and design a novel objective function that enforces point cloud vertices to remain outside of the naked body shape and tightly cling the personalized shape. Consolidating these three parts, our method integrates the advantages of free deformation method and model-based method. Our method is more effective than previous works in dealing with casual clothing situation. We evaluate the accuracy of estimated shape on noisy point cloud data captured by a commodity depth sensor.
Zongqing Lu 0001, Qingmin Liao
ICIP3
2019 Boundary Aware Multi-focus Image Fusion Using Deep Neural Network
abstract
Since it is usually difficult to capture an all-in-focus image of a 3D scene directly, various multi-focus image fusion methods are employed to generate it from several images focusing at different depths. However, the performance of existing methods is barely satisfactory and often degrades for areas near the focused/defocused boundary (FDB). In this paper, a boundary aware method using deep neural network is proposed to overcome this problem. (1) Aiming to acquire improved fusion images, a 2-channel deep network is proposed to better extract the relative defocus information of the two source images. (2) After analyzing the different situations for patches far away from and near the FDB, we use two networks to handle them respectively. (3) To simulate the reality more precisely, a new approach of dataset generation is designed. Experiments demonstrate that the proposed method outperforms the state-of-the-art methods, both qualitatively and quantitatively.
Juncheng Zhang, Qingmin Liao
ICME4
2019 A New Object Scene Flow Algorithm Based on Support Points Selection and Robust Moving Object Proposal
abstract
Recent algorithms of object scene flow estimation suffer from low computational efficiency or unstable moving object proposals. To tackle these two problems simultaneously, in this paper we propose a new, efficient and robust algorithm for object scene flow estimation, through making two technical contributions. Firstly to improve the efficiency, we propose to select only a few pixels termed support points for matching cost calculation rather than using all pixels. The support points are defined as those pixels with high confidence in feature matching. Secondly to attain stable moving object proposals, we propose a motion magnitude-adaptive thresholding scheme for ego-motion outlier detection, after patch matching on CNN-extracted high quality features. These two contributions, though simple, ensure a remarkable improvement in both efficiency and accuracy from the original object scene flow method, as well as making the proposed algorithm a strong practicable alternative to much more sophisticated state-of-the-art competitors.
Zhengyang Sun, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME4
2019 A New Approach to Automatic Clothing Matting from Mannequins
abstract
It is crucial to extract retail clothes from images of mannequins when building a database of clothing images for virtual try-on systems. However, clothes often have complex texture and translucent material, such as holes and laces. It is thus difficult to extract clothes as foreground by existing generic natural image matting methods. Hence in this paper, we present a novel approach to automatic clothing matting from mannequins, with auxiliary information from a rough background image of the mannequin only. Experiments show that we can achieve remarkable improvement on the alpha matte near challenging regions of complex texture and translucent material of clothes. Moreover, our approach can automatically generate trimaps to facilitate the development and evaluation of other image matting algorithms.
Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME4
2019 A New Rotation-Invariant Deep Network for 3D Object Recognition
abstract
When inputs are rotated, most 3D convolutional neural networks (CNNs) will have their performance much dropped, especially for those models with voxelized input of 3D objects. The newly proposed Spherical CNNS, with the concept of the rotation-equivariant spherical correlation, aims to achieve rotation invariance. Inspired by this, we propose a new rotation-invariant deep network to recognize rotated 3D objects. Specifically, we adopt the spherical representation and the spherical correlation S^2 layer of Spherical CNNs, for their capacity of representing 3D objects and rotation equivariance. In the meantime, we improve the computational efficiency and expressiveness of Spherical CNNs, by replacing its time-consuming and depth-limited SO(3) layer with a PointNet-style network architecture. Hence our proposed network can maintain the equivariance as the network grows deeper while substantially reducing its runtime, leading to a much better efficiency and expressiveness of rotation-invariant representation. Experimental results show that our network performs better than or comparable to the state-of-the-art methods in the ModelNet40 classification challenge.
Yachi Zhang, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME4
2019 Local polynomial contrast binary patterns for face recognition
Yinyan Jiang, Yicong Zhou, Weifeng Li 0001, Qingmin Liao
Neurocomputing6
2019 A hybrid finger identification pattern using Polarized depth-Weighted Binary Direction Coding
Wenming Yang, Wenyang Ji, Jing-Hao Xue, Qingmin Liao
Neurocomputing5
2019 Adaptive local-fitting-based active contour model for medical image segmentation
Qingmin Liao, Ziqin Chen, Ran Liao, Hui Ma 0003
Signal Process. Image Commun.2
2019 Lightweight Feature Fusion Network for Single Image Super-Resolution
abstract
Single image super-resolution (SISR) has witnessed great progress as convolutional neural network (CNN) gets deeper and wider. However, enormous parameters hinder its application to real world problems. In this letter, We propose a lightweight feature fusion network (LFFN) that can fully explore multi-scale contextual information and greatly reduce network parameters while maximizing SISR results. LFFN is built on spindle blocks and a softmax feature fusion module (SFFM). Specifically, a spindle block is composed of a dimension extension unit, a feature exploration unit. and a feature refinement unit. The dimension extension layer expands low dimension to high dimension and implicitly learns the feature maps which are suitable for the next unit. The feature exploration unit performs linear and nonlinear feature exploration aimed at different feature maps. The feature refinement layer is used to fuse and refine features. SFFM fuses the features from different modules in a self-adaptive learning manner with softmax function, making full use of hierarchical information with a small amount of parameter cost. Both qualitative and quantitative experiments on benchmark datasets show that LFFN achieves favorable performance against state-of-the-art methods with similar parameters.
Wenming Yang, Wei Wang 0194, Xuechen Zhang 0003, Shuifa Sun, Qingmin Liao
IEEE Signal Process. Lett.5
2019 $\alpha$ -Trimmed Weber Representation and Cross Section Asymmetrical Coding for Human Identification Using Finger Images
abstract
In this paper, a novel method that utilizes feature-level fusion of finger vein (FV) and finger dorsal texture (FDT) images is proposed for human identification. Motivated by Weber's law, we present α-trimmed Weber representation (α-TWR) to enhance the foreground lines (FLs), i.e., vessels underneath skin and line-like texture on skin. The proposed α-TWR is robust to illumination variation, as validated by a basic reflective and transmitted imaging model. Cross section asymmetrical coding (CSAC) is performed to extract features for each pixel. The coding value contains discriminative information on the orientation and internal point location of the FLs. The CSAC values of FV and FDT in each point are abreast in terms of binary representation. Local density weighted matching is developed to obtain the matching score between two feature maps. We experimentally show that the proposed method outperforms other unimodal and multimodal identification methods in terms of equal-error-rate.
Wenming Yang, Zhiquan Chen, Qingmin Liao
IEEE Trans. Inf. Forensics Secur.4
2019 FV-GAN: Finger Vein Representation Using Generative Adversarial Networks
abstract
In finger vein verification, the most important and challenging part is to robustly extract finger vein patterns from low-contrast infrared finger images with limited a priori knowledge. Although recent convolutional neural network (CNN)-based methods for finger vein verification have shown powerful capacity for feature representation and promising perspective in this area, they still have two critical issues to address. First, these CNN-based methods unexceptionally utilize fully connected layers, which restrict the size of finger vein images to process and increase the processing time. Second, the capacity of CNN for feature representation generally suffers from the low quality of finger vein ground-truth pattern maps for training, particularly due to outliers and vessel breaks. To address these issues, in this paper, we propose a novel approach termed FV-GAN to finger vein extraction and verification, based on generative adversarial network (GAN), as the first attempt in this area. Unlike the CNN-based methods, FV-GAN learns from the joint distribution of finger vein images and pattern maps rather than the direct mapping between them, with the aim at achieving stronger robustness against outliers and vessel breaks. Moreover, FV-GAN adopts fully convolutional networks as the basic architecture and discards fully connected layers, which relaxes the constraint on the input image size and reduces the computational expenditure for feature extraction. Furthermore, we design an adversarial training strategy and propose a hybrid loss function for FV-GAN. The experimental results on two public databases show significant improvement by FV-GAN in finger vein verification in terms of both verification accuracy and equal error rate.
Wenming Yang, Changqing Hui, Zhiquan Chen, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Inf. Forensics Secur.5
2019 Deep Learning for Single Image Super-Resolution: A Brief Review
abstract
Single image super-resolution (SISR) is a notoriously challenging ill-posed problem that aims to obtain a high-resolution output from one of its low-resolution versions. Recently, powerful deep learning algorithms have been applied to SISR and have achieved state-of-the-art performance. In this survey, we review representative deep learning-based SISR methods and group them into two categories according to their contributions to two essential aspects of SISR: The exploration of efficient neural network architectures for SISR and the development of effective optimization objectives for deep SISR learning. For each category, a baseline is first established, and several critical limitations of the baseline are summarized. Then, representative works on overcoming these limitations are presented based on their original content, as well as our critical exposition and analyses, and relevant comparisons are conducted from a variety of perspectives. Finally, we conclude this review with some current challenges and future trends in SISR that leverage deep learning algorithms.
Wenming Yang, Xuechen Zhang 0003, Yapeng Tian, Wei Wang 0194, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Multim.6
2018 Full-Reference Quality Assessment of Contrast Changed Images Based on Local Linear Model
abstract
This paper presents a new full-reference method to assess the quality of contrast changed images. In this method, we employ a linear model to describe the relationship between local patches of reference images and contrast changed images. With parameters of this model, three quality measures considering contrast comparison, structure variation, and luminance change are defined. Among them, the first measure produces larger quality scores for higher contrast, which is different from traditional forms of quality measures used in most existing full-reference methods. Experiments on four benchmark databases show that the proposed method is superior to state-of-the-art methods in assessing the quality of contrast changed images.
Wenming Yang, Fei Zhou 0001, Qingmin Liao
ICASSP4
2018 Tree-Shaped Sampling Based Hybrid Multi-Scale Feature Extraction for Texture Classification
abstract
Efficiency, distinctiveness and robustness are three main goals for feature extractors in application of texture classification. In this paper, a new feature extractor is designed which aims to achieve these three goals simultaneously. The contributions are threefold. Firstly, a tree-shaped multi-scale sampling structure is proposed to acquire points distributed along two circles and one octagon. Secondly, four histogram vectors are obtained by quantizing the sampling values through a hybrid strategy. In order to suppress the noise, mean filtering is used as a preprocessing step and the four vectors are concatenated to form the discriminant vector. Thirdly, experiments are conducted on different datasets with several well-known feature extractors. The results show that the proposed method improves the classification accuracy effectively and robustly, while has a moderate complexity. The source code is available at: https://github.com/madd2014/TSSHM.
Ziqin Chen, Qingmin Liao
ICIP3
2018 No-reference image quality assessment for photographic images based on robust statistics
Zhengda Zeng, Wenming Yang, Jing-Hao Xue, Qingmin Liao
Neurocomputing5
2018 Binarized features with discriminant manifold filters for robust single-sample face recognition
Wanping Zhang, Zongqing Lu 0001, Weifeng Li 0001, Qingmin Liao
Signal Process. Image Commun.6
2018 Margin Loss: Making Faces More Separable
abstract
The key point of face recognition is creating a discriminative feature representation to ensure intraclass compactness and interclass separability. Softmax loss is widely used in deep learning networks, but it is indirect for face verification. Center loss is effective to improve intraclass compactness, while interclass distances are ignored. In this letter, we propose a novel loss function, termed margin loss, to enlarge distances of interclass and reduce intraclass variations simultaneously. Margin loss aims to focus on samples hard to classify by a distance margin. Different from Softmax loss, margin loss is based on Euclidean distances that can directly measure face similarity. Experiments on different datasets have demonstrated the effectiveness of our method.
Riqiang Gao, Fuwei Yang, Wenming Yang, Qingmin Liao
IEEE Signal Process. Lett.4
2018 Discriminative Multidimensional Scaling for Low-Resolution Face Recognition
abstract
Face images captured by surveillance videos usually have limited resolution. Due to resolution mismatch, it is hard to match high-resolution (HR) faces with low-resolution (LR) faces directly. Recently, multidimensional scaling (MDS) has been employed to solve the problem. In this letter, we proposed a more discriminative MDS method to learn a mapping matrix, which projects the HR images and LR images to a common subspace. Our method is discriminative since both interclass distances and intraclass distances are taken into consideration. We add an interclass constraint to enlarge the distances of different subjects in the subspace to ensure discriminability. Besides, we consider not only the relationship of HR-LR images, but also the relationship of HR-HR images and LR-LR images in order to preserve local consistency. Experimental results on FERET, Multi-PIE, and SCface databases demonstrate the effectiveness of our proposed approach.
Fuwei Yang, Wenming Yang, Riqiang Gao, Qingmin Liao
IEEE Signal Process. Lett.4
2018 SPSIM: A Superpixel-Based Similarity Index for Full-Reference Image Quality Assessment
abstract
Full-reference image quality assessment algorithms usually perform comparisons of features extracted from square patches. These patches do not have any visual meanings. On the contrary, a superpixel is a set of image pixels that share similar visual characteristics and is thus perceptually meaningful. Features from superpixels may improve the performance of image quality assessment. Inspired by this, we propose a new superpixel-based similarity index by extracting perceptually meaningful features and revising similarity measures. The proposed method evaluates image quality on the basis of three measurements, namely, superpixel luminance similarity, superpixel chrominance similarity, and pixel gradient similarity. The first two measurements assess the overall visual impression on local images. The third measurement quantifies structural variations. The impact of superpixel-based regional gradient consistency on image quality is also analyzed. Distorted images showing high regional gradient consistency with the corresponding reference images are visually appreciated. Therefore, the three measurements are further revised by incorporating the regional gradient consistency into their computations. A weighting function that indicates superpixel-based texture complexity is utilized in the pooling stage to obtain the final quality score. Experiments on several benchmark databases demonstrate that the proposed method is competitive with the state-of-the-art metrics.
Qingmin Liao, Jing-Hao Xue, Fei Zhou 0001
IEEE Trans. Image Process.2
2017 Wavelet-based single image super-resolution with an overall enhancement procedure
abstract
In this paper, we address the problem of generating a super-resolution image based on a dictionary of low- and high-resolution exemplars from a single input image in wavelet domain with a overall enhancement procedure. Most methods extract different kinds of features in low-resolution image and high-resolution images to establish the mapping relation. But in this paper, we implement wavelet-transform to extract the same kind of feature to make the mapping more reasonable. Meanwhile we implement local Lipschitz regularity constraint and structure-keeping constraint to preserve the local singularity and edge in our method. Compared with current state-of-art methods on standard images, our method obtains both visual and PSNR improvement.
Zongqing Lu 0001, Quan Zou 0001, Fei Zhou 0001, Qingmin Liao
ICASSP4
2017 A robust feature descriptor based on multiple gradient-related features
abstract
In this paper, we propose a robust descriptor named as multiple gradient-related features (MGRF) in virtue of local and overall order encoding. Specifically, three types of features are introduced, including multidirectional gradient, gradient orientation, and first derivative of gradient orientation, each of which represents different aspect of region of interest (ROI). To extract these features, we also propose a novel sampling pattern of tree structure. Furthermore, each gradient-related feature is encoded with both local and overall order information of ROI, and the encoding results are respectively called local and overall gradient order code (GOC). Finally, our descriptor is formed by concatenating the respective feature vector of each type of feature, which is computed as a 2-D joint histogram of GOC and ordinal bin. The experiments conducted on Oxford dataset demonstrate that the proposed descriptor significantly outperforms other state-of-the-art descriptors.
Zhaomang Sun, Fei Zhou 0001, Qingmin Liao
ICASSP3
2017 Locality Sensitive Hashing based deepmatching for optical flow estimation
abstract
DeepMatching (DM) is one of the state-of-art matching algorithms to compute quasi-dense correspondences between images. Recent optical flow methods use DeepMatching to find initial image correspondences and achieves outstanding performance. However, the key building block of DeepMatching, the correlation map computation, is time-consuming. In this paper, we propose a new algorithm, LSHDM, which addresses the problem by employing Locality Sensitive Hashing (LSH) to DeepMatching. The computational complexity is greatly reduced for the correlation map computation step. Experiments show that image matching can be accelerated by our approach in ten times or more compared to DeepMatching, while retaining comparable accuracy for optical flow estimation.
Zongqing Lu 0001, Qingmin Liao, Danyi Li
ICASSP3
2017 Illumination-robust face recognition with Block-based Local Contrast Patterns
abstract
This paper proposes a novel facial image representation Block-based Local Contrast Patterns (BLCP) for illumination-robust face recognition. This method is based on an effective texture descriptor local contrast patterns (LCP). We use the directed and undirected difference masks to calculate three types of local intensity contrasts: directed, undirected, and maximum difference responses. These response images are divided into several nonoverlapping blocks. In each block these responses are quantized and encoded into specific patterns. A joint histogram of these patterns is computed for each block and then we concatenate all the blocks' histograms into an enhanced feature vector to be used as a face descriptor. The experimental results on Extended Yale-B and FERET databases illustrate the effectiveness of our proposed method in illumination-robust face recognition.
Weifeng Li 0001, Qingmin Liao
ICASSP4
2017 Learning adaptive local distance metric for face hallucination
abstract
In this paper, we propose a novel method for face hallucination by learning a new distance metric in the low-resolution (LR) patch space (source space). Local patch-based face hallucination methods usually assume that the two manifolds formed by LR and high-resolution (HR) image patches have similar local geometry. However, this assumption does not hold well in practice. Motivated by metric learning in machine learning, we propose to learn a new distance metric in the source space, under the supervision of the true local geometry in the target space (HR patch space). The learned new metric gives more freedom to the presentation of local geometry in the source space, and thus the local geometries of source and target space turn to be more consistent. Experiments conducted on two datasets demonstrate that the proposed method is superior to the state-of-the-art face hallucination and image super-resolution (SR) methods.
Yuanpeng Zou, Fei Zhou 0001, Qingmin Liao
ICASSP3
2017 Run-Based Connected Components Labeling Using Double-Row Scan
Qingmin Liao
ICIG (3)3
2017 Real-time 3D face reconstruction from one single image by displacement mapping
abstract
In this paper, we present a fast and robust method to reconstruct a plausible three-dimension (3D) face from one single frontal face image. In training phase, we classify the faces into several groups based on the facial structures and propose to learn a mapping, known as the displacement mapping (DM) in this paper, for each group. DM relates two displacements: One displacements, denoted as 2D displacements, represent the differences between the positions of feature points on the 2D training faces and those on the reference 2D face that has been pre-defined for the corresponding group; another displacements, denoted as 3D displacements, are the differences between the positions of vertices on the reconstructed 3D face and those on the reference 3D face that is also pre-defined. During the reconstruction phase, we first classify the input face as one of the groups and calculate the 2D displacements. Then we take advantage of the 2D displacements and the learned DM to estimate the 3D displacements. Subsequently, 3D displacements can be used to obtain the precise 3D face by shifting the 3D reference face. Experiments on Basel face model (BFM) database as well as some real-world 2D face images demonstrate the effectiveness and efficiency of the proposed method, in comparison with some state-of-arts methods.
Fei Zhou 0001, Qingmin Liao
ICIP3
2017 Microstructure analysis of silk samples using mueller matrix determination and sparse representation
abstract
In this paper, we propose to use Mueller matrix determination and sparse representation to classify silk samples washed in different detergents. Different detergents have different effects on the same silk samples after washing, and we distinguish their diversities in the Mueller matrix images(MMI) instead of visible light images(VLI). Compared with VLI, Mueller matrix, also known as polarization image, reflects the wavelength-scale microstructure and some optical properties of samples, and focuses on extracting the index to research the polarization property. To achieve a good performance with the microstructure analysis, we utilize the method of sparse representation which uses the reconstruction error for classification. Generally speaking, we introduce to combine Mueller matrix with sparse representation in the classification of the same silk samples washed in different detergents, and the high precision in experimental results indicates that our method works well.
Fei Zhou 0001, Hui Ma 0003, Qingmin Liao
ICIP5
2017 Face recognition via weighted sparse representation using metric learning
abstract
Face recognition methods utilizing Sparse Representation based Classification (SRC) and Collaborative Representation based Classification (CRC) have recently attracted a great deal of attention due to inherent simplicity and efficiency. In this paper, we introduce the Large Margin Nearest Neighbor (LMNN), which learns a Mahalanobis distance metric that is applied, to SRC and CRC as the locality constraint. Next, a locality LMNN Weighted Sparse Representation based Classification (LMNN-WSRC) and a locality LMNN Weighted Collaborative Representation based Classification (LMNN-WCRC) are proposed. Our methods utilize both linearity and data locality. For a query face image, our target is to exploit the appropriate distance metric as the locality constraint that could focus more on those truly related images in the code book. Experimental results on the Extended Yale B database and the AR database show that our methods are more effective than SRC, Weighted SRC (WSRC) and CRC.
Zongqing Lu 0001, Bokun Xu, Qingmin Liao
ICME4
2017 Weighted contourlet binary patterns and image-based fisher linear discriminant for face recognition
Weifeng Li 0001, Yinyan Jiang, Zongqing Lu 0001, Qingmin Liao
Neurocomputing6
2017 Scale the Internet routing table by generalized next hops of strict partial order
Qing Li 0006, Mingwei Xu 0001, Qi Li 0002, Dan Wang 0002, Yong Jiang 0001, Shutao Xia, Qingmin Liao
Inf. Sci.7
2017 MDID: A multiply distorted image database for image quality assessment
Fei Zhou 0001, Qingmin Liao
Pattern Recognit.3
2017 Cascaded Elastically Progressive Model for Accurate Face Alignment
abstract
While recently published face alignment algorithms mainly focused on occlusion, low image quality, and complex head poses, subtle variances of facial components were often overlooked. In this correspondence paper, we propose a new approach called cascaded elastically progressive model aiming for pixel-wise landmark localization. First of all, elastically progressive model (EPM) is designed to synthesize the prior knowledge of face shape and appearance of test image. More specifically, a novel framework referred to as inherent linear structure (ILS) is explored for capturing the characteristics of the shape, which is more plastic and flexible than extensively used principle component analysis-based modeling. A locally linear support vector machine (LL-SVM) is used as local expert for searching candidate feature points. In order to optimally integrate ILS with localization results of LL-SVM, we introduce Kalman filter (KF) to dynamically estimate the true shape in the sense of least mean square error. Two schemes are utilized based on our modeling of KF. First, we embedded heuristic line-like search strategy into the framework to guarantee and accelerate the convergence. Second, Kalman gain is manipulated adaptively in accordance with the confidence of the localizers so that poorly localized points are more subject to global constraint than well localized ones. To further improve robustness to initializations, two EPMs are cascaded, in which primary EPM detects the global structure and secondary EPM captures the details. Validation experiments are conducted on in-the-wild LFPW and HELEN databases. Our method shows advantages for accurate landmark localization compared with prevailing methods.
Wenming Yang, Xiang Sun 0003, Qingmin Liao
IEEE Trans. Syst. Man Cybern. Syst.3
2017 Single-Image Super-Resolution by Subdictionary Coding and Kernel Regression
abstract
In this paper, we present a new learning-based single-image super-resolution (SR) approach, inspired by existing sparse representation-based methods. As a promising image modeling theory, sparse representation has been effectively applied to solve the image SR problem, usually with the use of pretrained coupled or semi-coupled dictionaries. In our proposed method, we train independent dictionaries for high-resolution (HR) and low-resolution (LR) image patches to endow them more flexibility of expression. We use local subdictionaries to adaptively code image patches, which can characterize image local structures better and ensure the sparsity property of the image. Furthermore, we use kernel regression to relate HR and LR coding coefficients to capture and map the intrinsic nonlinear relationship between them. Such mapping is of central importance in the image SR problem, because high-order statistics play a significant role in the reconstruction of the detail structure of an HR image. The proposed model is generic for image SR in terms of two categories of blurring kernel. Experimental results show that our method can effectively reconstruct image details and outperform state-of-the-art algorithms in both quantitative and visual comparisons.
Wenming Yang, Tingrong Yuan, Wei Wang 0194, Fei Zhou 0001, Qingmin Liao
IEEE Trans. Syst. Man Cybern. Syst.5
2016 Saliency detection based on integration of central bias, reweighting and multi-scale for superpixels
abstract
Saliency detection has been a significant problem in computer vision and helpful to object detection. In this paper, we propose a new computational saliency detection model under the Bayesian framework. First, central bias and the reweighting of the salient regions in the convex hull are applied to guide the prior map. Then, multi-scale for superpixels is proposed to detect objects with various scales. At last, the Bayes formula is adopted to obtain the final saliency map. Experimental results on a standard database show that the proposed model outperforms state-of-the-art methods.
Xiaoling Hu 0002, Wenming Yang, Fei Zhou 0001, Qingmin Liao
ICASSP4
2016 An efficient anomaly detection approach in surveillance video based on oriented GMM
abstract
The detection and localization of abnormal activities are considered in this work. An efficient approach called oriented G-MM(OGMM) is proposed. The approach uses optical flow as low-level feature and quantizes the orientation of optical flow into 8 sections. In training stage, the approach will learn a GMM model at each orientation section and each position. In testing stage, the proposed approach estimates the probability of whether a position is abnormal using likelihood method. The proposed approach is a local method and can detect and locate anomaly. What's more, in the proposed approach, the same process is done to each position with little interaction between different positions. This makes the approach suit for parallel computing and can deal with large-scale tasks in Big Data times. The experiments verify that the proposed approach is efficient and effective.
Feiping Li, Wenming Yang, Qingmin Liao
ICASSP3
2016 Face recognition with local contourlet combined patterns
abstract
This paper proposes a novel face image descriptor called local contourlet combined patterns (LCCP), based on the Non-Subsampled Contourlet Transform (NSCT), for face recognition. NSCT is a multiresolution analysis tool and can capture image information at multiple scales, orientations, and frequency bands. To adapt to the NSCT filter bank, a new encoding method named mean-based contrast patterns (MCP) is presented. We apply LBP and MCP to different levels' NSCT coefficient images respectively and then combine them to obtain a robust representation. Futhermore, block-based kernel Fisher linear discriminant (BKFLD) is used to select the most discriminative feature sets. Face recognition experiments on FERET database demonstrate the effectiveness of our proposed approach.
Shilian Yu, Weifeng Li 0001, Longbiao Wang, Qingmin Liao
ICASSP5
2016 A fast 3D face reconstruction method from a single image using adjustable model
abstract
In this paper, we propose a fast and robust method which uses only a single frontal face image as input to reconstruct a plausible 3D face. Our method mainly consists of three stages: feature point detection, model adaptation in X-Y plane and model adjustment on Z-axis direction. At first stage, we detect some face regions such as face contour and facial components automatically. In these regions, we extract several feature points which can generally describe the structure of face. Subsequently, we apply several deformation processes and optimization procedures on an adjustable 3D face model in the X-Y plane based on these feature points. Finally, we present a method of insertion to obtain a dense and smooth model. Experimental results demonstrate the effectiveness and efficiency of our method as well as the robust adaptation to the complex imaging condition.
Fei Zhou 0001, Qingmin Liao
ICASSP3
2016 Anchored neighborhood regression based single image super-resolution from self-examples
abstract
In this paper, we present a novel self-learning single image super-resolution (SR) method, which restores a high-resolution (HR) image from self-examples extracted from the low-resolution (LR) input image itself without relying on extra external training images. In the proposed method, we directly use sampled image patches as the anchor points, and then learn multiple linear mapping functions based on anchored neighborhood regression to transform LR space into HR space. Moreover, we utilize the flipped and rotated versions of the self-examples to expand the internal patch space. Experimental comparison on standard benchmarks with state-of-the-art methods validates the effectiveness of the proposed approach.
Yapeng Tian, Fei Zhou 0001, Wenming Yang, Xuesen Shang, Qingmin Liao
ICIP5
2016 MSRT: Multi-Source Request and Transmission in Content-Centric Networks
abstract
In Content-Centric Networks (CCN), multiple routers may cache the same content, which makes it possible to retrieve the content chunks in parallel. In this paper, we propose Multi-Source Request and Transmission mechanism (MSRT) for CCN. We develop a MinMax problem to compute the optimal solution to retrieve all the chunks from multiple sources in the shortest time. We prove that the problem is NP complete and thus design a fully polynomial-time approximation algorithm to solve this problem. However, the previous works on multipath congestion control cannot be directly employed in MSRT. Therefore, we then propose the Half eXplicit Congestion Protocol (HXCP) to control the request/transmission pace in MSRT. To demonstrate the performance of MSRT, we construct comprehensive experiments. The results show that 1) our scheme reduces the content transmission time to at most 80%; 2) our multipath congestion control scheme HXCP effectively avoids congestion, improves the throughput and guarantees the fairness in the multi-source/multipath scenario.
Qing Li 0006, Bin Gan, Guangwu Hu, Yong Jiang 0001, Qingmin Liao, Mingwei Xu 0001
IWQoS5
2016 FICUS: Fast Incremental Consistent Update in SDN based on relation graph
abstract
In Software Defined Networking (SDN), the configuration inconsistency during updates is one main source of network instability. An efficient updating scheme with configuration consistency is required. In this paper, we propose the scheme of Fast Incremental Consistent Update for SDN (FICUS) based on the relation graph (RG). In our scheme, we analyse the relation between update operations, construct the relation graph and find a proper order of these update operations to avoid inconsistency. To solve the problem, we define two types of relations: the path dependency relation and the path rejection relation. We evaluate our scheme and algorithms by comprehensive experiments. The results show that our scheme needs only 10%–40% of the rules compared with the two-phase update scheme and speeds up the update process by 40% in average.
Qing Li 0006, Lei Wang 0071, Yong Jiang 0001, Guangwu Hu, Mingwei Xu 0001, Qingmin Liao
IWQoS6
2016 Visual domain adaptation using weighted subspace alignment
abstract
Domain Adaptation (DA) has attracted a lot of attention in recent years. DA aims at overcoming the covariate shift in dataset and aligning multiple existing but partially related data collections. In this paper, we propose a new DA algorithm which aligns the weighted subspaces generated from source samples and target samples. The weighted subspaces of source samples are generated using weighted Principal Component Analysis (PCA). Specifically, the source samples closer to the target domain are given higher weights during the construction of subspaces, which is definitely beneficial for building an adaptable classifier. Subsequently, the weighted subspaces of source samples and the subspaces of target samples are aligned to achieve domain adaptation. Experimental results on standard datasets demonstrate the advantages of our approach over state-of-the-art DA approaches.
Shuo Chen 0010, Fei Zhou 0001, Qingmin Liao
VCIP3
2016 Two-stage patch-based sparse multi-value descriptor for face recognition
abstract
In this paper, we propose Two-stage Patch-based Sparse Multi-value Descriptor (TPSMD), a generalization of Sparse Linear Regression Binary method. The TPSMD makes two contributions. First, the multi-value strategy introduces user-specified parameters to improve the binarization, which makes our method more discriminant and less sensitive to noise. The multi-value strategy is a comprise between the simplification and discrimination. Second, the two-stage patch-based strategy contains two independent patch-segmentations for the face image. In the first stage, according to the Multi-value strategy we obtain the discriminative local descriptor based on small patches. In the second stage, we calculate weights for larger patches, and the discriminative face regions, such as eyes and month, are strengthened by the weights. The Two-stage strategy considers local similarity in the first stage and global differences in the second one. Extensive experiments on Extended Yale B and FERET show that our method outperforms state-of-the-art methods.
Riqiang Gao, Wenming Yang, Xiaoling Hu 0002, Qingmin Liao
VCIP4
2016 Latent variable pictorial structure for human pose estimation on depth images
Guijin Wang, Qingmin Liao, Jing-Hao Xue
Neurocomputing3
2016 Two strategies to optimize the decisions in signature verification with the presence of spoofing attacks
Shilian Yu, Ye Ai, Yicong Zhou, Weifeng Li 0001, Qingmin Liao, Norman Poh
Inf. Sci.6
2016 Defocus Map Estimation From a Single Image Based on Two-Parameter Defocus Model
abstract
Defocus map estimation (DME) is highly important in many computer vision applications. Nearly, all existing approaches for DME from a single image are based on a one-parameter defocus model, which does not allow for the variation of depth over edges. In this paper, a novel two-parameter model of defocused edges is proposed for DME from a single image. We can estimate the defocus amounts for each side of the edges through this proposed model, and the confidence that the edge is a pattern edge, where the depth remains the same over the edge, can be generated. Then, we modify the TV-L1 algorithm for structure-texture decomposition by taking advantage of this confidence to eliminate pattern edges while preserving structural ones. Finally, the defocus amounts estimated at the edge positions are used as initial values, and the structure component is employed as a guidance in the following Laplacian matting procedure to avoid the influence of pattern edges on the final defocus map. Experiment results show that the proposed method can effectively eliminate the influence of pattern edges compared with the state-of-art method. Furthermore, the estimated defocus map is feasible in applications of depth estimation and foreground/background segmentation.
Fei Zhou 0001, Qingmin Liao
IEEE Trans. Image Process.3
2016 Consistent Coding Scheme for Single-Image Super-Resolution Via Independent Dictionaries
abstract
In this paper, we present a unified frame based on collaborative representation (CR) for single-image super-resolution (SR), which learns low-resolution (LR) and high-resolution (HR) dictionaries independently in the training stage and adopts a consistent coding scheme (CCS) to guarantee the prediction accuracy of HR coding coefficients during SR reconstruction. The independent LR and HR dictionaries are learned based on CR with l2-norm regularization, which can well describe the corresponding LR and HR patch space, respectively. Furthermore, a mapping function is learned to map LR coding coefficients onto the corresponding HR coding coefficients. Propagation filtering can achieve smoothing over an image while preserving image context like edges or textural regions. Moreover, to preserve the edge structures of a super-resolved image and suppress artifacts, a propagation filtering-based constraint and image nonlocal self-similarity regularization are introduced into the SR reconstruction framework. Experimental comparison with state-of-the-art single image SR algorithms validates the effectiveness of proposed approach.
Wenming Yang, Yapeng Tian, Fei Zhou 0001, Qingmin Liao, Hai Chen, Chenglin Zheng
IEEE Trans. Multim.4
2015 A Fast and Accurate Iris Segmentation Approach
Guojun Cheng, Wenming Yang, Qingmin Liao
ICIG (1)4
2015 Texture classification using uniform rotation invariant gradient
abstract
In this paper, we present a novel descriptor called uniform rotation invariant gradient(URIG) aiming at texture classification under variant rotation and illumination condition. Instead of using URIG directly, a 2D descriptor can be formulated combining URIG with average of local pixels. Given a texture image, such 2D descriptors are extracted from every pixel followed by clustering. The centers of clustering can be viewed as a texton dictionary over which a histogram is computed as the representation of given texture image. Experiments are carried out on Outex and CUReT databases comparing to state-of-the-art approaches. Our proposed method achieved promising performance against illumination and rotation changes with least cost for representing histogram dimension.
Wenteng Zhao, Zongqing Lu 0001, Qingmin Liao
ICIP3
2015 Single-frame image super-resolution inspired by perceptual criteria
abstract
In this study, the authors consider the problem of image super‐resolution (SR) in terms of the perceptual criteria. Existing SR methods treat the traditional mean‐squared error (MSE) as an irreplaceable objective function. However, MSE has been widely criticised since it is inconsistent with visual perception of human beings. The perceptual criteria, including the structural similarity (SSIM) index and feature similarity (FSIM) index, have been reported to be more effective in assessing image quality. Therefore SSIM and FSIM are included for the SR task in this study. Specifically, the authors first propose to reform principal component analysis (PCA), which is named as visual perceptual PCA (VP‐PCA), by adopting SSIM as the object function. Subsequently, to accomplish the SR task, the authors cluster the training data and perform VP‐PCA on each cluster to calculate the coefficients. Finally, based on the principle of FSIM, the traditional SR results and the SR results using VP‐PCA are combined to form our fused results. Experimental results are provided to show the superiority of the proposed method over several state‐of‐the‐art methods in both quantitative and visual comparisons.
Fei Zhou 0001, Qingmin Liao
IET Image Process.2
2015 Depth-images-based pose estimation using regression forests and graphical models
Guijin Wang, Qingmin Liao, Jing-Hao Xue
Neurocomputing3
2015 Patterns of Weber magnitude and orientation for uncontrolled face representation and recognition
Yinyan Jiang, Yicong Zhou, Weifeng Li 0001, Qingmin Liao
Neurocomputing5
2015 Single-Image Super-Resolution Based on Compact KPCA Coding and Kernel Regression
abstract
In this letter, we propose a novel approach for single-image super-resolution (SR). Our method is based on the idea of learning a dictionary which can capture the high-order statistics of high-resolution (HR) images. It is of central importance in image SR application, since the high-order statistics play a significant role in the reconstruction of HR image structure. Kernel principal component analysis (KPCA) is adopted to learn such a dictionary. A compact solution is adopted to reduce the time complexity of learning and testing for KPCA. Meanwhile, kernel ridge regression is employed to connect the input low-resolution (LR) image patches with the HR coding coefficients. Experimental results show that the proposed method is effective and efficient in comparison with state-of-art algorithms.
Fei Zhou 0001, Tingrong Yuan, Wenming Yang, Qingmin Liao
IEEE Signal Process. Lett.4
2014 Removal of bleed-through effect via nonnegative least-correlation
abstract
Bleed-through effect is one of the most common degradations in old documents, even in today's newspapers. This effect must be removed for hancing human and automatic readability. The two images scanned from the recto and verso pages of a document can be treated as a linear combination of clean text images from two sides. In this paper, we introduce a nonnegative least-correlation approach to demix bleed-through text images. Experiments have been conducted on both synthetic and real world images. In synthetic case, our method can recover the source text images exactly. Under real world conditions, our approach also performs well. In addition, our approach is computationally efficient and does not need any postprocessing task.
Ye Ai, Weifeng Li 0001, Tsung-Han Chan, Qingmin Liao
ICASSP4
2014 Log-domain polynomial filters for illumination-robust face recognition
abstract
This paper proposes a novel face image descriptor local surface pattern (LSP) for illumination-robust face recognition. It is assumed that the discrete array of pixel values comes about by sampling an underlying smooth surface on the domain of the image. The proposed method efficiently estimates the underlying local surface information, which is approximately represented as linear projection coefficients of the pixels in a local patch. Thus, by filtering local image patches using the polynomial filters and binarizing the filter responses via thresholding, the method can compute a binary code for each pixel in the face image. Then the distribution of the code over suitable image regions is used for face representation. Furthermore, we prove that applying zero-mean filters in logdomain may enable the responses to be more robust to illumination variations. The experimental results on Extended Yale-B and FERET fc databases illustrate the effectiveness of our proposed method in illumination-robust face recognition.
Yinyan Jiang, Yong Wu 0003, Weifeng Li 0001, Longbiao Wang, Qingmin Liao
ICASSP5
2014 Image super-resolution via Kernel regression of sparse coefficients
abstract
In this paper, we present a sparse coding (SC) inspired method to reconstruct a high-resolution (HR) image from one single low-resolution (LR) image. Instead of restricting the coding coefficients of LR and HR image patches to be equal or linearly mapped, we introduce kernel regression to nonlinearly relate the coding coefficients of LR patches and those of corresponding HR ones in an implicit fashion. Meanwhile, principal component analysis (PCA) is employed to train independent dictionaries which can well express image geometrical structure and ensure image sparse property. Experimental results show that the proposed method can effectively reconstruct image details and outperforms state-of-the-art algorithms in both quantitative and visual comparisons.
Tingrong Yuan, Fei Zhou 0001, Wenming Yang, Qingmin Liao
ICASSP4
2014 Image amplification based on pixel-splitting
abstract
In this paper, we propose a pixel-splitting based image amplification method, which involves two main operations: edge-keeping and mean-keeping. Our main idea is to splitting each parent-pixel into multiple sub-pixels. Based on the edge detection and image gradient, we separate edge into vertical and horizontal short rods. The main orientation of each rod is estimated. Then the intensities of sub-pixels around the rod are calculated according to the intensities and edge orientation of parent-pixels. The rest sub-pixels are estimated based on the principle of mean-keeping. The experimental results demonstrate that our method is effective in zigzagging artifacts reduction, de-blurring, and contrast enhancement.
Xiangyi Fu, Fei Zhou 0001, Qingmin Liao
ICIP3
2014 Vanishing point estimation for challenging road images
abstract
In this paper, we present an efficient vanishing point detection method for challenging road images. This detection process is based on the geometrical features of the roads. The slope distribution of the line segments is analyzed to reduce the spurious lines. A distance-based weighting scheme is also utilized to eliminate the voting noise in the voting stage. The proposed algorithm has been tested on a natural data set from Defense Advanced Research Projects Agency (DARPA). Experimental results with both quantitative and qualitative analyses are provided, which demonstrate the superiority of the proposed method over some state-of-the-art methods.
Qingyun She, Zongqing Lu 0001, Qingmin Liao
ICIP3
2014 Local texture based optical flow for complex brightness variations
abstract
In real-world scenarios, complex brightness variations are commonly seen, due to shadows, global illumination changes and nonlinear camera responses, etc. Classical optical flow methods based on brightness or gradient constancy assumption tends to fail under these circumstances. This work proposes an image texture descriptor called LSOT, based on the local spatial structure of a pixel and the relative ordinal information. Then a texture constancy assumption is embedded into a variational optical flow estimation framework as a data term, in order to cope with complex brightness variations. In addition, a non-local regularization term is used to improve the accuracy of the obtained flow fields. The energy functional is optimized using a primal-dual algorithm in a coarse-to-fine warping fashion. Experimental results on synthetic and real image sequences demonstrate the superior performance of the proposed method.
Zongqing Lu 0001, Qingmin Liao
ICIP3
2014 Single image super-resolution via sparse KPCA and regression
abstract
In this paper, we present a new approach to single image super-resolution (SR). The basic idea is to learn a dictionary which can capture the high-order statistics of high-resolution (HR) images. This is of central importance in image SR application, since the high-order statistics play a significant role in the reconstruction of HR image structure. Kernel principal component analysis (KPCA) is used to learn such a dictionary. To reduce the time complexity of learning and testing for KPCA, a sparse solution is adopted. Meanwhile, kernel ridge regression is employed to relate the input low-resolution (LR) image patches and the HR coding coefficients. Experimental results show that the proposed method can effectively reconstruct image details and outperform state-of-the-art algorithms in both quantitative and visual comparisons.
Tingrong Yuan, Wenming Yang, Fei Zhou 0001, Qingmin Liao
ICIP4
2014 Multi-channel speech enhancement using sparse coding on local time-frequency structures
Zhaogui Ding, Weifeng Li 0001, Zhiyong Wu 0001, Longbiao Wang, Qingmin Liao
INTERSPEECH6
2014 Evaluation of PM2.5 and PM10 using normalized first-order absolute sum of high-frequency spectrum
abstract
A new method for air quality evaluation using only visible image analysis is introduced in this paper. Based on the fact that suspended particles in air are visible, we attempted to use visible images to develop an appropriate measure which can be closely related to the density of suspended particles (namely the values of PM2.5 and PM10). Furthermore, using this measure, we can evaluate the values of PM2.5 and PM10 via digital image processing. Combined with water droplets, suspended particles in air can form fog or haze. Based on the monochrome atmospheric scattering model, which has been widely used to describe the formation of a haze image, we propose a measure, normalized first-order absolute sum of high-frequency spectrum (NFAS) and attempt to investigate its relationship with the values of PM2.5 and PM10. The experimental results showed the proposed measure is closely related to PM2.5 and has a relation with PM10.
Wenming Yang, Qingmin Liao
SMARTCOMP3
2014 Face hallucination via position-based dictionaries coding in kernel feature space
abstract
In this paper, we present a new method to reconstruct a high-resolution (HR) face image from a low-resolution (LR) observation. Inspired by position-patch based face hallucination approach, we design position-based dictionaries to code image patches, and recovery HR patch using the coding coefficients as reconstruction weights. In order to capture nonlinear similarity of face features, we implicitly map the data into a high dimensional feature space. By applying kernel principal analysis (KPCA) on the mapped data in the high dimensional feature space, we can obtain reconstruction coefficients in a reduced subspace. Experimental results show that the proposed method can effectively reconstruct details of face images and outperform state-of-the-art algorithms in both quantitative and visual comparisons.
Wenming Yang, Tingrong Yuan, Fei Zhou 0001, Qingmin Liao
SMARTCOMP4
2014 Generalized Weber-face for illumination-robust face recognition
Yong Wu 0003, Yinyan Jiang, Yicong Zhou, Weifeng Li 0001, Zongqing Lu 0001, Qingmin Liao
Neurocomputing6
2014 Super-resolution for facial image using multilateral affinity function
Fei Zhou 0001, Qingmin Liao
Neurocomputing3
2014 Comparative competitive coding for personal identification by using finger vein and finger dorsal texture fusion
Wenming Yang, Xiaola Huang, Fei Zhou 0001, Qingmin Liao
Inf. Sci.4
2014 Feature mapping of multiple beamformed sources for robust overlapping speech recognition using a microphone array
abstract
This paper introduces a nonlinear vector-based feature mapping approach to extract robust features for automatic speech recognition (ASR) of overlapping speech using a microphone array. We explore different configurations and additional sources of information to improve the effectiveness of the feature mapping. First, we investigate the full-vector based mapping of different sources in a log mel-filterbank energy (log MFBE) domain, and demonstrate that retraining the acoustic model using the generated training data can help improve the recognition performance. Then we investigate the feature mapping between different domains. Finally in order to improve the qualities of the mapping inputs we propose a nonlinear mapping of the features from multiple beamformed sources, which are directed at the target and interfering speakers, respectively. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes non-overlapping single speaker and overlapping multi-speaker conditions.
Weifeng Li 0001, Longbiao Wang, Yicong Zhou, John Dines, Mathew Magimai-Doss, Hervé Bourlard, Qingmin Liao
IEEE ACM Trans. Audio Speech Lang. Process.7
2014 Nonlocal Pixel Selection for Multisurface Fitting-Based Super-Resolution
abstract
In this paper, we address a super-resolution (SR) problem that constructs a high-resolution (HR) frame/image from a short sequence of low-resolution (LR) frames/images. It is well known that SR is a difficult problem, especially when the number of LR inputs is small. In particular, our previous work involving multisurface fitting-based SR exhibits relatively poor performance in the above case. To cope with this problem, we take advantage of nonlocal pixels to fit local surfaces. The pixels from nonlocal spatial-temporal positions are selected and weighted based on patch similarity and outlier removal. With this method, the fitted surfaces become more elaborate so that more details can be retrieved in SR results. Experiments demonstrate that the proposed method is very effective in producing HR frames through a small number of LR inputs when compared with some state-of-the-art methods.
Fei Zhou 0001, Shutao Xia, Qingmin Liao
IEEE Trans. Circuits Syst. Video Technol.3
2013 Joint sparse representation based cepstral-domain dereverberation for distant-talking speech recognition
abstract
In this paper we address reducing the mismatch between training and testing conditions for robust distant-talking speech recognition under realistic reverberant environments. It is well known that the distortions caused by reverberation, background noise, etc., are highly nonlinear in the cepstral domain. In this paper we propose to capture the complex relationships between clean and reverberant speech via joint dictionary learning. Given a test reverberant speech with a sequence of feature vectors we first find their sparse representations, and then estimate the underlying clean feature vectors using the dictionary of clean speech. Based on speech recognition experiments conducted under realistic reverberation conditions, the proposed method is shown to perform very well, resulting in an average relative improvement of 59.1% compared with the baseline front-ends.
Weifeng Li 0001, Longbiao Wang, Fei Zhou 0001, Qingmin Liao
ICASSP4
2013 Kernel collaborative representation-based classifier for face recognition
abstract
Recent research has shown that collaborative representation-based classifier (CRC) can lead to promising results for the classification of face images. However, CRC is conducted in the original image space rather than the nonlinear high dimensional feature space in which features belonging to the same class are better grouped together and thus can be easily separable. To address this problem, this paper presents a novel classifier, Kernel Collaborative Representation-based Classifier (KCRC), by incorporating the kernel trick into the framework of CRC. Extensive experiments on both the AT&T and the FERET face databases demonstrate the priority of KCRC to CRC and several state-of-the-art methods.
Weifeng Li 0001, Norman Poh, Qingmin Liao
ICASSP4
2013 A Space Carving Based Reconstruction Method Using Discrete Viewing Edges
abstract
In this paper, we consider the problem of reconstructing a 3D model from a set of pictures taken from calibrated and arbitrarily placed cameras. Our method is based on existing space carving algorithm which considers photo hull as the final result. Our goal is to solve the visibility problem during carving the visual hull. A new concept Discrete Viewing Edge (DVE) is proposed to represent the visual hull instead of a 3D array. DVE is based on voxels and is simple but effective. With models represented by DVEs, we present a surface extraction algorithm and a carve procedure, during which the visibility of a voxel can be determined rapidly and easily. Our way of determining the visibility of a voxel is global, i.e., we take all possible cameras to which this voxel is visible into account. We apply our method to a set of synthetic pictures and provide arbitrary views of target model which are different from existing cameras.
Wenming Yang, Qingmin Liao
ICIG3
2013 Illumination Variation Dictionary Designing for Single-Sample Face Recognition via Sparse Representation
Weifeng Li 0001, Qingmin Liao
MMM (2)3
2013 Iterative Super-Resolution for Facial Image by Local and Global Regression
Fei Zhou 0001, Wenming Yang, Qingmin Liao
MMM (1)4
2013 Adaptive linear regression for single-sample face recognition
Weifeng Li 0001, Qingmin Liao
Neurocomputing4
2013 Active contours driven by local and global probability distributions
Danyi Li, Weifeng Li 0001, Qingmin Liao
J. Vis. Commun. Image Represent.3
2013 Part template: 3D representation for multiview human pose estimation
Jianfeng Shen, Wenming Yang, Qingmin Liao
Pattern Recognit.3
2013 Feature Denoising Using Joint Sparse Representation for In-Car Speech Recognition
abstract
We address reducing the mismatch between training and testing conditions for hands-free in-car speech recognition. It is well known that the distortions caused by background noise, channel effects, etc., are highly nonlinear in the log-spectral or cepstral domain. This letter introduces a joint sparse representation (JSR) to estimate the underlying clean feature vector from a noisy feature vector. Performing a joint dictionary learning by sharing the same representation coefficients, the proposed method intends to capture the complex relationships (or mapping functions) between clean and noisy speech. Speech recognition experiments on realistic in-car data demonstrate that the proposed method shows excellent recognition performance with a relative improvement of 39.4% compared with the “baseline” frontends.
Weifeng Li 0001, Yicong Zhou, Norman Poh, Fei Zhou 0001, Qingmin Liao
IEEE Signal Process. Lett.5
2013 Robust Log-Energy Estimation and its Dynamic Change Enhancement for In-car Speech Recognition
abstract
The log-energy parameter, typically derived from a full-band spectrum, is a critical feature commonly used in automatic speech recognition (ASR) systems. However, log-energy is difficult to estimate reliably in the presence of background noise. In this paper, we theoretically show that background noise affects the trajectories of not only the “conventional” log-energy, but also its delta parameters. This results in a poor estimation of the actual log-energy and its delta parameters, which no longer describe the speech signal. We thus propose a new method to estimate log-energy from a sub-band spectrum, followed by dynamic change enhancement and mean smoothing. We demonstrate the effectiveness of the proposed log-energy estimation and its post-processing steps through speech recognition experiments conducted on the in-car CENSREC-2 database. The proposed log-energy (together with its corresponding delta parameters) yields an average improvement of 32.8% compared with the baseline front-ends. Moreover, it is also shown that further improvement can be achieved by incorporating the new Mel-Frequency Cepstral Coefficients (MFCCs) obtained by non-linear spectral contrast stretching.
Weifeng Li 0001, Longbiao Wang, Yicong Zhou, Hervé Bourlard, Qingmin Liao
IEEE Trans. Speech Audio Process.5
2012 Water droplets segmentation for hydrophobicity classification
abstract
In this paper, we propose an effective water droplets segmentation algorithm based on HSV color space and watershed method. Water droplets segmentation is the key issue to design an automatic hydrophobicity classification algorithm, and the challenge is two-fold: highlight spots on water droplets and the transparency of water. By decomposing the color images into HSV color space, water droplets are easy to be separated in the saturation channel, and watershed method is incorporated to reduce the side-effect of highlight spots. Experimental results on real images demonstrate the advantage of our proposed method.
Wenming Yang, Qingmin Liao
ICASSP3
2012 Face recognition based on nonsubsampled contourlet transform and block-based kernel Fisher linear discriminant
abstract
Face representation, including both feature extraction and feature selection, is the key issue for a successful face recognition system. In this paper, we propose a novel face representation scheme based on nonsubsampled contourlet transform (NSCT) and block-based kernel Fisher linear discriminant (BKFLD). NSCT is a newly developed multiresolution analysis tool and has the ability to extract both intrinsic geometrical structure and directional information in images, which implies its discriminative potential for effective feature extraction of face images. By encoding the the NSCT coefficient images with the local binary pattern (LBP) operator, we could obtain a robust feature set. Furthermore, kernel Fisher linear discriminant is introduced to select the most discriminative feature sets, and the block-based scheme is incorporated to address the small sample size problem. Face recognition experiments on FERET database demonstrate the effectiveness of our proposed approach.
Weifeng Li 0001, Qingmin Liao
ICASSP3
2012 Patterns of weber magnitude and orientation for face recognition
abstract
Feature extraction is vital for a successful face recognition system. In this paper, we propose a computationally efficient, discriminative and robust feature descriptor for face images, named Patterns of Weber magnitude and orientation (PWMO), which encodes Weber magnitude and orientation with patch-based local binary pattern (p-LBP) and patch-based local XOR pattern (p-LXP), respectively. Furthermore, whitened PCA is introduced to reduce the feature dimensionality and select the most discriminative feature sets, and the block-based scheme is incorporated to address the small sample size problem. The effectiveness and robustness of our proposed approach has been demonstrated experimentally on the well-known FERET database.
Weifeng Li 0001, Qingmin Liao
ICIP4
2012 Fusion of finger vein and finger dorsal texture for personal identification based on Comparative Competitive Coding
abstract
In this paper, we present a multimodal personal identification system using finger vein and finger dorsal images with their fusion applied at the feature level. A scheme which combines the registration of image pairs with the region-of-interest (ROI) segmentation, is explored on simultaneously captured finger ventral vein and finger dorsal images. We developed a “Comparative Competitive Coding” (C2Code) fusion scheme. It is capable of discarding undesired information in unimodal feature extraction stage. And only discriminative information can be preserved. Furthermore, the C2Code contains new feature of junction points from the finger vein and finger dorsal image pairs. Experimentally, we establish a dataset of finger vein and finger dorsal images. Comparing the performance of proposed fusion scheme with unimodal methods, higher identification accuracy and lower Equal-Error-Rate (EER) are achieved.
Wenming Yang, Xiaola Huang, Qingmin Liao
ICIP3
2012 A Coarse-to-Fine Subpixel Registration Method to Recover Local Perspective Deformation in the Application of Image Super-Resolution
abstract
In this paper, a coarse-to-fine framework is proposed to register accurately the local regions of interest (ROIs) of images with independent perspective motions by estimating their deformation parameters. A coarse registration approach based on control points (CPs) is presented to obtain the initial perspective parameters. This approach exploits two constraints to solve the problem with a very limited number of CPs. One is named the point-point-line topology constraint, and the other is named the color and intensity distribution of segment constraint. Both of the constraints describe the consistency between the reference and sensed images. To obtain a finer registration, we have converted the perspective deformation into affine deformations in local image patches so that affine refinements can be used readily. Then, the local affine parameters that have been refined are utilized to recover precise perspective parameters of a ROI. Moreover, the location and dimension selections of local image patches are discussed by mathematical demonstrations to avoid the aperture effect. Experiments on simulated data and real-world sequences demonstrate the accuracy and the robustness of the proposed method. The experimental results of image super-resolution are also provided, which show a possible practical application of our method.
Fei Zhou 0001, Wenming Yang, Qingmin Liao
IEEE Trans. Image Process.3
2012 Interpolation-Based Image Super-Resolution Using Multisurface Fitting
abstract
In this paper, we propose a new interpolation-based method of image super-resolution reconstruction. The idea is using multisurface fitting to take full advantage of spatial structure information. Each site of low-resolution pixels is fitted with one surface, and the final estimation is made by fusing the multisampling values on these surfaces in the maximum a posteriori fashion. With this method, the reconstructed high-resolution images preserve image details effectively without any hypothesis on image prior. Furthermore, we extend our method to a more general noise model. Experimental results on the simulated and real-world data show the superiority of the proposed method in both quantitative and visual comparisons.
Fei Zhou 0001, Wenming Yang, Qingmin Liao
IEEE Trans. Image Process.3
2011 Fast single image fog removal using edge-preserving smoothing
abstract
Imaging in poor weather is often severely degraded by scattering due to suspended particles in the atmosphere such as haze and fog. In this paper, we propose a novel fast defogging method from a single image of a scene based on the atmospheric scattering model. In the inference process of the atmospheric veil, the coarser estimate is refined using a fast edge-preserving smoothing approach. The complexity of the proposed method is only a linear function of the number of image pixels and this thus allows a very fast implementation. Results on a variety of outdoor foggy images demonstrate that the proposed method achieves good restoration for contrast and color fidelity resulting in a great improvement in image visibility.
Jing Yu 0005, Qingmin Liao
ICASSP2
2011 Kernel feature selection to fuse multi-spectral MRI images for brain tumor segmentation
Su Ruan, Stéphane Lebonvallet, Qingmin Liao, Yue Min Zhu
Comput. Vis. Image Underst.4
2011 Multiview human pose estimation with unconstrained motions
Jianfeng Shen, Wenming Yang, Qingmin Liao
Pattern Recognit. Lett.3
2011 Illumination Normalization Based on Weber's Law With Application to Face Recognition
abstract
Weber's law suggests that for a stimulus, the ratio between the smallest perceptual change and the background is a constant, which implies stimuli are perceived not in absolute terms but in relative terms. Inspired from this, we exploit and analyze a novel illumination insensitive representation of face images under varying illuminations via a ratio image, called “Weber-face,” where a ratio between local intensity variation and the background is computed. Experimental results on both CMU-PIE and Yale B face databases show that Weber-face performs better than the existing representative approaches.
Weifeng Li 0001, Wenming Yang, Qingmin Liao
IEEE Signal Process. Lett.4
2010 Object Tracking and Local Appearance Capturing in a Remote Scene Video Surveillance System with Two Cameras
Wenming Yang, Fei Zhou 0001, Qingmin Liao
MMM3
2010 Locating human hands for real-time pose estimation from monocular video
abstract
This paper presents a real-time system to detect and estimate the pose of human upper body from a monocular video. A novel approach to locate the hands is proposed, which is designed to cope with the complicated situations such as short sleeves, fast motion and occlusion. Human silhouette and skin color blobs are extracted from the frames of the video; then candidate locations of head, hands, and elbows are chosen and evaluated by an inverse kinematics based strategy. Experiments demonstrate the efficacy and robustness of this approach. The algorithm is developed for a camera-based tennis game, in which poses of a player have to be estimated in real time (for avatar animation, action recognition, etc). It can also be applied in other human-computer interaction applications.
Xin Lian, Qingmin Liao
VRST2
2009 A Novel System of Stereoscopic Video Based on TMS320DM642 DSP
abstract
In this paper, we propose a novel stereoscopic video display system based on Texas Instruments (TI) company's Multimedia processor DM642. The system uses a common stereoscopic vision display method by which same scenes captured in different angles bring stereo depth information and therefore a stereoscopic perception in human's brain. The system is a time-sequential stereoscopic display system. By using Liquid Crystal Shutter (LCS) glasses, images captured by two NTSC cameras are respectively sent to the observer's left and right eyes. In order to enhance the quality of video, the frame rate for each eye is designed to 60 Hz, and the image's vertical resolution is doubled compared with the input video of the camera. A high computational ability and powerful platform for real-time compression and transmission of the stereoscopic video can be achieved in our system.
Haijin Fan, Qingmin Liao
ICIG3
2009 An Effective Method for Foreground Segmentation of Video
abstract
In this paper, we propose a novel foreground segmentation approach for applications using static cameras. The foreground segmentation is modeled as an energy function optimum process, where energy function is based on Markov Random Field (MRF) and efficiently optimized by Gibbs sampling. The essence of our method is that we fuse four foreground/background models based on color and texture. This allows composing a robust likelihood term that not only reflects the appearance of foreground/background, but also models the shadow removal process, together with a spatial contrast term and a better temporal persistence term, which achieves a more accurate segmentation. This method has been run on both indoor and outdoor sequences, and the results have proved its effectiveness.
Jianfeng Shen, Zongqing Lu 0001, Qingmin Liao
ICIG3
2009 A variational approach to automatic segmentation of RNFL on OCT data sets of the retina
abstract
Optical coherence tomography (OCT) as a new imaging technology is gaining popularity in the diagnosis of ocular diseases. It enable clinicians to perform accurate, objective, and reproducible measurements of the retinal nerve fibre layer (RNFL) whose thickness is closely related to many ocular diseases. Automatic segmenting RNFL is a challenging image processing problem, which is a critical job for final thickness estimation. We modeled the OCT data sets as probability density fields and introduced a level set model to outline the RNFL region within the retina. We also introduced the symmetrized Kullback-Leibler distance to describe the difference of two density functions. The new approach can deal with the typical problems of OCT image analysis: speckle noise and faint structure in an efficient way.
Zongqing Lu 0001, Qingmin Liao
ICIP2
2009 Multi-kernel SVM based classification for brain tumor segmentation of MRI multi-sequence
abstract
In this paper, the multi-kernel SVM (Support Vector Machine) classification, integrated with a fusion process, is proposed to segment brain tumor from multi-sequence MRI images (T2, PD, FLAIR). The objective is to quantify the evolution of a tumor during a therapeutic treatment. As the procedure develops, a manual learning process about the tumor is carried out just on the first MRI examination. Then the follow-up on coming examinations adapts the learning automatically and delineates the tumor. Our method consists of two steps. The first one classifies the tumor region using a multi-kernel SVM which performs on multi-image sources and obtains relative multi-result. The second one ameliorates the contour of the tumor region using both the distance and the maximum likelihood measures. Our method has been tested on real patient images. The quantification evaluation proves the effectiveness of the proposed method.
Su Ruan, Stéphane Lebonvallet, Qingmin Liao, Yue Min Zhu
ICIP4
2009 Personal authentication using finger vein pattern and finger-dorsa texture fusion
abstract
Personal authentication has attracted great attention due to its large potential of security application, and many researches have shown that fusion of features or decisions obtained from various single-modal biometrics verification systems can enhance the overall performance of system. In this paper, we proposed a novel multimodal biometric approach fusing finger vein pattern with finger-dorsa texture. Firstly, Finger Vein image and finger-dorsa image from the same finger are captured simultaneously, and a method is designed to segment Regions Of Interest(ROI) of vein image and dorsal image. Secondly, two strategies are designed to extract finger vein pattern and finger-dorsa texture respectively. Vein extraction strategy consists of four steps: local thresholding, modified line tracking, thorough probability map creating and directional neighbor analysis. Gray normalization is performed on finger-dorsa image to extract main finger-dorsa texture. Thirdly, the binarized vein pattern and normalized dorsal texture are fused into one feature image. Finally, a block-based texture feature is proposed for personal authentication. Experimental results showed that the proposed fusion method outperforms any one of finger-dorsa and finger vein methods.
Wenming Yang, Qingmin Liao
ACM Multimedia3
2004 Possibilistic-clustering-based MR brain image segmentation with accurate initialization
abstract
Magnetic resonance image analysis by computer is useful to aid diagnosis of malady. We present in this paper a automatic segmentation method for principal brain tissues. It is based on the possibilistic clustering approach, which is an improved fuzzy c-means clustering method. In order to improve the efficiency of clustering process, the initial value problem is discussed and solved by combining with a histogram analysis method. Our method can automatically determine number of classes to cluster and the initial values for each class. It has been tested on a set of forty MR brain images with or without the presence of tumor. The experimental results showed that it is simple, rapid and robust to segment the principal brain tissues.
Qingmin Liao, Yingying Deng, Weibei Dou, Su Ruan, Daniel Bloyet
VCIP1
2002 Virtual face rendering based on gradient features in VLBR networks
Yang Ran, Qingmin Liao, Xinggang Lin
VCIP2
2002 Rate-distortion model based rate control for real-time VBR video coding and low-delay communications
Junfeng Bai, Qingmin Liao, Xinggang Lin, Xinhua Zhuang
Signal Process. Image Commun.2
2001 Accurate estimation of R-D characteristics for rate control in real-time video encoding
abstract
In real-time video communications, the rate control strategies must be utilized to satisfy the end-to-end delay and prevent the encoding buffer from over/underflow. In other words, to acquire the best possible video quality with a minimal quality variation in the playback video, an accurate rate-distortion (R-D) model of the video source is critical in optimizing the bit allocation for video coding. In the paper, an exponential functions based R-D model is proposed for intra-coded frames in video coding. Numerous experiments have consistently shown that the proposed model outperforms other popular R-D models in terms of both the estimation accuracy and computation complexity, making it suitable for rate control in real-time video coding.
Junfeng Bai, Chang Feng, Qingmin Liao, Xinggang Lin, Xinhua Zhuang
ICASSP3
2000 A rate control algorithm for VBR video encoding and transmission
abstract
To satisfy the rigorous limitation of delay in real time transmission of high quality video, the coding rate and transmission rate should be selected judiciously. For the ATM variable bit rate (VBR) channel, the channel rate is tightly restrained by the traffic contract negotiated before communication. The leaky-bucket based rate control algorithm proposed by us could meet the delay and traffic contract, so as to abstain the cell loss over the user-network interface (UNI). Furthermore, in order to benefit the statistical multiplexing and improve the overall decoded video quality, the burstiness of the MPEG encoded stream is reduced dramatically in our system. Our experiment shows that the algorithm is a good compromise between CBR (constant bit rate) and VBR video coding. And its implementation is so simple that it could be easily utilized in on-line rate control for real-time video communication.
Junfeng Bai, Qingmin Liao, Xinggang Lin
ICASSP2
2000 Rate-distortion-model-based rate control algorithm for real-time VBR video encoding
Junfeng Bai, Qingmin Liao, Xinggang Lin
VCIP2
1997 An Object-Oriented System Architecture of General Image Processing Systems
abstract
An object-oriented system architecture for general image processing systems is studied. Centering around functional requirements of general image processing systems, the system architecture is designed based on object-oriented technology. Such a system architecture has four main advantages over traditional structured design, such as openness, extendibility and development platforms independence etc. By adopting the architecture, a prototypical system of general image processing is implemented based on the Document/View structure with C++. Implementation of the system testifies that such an object-oriented architecture is effective for developing general image processing systems.
Bin Zhang 0002, Xinggang Lin, Qingmin Liao
ICIP (2)3