Zhongxuan Luo

dblp:17/6264 · DBLP profile ↗
← Back
160ranked-venue papers
6as first author
74since 2021 · last 2026
0000-0001-5997-2646ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 109 · 5 first-author · 52 since 2021Artificial intelligence and machine learning · 52 · 26 since 2021Computer networks · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 2 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2026 RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels
abstract
Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their inability to provide precise annotation data for sonar images. Therefore, designing effective object detection methods for sonar images with extremely limited labels is particularly important. To address this, we propose a teacher-student framework called RSOD, which aims to fully learn the characteristics of sonar images and develop a pseudo-label strategy suitable for these images to mitigate the impact of limited labels. First, RSOD calculates a reliability score by assessing the consistency of the teacher's predictions across different views. To leverage this score, we introduce an object mixed pseudo-label method to tackle the shortage of labeled data in sonar images. Finally, we optimize the performance of the student by implementing a reliability-guided adaptive constraint. By taking full advantage of unlabeled data, the student can perform well even in situations with extremely limited labels. Notably, on the UATD dataset, our method, using only 5% of labeled data, achieves results that can compete against those of our baseline algorithm trained on 100% labeled data. We also collected a new dataset to provide more valuable data for research in the field of sonar.
Chengzhou Li, Guanchen Meng, Qi Jia 0001, Jinyuan Liu 0001, Zhu Liu 0004, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
AAAI9
2026 OT-ALD: Aligning Latent Distributions with Optimal Transport for Accelerated Image-to-Image Translation
Zhanpeng Wang, Shuting Cao, Na Lei, Zhongxuan Luo
AAAI6
2026 Underwater Data Collection Scheme based on LLMs
Kunhong Ji, Chi Lin 0001, Jiankang Ren, Xin Fan 0001, Zhongxuan Luo
INFOCOM6
2026 Topo-GenMeta: Generative design of metamaterials based on diffusion model with attention to topology
Jiangbei Hu, Shengfa Wang, Yu Jiang 0019, Na Lei, Ying He 0001, Zhongxuan Luo
Comput. Aided Des.7
2026 High-connectivity polycube-maps: Solvable space expansion through validity-augmented topological conditions
abstract
Polycube-maps play a critical role in computer graphics, especially for generating high-quality hexahedral meshes. Existing polycube validity conditions, primarily based on Steinitz and Eppstein’s approach, are limited to 3-connected graphs. Extending polycube-maps to handle higher connectivity graphs is crucial for practical applications. In this work, we introduce Validity-Augmented Topological Conditions (VAT conditions) based on the Gauss–Bonnet theorem. These conditions offer both global and local topological criteria, enabling the solvability of k-connected graphs, non-manifold structures, and meshes with voids. Our VAT conditions allow models that do not meet traditional polycube validity criteria but are still valid polycube polyhedra in practice. Additionally, we propose an Immune Genetic Algorithm (ImGA) tailored to our VAT conditions to enhance the robustness of polycube-map generation. We evaluate our method using the Thingi10k and ABC datasets. Results demonstrate that our VAT conditions expands the solvable space of polycubes and achieves higher quality all-hexahedral meshing for higher-connectivity or more complex models. Furthermore, we discuss the limitations associated with our proposed method.
Na Lei, Xiaopeng Zheng, Zhongxuan Luo
Comput. Aided Des.6
2026 Gen-Porous: An INR-based generative framework for multiscale TPMS-like porous structure design and optimization
Shengfa Wang, Jiangbei Hu, Yu Jiang 0019, Na Lei, Zhongxuan Luo
Comput. Aided Des.8
2026 Quadrilateral mesh generation based on foliation and meromorphic quadratic differential
Xiaopeng Zheng, Na Lei, Zhongxuan Luo
Comput. Aided Des.4
2026 Energy-Aware Adaptive Topology Control for UOWSNs
abstract
Underwater Optical Wireless Sensor Network (UOWSN) is a promising technology as it can achieve high-speed communication in underwater environment. However, affected by the uncertainty of complex underwater environment, the network topology of UOWSN is highly dynamic, making it difficult to quantify flexibility or further optimize the topological structure. Additionally, node mobility and energy constraints pose significant challenges to reliable communication. In this paper, we propose a mobility-aware and energy-efficient flexibility-based network topology evaluation model (ME-FEM) for UOWSNs. Then, a reinforcement learning model, termed ME-FEM-DRL, for optimizing the network topology based on ME-FEM is developed, which enables UOWSN to maintain an optimal topology when working in harsh underwater environments. Theoretical analysis proves the NP-hardness of the optimization problem and demonstrates that our algorithm achieves an approximation ratio of$O(\log N)$with optimal parameter boundaries. Simulation results demonstrate that the proposed method can significantly improve the network flexibility. Compared with the five baseline algorithms in simulations, ME-FEM-DRL reduces normalized topology optimization time cost by 64% and extends network lifetime by 95% on average. Test-bed experiments verify the applicability and effectiveness in practical applications for detecting emergent events.
Yang Chi, Chi Lin 0001, Haipeng Dai 0001, Yu Tian 0014, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Mob. Comput.6
2026 Adaptive Interference Alignment for Underwater Optical Wireless Sensor Networks
abstract
Underwater Optical Wireless Sensor Networks (UOWSNs) have emerged as a promising solution for high-speed underwater communication. However, these networks face a critical challenge of mutual interference among optical nodes, which occurs when the directional optical beams intersect or coverage areas overlap due to node mobility in dynamic underwater environments. Existing interference management approaches demonstrate limited effectiveness due to their reliance on simplified channel models and inability to handle rapid topology changes, resulting in significant network performance degradation. This paper presents a novel framework that systematically addresses interference management in UOWSNs through two key innovations. First, we propose a Sparse Bayesian Learning-based Interference Detection (SBL-ID) algorithm that enables real-time identification and characterization of interference patterns under complex underwater channel conditions. Second, we develop an Adaptive Interference Alignment and Delay Compensation (AIADC) algorithm that projects interference signals into a reduced-dimensional subspace, thereby enhancing the signal-to-interference ratio and facilitating accurate detection of desired signals amid interference. Our framework transforms the NP-hard interference management problem into tractable optimizations, achieving near-optimal solutions with polynomial time complexity. Extensive simulations demonstrate that our approach reduces BER by 95% and improves network throughput by 67% compared to state-of-the-art techniques. Testbed experiments conducted in both pool and lake further validate our framework's effectiveness, maintaining consistent performance improvements under diverse underwater conditions.
Yang Chi, Chi Lin 0001, Fengqi Li, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Mob. Comput.6
2026 Beyond Implicit Mapping: Advancing Generative Models Through Smoothed Optimal Transport
abstract
Optimal transport (OT) has gained significant attention in deep learning as a powerful mathematical tool for transforming distributions. Specifically, in deep generative models, the incorporation of OT helps address issues such as training instability, vanishing gradients, and mode collapse. However, in these models, most of the OT mappings learned by neural networks are typically implicit, making it difficult to explicitly model the relationship between the source and target domains. This limitation reduces the interpretability of the model and hinders its applicability in conditional generation tasks. To address this issue, we introduce Nesterov's smoothing technique to smooth the Brenier potential, enabling the derivation of an explicit OT mapping that serves as the foundation for constructing an advanced generative model. The proposed model offers the following advantages. First, it explicitly captures the mapping between the source and target domains, thereby enhancing the interpretability of the generative process and enabling a novel pathway for conditional sample generation based on a smoothed approximation of OT mapping. Second, the model can generate new samples directly through an explicit OT mapping, eliminating the need for interpolation and rejection sampling commonly seen in traditional methods, thereby improving generation efficiency. Moreover, extensive experiments show that our proposed model achieves superior performance in both unconditional and conditional generation tasks.
Lianbao Jin, Zhanpeng Wang, Zebin Xu, Na Lei, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.6
2025 TNT-GS: Truncated and Tailored Gaussian Splatting
abstract
Gaussian Splatting (GS) is widely used for efficient 3D scene representation and rendering by modeling scenes as continuous Gaussian distributions. However, GS struggles with high-frequency details and sharp transitions due to its low-pass filtering effect, often requiring multiple Gaussian stacking, which increases computational and memory costs. To overcome these limitations, we propose Truncated and Tailored Gaussian Splatting (TNT-GS), a novel approach that enhances shape complexity and preserves sharp boundaries. Our method truncates Gaussians to generate sharp edges and flexible shapes without excessive stacking, improving efficiency. We also introduce learnable parameters to dynamically tailor the receptive field of the primitives, optimizing the balance between high-frequency details and smooth regions. Furthermore, we employ specialized densification strategies to further improve efficiency during tile computation. Experimental results show that TNT-GS outperforms state-of-the-art methods in storage efficiency and rendering speed, offering a robust solution for real-time rendering. The code of TNT-GS is available at https://github.com/GoogolplexGoodenough/TNT-GS.
Xiaofeng Liu 0001, Guanchen Meng, Chongyang Feng, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
ACM Multimedia5
2025 Physics-Guided Sonar Image Fine-grained Recognition under Scarce Annotations
abstract
Sonar image recognition is a key technology in underwater exploration systems. Compared with natural images, sonar images have fewer texture details and are easily affected by heavy noise, making it more challenging for specialists to distinguish the subtle differences among classes. In view of this, studying fine-grained classification methods for sonar images with scarce annotations is of significant importance. To address this issue, we propose a Physics-Guided Teacher-Student (PGTS) framework to explore the unique physical information of sonar images while simultaneously mitigating the effects of limited annotations. First, PGTS reconstructs sonar signals through physical simulation and a specially designed physics-guided feature generation module, which allows it to bypass the time-consuming physical simulation during inference. Then, we design a multi-modal teacher model combines the reconstructed sonar signals and sonar images to extract discriminative features to generate robust pseudo labels for fine-grained target categories. Finally, the knowledge is transferred to a single-modal student model through consistency loss. Under the joint constraints of the teacher model and the reconstructed sonar physical signals, the student model continuously improves its performance in annotation-scarce scenarios. Notably, when merely 1% of the data is labeled, our method outperforms other state-of-the-art approaches by 12.46% in terms of accuracy.
Chengzhou Li, Qi Jia 0001, Jinyuan Liu 0001, Zhiying Jiang, Longhan Feng, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
ACM Multimedia8
2025 Feature-aware Singularity Structure Optimization for Hex Mesh
Xiaopeng Zheng, Junyi Duan, Na Lei, Zhongxuan Luo
Comput. Aided Des.4
2025 Bilevel Fast Scene Adaptation for Low-Light Image Enhancement
Long Ma 0002, Dian Jin 0003, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
Int. J. Comput. Vis.6
2025 An optimal transport-guided diffusion framework with mitigating mode mixture
Zhanpeng Wang, Zhongxuan Luo, Na Lei
Neurocomputing3
2025 Semi-Discrete Optimal Transport for Long-Tailed Classification
Lianbao Jin, Na Lei, Zhongxuan Luo, Chao Ai, Xianfeng Gu
J. Comput. Sci. Technol.3
2025 Anomaly-aware symmetric non-negative matrix factorization for short text clustering
Ximing Li 0002, Yuanyuan Guan, Bo Fu 0001, Zhongxuan Luo
Knowl. Inf. Syst.4
2025 Learning With Self-Calibrator for Fast and Robust Low-Light Image Enhancement
abstract
Convolutional Neural Networks (CNNs) have shown significant success in the low-light image enhancement task. However, most of existing works encounter challenges in balancing quality and efficiency simultaneously. This limitation hinders practical applicability in real-world scenarios and downstream vision tasks. To overcome these obstacles, we propose a Self-Calibrated Illumination (SCI) learning scheme, introducing a new perspective to boost the model's capability. Based on a weight-sharing illumination estimation process, we construct an embedded self-calibrator to accelerate stage-level convergence, yielding gains that utilize only a single basic block for inference, which drastically diminishes computation cost. Additionally, by introducing the additivity condition on the basic block, we acquire a reinforced version dubbed SCI++, which disentangles the relationship between the self-calibrator and illumination estimator, providing a more interpretable and effective learning paradigm with faster convergence and better stability. We assess the proposed enhancers on standard benchmarks and in-the-wild datasets, confirming that they can restore clean images from diverse scenes with higher quality and efficiency. The verification on different levels of low-light vision tasks shows our applicability against other methods.
Long Ma 0002, Tengyu Ma 0004, Chengpei Xu, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Bi-level Learning of Task-Specific Decoders for Joint Registration and One-Shot Medical Image Segmentation
abstract
One-shot medical image segmentation (MIS) aims to cope with the expensive, time-consuming, and inherent human bias annotations. One prevalent method to address one-shot MIS is joint registration and segmentation (JRS) with a shared encoder, which mainly explores the voxel-wise correspondence between the labeled data and unlabeled data for better segmentation. However, this method omits underlying connections between task-specific decoders for segmentation and registration, leading to unstable training. In this paper, we propose a novel Bi-level Learning of Task-Specific Decoders for one-shot MIS, employing a pretrained fixed shared encoder that is proved to be more quickly adapted to brand-new datasets than existing JRS without fixed shared encoder paradigm. To be more specific, we introduce a bi-level optimization training strategy considering registration as a major objective and segmentation as a learnable constraint by leveraging inter-task coupling dependencies. Furthermore, we design an appearance conformity constraint strategy that learns the backward transformations generating the fake labeled data used to perform data augmentation instead of the labeled image, to avoid performance degradation caused by inconsistent styles between unlabeled data and labeled data in previous methods. Extensive experiments on the brain MRI task across ABIDE, ADNI, and PPMI datasets demonstrate that the proposed Bi-JROS outperforms state-of-the-art one-shot MIS methods for both segmentation and registration tasks. The code will be available at https://github.com/Coradlut/Bi-JROS.
Xin Fan 0001, Jiaxin Gao 0001, Jia Wang 0036, Zhongxuan Luo, Risheng Liu
CVPR5
2024 UWBeacon: Lighting up Centimeter-Level Underwater Positioning
abstract
Underwater positioning plays a key role in many underwater operations. This paper presents the design, implementation, and evaluation of UWBeacon, a centimeter-level visible light-based underwater positioning system. UWBeacon consists of LED beacons as the light signal transmitter and a camera-based receiver as the target. To address unique challenges in underwater environment such as limited visibility and strong ambient interference, we exploit a novel design that utilizes polarized lights of different colors with different polarization angles for background subtraction. UWBeacon is implemented with commercial-off-the-shelf LEDs and cameras. Comprehensive experiments conducted in various real underwater environments show that UWBeacon can achieve a mean positioning error below 6 cm and an orientation error below 1.5° at a distance of 10 meters.
Chi Lin 0001, Jie Xiong 0001, Lei Wang 0005, Guowei Wu 0001, Xin Fan 0001, Zhongxuan Luo
MobiCom8
2024 Motion-Driven Neural Optimizer for Prophylactic Braces Made by Distributed Microstructures
abstract
Joint injuries, and their long-term consequences, present a substantial global health burden. Wearable prophylactic braces are an attractive potential solution to reduce the incidence of joint injuries by limiting joint movements that are related to injury risk. Given human motion and ground reaction forces, we present a computational framework that enables the design of personalized braces by optimizing the distribution of microstructures and elasticity. As varied brace designs yield different reaction forces that influence kinematics and kinetics analysis outcomes, the optimization process is formulated as a differentiable end-to-end pipeline in which the design domain of microstructure distribution is parameterized onto a neural network. The optimized distribution of microstructures is obtained via a self-learning process to determine the network coefficients according to a carefully designed set of losses and the integrated biomechanical and physical analyses. Since knees and ankles are the most commonly injured joints, we demonstrate the effectiveness of our pipeline by designing, fabricating, and testing prophylactic braces for the knee and ankle to prevent potentially harmful joint movements.
Xingjian Han, Yu Jiang 0019, Weiming Wang 0003, Guoxin Fang, Simeon Gill, Zhiqiang Zhang 0001, Shengfa Wang, Jun Saito, Zhongxuan Luo, Emily Whiting, Charlie C. L. Wang
SIGGRAPH Asia10
2024 Singularity structure simplification for hex mesh via integer linear program
Junyi Duan, Xiaopeng Zheng, Na Lei, Zhongxuan Luo
Comput. Aided Des.4
2024 IF-TONIR: Iteration-free Topology Optimization based on Implicit Neural Representations
Jiangbei Hu, Ying He 0001, Baixin Xu, Shengfa Wang, Na Lei, Zhongxuan Luo
Comput. Aided Des.6
2024 CoCoNet: Coupled Contrastive Learning Network with Multi-level Feature Ensemble for Multi-modality Image Fusion
Jinyuan Liu 0001, Runjia Lin, Guanyao Wu, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
Int. J. Comput. Vis.5
2024 OT-net: a reusable neural optimal transport solver
Zezeng Li, Lianbao Jin, Na Lei, Zhongxuan Luo
Mach. Learn.5
2024 A Task-Guided, Implicitly-Searched and Meta-Initialized Deep Model for Image Fusion
abstract
Image fusion plays a key role in a variety of multi-sensor-based vision systems, especially for enhancing visual quality and/or extracting aggregated features for perception. However, most existing methods just consider image fusion as an individual task, thus ignoring its underlying relationship with these downstream vision problems. Furthermore, designing proper fusion architectures often requires huge engineering labor. It also lacks mechanisms to improve the flexibility and generalization ability of current fusion approaches. To mitigate these issues, we establish a Task-guided, Implicit-searched and Meta-initialized (TIM) deep model to address the image fusion problem in a challenging real-world scenario. Specifically, we first propose a constrained strategy to incorporate information from downstream tasks to guide the unsupervised learning process of image fusion. Within this framework, we then design an implicit search scheme to automatically discover compact architectures for our fusion model with high efficiency. In addition, a pretext meta initialization technique is introduced to leverage divergence fusion data to support fast adaptation for different kinds of image fusion tasks. Qualitative and quantitative experimental results on different categories of image fusion problems and related downstream tasks (e.g., visual enhancement and semantic understanding) substantiate the flexibility and effectiveness of our TIM.
Risheng Liu, Zhu Liu 0004, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Measure-Driven Neural Solver for Optimal Transport Mapping
abstract
Optimal transport (OT) studies the most economical transformation of one probability measure into another, attracting attention across diverse fields and inspiring various OT-solving algorithms. However, adjusting the probability measure according to specific application requirements, such as achieving unbiased generated images or generating images with specific attributes, necessitates recalculating the OT mapping. This process may result in inefficiency and limited usage flexibility of existing algorithms. To address this, we propose a measure-driven neural solver for OT, the key of which is to construct a network module to learn Brenier’s height representation, and then compute the gradient of Brenier’s potential to derive the OT mapping. Our algorithm has two main advantages: i) It enables direct calculation or fine-tuning of the OT mapping when the target sample measure changes, enhancing efficiency. ii) For unbiased image generation or attribute-specific face generation, adjusting the posterior probability measure of the latent space in the pre-trained model suffices, without the need for additional auxiliary components, this highlights the flexibility of our algorithm. Extensive experiments demonstrate the excellent performance of our algorithm in debiased generation and controllable generation, and its flexibility and efficiency. In addition, both of these generation ways can enhance the classification performance of minority groups.
Zezeng Li, Zhanpeng Wang, Zebin Xu, Na Lei, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.6
2024 A Parametric Design Method for Engraving Patterns on Thin Shells
abstract
Designing thin-shell structures that are diverse, lightweight, and physically viable is a challenging task for traditional heuristic methods. To address this challenge, we present a novel parametric design framework for engraving regular, irregular, and customized patterns on thin-shell structures. Our method optimizes pattern parameters such as size and orientation, to ensure structural stiffness while minimizing material consumption. Our method is unique in that it works directly with shapes and patterns represented by functions, and can engrave patterns through simple function operations. By eliminating the need for remeshing in traditional FEM methods, our method is more computationally efficient in optimizing mechanical properties and can significantly increase the diversity of shell structure design. Quantitative evaluation confirms the convergence of the proposed method. We conduct experiments on regular, irregular, and customized patterns and present 3D printed results to demonstrate the effectiveness of our approach.
Jiangbei Hu, Shengfa Wang, Ying He 0001, Zhongxuan Luo, Na Lei, Ligang Liu 0001
IEEE Trans. Vis. Comput. Graph.4
2023 DPM-OT: A New Diffusion Probabilistic Model Based on Optimal Transport
abstract
Sampling from diffusion probabilistic models (DPMs) can be viewed as a piecewise distribution transformation, which generally requires hundreds or thousands of steps of the inverse diffusion trajectory to get a high-quality image. Recent progress in designing fast samplers for DPMs achieves a trade-off between sampling speed and sample quality by knowledge distillation or adjusting the variance schedule or the denoising equation. However, it can’t be optimal in both aspects and often suffer from mode mixture in short steps. To tackle this problem, we innovatively regard inverse diffusion as an optimal transport (OT) problem between latents at different stages and propose the DPM-OT, a unified learning framework for fast DPMs with a direct expressway represented by OT map, which can generate high-quality samples within around 10 function evaluations. By calculating the semi-discrete optimal transport map between the data latents and the white noise, we obtain an expressway from the prior distribution to the data distribution, while significantly alleviating the problem of mode mixture. In addition, we give the error bound of the proposed method, which theoretically guarantees the stability of the algorithm. Extensive experiments validate the effectiveness and advantages of DPM-OT in terms of speed and quality (FID and mode mixture), thus representing an efficient solution for generative modeling. Source codes are available at https://github.com/cognaclee/DPM-OT.
Zezeng Li, Zhanpeng Wang, Na Lei, Zhongxuan Luo, Xianfeng Gu
ICCV5
2023 Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation
abstract
Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach ‘Best of Both Worlds’. To overcome this issue, in this paper, we propose a Multi-interactive Feature learning architecture for image fusion and Segmentation, namely SegMiF, and exploit dual-task correlation to promote the performance of both tasks. The SegMiF is of a cascade structure, containing a fusion sub-network and a commonly used segmentation sub-network. By slickly bridging intermediate features between two components, the knowledge learned from the segmentation task can effectively assist the fusion task. Also, the benefited fusion network supports the segmentation one to perform more pretentiously. Besides, a hierarchical interactive attention block is established to ensure fine-grained mapping of all the vital information between two tasks, so that the modality/semantic features can be fully mutual-interactive. In addition, a dynamic weight factor is introduced to automatically adjust the corresponding weights of each task, which can balance the interactive feature correspondence and break through the limitation of laborious tuning. Furthermore, we construct a smart multi-wave binocular imaging system and collect a full-time multi-modality benchmark with 15 annotated pixel-level categories for image fusion and segmentation. Extensive experiments on several public datasets and our benchmark demonstrate that the proposed method outputs visually appealing fused images and perform averagely 7.66% higher segmentation mIoU in the real-world scene than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/JinyuanLiu-CV/SegMiF.
Jinyuan Liu 0001, Zhu Liu 0004, Guanyao Wu, Long Ma 0002, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
ICCV7
2023 Bilevel Generative Learning for Low-Light Vision
abstract
Recently, there has been a growing interest in constructing deep learning schemes for Low-Light Vision (LLV). Existing techniques primarily focus on designing task-specific and data-dependent vision models on the standard RGB domain, which inherently contain latent data associations. In this study, we propose a generic low-light vision solution by introducing a generative block to convert data from the RAW to the RGB domain. This novel approach connects diverse vision problems by explicitly depicting data generation, which is the first in the field. To precisely characterize the latent correspondence between the generative procedure and the vision task, we establish a bilevel model with the parameters of the generative block defined as the upper level and the parameters of the vision task defined as the lower level. We further develop two types of learning strategies targeting different goals, namely low cost and high accuracy, to acquire a new bilevel generative learning paradigm. The generative blocks embrace a strong generalization ability in other low-light vision tasks through the bilevel optimization on enhancement tasks. Extensive experimental evaluations on three representative low-light vision tasks, namely enhancement, detection, and segmentation, fully demonstrate the superiority of our proposed approach. The code will be available at https://github.com/Yingchi1998/BGL.
Yingchi Liu, Zhu Liu 0004, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
ACM Multimedia6
2023 Meshless Optimization of Triply Periodic Minimal Surface Based Two-Fluid Heat Exchanger
Yu Jiang 0019, Jiangbei Hu, Shengfa Wang, Na Lei, Zhongxuan Luo, Ligang Liu 0001
Comput. Aided Des.5
2023 Differentiable Channel Design for Enhancing Manufacturability of Enclosed Cavities
Jiangbei Hu, Shengfa Wang, Na Lei, Zhongxuan Luo
Comput. Aided Des.5
2023 An Efficient Self-supporting Infill Structure for Computational Fabrication
abstract
Abstract Efficiently optimizing the internal structure of 3D printing models is a critical focus in the field of industrial manufacturing, particularly when designing self‐supporting structures that offer high stiffness and lightweight characteristics. To tackle this challenge, this research introduces a novel approach featuring a self‐supporting polyhedral structure and an efficient optimization algorithm. Specifically, the internal space of the model is filled with a combination of self‐supporting octahedrons and tetrahedrons, strategically arranged to maximize structural integrity. Our algorithm optimizes the wall thickness of the polyhedron elements to satisfy specific stiffness requirements, while ensuring efficient alignment of the filled structures in finite element calculations. Our approach results in a considerable decrease in optimization time. The optimization process is stable, converges rapidly, and consistently delivers effective results. Through a series of experiments, we have demonstrated the effectiveness and efficiency of our method in achieving the desired design objectives.
Shengfa Wang, Jiangbei Hu, Na Lei, Zhongxuan Luo
Comput. Graph. Forum5
2023 Rethinking general underwater object detection: Datasets, challenges, and solutions
Chenping Fu, Risheng Liu, Xin Fan 0001, Puyang Chen, Hao Fu 0004, Wanqi Yuan, Ming Zhu 0001, Zhongxuan Luo
Neurocomputing8
2023 Learning With Nested Scene Modeling and Cooperative Architecture Search for Low-Light Vision
abstract
Images captured from low-light scenes often suffer from severe degradations, including low visibility, color casts, intensive noises, etc. These factors not only degrade image qualities, but also affect the performance of downstream Low-Light Vision (LLV) applications. A variety of deep networks have been proposed to enhance the visual quality of low-light images. However, they mostly rely on significant architecture engineering and often suffer from the high computational burden. More importantly, it still lacks an efficient paradigm to uniformly handle various tasks in the LLV scenarios. To partially address the above issues, we establish Retinex-inspired Unrolling with Architecture Search (RUAS), a general learning framework, that can address low-light enhancement task, and has the flexibility to handle other challenging downstream vision tasks. Specifically, we first establish a nested optimization formulation, together with an unrolling strategy, to explore underlying principles of a series of LLV tasks. Furthermore, we design a differentiable strategy to cooperatively search specific scene and task architectures for RUAS. Last but not least, we demonstrate how to apply RUAS for both low- and high-level LLV applications (e.g., enhancement, detection and segmentation). Extensive experiments verify the flexibility, effectiveness, and efficiency of RUAS.
Risheng Liu, Long Ma 0002, Tengyu Ma 0004, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Learning Heavily-Degraded Prior for Underwater Object Detection
abstract
Underwater object detection suffers from low detection performance because the distance and wavelength dependent imaging process yield evident image quality degradations such as haze-like effects, low visibility, and color distortions. Therefore, we commit to resolving the issue of underwater object detection with compounded environmental degradations. Typical approaches attempt to develop sophisticated deep architecture to generate high-quality images or features. However, these methods are only work for limited ranges because imaging factors are either unstable, too sensitive, or compounded. Unlike these approaches catering for high-quality images or features, this paper seeks transferable prior knowledge from detector-friendly images. The prior guides detectors removing degradations that interfere with detection. It is based on statistical observations that, the heavily degraded regions of detector-friendly (DFUI) and underwater images have evident feature distribution gaps while the lightly degraded regions of them overlap each other. Therefore, we propose a residual feature transference module (RFTM) to learn a mapping between deep representations of the heavily degraded patches of DFUI- and underwater-images, and make the mapping as a heavily degraded prior (HDP) for underwater detection. Since the statistical properties are independent to image content, HDP can be learned without the supervision of semantic labels and plugged into popular CNN-based feature extraction networks to improve their performance on underwater object detection. Without bells and whistles, evaluations on URPC2020 and UODD show that our methods outperform CNN-based detectors by a large margin. Our method with higher speeds and less parameters still performs better than transformer-based detectors. Our code and DFUI dataset can be found inhttps://github.com/xiaoDetection/Learning-Heavily-Degraed-Prior.
Chenping Fu, Xin Fan 0001, Jiewen Xiao, Wanqi Yuan, Risheng Liu, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.6
2023 Automated Learning for Deformable Medical Image Registration by Jointly Optimizing Network Architectures and Objective Functions
abstract
Deformable image registration plays a critical role in various tasks of medical image analysis. A successful registration algorithm, either derived from conventional energy optimization or deep networks, requires tremendous efforts from computer experts to well design registration energy or to carefully tune network architectures with respect to medical data available for a given registration task/scenario. This paper proposes an automated learning registration algorithm (AutoReg) that cooperatively optimizes both architectures and their corresponding training objectives, enabling non-computer experts to conveniently find off-the-shelf registration algorithms for various registration scenarios. Specifically, we establish a triple-level framework to embrace the searching for both network architectures and objectives with a cooperating optimization. Extensive experiments on multiple volumetric datasets and various registration scenarios demonstrate that AutoReg can automatically learn an optimal deep registration network for given volumes and achieve state-of-the-art performance. The automatically learned network also improves computational efficiency over the mainstream UNet architecture from 0.558 to 0.270 seconds for a volume pair on the same configuration.
Xin Fan 0001, Risheng Liu, Zhongxuan Luo, Hao Huang 0016
IEEE Trans. Image Process.6
2023 Characteristic Mapping for Ellipse Detection Acceleration
abstract
It is challenging to characterize the intrinsic geometry of high-degree algebraic curves with lower-degree algebraic curves. The reduction in the curve's degree implies lower computation costs, which is crucial for various practical computer vision systems. In this paper, we develop a characteristic mapping (CM) to recursively degenerate 3n points on a planar curve of n th order to 3(n-1) points on a curve of (n-1) th order. The proposed characteristic mapping enables curve grouping on a line, a curve of the lowest order, that preserves the intrinsic geometric properties of a higher-order curve (ellipse). We prove a necessary condition and derive an efficient arc grouping module that finds valid elliptical arc segments by determining whether the mapped three points are colinear, invoking minimal computation. We embed the module into two latest arc-based ellipse detection methods, which reduces their running time by 25% and 50% on average over five widely used data sets. This yields faster detection than the state-of-the-art algorithms while keeping their precision comparable or even higher. Two CM embedded methods also significantly surpass a deep learning method on all evaluation metrics.
Qi Jia 0001, Xin Fan 0001, Yang Yang 0120, Xuxu Liu, Zhongxuan Luo, Xinchen Zhou, Longin Jan Latecki
IEEE Trans. Image Process.5
2023 Optimization-Inspired Learning With Architecture Augmentations and Control Mechanisms for Low-Level Vision
abstract
In recent years, there has been a growing interest in combining learnable modules with numerical optimization to solve low-level vision tasks. However, most existing approaches focus on designing specialized schemes to generate image/feature propagation. There is a lack of unified consideration to construct propagative modules, provide theoretical analysis tools, and design effective learning mechanisms. To mitigate the above issues, this paper proposes a unified optimization-inspired learning framework to aggregate Generative, Discriminative, and Corrective (GDC for short) principles with strong generalization for diverse optimization models. Specifically, by introducing a general energy minimization model and formulating its descent direction from different viewpoints (i.e., in a generative manner, based on the discriminative metric and with optimality-based correction), we construct three propagative modules to effectively solve the optimization models with flexible combinations. We design two control mechanisms that provide the non-trivial theoretical guarantees for both fully- and partially-defined optimization formulations. Under the support of theoretical guarantees, we can introduce diverse architecture augmentation strategies such as normalization and search to ensure stable propagation with convergence and seamlessly integrate the suitable modules into the propagation respectively. Extensive experiments across varied low-level vision tasks validate the efficacy and adaptability of GDC.
Risheng Liu, Zhu Liu 0004, Pan Mu, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.5
2023 Low-Light Image Enhancement via Self-Reinforced Retinex Projection Model
abstract
Low-light image enhancement aims to improve the quality of images captured under low-lightening conditions, which is a fundamental problem in computer vision and multimedia areas. Although many efforts have been invested over the years, existing illumination-based models tend to generate unnatural-looking results (e.g., over-exposure). It is because that the widely-adopted illumination adjustment (e.g., Gamma Correction) breaks down the favorable smoothness property of the original illumination derived from the well-designed illumination estimation model. To settle this issue, a great-efficiency and high-quality Self-Reinforced Retinex Projection (SRRP) model is developed in this paper, which contains optimization modules of both illumination and reflectance layers. Specifically, we construct a new fidelity term with the self-reinforced function for the illumination optimization to eliminate the dependence of the illumination adjustment to obtain a desired illumination with the excellent smoothing property. By introducing a flexible feasible constraint, we obtain a reflectance optimization module with projection. Owing to its flexibility, we can extend our model to an enhanced version by integrating a data-driven denoising mechanism as the projection, which is able to effectively handle the generated noises/artifacts in the enhanced procedure. In the experimental part, on one side, we make ample comparative assessments on multiple benchmarks with considerable state-of-the-art methods. These evaluations fully verify the outstanding performance of our method, in terms of the qualitative and quantitative analyses and execution efficiency. On the other side, we also conduct extensive analytical experiments to indicate the effectiveness and advantages of our proposed model.
Long Ma 0002, Risheng Liu, Yiyang Wang 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Multim.5
2022 Segment, Magnify and Reiterate: Detecting Camouflaged Objects the Hard Way
abstract
It is challenging to accurately detect camouflaged objects from their highly similar surroundings. Existing methods mainly leverage a single-stage detection fashion, while neglecting small objects with low-resolution fine edges requires more operations than the larger ones. To tackle camouflaged object detection (COD), we are inspired by humans attention coupled with the coarse-to-fine detection strategy, and thereby propose an iterative refinement framework, coined SegMaR, which integrates Segment, Magnify and Reiterate in a multi-stage detection fashion. Specifically, we design a new discriminative mask which makes the model attend on the fixation and edge regions. In addition, we leverage an attention-based sampler to magnify the object region progressively with no need of enlarging the image size. Extensive experiments show our SegMaR achieves remarkable and consistent improvements over other state-of-the-art methods. Especially, we surpass two competitive methods 7.4% and 20.0% respectively in average over standard evaluation metrics on small camouflaged objects. Additional studies provide more promising insights into Seg-MaR, including its effectiveness on the discriminative mask and its generalization to other network architectures. Code is available at https://github.com/dlut-dimt/SegMaR.
Qi Jia 0001, Shuilian Yao, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Zhongxuan Luo
CVPR6
2022 Toward Fast, Flexible, and Robust Low-Light Image Enhancement
abstract
Existing low-light image enhancement techniques are mostly not only difficult to deal with both visual quality and computational efficiency but also commonly invalid in unknown complex scenarios. In this paper, we develop a new Self-Calibrated Illumination (SCI) learning framework for fast, flexible, and robust brightening images in real-world low-light scenarios. To be specific, we establish a cascaded illumination learning process with weight sharing to handle this task. Considering the computational burden of the cascaded pattern, we construct the self-calibrated module which realizes the convergence between results of each stage, producing the gains that only use the single basic block for inference (yet has not been exploited in previous works), which drastically diminishes computation cost. We then define the unsupervised training loss to elevate the model capability that can adapt general scenes. Further, we make comprehensive explorations to excavate SCI's inherent properties (lacking in existing works) including operation-insensitive adaptability (acquiring stable performance under the settings of different simple operations) and model-irrelevant generality (can be applied to illumination-based existing works to improve performance). Finally, plenty of experiments and ablation studies fully indicate our superiority in both quality and efficiency. Applications on low-light face detection and nighttime semantic segmentation fully reveal the latent practical values for SCI. The source code is available at https://github.com/vis-opt-group/SCI.
Long Ma 0002, Tengyu Ma 0004, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
CVPR5
2022 Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection
abstract
This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization or deep networks. These approaches neglect that modality differences implying the complementary information are extremely important for both fusion and subsequent detection task. This paper proposes a bilevel optimization formulation for the joint problem of fusion and detection, and then unrolls to a target-aware Dual Adversarial Learning (TarDAL) network for fusion and a commonly used detection network. The fusion network with one generator and dual discriminators seeks commons while learning from differences, which preserves structural information of targets from the infrared and textural details from the visible. Furthermore, we build a synchronized imaging system with calibrated infrared and optical sensors, and collect currently the most comprehensive benchmark covering a wide range of scenarios. Extensive experiments on several public datasets and our benchmark demonstrate that our method outputs not only visually appealing fusion but also higher detection mAP than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/dlut-dimt/TarDAL.
Jinyuan Liu 0001, Xin Fan 0001, Zhanbo Huang, Guanyao Wu, Risheng Liu, Zhongxuan Luo
CVPR7
2022 ReCoNet: Recurrent Correction Network for Fast and Efficient Multi-modality Image Fusion
Zhanbo Huang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu, Zhongxuan Luo
ECCV (18)6
2022 Blindfold Attention: Novel Mask Strategy for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) is a basic and crucial computer vision task of classifying emotional expressions from human faces images into various emotion categories such as happy, sad, surprised, scared, angry, etc. Recently, facial expression recognition based on deep learning has made great progress. However, no matter the weight initialization technology or the attention mechanism, the face recognition method based on deep learning hard to capture those visually insignificant but semantically important features. To aid above question, in this paper we present a novel Facial Expression Recognition training strategy consisting of two components: Memo Affinity Loss (MAL) and Mask Attention Fine Tuning (MAFT). MAL is a variant of center loss, which uses memory bank strategy as well as discriminative center. MAL widens the distance between different clusters and narrows the distance within each cluster. Therefore, the features extracted by CNN were comprehensive and independent, which produced a more robust model. MAFT is a strategy that blindfolds attention parts temporarily and forces the model to learn from other important regions of the input image. It's not only an augmenting technique, but also a novel fine-tuning approach. As we know, we are the first to apply the mask strategy to the attention part and use this strategy to fine-tune the models. Finally, to implement our ideas, we constructed a new network named Architecture Attention ResNet based on ResNet-18. Our methods are conceptually and practically simple, but receives superior results on popular public facial expression recognition benchmarks with 88.75% on RAF-DB, 65.17% on AffectNet-7, 60.72% on AffectNet-8. The code will open source soon.
Bo Fu 0001, Yuanxin Mao, Shilin Fu, Yonggong Ren, Zhongxuan Luo
ICMR5
2022 PIA: Parallel Architecture with Illumination Allocator for Joint Enhancement and Detection in Low-Light
abstract
Visual perception in low-light conditions (e.g., nighttime) plays an important role in various multimedia-related applications (e.g., autonomous driving). The enhancement (provides a visual-friendly appearance) and detection (detects the instances of objects) in low-light are two fundamental and crucial visual perception tasks. In this paper, we make efforts on how to simultaneously realize low-light enhancement and detection from two aspects. First, we define a parallel architecture to satisfy the task demand for both two tasks. In which, a decomposition-type warm-start acting on the entrance of parallel architecture is developed to narrow down the adverse effects brought by low-light scenes to some extent. Second, a novel illumination allocator is designed by encoding the key illumination component (the inherent difference between normal-light and low-light) to extract hierarchical features for assisting in enhancement and detection. Further, we make a substantive discussion for our proposed method. That is, we solve enhancement in a coarse-to-fine manner and handle detection in a decomposed-to-integrated fashion. Finally, multidimensional analytical and evaluated experiments are performed to indicate our effectiveness and superiority. The code is available at \urlhttps://github.com/tengyu1998/PIA
Tengyu Ma 0004, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
ACM Multimedia4
2022 DeepEnReg: Joint Enhancement and Affine Registration for Low-contrast Medical Images
Yun Peng 0005, Huanyu Luo, Huanjie Li, Zhongxuan Luo, Xin Fan 0001
PRCV (2)8
2022 Efficient Representation and Optimization of TPMS-Based Porous Structures for 3D Heat Dissipation
Shengfa Wang, Yu Jiang 0019, Jiangbei Hu, Xin Fan 0001, Zhongxuan Luo, Ligang Liu 0001
Comput. Aided Des.5
2022 Function Representation Based Analytic Shape Hollowing Optimization
Shengfa Wang, Baojun Li, Yi Wang 0037, Zhongxuan Luo, Ligang Liu 0001
Comput. Aided Des.5
2022 Learning Deformable Image Registration From Optimization: Perspective, Modules, Bilevel Training and Beyond
abstract
Conventional deformable registration methods aim at solving an optimization model carefully designed on image pairs and their computational costs are exceptionally high. In contrast, recent deep learning-based approaches can provide fast deformation estimation. These heuristic network architectures are fully data-driven and thus lack explicit geometric constraints which are indispensable to generate plausible deformations, e.g., topology-preserving. Moreover, these learning-based approaches typically pose hyper-parameter learning as a black-box problem and require considerable computational and human effort to perform many training runs. To tackle the aforementioned problems, we propose a new learning-based framework to optimize a diffeomorphic model via multi-scale propagation. Specifically, we introduce a generic optimization model to formulate diffeomorphic registration and develop a series of learnable architectures to obtain propagative updating in the coarse-to-fine feature space. Further, we propose a new bilevel self-tuned training strategy, allowing efficient search of task-specific hyper-parameters. This training strategy increases the flexibility to various types of data while reduces computational and human burdens. We conduct two groups of image registration experiments on 3D volume datasets including image-to-atlas registration on brain MRI data and image-to-image registration on liver CT data. Extensive results demonstrate the state-of-the-art performance of the proposed method with diffeomorphic guarantee and extreme efficiency. We also apply our framework to challenging multi-modal image registration, and investigate how our registration to support the down-streaming tasks for medical image analysis including multi-modal fusion and image segmentation.
Risheng Liu, Xin Fan 0001, Chenying Zhao, Hao Huang 0016, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Discriminative information restoration and extraction for weakly supervised low-resolution fine-grained image recognition
Tiantian Yan, Zhongxuan Luo, Zhihui Wang 0001
Pattern Recognit.4
2022 Learning a Deep Multi-Scale Feature Ensemble and an Edge-Attention Guidance for Image Fusion
abstract
Image fusion integrates a series of images acquired from different sensors,e.g., infrared and visible, outputting an image with richer information than either one. Traditional and recent deep-based methods have difficulties in preserving prominent structures and recovering vital textural details for practical applications. In this article, we propose a deep network for infrared and visible image fusion cascading a feature learning module with a fusion learning mechanism. Firstly, we apply a coarse-to-fine deep architecture to learn multi-scale features for multi-modal images, which enables discovering prominent common structures for later fusion operations. The proposed feature learning module requires no well-aligned image pairs for training. Compared with the existing learning-based methods, the proposed feature learning module can ensemble numerous examples from respective modals for training, increasing the ability of feature representation. Secondly, we design an edge-guided attention mechanism upon the multi-scale features to guide the fusion focusing on common structures, thus recovering details while attenuating noise. Moreover, we provide a new aligned infrared and visible image fusion dataset, RealStreet, collected in various practical scenarios for comprehensive evaluation. Extensive experiments on two benchmarks, TNO and RealStreet, demonstrate the superiority of the proposed method over the state-of-the-art in terms of both visual inspection and objective analysis on six evaluation metrics. We also conduct the experiments on the FLIR and NIR datasets, containing foggy weather and poor light conditions, to verify the generalization and robustness of the proposed method.
Jinyuan Liu 0001, Xin Fan 0001, Ji Jiang, Risheng Liu, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.5
2022 Discriminative Feature Mining and Enhancement Network for Low-Resolution Fine-Grained Image Recognition
abstract
Existing fine-grained image recognition methods are difficult to learn complete discriminative features from low-resolution (LR) data, because the original subtle inter-class distinctions become slimmer with the reduction of the image resolution. Besides, existing methods of LR fine-grained image recognition and general LR image recognition only consider the restoration and extraction of global discriminative features, ignoring unreliable local fine-grained details can be detrimental to final recognition. To address the above problems, we propose a multi-tasking framework, discriminative feature mining and enhancement network (DME-Net), for the LR fine-grained image recognition task, which aims to capture the reliable object descriptions from macro and micro perspectives, respectively. Macroscopically, we train the framework’s ability to recover and extract global discriminative features based on the whole images. Microscopically, we purposefully reinforce the framework’s ability to repair and capture the local discriminative details on the mined informative parts. To precisely excavate the most potential parts, we design an informative part mining (IPM) module, in which we firstly employ a part generation layer to predict several part masks that focus on different discriminative parts under the guidance of discrepancy loss and discriminant loss. Then we introduce a part selection (PS) submodule to further screen out a group of most informative parts from the predicted part masks according to their corresponding scores, which measure the semantic correlation degree of each part to the others. Experimental results on three benchmark datasets and one retail product dataset consistently show that our proposed framework can significantly boost the performance of the baseline model. Besides, extensive ablation studies are conducted, which further prove the effectiveness of each component of our designs.
Tiantian Yan, Baoli Sun, Zhihui Wang 0001, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.5
2022 Underexposed Image Correction via Hybrid Priors Navigated Deep Propagation
abstract
Enhancing visual quality for underexposed images is an extensively concerning task that plays an important role in various areas of multimedia and computer vision. Most existing methods often fail to generate high-quality results with appropriate luminance and abundant details. To address these issues, we develop a novel framework, integrating both knowledge from physical principles and implicit distributions from data to address underexposed image correction. More concretely, we propose a new perspective to formulate this task as an energy-inspired model with advanced hybrid priors. A propagation procedure navigated by the hybrid priors is well designed for simultaneously propagating the reflectance and illumination toward desired results. We conduct extensive experiments to verify the necessity of integrating both underlying principles (i.e., with knowledge) and distributions (i.e., from data) as navigated deep propagation. Plenty of experimental results of underexposed image correction demonstrate that our proposed method performs favorably against the state-of-the-art methods on both subjective and objective assessments. In addition, we execute the task of face detection to further verify the naturalness and practical value of underexposed image correction. What is more, we apply our method to solve single-image haze removal whose experimental results further demonstrate our superiorities.
Risheng Liu, Long Ma 0002, Yuxi Zhang 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.5
2022 Learning Deep Context-Sensitive Decomposition for Low-Light Image Enhancement
abstract
Enhancing the quality of low-light (LOL) images plays a very important role in many image processing and multimedia applications. In recent years, a variety of deep learning techniques have been developed to address this challenging task. A typical framework is to simultaneously estimate the illumination and reflectance, but they disregard the scene-level contextual information encapsulated in feature spaces, causing many unfavorable outcomes, e.g., details loss, color unsaturation, and artifacts. To address these issues, we develop a new context-sensitive decomposition network (CSDNet) architecture to exploit the scene-level contextual dependencies on spatial scales. More concretely, we build a two-stream estimation mechanism including reflectance and illumination estimation network. We design a novel context-sensitive decomposition connection to bridge the two-stream mechanism by incorporating the physical principle. The spatially varying illumination guidance is further constructed for achieving the edge-aware smoothness property of the illumination component. According to different training patterns, we construct CSDNet (paired supervision) and context-sensitive decomposition generative adversarial network (CSDGAN) (unpaired supervision) to fully evaluate our designed architecture. We test our method on seven testing benchmarks [including massachusetts institute of technology (MIT)-Adobe FiveK, LOL, ExDark, and naturalness preserved enhancement (NPE)] to conduct plenty of analytical and evaluated experiments. Thanks to our designed context-sensitive decomposition connection, we successfully realized excellent enhanced results (with sufficient details, vivid colors, and few noises), which fully indicates our superiority against existing state-of-the-art approaches. Finally, considering the practical needs for high efficiency, we develop a lightweight CSDNet (named LiteCSDNet) by reducing the number of channels. Furthermore, by sharing an encoder for these two components, we obtain a more lightweight version (SLiteCSDNet for short). SLiteCSDNet just contains 0.0301M parameters but achieves the almost same performance as CSDNet. Code is available at https://github.com/KarelZhang/CSDNet-CSDGAN.
Long Ma 0002, Risheng Liu, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.5
2022 Efficient Representation and Optimization for TPMS-Based Porous Structures
abstract
In this approach, we present an efficient topology and geometry optimization of triply periodic minimal surfaces (TPMS) based porous shell structures, which can be represented, analyzed, optimized and stored directly using functions. The proposed framework is directly executed on functions instead of remeshing (tetrahedral/hexahedral), and this framework substantially improves the controllability and efficiency. Specifically, a valid TPMS-based porous shell structure is first constructed by function expressions. The porous shell permits continuous and smooth changes of geometry (shell thickness) and topology (porous period). The porous structures also inherit several of the advantageous properties of TPMS, such as smoothness, full connectivity (no closed hollows), and high controllability. Then, the problem of filling an object's interior region with porous shell can be formulated into a constraint optimization problem with two control parameter functions. Finally, an efficient topology and geometry optimization scheme is presented to obtain optimized scale-varying porous shell structures. In contrast to traditional heuristic methods for TPMS, our work directly optimize both the topology and geometry of TPMS-based structures. Various experiments have shown that our proposed porous structures have obvious advantages in terms of efficiency and effectiveness.
Jiangbei Hu, Shengfa Wang, Baojun Li, Fengqi Li, Zhongxuan Luo, Ligang Liu 0001
IEEE Trans. Vis. Comput. Graph.5
2021 Retinex-Inspired Unrolling With Cooperative Prior Architecture Search for Low-Light Image Enhancement
abstract
Low-light image enhancement plays very important roles in low-level vision areas. Recent works have built a great deal of deep learning models to address this task. However, these approaches mostly rely on significant architecture engineering and suffer from high computational burden. In this paper, we propose a new method, named Retinex-inspired Unrolling with Architecture Search (RUAS), to construct lightweight yet effective enhancement network for low-light images in real-world scenario. Specifically, building upon Retinex rule, RUAS first establishes models to characterize the intrinsic underexposed structure of low-light images and unroll their optimization processes to construct our holistic propagation structure. Then by designing a cooperative reference-free learning strategy to discover low-light prior architectures from a compact search space, RUAS is able to obtain a top-performing image enhancement network, which is with fast speed and requires few computational resources. Extensive experiments verify the superiority of our RUAS framework against recently proposed state-of-the-art methods. The project page is available at http://dutmedia.org/RUAS/.
Risheng Liu, Long Ma 0002, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
CVPR5
2021 NASA: A Noise-Adaptive and Structure-Aware Learning Framework for Image Deblurring
abstract
Image deblurring is a classical low-level visual processing task, which aims to recover a potentially noise-free sharp image from the blurred image. Existing prior-based and learning-based methods usually need to manually set some vital auxiliary components (e.g., noise level). It brings about extremely weak adaptability and flexibility. To settle this issue, we develop a Noise-Adaptive Structure-Aware learning framework (NASA) to achieve fully intelligent manufacturing. Concretely, by introducing a new task-assisted module, we define a novel robust image deblurring model derived from a MAP-based energy function. Consequently, we establish the NASA which consists of three basic modules including the task-assisted, fidelity-term, and regularization-term modules, to solve our designed model. The task-assisted module generates the noise-adaptive and structure-aware maps, which are fed to the other two modules. By end-to-end training our NASA, we successfully avoid the cumbersome manually parameters-adjustment process. Quantitative and qualitative experiments demonstrate our superiority compared to the state-of-the-art methods, both in visual effect and numerical scores. A series of ablation study also verify the effectiveness and necessity of our designed mechanism.
Xiaokun Liu, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICASSP6
2021 Dynamic Context-Sensitive Filtering Network for Video Salient Object Detection
abstract
The ability to capture inter-frame dynamics has been critical to the development of video salient object detection (VSOD). While many works have achieved great success in this field, a deeper insight into its dynamic nature should be developed. In this work, we aim to answer the following questions: How can a model adjust itself to dynamic variations as well as perceive fine differences in the real-world environment; How are the temporal dynamics well introduced into spatial information over time? To this end, we propose a dynamic context-sensitive filtering network (DCFNet) equipped with a dynamic context-sensitive filtering module (DCFM) and an effective bidirectional dynamic fusion strategy. The proposed DCFM sheds new light on dynamic filter generation by extracting location-related affinities between consecutive frames. Our bidirectional dynamic fusion strategy encourages the interaction of spatial and temporal information in a dynamic manner. Experimental results demonstrate that our proposed method can achieve state-of-the-art performance on most VSOD datasets while ensuring a real-time speed of 28 fps. The source code is publicly available at https://github.com/OIPLab-DUT/DCFNet.
Miao Zhang 0004, Jie Liu 0044, Yongri Piao, Shunyu Yao 0004, Wei Ji 0011, Huchuan Lu, Zhongxuan Luo
ICCV9
2021 Robust Image Denoising with Texture-Aware Neural Network
abstract
Image denoising is a well-studied yet still hot research topic in the image processing community. Recently, image denoising with deep neural networks has achieved superior performance, however they can not recover tiny details from noisy images. Motivated by this problem, we propose a Texture-Aware Neural Network named TANet, which is composed of main network part with attention mechanism, residual structure and Texture-Aware Modular. Proposed Texture-Aware Modular owns dual paths, denoised image from main denoising network and clean image are input different path respectively. From Texture-Aware Modular, we get two sets intermediate codes and calculate corresponding perceptual loss. This perceptual loss is designed to generate auxiliary super-vision for tiny detail recovery from mixed residual details and noise set. Extensive experimental results demonstrate that the proposed TANet is on a par with the state-of-the-art denoising methods.
Bo Fu 0001, Zhongxuan Luo
ICME3
2021 Multiple Task-Oriented Encoders for Unified Image Fusion
abstract
Image fusion methods have achieved incredible progress, but they are vulnerable to handling a certain type of fusion task rather than considering deeper relations between cross-realm task correlations. To achieve this, we integrate different image fusion tasks into a unified network. Our method is accomplished through multiple task-oriented encoders and a generic decoder, in addition to a self-adapting loss function. The taskoriented encoders are trained to learn task-specific features, while the generic decoder reconstructs the fused features to generate a comprehensive image. Subsequently, by introducing the self-adapting loss in our method, it can automatically adjust itself to source data characteristics on different tasks. Besides, we formulate a training strategy based on bilevel optimization to update the multi-encoder and generic decoder in an alternative manner. Extensive experimental results demonstrate the superior performance of our method over the stateof-the-art methods.
Zhuoxiao Li, Jinyuan Liu 0001, Risheng Liu, Xin Fan 0001, Zhongxuan Luo, Wen Gao 0001
ICME5
2021 Spatial-Temporal Integration Network with Self-Guidance for Robust Video Deraining
abstract
Recently, video deraining has become a research focus. Network-based approaches are continuously showing extrusive performance. However, they lack precise control over the motion consistency in temporal information and characterize spatial distribution, so that their results are unsatisfying, especially in some real-world scenarios. To settle them, we develop a spatial-temporal integration network with self-guidance. It contains flow-induced alignment, self-guidance generation, and spatial-temporal integration modules. The alignment module not only preliminarily removes rain to provide more effective temporal correlation but also accurately keeps motion consistency between frames. The self-guidance map characterizes the pixel-level spatial distribution for the target to avoid injuring the background. Finally, we concatenate adjacent aligned frames, self-guidance map, and original current rain frame into the integration module to progressively fuse them in a coarse-to-fine way. Extensive evaluations demonstrate our superiority against other state-of-the-art methods qualitatively and quantitatively.
Xiaokun Liu, Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICME5
2021 Hardware-Aware Low-Light Image Enhancement via One-Shot Neural Architecture Search with Shrinkage Sampling
abstract
Low-light image enhancement has traditionally been tackled by training a heuristically designed neural network architecture. Despite the success of these approaches, the heuristic design pattern inherently not only hinders further optimization of network architectures, but also limits the factors that the designer can take into consideration. As a result, these methods are difficult to achieve a balance between enhancing performance and hardware related performance. In this paper, we equip a basic enhancing algorithm with a neural architecture search technique. This technique helps to automatically search an optimal hardware-aware architecture while also increases neglectable computation burden. In this work, we propose a shrinkage sampling strategy to drastically decrease the computation cost of neural architecture search while improving the quality of search. Extensive experiments on various benchmarks demonstrate that our algorithm achieves state-of-the-art performance with higher speed.
Yuansheng Yao, Risheng Liu, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
ICME6
2021 Star-Net: Spatial-Temporal Attention Residual Network for Video Deraining
abstract
Learning-based video deraining has recently drawn increasing attention. They tend to directly package aligned frames to input a fully end-to-end network. However, the network is generally object-driven and cannot recognize how to utilize temporal information so that the results are unsatisfied. In this work, we design a novel Spatial-Temporal Attention Network (STAR-Net) to explicitly utilize the temporal information. Concretely, we define the self-spatial attention to characterizing the rain region of the target frame, and the temporal-spatial attention to learn the profitable information for remedying the rain region of the target frame from the adjacent frame. We also introduce a simple residual network to further strengthen the relationship between the target and the adjacent frame. These addressed frames are fused by a three-layers convolutional module to further improve the capability. Extensive evaluations indicate our superiority against state-of-the-art methods.
Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME6
2021 Collaborative Reflectance-And-Illumination Learning For High-Efficient Low-Light Image Enhancement
abstract
In this paper, we settle the low-light image enhancement problem by developing a collaborative learning framework, which not only improves lightness and suppresses noises simultaneously but also with fast speed and requires few computational resources. The approach is inspired by the fact that reflectance and illumination are highly correlated to satisfy the well-known Retinex decomposition principle. With this in mind, we establish a Reflectance-and-Illumination Collaborative (RIC) block to depict the compact physical relationship between reflectance and illumination. By cascading multiple RIC blocks, we obtain an end-to-end RICNet to interactively optimize these two components in a collaborative manner. Benefiting from the RIC block that integrates powerful task cues, RICNet just needs few parameters to simultaneously improve brightness and remove noises. Extensive experiments demonstrate our superiority against existing state-of-the-art methods. We also make meticulous analysis for the RIC block. The results reveal the rationality and effectiveness of our built mechanism.
Guijing Zhu, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME5
2021 Shrimp: a robust underwater visible light communication system
abstract
This paper presents the design, implementation, and evaluation of Shrimp, an underwater visible light communication (VLC) system. To address the unique issues in underwater environment such as water flow and scattered sunlight interference, we exploit the circularly polarized light (CPL) and double links for underwater VLC transmission. A coding scheme tailored for underwater communication based on double CPL design is developed. We prototype Shrimp on commercial-off-the-shelf (COTS) LEDs with fabricated printed circuit boards (PCBs). Extensive experiments conducted in an indoor water pool, a lake, and the sea demonstrate that Shrimp can combat against environmental interference and achieve robust communication in underwater environments. The communication distance can be up to 3 m in sea/lake water using a 3 W commodity LED, outperforming the VLC schemes designed for in-air communication.
Chi Lin 0001, Yongda Yu, Jie Xiong 0001, Lei Wang 0005, Guowei Wu 0001, Zhongxuan Luo
MobiCom7
2021 Latency-Constrained Spatial-Temporal Aggregated Architecture Search for Video Deraining
Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo, Yuduo Zhang
PRCV (3)5
2021 Semantic-Driven Context Aggregation Network for Underwater Image Enhancement
Dongxiang Shi, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
PRCV (3)5
2021 Novelty Detection and Online Learning for Chunk Data Streams
abstract
Datastream analysis aims at extracting discriminative information for classification from continuously incoming samples. It is extremely challenging to detect novel data while incrementally updating the model efficiently and stably, especially for high-dimensional and/or large-scale data streams. This paper proposes an efficient framework for novelty detection and incremental learning for unlabeled chunk data streams. First, an accurate factorization-free kernel discriminative analysis (FKDA-X) is put forward through solving a linear system in the kernel space. FKDA-X produces a Reproducing Kernel Hilbert Space (RKHS), in which unlabeled chunk data can be detected and classified by multiple known-classes in a single decision model with a deterministic classification boundary. Moreover, based on FKDA-X, two optimal methods FKDA-CX and FKDA-C are proposed. FKDA-CX uses the micro-cluster centers of original data as the input to achieve excellent performance in novelty detection. FKDA-C and incremental FKDA-C (IFKDA-C) using the class centers of original data as their input have extremely fast speed in online learning. Theoretical analysis and experimental validation on under-sampled and large-scale real-world datasets demonstrate that the proposed algorithms make it possible to learn unlabeled chunk data streams with significantly lower computational costs and comparable accuracies than the state-of-the-art approaches.
Yi Wang 0037, Xiangjian He, Xin Fan 0001, Chi Lin 0001, Fengqi Li, Tianzhu Wang, Zhongxuan Luo, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2021 Learning Hadamard-Product-Propagation for Image Dehazing and Beyond
abstract
Image dehazing has evolved into an attractive research field in the computer vision community in the past few decades. Previous traditional approaches attempt to design energy-based objective functions. However, they cannot accurately express the intrinsic characteristics of the images, posing weak adaptation ability for real-world complex scenarios. More recently, deep learning techniques for image dehazing have matured and become more reliable, showing outstanding performance. Nevertheless, these methods heavily depend on training data, restricting their application ranges. More importantly, both traditional and deep learning approaches all ignore a common issue, noises/artifacts always appear in the recovery process. To this end, a new Hadamard-Product (HP) model is proposed, which consists of a series of data-driven priors. Based on this model, we derive a Learnable Hadamard-Product-Propagation (LHPP) by cascading a series of principle-inspired guidance and recovery modules. In which, the principle-inspired guidance related to transmission is endowed the smoothness property, the other recovery module satisfies the distribution of natural images. The Hadamard-product-based propagations is generated in our developed learnable framework for the task of image dehazing. In this way, we can eliminate noises/artifacts in the recovery procedure to obtain the ideal outputs. Subsequently, since the generality of our HP model, we successfully extend our LHPP to settle low-light image enhancement and underwater image enhancement problems. A series of analytical experiments are performed to verify our effectiveness. Plenty of performance evaluations on three complex tasks fully reveal our superiority against multiple state-of-the-art methods.
Risheng Liu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.6
2021 A Bilevel Integrated Model With Data-Driven Layer Ensemble for Multi-Modality Image Fusion
abstract
Image fusion plays a critical role in a variety of vision and learning applications. Current fusion approaches are designed to characterize source images, focusing on a certain type of fusion task while limited in a wide scenario. Moreover, other fusion strategies (i.e., weighted averaging, choose-max) cannot undertake the challenging fusion tasks, which furthermore leads to undesirable artifacts facilely emerged in their fused results. In this paper, we propose a generic image fusion method with a bilevel optimization paradigm, targeting on multi-modality image fusion tasks. Corresponding alternation optimization is conducted on certain components decoupled from source images. Via adaptive integration weight maps, we are able to get the flexible fusion strategy across multi-modality images. We successfully applied it to three types of image fusion tasks, including infrared and visible, computed tomography and magnetic resonance imaging, and magnetic resonance imaging and single-photon emission computed tomography image fusion. Results highlight the performance and versatility of our approach from both quantitative and qualitative aspects.
Risheng Liu, Jinyuan Liu 0001, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.5
2021 Dual Neural Networks Coupling Data Regression With Explicit Priors for Monocular 3D Face Reconstruction
abstract
We address the challenging issue of reconstructing a 3D face from one single image under various expressions and illuminations, which is widely applied in multimedia tasks. Methods built upon classical parametric morphable models (3DMMs) gain success on reconstructing the global geometry of a 3D face, but fail to precisely characterize local facial details. Recently, deep neural networks (DNN) have been applied to the reconstruction that directly predicts depth maps, showing compelling performance on detail recovery. Unfortunately, their reconstruction is prone to structural distortions owing to the lack of explicit prior constraints. In this paper, we propose dual neural networks that optimize one energy coupling data fitting with local explicit geometric prior. Specifically, we build one residual network upon traditional convolution layers in order to directly predict 3D structures by fitting an input image. Meanwhile, we devise a novel architecture stacking shallow networks to refine 3D clouds with geometric priors given by Markov random fields (MRFs). Quantitative evaluations demonstrate the superior performance of the dual networks over either end-to-end DNNs or parametric models. Comparisons with the state-of-the-art also show competitive reconstruction quality on various conditions.
Xin Fan 0001, Shichao Cheng, Kang Huyan, Minjun Hou, Risheng Liu, Zhongxuan Luo
IEEE Trans. Multim.6
2021 Location-Aware and Regularization-Adaptive Correlation Filters for Robust Visual Tracking
abstract
Correlation filter (CF) has recently been widely used for visual tracking. The estimation of the search window and the filter-learning strategies is the key component of the CF trackers. Nevertheless, prevalent CF models separately address these issues in heuristic manners. The commonly used CF models directly set the estimated location in the previous frame as the search center for the current one. Moreover, these models usually rely on simple and fixed regularization for filter learning, and thus, their performance is compromised by the search window size and optimization heuristics. To break these limits, this article proposes a location-aware and regularization-adaptive CF (LRCF) for robust visual tracking. LRCF establishes a novel bilevel optimization model to address simultaneously the location-estimation and filter-training problems. We prove that our bilevel formulation can successfully obtain a globally converged CF and the corresponding object location in a collaborative manner. Moreover, based on the LRCF framework, we design two trackers named LRCF-S and LRCF-SA and a series of comparisons to prove the flexibility and effectiveness of the LRCF framework. Extensive experiments on different challenging benchmark data sets demonstrate that our LRCF trackers perform favorably against the state-of-the-art methods in practice.
Risheng Liu, Qianru Chen, Yuansheng Yao, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.5
2020 Image Restoration Via Data-Dependent Proximal Averaged Optimization
abstract
Maximum A Posterior (MAP) acts as one of the most popular modeling scheme in image restoration and is usually reduced to a separable optimization model. Unfortunately, it is challenging to establish exact regularization term and the model with complex priors is hard to optimize. In additionally, it is still hard to incorporate different domain knowledge and data-dependent information into MAP model without changing the property of the objective. To partially address the above issues, we develop a Data-dependent Proximal Averaged (DPA) paradigm through optimizing objective and data-dependent feasibility constraint for the challenging Image Restoration (IR) tasks. Both visual and quantitative comparison results demonstrate that our method outperforms the state of the art.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICASSP6
2020 Principle-Inspired Multi-Scale Aggregation Network for Extremely Low-Light Image Enhancement
abstract
The under-exposure and low-light environments are common to degrade the image-quality with invisible information. To ameliorate this case, a copious of low-light image enhancement methods are developed. However, these existing works are hard to handle extremely low-light conditions with noises, even well-known network-based methods. To address this issue, we develop a Principle-inspired Multi-scale Aggregation Network (PMA-Net) to simultaneously achieve the exposure enhancement and noises removal. Specifically, we establish a pioneering principle-inspired connection to present the physical principle in the inside of the network, to strengthen the structural depict. Subsequently, we propose a multi-scale aggregation strategy to eliminate the noises in the enhanced results. Sufficient ablation studies manifest the effectiveness of our PMA-Net. Extensive qualitative and quantitative comparisons with other state-of-the-art methods are conducted to fully indicates our outstanding performance.
Jiaao Zhang, Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP6
2020 An Efficient Ellipse Detector Based On Region Detection And Arc Pruning
abstract
Detecting ellipses accurately and efficiently for real-world images is crucial for various visual-based applications. Most existing methods employ detection strategies throughout images, while most time is spent on the non-ellipse region. Meanwhile, small ellipses are often miss-detected due to the low resolution and fixed parameters of detectors. In this paper, we proposed an effective ellipse detector benefiting from the region detection method, which provides a basic estimation on the region and size of ellipses. Then, a two-level arc pruning strategy is proposed to detect ellipses efficiently while limiting false-positive and false-negative results. Furthermore, for the pre-estimated region without detected ellipses, interpolation method is employed to enlarge the target region, which makes small and blur ellipses to be detected. Experimental results demonstrate that the proposed method achieves competitive accuracy compared with the state-of-the-art methods.
Ruike Zhang, Jingchao Liang, Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo
ICIP6
2020 Ae-OT: a New Generative Model based on Extended Semi-discrete Optimal transport
Dongsheng An, Na Lei, Zhongxuan Luo, Shing-Tung Yau, Xianfeng Gu
ICLR4
2020 Flexible Bilevel Image Layer Modeling For Robust Deraining
abstract
Visual quality degradation by rain streaks in images/videos is a significant factor that makes many computer vision systems fail to function properly. However, existing rain removal methods tend to remove a specific type of rain streaks while cannot deal with diverse real rainy images. In this paper, we formulate a novel rain model collectively with two contrasting rain streaks and a weighting map. To self-adaptively handle the rain removal problem in the presence of various types of rain streaks, we further propose a bilevel optimization learning framework. Then, we synthesize a new dataset to evaluate the ability of our method to deal with diverse rain streaks. Extensive experiments show that our method can make better performance on both synthesized and real rainy images.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME5
2020 Bi-level Probabilistic Feature Learning for Deformable Image Registration
abstract
We address the challenging issue of deformable registration that robustly and efficiently builds dense correspondences between images. Traditional approaches upon iterative energy optimization typically invoke expensive computational load. Recent learning-based methods are able to efficiently predict deformation maps by incorporating learnable deep networks. Unfortunately, these deep networks are designated to learn deterministic features for classification tasks, which are not necessarily optimal for registration. In this paper, we propose a novel bi-level optimization model that enables jointly learning deformation maps and features for image registration. The bi-level model takes the energy for deformation computation as the upper-level optimization while formulates the maximum \emph{a posterior} (MAP) for features as the lower-level optimization. Further, we design learnable deep networks to simultaneously optimize the cooperative bi-level model, yielding robust and efficient registration. These deep networks derived from our bi-level optimization constitute an unsupervised end-to-end framework for learning both features and deformations. Extensive experiments of image-to-atlas and image-to-image deformable registration on 3D brain MR datasets demonstrate that we achieve state-of-the-art performance in terms of accuracy, efficiency, and robustness.
Risheng Liu, Yuxi Zhang 0001, Xin Fan 0001, Zhongxuan Luo
IJCAI5
2020 A Link Scheduling Algorithm for Underwater Optical Wireless Networks
Zhengxin Fan, Lei Wang 0005, Bingxian Lu, Yongda Yu, Chi Lin 0001, Zhongxuan Luo, Zhenquan Qin, Ming Zhu 0001
Networking6
2020 Learning Multi-scale Retinex with Residual Network for Low-Light Image Enhancement
Long Ma 0002, Jingjie Shang, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
PRCV (1)6
2020 Non-rigid 3D shape retrieval based on multi-scale graphical image and joint Bayesian
Haohao Li, Zhixun Su, Nannan Li 0002, Ximin Liu, Shengfa Wang, Zhongxuan Luo
Comput. Aided Geom. Des.6
2020 Blind image deblurring via hybrid deep priors modeling
Shichao Cheng, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
Neurocomputing5
2020 On the Convergence of Learning-Based Iterative Methods for Nonconvex Inverse Problems
abstract
Numerous tasks at the core of statistics, learning and vision areas are specific cases of ill-posed inverse problems. Recently, learning-based (e.g., deep) iterative methods have been empirically shown to be useful for these problems. Nevertheless, integrating learnable structures into iterations is still a laborious process, which can only be guided by intuitions or empirical insights. Moreover, there is a lack of rigorous analysis about the convergence behaviors of these reimplemented iterations, and thus the significance of such methods is a little bit vague. This paper moves beyond these limits and proposes Flexible Iterative Modularization Algorithm (FIMA), a generic and provable paradigm for nonconvex inverse problems. Our theoretical analysis reveals that FIMA allows us to generate globally convergent trajectories for learning-based iterative methods. Meanwhile, the devised scheduling policies on flexible modules should also be beneficial for classical numerical methods in the nonconvex scenario. Extensive experiments on real applications verify the superiority of FIMA.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhouchen Lin, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 Progressive learning for weakly supervised fine-grained classification
Tiantian Yan, Shijie Wang 0003, Zhihui Wang 0001, Zhongxuan Luo
Signal Process.5
2020 Joint Over and Under Exposures Correction by Aggregated Retinex Propagation for Image Enhancement
abstract
Since the interference of ambient light and the limitation of physical devices, it is quite a common phenomenon that images taken in real-world scenarios turn out to be incorrectly exposed. Most existing techniques emphasize underexposed image correction. On one hand, these works ignore the correction of over-exposure regions in the original input. On the other hand, it is likely to generate over-exposure images. To mitigate these issues, we have developed a novel aggregated Retinex propagations to simultaneously correct over and under-exposure correction of a single image. Concretely, we first manifest the necessity of concurrently correcting under and over-exposure appearances. We establish a Retinex image propagation framework with shared weights to correct different levels of exposure. Then by introducing the fusion computational module, we achieve the accurate exposure correction for a single image. Plenty of quantitative and qualitative comparisons are conducted to fully indicate our superiority against other state-of-the-art algorithms. The elaborated algorithmic analyses show our effectiveness. Experiments on face detection further verify our practicability.
Long Ma 0002, Dian Jin 0003, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
IEEE Signal Process. Lett.5
2020 Real-World Underwater Enhancement: Challenges, Benchmarks, and Solutions Under Natural Light
abstract
Underwater image enhancement is such an important low-level vision task with many applications that numerous algorithms have been proposed in recent years. These algorithms developed upon various assumptions demonstrate successes from various aspects using different data sets and different metrics. In this work, we setup an undersea image capturing system, and construct a large-scale Real-world Underwater Image Enhancement (RUIE) data set divided into three subsets. The three subsets target at three challenging aspects for enhancement, i.e., image visibility quality, color casts, and higher-level detection/classification, respectively. We conduct extensive and systematic experiments on RUIE to evaluate the effectiveness and limitations of various algorithms to enhance visibility and correct color casts on images with hierarchical categories of degradation. Moreover, underwater image enhancement in practice usually serves as a preprocessing step for mid-level and high-level vision tasks. We thus exploit the object detection performance on enhanced images as a brand new task-specific evaluation criterion. The findings from these evaluations not only confirm what is commonly believed, but also suggest promising solutions and new directions for visibility enhancement, color correction, and object detection on real-world underwater images. The benchmark is available at: https://github.com/dlut-dimt/Realworld-Underwater-Image-Enhancement-RUIE-Benchmark.
Risheng Liu, Xin Fan 0001, Ming Zhu 0001, Minjun Hou, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.5
2020 Investigating Task-Driven Latent Feasibility for Nonconvex Image Modeling
abstract
Properly modeling latent image distributions plays an important role in a variety of image-related vision problems. Most exiting approaches aim to formulate this problem as optimization models (e.g., Maximum A Posterior, MAP) with handcrafted priors. In recent years, different CNN modules are also considered as deep priors to regularize the image modeling process. However, these explicit regularization techniques require deep understandings on the problem and elaborately mathematical skills. In this work, we provide a new perspective, named Task-driven Latent Feasibility (TLF), to incorporate specific task information to narrow down the solution space for the optimization-based image modeling problem. Thanks to the flexibility of TLF, both designed and trained constraints can be embedded into the optimization process. By introducing control mechanisms based on the monotonicity and boundedness conditions, we can also strictly prove the convergence of our proposed inference process. We demonstrate that different types of image modeling problems, such as image deblurring and rain streaks removals, can all be appropriately addressed within our TLF framework. Extensive experiments also verify the theoretical results and show the advantages of our method against existing state-of-the-art approaches.
Risheng Liu, Pan Mu, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.5
2020 A Deep Framework Assembling Principled Modules for CS-MRI: Unrolling Perspective, Convergence Behaviors, and Practical Modeling
abstract
Compressed Sensing Magnetic Resonance Imaging (CS-MRI) significantly accelerates MR acquisition at a sampling rate much lower than the Nyquist criterion. A major challenge for CS-MRI lies in solving the severely ill-posed inverse problem to reconstruct aliasing-free MR images from the sparse k -space data. Conventional methods typically optimize an energy function, producing restoration of high quality, but their iterative numerical solvers unavoidably bring extremely large time consumption. Recent deep techniques provide fast restoration by either learning direct prediction to final reconstruction or plugging learned modules into the energy optimizer. Nevertheless, these data-driven predictors cannot guarantee the reconstruction following principled constraints underlying the domain knowledge so that the reliability of their reconstruction process is questionable. In this paper, we propose a deep framework assembling principled modules for CS-MRI that fuses learning strategy with the iterative solver of a conventional reconstruction energy. This framework embeds an optimal condition checking mechanism, fostering efficient and reliable reconstruction. We also apply the framework to three practical tasks, i.e., complex-valued data reconstruction, parallel imaging and reconstruction with Rician noise. Extensive experiments on both benchmark and manufacturer-testing images demonstrate that the proposed method reliably converges to the optimal solution more efficiently and accurately than the state-of-the-art in various scenarios.
Risheng Liu, Yuxi Zhang 0001, Shichao Cheng, Zhongxuan Luo, Xin Fan 0001
IEEE Trans. Medical Imaging4
2020 Knowledge-Driven Deep Unrolling for Robust Image Layer Separation
abstract
Single-image layer separation targets to decompose the observed image into two independent components in terms of different application demands. It is known that many vision and multimedia applications can be (re)formulated as a separation problem. Due to the fundamentally ill-posed natural of these separations, existing methods are inclined to investigate model priors on the separated components elaborately. Nevertheless, it is knotty to optimize the cost function with complicated model regularizations. Effectiveness is greatly conceded by the settled iteration mechanism, and the adaption cannot be guaranteed due to the poor data fitting. What is more, for a universal framework, the most taxing point is that one type of visual cue cannot be shared with different tasks. To partly overcome the weaknesses mentioned earlier, we delve into a generic optimization unrolling technique to incorporate deep architectures into iterations for adaptive image layer separation. First, we propose a general energy model with implicit priors, which is based on maximum a posterior, and employ the extensively accepted alternating direction method of multiplier to determine our elementary iteration mechanism. By unrolling with one general residual architecture prior and one task-specific prior, we attain a straightforward, flexible, and data-dependent image separation framework successfully. We apply our method to four different tasks, including single-image-rain streak removal, high-dynamic-range tone mapping, low-light image enhancement, and single-image reflection removal. Extensive experiments demonstrate that the proposed method is applicable to multiple tasks and outperforms the state of the arts by a large margin qualitatively and quantitatively.
Risheng Liu, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.4
2019 A Theoretically Guaranteed Deep Optimization Framework for Robust Compressive Sensing MRI
abstract
Magnetic Resonance Imaging (MRI) is one of the most dynamic and safe imaging techniques available for clinical applications. However, the rather slow speed of MRI acquisitions limits the patient throughput and potential indications. Compressive Sensing (CS) has proven to be an efficient technique for accelerating MRI acquisition. The most widely used CS-MRI model, founded on the premise of reconstructing an image from an incompletely filled k-space, leads to an ill-posed inverse problem. In the past years, lots of efforts have been made to efficiently optimize the CS-MRI model. Inspired by deep learning techniques, some preliminary works have tried to incorporate deep architectures into CS-MRI process. Unfortunately, the convergence issues (due to the experience-based networks) and the robustness (i.e., lack real-world noise modeling) of these deeply trained optimization methods are still missing. In this work, we develop a new paradigm to integrate designed numerical solvers and the data-driven architectures for CS-MRI. By introducing an optimal condition checking mechanism, we can successfully prove the convergence of our established deep CS-MRI optimization scheme. Furthermore, we explicitly formulate the Rician noise distributions within our framework and obtain an extended CS-MRI network to handle the real-world nosies in the MRI process. Extensive experimental results verify that the proposed paradigm outperforms the existing state-of-theart techniques both in reconstruction accuracy and efficiency as well as robustness to noises in real scene.
Risheng Liu, Yuxi Zhang 0001, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
AAAI5
2019 Compounded Layer-Prior Unrolling: A Unified Transmission-Based Image Enhancement Framework
abstract
Improving the quality of images degraded by various transmission media has important practical significance. Such enhancement tasks involve resolving both transmission degradation and residual contamination including imaging noise, color distortion, and occlusions. Existing methods typically develop the priors on natural scenes to resolve ill-posed problems separately. However, the solutions derived from hand-crafted priors may fail on specific regions where a priori assumptions break, and recent data-driven methods highly depend on training data owing to the absence of effective priors. Based on a unified formulation for transmission-based image enhancement tasks, we develop a compounded unrolling framework to generate hybrid image layer propagations. Specifically, as multiple deeply-trained priors are integrated into the iterative propagation scheme, the deep model can recognize specific task properties and data distributions for different applications. Both quantitative and qualitative experiments demonstrate the superior performance of the proposed framework on various transmission-based tasks (haze removal, underwater image enhancement and rain removal).
Risheng Liu, Minjun Hou, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo
ICME5
2019 Enhanced Residual Dense Intrinsic Network for Intrinsic Image Decomposition
abstract
Intrinsic image decomposition is a challenging task, which aims at recovering intrinsic components from the observation. Hand-crafted priors have been widely used in traditional methods, yet with unsatisfactory performance of quality and runtime. Recently, network-based approaches have been greatly developed, but the physical imaging principle is ignored causing the multiplication of estimated components is hard to reconstruct the observation. To overcome these limitations, we develop an enhanced residual dense intrinsic network (ERDIN) for intrinsic decomposition. Specifically, we construct the basic module (i.e., enhanced residual dense block (ERDB)) to fully exploit the hierarchical features. The physical imaging principle is designed as the reconstruction loss to ensure the consistency between the observation and the multiplication of estimated components, which is of equal importance with the data loss. Extensive experimental results illustrate our excellent performance compared with other state-of-the-art methods.
Risheng Liu, Long Ma 0002, Miao Zhang 0004, Xin Fan 0001, Zhongxuan Luo
ICME6
2019 Continuous Scale Adaption for Efficient Box-Based Scene Text Detection
abstract
Due to the diversity of text size in scene images, the current box-based methods employ a large amount of fixed-size anchors with different scales to match texts, thus leading to high computational cost. In this paper, we propose to learn the scales of texts and adjust the sizes of anchors accordingly, which can largely reduce the numbers of anchors and therefore significantly reduces the time cost. Moreover, compared to discrete scales used in previous methods, the learned scales are continuous and more reliable. Additionally, we propose Anchor convolution to exploit scaled feature for each anchor by dynamically adjusting the sizes of receptive fields according to the learned scales. Experimental results show that the proposed method significantly improves the computational efficiency of box-based framework(reduce the running time from 0.73s to 0.28s) and enhances its robustness against small texts, while achieving competitive performance with other methods.
Bingwang Zhang, Zhihui Wang 0001, Zhongxuan Luo
ICME5
2019 Learning diffusion on global graph: A PDE-directed approach for feature detection on geometric shapes
Nannan Li 0002, Shengfa Wang, Risheng Liu, Ziqiao Guan, Zhixun Su, Zhongxuan Luo, Hong Qin 0001
Comput. Aided Geom. Des.6
2019 Learning Bilevel Layer Priors for Single Image Rain Streaks Removal
abstract
Rain streaks removal is an important issue of the outdoor vision system and recently has been investigated extensively. In the past decades, maximum a posterior and network-based architecture have been attracting considerable attention for this problem. However, it is challenging to establish effective regularization priors and the cost function with complex prior is hard to optimize. On the other hand, it is still hard to incorporate data-dependent information into conventional numerical iterations. To partially address the above limits and inspired by the leader-follower gaming perspective, we introduce an unrolling strategy to incorporate data-dependent network architectures into the established iterations, i.e., a learning bilevel layer priors method to jointly investigate the learnable feasibility and optimality of rain streaks removal problem. Both visual and quantitative comparison results demonstrate that our method outperforms the state of the art.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
IEEE Signal Process. Lett.5
2019 Deep Proximal Unrolling: Algorithmic Framework, Convergence Analysis and Applications
abstract
Deep learning models have gained great success in many real-world applications. However, most existing networks are typically designed in heuristic manners, thus these approaches lack of rigorous mathematical derivations and clear interpretations. Several recent studies try to build deep models by unrolling a particular optimization model that involves task information. Unfortunately, due to the dynamic nature of network parameters, their resultant deep propagations do not possess the nice convergence property as the original optimization scheme does. In this work, we develop a generic paradigm to unroll nonconvex optimization for deep model design. Different from most existing frameworks, which just replace the iterations by network architectures, we prove in theory that the propagation generated by our proximally unrolled deep model can globally converge to the critical-point of the original optimization model. Moreover, even if the task information is only partially available (e.g., no prior regularization), we can still train a convergent deep propagations. We also extend these theoretical investigations on the more general multi-block models and thus a lot of real-world applications can be successfully handled by the proposed framework. Finally, we conduct experiments on various low-level vision tasks (i.e., non-blind deconvolution, dehazing, and low-light image enhancement) and demonstrate the superiority of our proposed framework, compared with existing state-of-the-art approaches.
Risheng Liu, Shichao Cheng, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.5
2019 Learning Aggregated Transmission Propagation Networks for Haze Removal and Beyond
abstract
Single-image dehazing is an important low-level vision task with many applications. Early studies have investigated different kinds of visual priors to address this problem. However, they may fail when their assumptions are not valid on specific images. Recent deep networks also achieve a relatively good performance in this task. But unfortunately, due to the disappreciation of rich physical rules in hazes, a large amount of data are required for their training. More importantly, they may still fail when there exist completely different haze distributions in testing images. By considering the collaborations of these two perspectives, this paper designs a novel residual architecture to aggregate both prior (i.e., domain knowledge) and data (i.e., haze distribution) information to propagate transmissions for scene radiance estimation. We further present a variational energy-based perspective to investigate the intrinsic propagation behavior of our aggregated deep model. In this way, we actually bridge the gap between prior-driven models and data-driven networks and leverage advantages but avoid limitations of previous dehazing approaches. A lightweight learning framework is proposed to train our propagation network. Finally, by introducing a task-aware image separation formulation with a flexible optimization scheme, we extend the proposed model for more challenging vision tasks, such as underwater image enhancement and single-image rain removal. Experiments on both synthetic and real-world images demonstrate the effectiveness and efficiency of the proposed framework.
Risheng Liu, Xin Fan 0001, Minjun Hou, Zhiying Jiang, Zhongxuan Luo, Lei Zhang 0006
IEEE Trans. Neural Networks Learn. Syst.5
2019 Recommending New Features from Mobile App Descriptions
abstract
The rapidly evolving mobile applications (apps) have brought great demand for developers to identify new features by inspecting the descriptions of similar apps and acquire missing features for their apps. Unfortunately, due to the huge number of apps, this manual process is time-consuming and unscalable. To help developers identify new features, we propose a new approach named SAFER. In this study, we first develop a tool to automatically extract features from app descriptions. Then, given an app, we leverage the topic model to identify its similar apps based on the extracted features and API names of apps. Finally, we design a feature recommendation algorithm to aggregate and recommend the features of identified similar apps to the specified app. Evaluated over a collection of 533 annotated features from 100 apps, SAFER achieves a Hit@15 score of up to 78.68% and outperforms the baseline approach KNN+ by 17.23% on average. In addition, we also compare SAFER against a typical technique of recommending features from user reviews, i.e., CLAP. Experimental results reveal that SAFER is superior to CLAP by 23.54% in terms of Hit@15.
He Jiang 0001, Zhilei Ren, David Lo 0001, Xindong Wu 0001, Zhongxuan Luo
ACM Trans. Softw. Eng. Methodol.7
2019 Fast example searching for input-adaptive data-driven dehazing with Gaussian process regression
Xin Fan 0001, Xianxuan Tang, Minjun Hou, Zhongxuan Luo
Vis. Comput.4
2019 A lightweight methodology of 3D printed objects utilizing multi-scale porous structures
Jiangbei Hu, Shengfa Wang, Yi Wang 0037, Fengqi Li, Zhongxuan Luo
Vis. Comput.5
2018 Self-Reinforced Cascaded Regression for Face Alignment
Xin Fan 0001, Risheng Liu, Kang Huyan, Yuyao Feng, Zhongxuan Luo
AAAI5
2018 Proximal Alternating Direction Network: A Globally Converged Deep Unrolling Framework
abstract
Deep learning models have gained great success in many real-world applications. However, most existing networks are typically designed in heuristic manners, thus lack of rigorous mathematical principles and derivations. Several recent studies build deep structures by unrolling a particular optimization model that involves task information. Unfortunately, due to the dynamic nature of network parameters, their resultant deep propagation networks do not possess the nice convergence property as the original optimization scheme does. This paper provides a novel proximal unrolling framework to establish deep models by integrating experimentally verified network architectures and rich cues of the tasks. More importantly,we prove in theory that 1) the propagation generated by our unrolled deep model globally converges to a critical-point of a given variational energy, and 2) the proposed framework is still able to learn priors from training data to generate a convergent propagation even when task information is only partially available. Indeed, these theoretical results are the best we can ask for, unless stronger assumptions are enforced. Extensive experiments on various real-world applications verify the theoretical convergence and demonstrate the effectiveness of designed deep models.
Risheng Liu, Xin Fan 0001, Shichao Cheng, Zhongxuan Luo
AAAI5
2018 Deep Layer Prior Optimization for Single Image Rain Streaks Removal
abstract
Visible distortions caused by rain streaks have significant negative effects on the performance of many vision and learning algorithms. Most of the existing deraining approaches propose to build complex prior models to formulate the appearance of rain streaks. Unfortunately, these human-designed priors tend to over-smooth the background and leave too many rain streaks since the distribution of rain streaks is complex and disordered. In this work, we exploit a deep layer prior under the maximum a posterior framework to recover the intrinsic rain structure. The optimization of the resulted variational energy can be understood as simultaneously performing rain and image propagations based on data-dependent residual networks and task cues (e.g., total variation regularization), respectively. Experimental results on both synthetic and real test images demonstrate the effectiveness of our approach against both designed priors and fully data-dependent convolutional neural networks.
Risheng Liu, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP6
2018 Robust Haze Removal Via Joint Deep Transmission and Scene Propagation
abstract
Haze is one of the most important factors which reduce the outdoor image quality. Existing approaches often aim to design their models based on principles of hazes. However, even with exactly modeled haze distribution, it is still a challenging task due to factors in real scenario, such as noises, halos and artifacts. To address limitations of existing approaches for real-world hazy removal problem, this paper proposes a novel framework to incorporate deep residual architectures into a propagation scheme to jointly estimate transmission and clean scene. We evaluate the proposed framework on both widely used benchmarks and real-world low-quality hazy images. Extensive experimental results demonstrate that our method performs favorably against approaches designed only based on haze cues and achieves the state-of-the-art results, compared with both conventional shallow models and deep dehzaing networks.
Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP6
2018 Joint Residual Learning for Underwater Image Enhancement
abstract
Improving the quality of underwater image has a significant impact on many signal processing and computer vision applications, while haze-effect and color shift are main handicaps need to be surmounted. Due to the complexity of the underwater environmental factors, most existing image enhancement techniques cannot be directly applied to address this task. In this work, we develop a novel framework to jointly performing residual learning on transmission and image domains for underwater scene entrenchment. Indeed, our deep model consists of a data-driven residual architecture for transmission estimation and a knowledge-driven scene residual formulation for underwater illumination balance. Therefore, we can aggregate the prior knowledge and data information to investigate the underlying underwater image distribution. Moreover, by introducing adaptive exposure map, image colors will also be corrected accordingly. Experimentally, both quantitative and qualitative analysis can indicate outstanding effectiveness of the proposed algorithm, against state-of-the-art approaches.
Minjun Hou, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICIP4
2018 Single Image Layer Separation via Deep Admm Unrolling
abstract
Single image layer separation aims to divide the observed image into two independent components according to special task requirements and has been widely used in many vision and multimedia applications. Because this task is fundamentally ill-posed, most existing approaches tend to design complex priors on the separated layers. However, the cost function with complex prior regularization is hard to optimize. The performance is also compromised by fixed iteration schemes and less data fitting ability. More importantly, it is also challenging to design a unified framework to separate image layers for different applications. To partially mitigate the above limitations, we develop a flexible optimization unrolling technique to incorporate deep architectures into iterations for adaptive image layer separation. Specifically, we first design a general energy model with implicit priors and adopt the widely used alternating direction method of multiplier (ADMM) to establish our basic iteration scheme. By unrolling with residual convolution architectures, we successfully obtain a simple, flexible, and data-dependent image separation method. Extensive experiments on the tasks of rain streak removal and reflection removal validate the effectiveness of our approach.
Risheng Liu, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo
ICME5
2018 Toward Designing Convergent Deep Operator Splitting Methods for Task-specific Nonconvex Optimization
abstract
Operator splitting methods have been successfully used in computational sciences, statistics, learning and vision areas to reduce complex problems into a series of simpler subproblems. However, prevalent splitting schemes are mostly established only based on the mathematical properties of some general optimization models. So it is a laborious process and often requires many iterations of ideation and validation to obtain practical and task-specific optimal solutions, especially for nonconvex problems in real-world scenarios. To break through the above limits, we introduce a new algorithmic framework, called Learnable Bregman Splitting (LBS), to perform deep-architecture-based operator splitting for nonconvex optimization based on specific task model. Thanks to the data-dependent (i.e., learnable) nature, our LBS can not only speed up the convergence, but also avoid unwanted trivial solutions for real-world tasks. Though with inexact deep iterations, we can still establish the global convergence and estimate the asymptotic convergence rate of LBS only by enforcing some fairly loose assumptions. Extensive experiments on different applications (e.g., image completion and deblurring) verify our theoretical results and show the superiority of LBS against existing methods.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
IJCAI5
2018 Fast Factorization-free Kernel Learning for Unlabeled Chunk Data Streams
abstract
Data stream analysis aims at extracting discriminative information for classification from continuously incoming samples. It is extremely challenging to detect novel data while updating the model in an efficient and stable fashion, especially for the chunk data. This paper proposes a fast factorization-free kernel learning method to unify novelty detection and incremental learning for unlabeled chunk data streams in one framework. The proposed method constructs a joint reproducing kernel Hilbert space from known class centers by solving a linear system in kernel space. Naturally, unlabeled data can be detected and classified among multi-classes by a single decision model. And projecting samples into the discriminative feature space turns out to be the product of two small-sized kernel matrices without needing such time-consuming factorization like QR-decomposition or singular value decomposition. Moreover, the insertion of a novel class can be treated as the addition of a new orthogonal basis to the existing feature space, resulting in fast and stable updating schemes. Both theoretical analysis and experimental validation on real-world datasets demonstrate that the proposed methods learn chunk data streams with significantly lower computational costs and comparable or superior accuracy than the state of the art.
Yi Wang 0037, Nan Xue 0004, Xin Fan 0001, Jiebo Luo 0001, Risheng Liu, Zhongxuan Luo
IJCAI8
2018 User-Guided Deep Anime Line Art Colorization with Conditional Adversarial Networks
abstract
Scribble colors based line art colorization is a challenging computer vision problem since neither greyscale values nor semantic information is presented in line arts, and the lack of authentic illustration-line art training pairs also increases difficulty of model generalization. Recently, several Generative Adversarial Nets (GANs) based methods have achieved great success. They can generate colorized illustrations conditioned on given line art and color hints. However, these methods fail to capture the authentic illustration distributions and are hence perceptually unsatisfying in the sense that they often lack accurate shading. To address these challenges, we propose a novel deep conditional adversarial architecture for scribble based anime line art colorization. Specifically, we integrate the conditional framework with WGAN-GP criteria as well as the perceptual loss to enable us to robustly train a deep network that makes the synthesized images more natural and real. We also introduce a local features network that is independent of synthetic data. With GANs conditioned on features from such network, we notably increase the generalization capability over "in the wild" line arts. Furthermore, we collect two datasets that provide high-quality colorful illustrations and authentic line arts for training and benchmarking. With the proposed model trained on our illustration dataset, we demonstrate that images synthesized by the presented approach are considerably more realistic and precise than alternative approaches.
Yuanzheng Ci, Xinzhu Ma, Zhihui Wang 0001, Zhongxuan Luo
ACM Multimedia5
2018 Learning Collaborative Generation Correction Modules for Blind Image Deblurring and Beyond
abstract
Blind image deblurring plays a very important role in many vision and multimedia applications. Most existing works tend to introduce complex priors to estimate the sharp image structures for blur kernel estimation. However, it has been verified that directly optimizing these models is challenging and easy to fall into degenerate solutions. Although several experience-based heuristic inference strategies, including trained networks and designed iterations, have been developed, it is still hard to obtain theoretically guaranteed accurate solutions. In this work, a collaborative learning framework is established to address the above issues. Specifically, we first design two modules, named Generator and Corrector, to extract the intrinsic image structures from the data-driven and knowledge-based perspectives, respectively. By introducing a collaborative methodology to cascade these modules, we can strictly prove the convergence of our image propagations to a deblurring-related optimal solution. As a nontrivial byproduct, we also apply the proposed method to address other related tasks, such as image interpolation and edge-preserved smoothing. Plenty of experiments demonstrate that our method can outperform the state-of-the-art approaches on both synthetic and real datasets.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
ACM Multimedia5
2018 A Bridging Framework for Model Optimization and Deep Propagation
abstract
Optimizing task-related mathematical model is one of the most fundamental methodologies in statistic and learning areas. However, generally designed schematic iterations may hard to investigate complex data distributions in real-world applications. Recently, training deep propagations (i.e., networks) has gained promising performance in some particular tasks. Unfortunately, existing networks are often built in heuristic manners, thus lack of principled interpretations and solid theoretical supports. In this work, we provide a new paradigm, named Propagation and Optimization based Deep Model (PODM), to bridge the gaps between these different mechanisms (i.e., model optimization and deep propagation). On the one hand, we utilize PODM as a deeply trained solver for model optimization. Different from these existing network based iterations, which often lack theoretical investigations, we provide strict convergence analysis for PODM in the challenging nonconvex and nonsmooth scenarios. On the other hand, by relaxing the model constraints and performing end-to-end training, we also develop a PODM based strategy to integrate domain knowledge (formulated as models) and real data distributions (learned by networks), resulting in a generic ensemble framework for challenging real-world applications. Extensive experiments verify our theoretical results and demonstrate the superiority of PODM against these state-of-the-art approaches.
Risheng Liu, Shichao Cheng, Xiaokun Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
NeurIPS6
2018 Disparity-Based Robust Unstructured Terrain Segmentation
Xinzhu Ma, Zhihui Wang 0001, Zhongxuan Luo
PRCV (4)5
2018 Designing a stable feedback control system for blind image deconvolution
Shichao Cheng, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
Neural Networks4
2018 Line matching based on line-points invariant and local homography
Qi Jia 0001, Xin Fan 0001, Xinkai Gao, Meiyu Yu, Zhongxuan Luo
Pattern Recognit.6
2018 Explicit Shape Regression With Characteristic Number for Facial Landmark Localization
abstract
Robustly localizing facial landmarks plays a very important role in many multimedia and vision applications. Most recently proposed regression-based methods prevailing in the community lack explicit shape constraints for faces and require a large number of facial images to cover great appearance variations. To address these limitations, this paper introduces a novel projective invariant called characteristic number (CN) to explicitly characterize the intrinsic geometries of facial points shared by human faces. It can be verified that the shape priors from CN are inherently invariant to pose changes. By further developing a shape-to-gradient regression framework, we provide a robust and efficient landmark detector for facial images in the wild. The computation of our model can be successfully addressed by learning the descent directions using point-CN pairs without the need for large collections for appearance training. As a nontrivial byproduct, this paper also builds a face dataset, where each face has 15 well-defined viewpoints (poses) to quantitatively analyze the effects of different poses on localization methods. Extensive experiments on challenging benchmarks and our newly built dataset demonstrate the effectiveness of our proposed detector against other state-of-the-art approaches.
Xin Fan 0001, Risheng Liu, Zhongxuan Luo, Yuyao Feng
IEEE Trans. Multim.3
2017 Fast Online Incremental Learning on Mixture Streaming Data
abstract
The explosion of streaming data poses challenges to feature learning methods including linear discriminant analysis (LDA). Many existing LDA algorithms are not efficient enough to incrementally update with samples that sequentially arrive in various manners. First, we propose a new fast batch LDA (FLDA/QR) learning algorithm that uses the cluster centers to solve a lower triangular system that is optimized by the Cholesky-factorization. To take advantage of the intrinsically incremental mechanism of the matrix, we further develop an exact incremental algorithm (IFLDA/QR). The Gram-Schmidt process with reorthogonalization in IFLDA/QR significantly saves the space and time expenses compared with the rank-one QR-updating of most existing methods. IFLDA/QR is able to handle streaming data containing 1) new labeled samples in the existing classes, 2) samples of an entirely new (novel) class, and more significantly, 3) a chunk of examples mixed with those in 1) and 2). Both theoretical analysis and numerical experiments have demonstrated much lower space and time costs (2~10 times faster) than the state of the art, with comparable classification accuracy.
Yi Wang 0037, Xin Fan 0001, Zhongxuan Luo, Tianzhu Wang, Maomao Min, Jiebo Luo 0001
AAAI3
2017 Leveraging geometric correlation for input-adaptive facial landmark regression
abstract
Facial analysis plays very important role in many vision applications, such as authentication and entertainments. The very early works in the 1990s mostly focus on estimating geometric deformations of facial landmarks to address this task. While in the past several years, more and more efforts have been made to directly learn an appearance regression for facial analysis. Though training regressions on controlled facial images can successfully capture the appearance variations, the performance of these appearance-based models are tightly related to the quantity and quality of the training data. In this paper, we develop a novel framework, named geometric correlated landmark regression (GCLR), to inherit the advantages but overcome limitations of these two categories of methods. Specifically, we first establish a landmark-to-landmark regression to estimate the geometry of facial images. By further incorporating a sparse coding term into the regression framework, we can successfully leverage the geometric correlations between the test image and the shape dictionary, thus significantly enhance the geometry regression performance. Experimental results on various challenging facial data sets verify the effectiveness and efficiency of GCLR.
Yuyao Feng, Risheng Liu, Xin Fan 0001, Kang Huyan, Zhongxuan Luo
ICME5
2017 Blind image deblurring via adaptive dynamical system learning
abstract
Blind image deblurring is one of the main phases in most media analysis tasks. Many existing works aim to simultaneously estimate the latent image and the blur kernel under a MAP framework. However, it has been demonstrated that such joint estimation strategies may lead to the undesired trivial solution. In this paper, we propose a learnable nonlinear dynamical system to formulate the image propagation so that the blur kernel estimation can be efficiently controlled by both cues and training data. Our analysis also indicates that the proposed dynamical system is feasible on image modeling socialities. Experimental results on different benchmark image sets evaluate the effectiveness of our proposed approach.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
ICME4
2017 Deep hybrid residual learning with statistic priors for single image super-resolution
abstract
This paper considers single image super-resolution (SISR), which is an important low-level vision task and has various applications in multimedia society. Recently, deep neural networks have archived good performance on this field. But most of existing deep models are based on the fully data-dependent network architecture, thus missing majority of domain-knowledge of the super-resolution task. To address this limitation, we develop a new hybrid residual learning approach to leverage priors of SISR within the maximum a posteriori framework for network architecture design. We demonstrate that it can incorporate both image priors and data fidelity into the network, leading to a novel cascaded residual learning system for SISR process. Extensive experimental results on real-world images show that the proposed algorithm performs favorably against state-of-the-art methods.
Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME5
2017 Compact CNN Based Video Representation for Efficient Video Copy Detection
Xin Fan 0001, Zhongxuan Luo
MMM (1)5
2017 Feature based problem hardness understanding for requirements engineering
Zhilei Ren, He Jiang 0001, Jifeng Xuan, Shuwei Zhang, Zhongxuan Luo
Sci. China Inf. Sci.5
2017 Adaptive low-rank subspace learning with online optimization for robust visual tracking
Risheng Liu, Di Wang 0018, Yuzhuo Han, Xin Fan 0001, Zhongxuan Luo
Neural Networks5
2017 Two-Layer Gaussian Process Regression With Example Selection for Image Dehazing
abstract
Researchers have devoted great efforts to image dehazing with prior assumptions in the past decade. Recently developed example-based approaches typically lack elegant models for the hazy process and meanwhile demand synthetic hazy images by manual selection. The priors from observations, and those trained from synthetic images cannot always reflect true structural information of natural images in practice. In this paper, we present a learning model for haze removal by using two-layer Gaussian process regression (GPR). By using training examples, the two-layer GPR establishes a direct relationship from the input image to the depth-dependent transmission, and learns local image priors to further improve the estimation. We also provide a systematic scheme to automatically collect suitable training pairs, which works for both simulated examples and images of natural scenes. Both qualitative and quantitative comparisons on real-world and synthetic hazy images demonstrate the effectiveness of the proposed approach, especially for white or bright objects and heavy haze regions in which traditional methods may fail.
Xin Fan 0001, Yi Wang 0037, Xianxuan Tang, Renjie Gao, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.5
2017 A Fast Ellipse Detector Using Projective Invariant Pruning
abstract
Detecting elliptical objects from an image is a central task in robot navigation and industrial diagnosis, where the detection time is always a critical issue. Existing methods are hardly applicable to these real-time scenarios of limited hardware resource due to the huge number of fragment candidates (edges or arcs) for fitting ellipse equations. In this paper, we present a fast algorithm detecting ellipses with high accuracy. The algorithm leverages a newly developed projective invariant to significantly prune the undesired candidates and to pick out elliptical ones. The invariant is able to reflect the intrinsic geometry of a planar curve, giving the value of -1 on any three collinear points and +1 for any six points on an ellipse. Thus, we apply the pruning and picking by simply comparing these binary values. Moreover, the calculation of the invariant only involves the determinant of a 3×3 matrix. Extensive experiments on three challenging data sets with 648 images demonstrate that our detector runs 20%-50% faster than the state-of-the-art algorithms with the comparable or higher precision.
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Lianbo Song, Tie Qiu 0001
IEEE Trans. Image Process.3
2016 Novel Coplanar Line-Points Invariants for Robust Line Matching Across Views
Qi Jia 0001, Xinkai Gao, Xin Fan 0001, Zhongxuan Luo, Ziyao Chen
ECCV (8)4
2016 Discriminative Feature Learning with an Optimal Pattern Model for Image Classification
Xin Fan 0001, Zhongxuan Luo
MMM (1)5
2016 1D Barcode Region Detection Based on the Hough Transform and Support Vector Machine
Zhihui Wang 0001, Ai Chen, Jianjun Li 0007, Zhongxuan Luo
MMM (2)5
2016 An efficient mesh-based face beautifier on mobile devices
Xin Fan 0001, Yuyao Feng, Yi Wang 0037, Shengfa Wang, Zhongxuan Luo
Neurocomputing6
2016 ARAP++: an extension of the local/global approach to mesh parameterization
abstract
Mesh parameterization is one of the fundamental operations in computer graphics (CG) and computeraided design (CAD). In this paper, we propose a novel local/global parameterization approach, ARAP++, for singleand multi-boundary triangular meshes. It is an extension of the as-rigid-as-possible (ARAP) approach, which stitches together 1-ring patches instead of individual triangles. To optimize the spring energy, we introduce a linear iterative scheme which employs convex combination weights and a fitting Jacobian matrix corresponding to a prescribed family of transformations. Our algorithm is simple, efficient, and robust. The geometric properties (angle and area) of the original model can also be preserved by appropriately prescribing the singular values of the fitting matrix. To reduce the area and stretch distortions for high-curvature models, a stretch operator is introduced. Numerical results demonstrate that ARAP++ outperforms several state-of-the-art methods in terms of controlling the distortions of angle, area, and stretch. Furthermore, it achieves a better visualization performance for several applications, such as texture mapping and surface remeshing.
Zhongxuan Luo, Jielin Zhang, Emil Saucan
Frontiers Inf. Technol. Electron. Eng.2
2016 On the tag localization of web video
Bin Liu 0040, Lei Yi, Yue Guan 0002, Zhongxuan Luo
Multim. Syst.5
2016 Cross-view action matching using a novel projective invariant on non-coplanar space-time points
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Kang Huyan, Zezhou Li
Multim. Tools Appl.3
2016 Learning to Diffuse: A New Perspective to Design PDEs for Visual Analysis
abstract
Partial differential equations (PDEs) have been used to formulate image processing for several decades. Generally, a PDE system consists of two components: the governing equation and the boundary condition. In most previous work, both of them are generally designed by people using mathematical skills. However, in real world visual analysis tasks, such predefined and fixed-form PDEs may not be able to describe the complex structure of the visual data. More importantly, it is hard to incorporate the labeling information and the discriminative distribution priors into these PDEs. To address above issues, we propose a new PDE framework, named learning to diffuse (LTD), to adaptively design the governing equation and the boundary condition of a diffusion PDE system for various vision tasks on different types of visual data. To our best knowledge, the problems considered in this paper (i.e., saliency detection and object tracking) have never been addressed by PDE models before. Experimental results on various challenging benchmark databases show the superiority of LTD against existing state-of-the-art methods for all the tested visual analysis tasks.
Risheng Liu, Guangyu Zhong, Junjie Cao 0001, Zhouchen Lin, Shiguang Shan, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.6
2016 Hierarchical projective invariant contexts for shape recognition
Qi Jia 0001, Xin Fan 0001, Yu Liu 0012, Zhongxuan Luo, He Guo 0001
Pattern Recognit.5
2016 3D facial landmark localization using texture regression via conformal mapping
Xin Fan 0001, Qi Jia 0001, Kang Huyan, Xianfeng Gu, Zhongxuan Luo
Pattern Recognit. Lett.5
2016 Image morphing with conformal welding
Xin Fan 0001, Yuyao Feng, Xianfeng Gu, Zhongxuan Luo
Vis. Comput.5
2016 Haze editing with natural transmission
Xin Fan 0001, Yi Wang 0037, Renjie Gao, Zhongxuan Luo
Vis. Comput.4
2016 Generalized rational Bézier curves for the rigid body motion design
Zhongxuan Luo, Xin Fan 0001, Yaqi Gao, Panpan Shui
Vis. Comput.1
2015 Sparse concept discriminant matrix factorization for image representation
abstract
Over the past few decades, matrix factorization has attracted considerable attention for image representation. It is desired for a matrix factorization technique to find the basis that is able to capture highly discriminant information as well as to preserve the intrinsic manifold structure. Besides, the basis has to generate a sparse representation for a given image. In this paper, we propose a matrix factorization method called Sparse concept Discriminant Matrix Factorization (SDMF) by combining a novel fisher-like criterion with the sparse coding. The criterion is discriminant enough across different feature spaces, and meanwhile maintains locally neighboring structures. The proposed method is general for both cases with and without class labels, hence yielding supervised and un-supervised SDMFs. Experimental results show that SDMF provides better representation with higher performance on two tasks (image recognition and clustering) compared with the existing matrix factorization methods.
Chuang Lin 0001, Risheng Liu, Xin Fan 0001, Jifeng Jiang, Zhongxuan Luo
ICIP6
2015 Characteristic number regression for facial feature extraction
abstract
Facial feature extraction plays an important role in many multimedia and vision applications. Recent regression methods for extraction lack the explicit shape constraints for faces, and require a large number of facial images covering great appearance variations. This paper introduces a novel projective invariant, named characteristic number (CN), to explicitly characterize the intrinsic geometries of facial points shared by human faces, which is inherently invariant to pose changes. By further developing a shape-to-gradient regression framework, we provide a robust and efficient feature extractor for facial images in the wild. The computation of our model can be successfully addressed by learning the descent directions using point-CN pairs without the need of large collections for appearance training. Extensive experiments on challenging benchmark data sets demonstrate the effectiveness of our proposed detector against other state-of-the-art approaches.
Xin Fan 0001, Risheng Liu, Yuyao Feng, Zhongxuan Luo, Zezhou Li
ICME5
2015 Multi-scale mesh saliency based on low-rank and sparse analysis in shape feature space
Shengfa Wang, Nannan Li 0002, Shuai Li 0001, Zhongxuan Luo, Zhixun Su, Hong Qin 0001
Comput. Aided Geom. Des.4
2015 Community-Based Event Dissemination with Optimal Load Balancing
abstract
Distributed publish/subscribe systems are poised with challenges of performance degradation and poor scalability. This is typically caused by an uneven load distribution of real-world applications and the susceptibility of link failure in networks. Partitioning and replication techniques have been implemented by exploring community-based load balancing to cope with such issues. The novel approach herein exploits offloading at the inter-community level as well as filter replication at the intra-community level. This results in the dynamic distribution and forwarding of publication and subscription services among brokers during run time. The proposed method, Co-Lab (COmmunity-based LoAd Balancing), seeks to improve the network performance by clustering brokers in a community by taking into consideration interest similarity and filter replication. It attempts to effectively achieve a more consistent and uniform load distribution among brokers and to circumvent the occurrence of highly overloaded brokers. Performance evaluations indicate that Co-Lab has promising advantages by achieving relatively better load balance, reduced overall load and robustness against failures.
Feng Xia 0001, Ahmedin Mohammed Ahmed, Laurence T. Yang, Zhongxuan Luo
IEEE Trans. Computers4
2015 A unified approach to computing the nearest complex polynomial with a given zero
Xingjun Luo, Zhongxuan Luo
Theor. Comput. Sci.3
2015 Fiducial Facial Point Extraction Using a Novel Projective Invariant
abstract
Automatic extraction of fiducial facial points is one of the key steps to face tracking, recognition, and animation.Great facial variations, especially pose or viewpoint changes,typically degrade the performance of classical methods. Recent learning or regression-based approaches highly rely on the availability of a training set that covers facial variations as wide as possible. In this paper, we introduce and extend a novel projective invariant, named the characteristic number (CN), which unifies the collinearity, cross ratio, and geometrical characteristics given by more (6) points. We derive strong shape priors from CN statistics on a moderate size (515) of frontal upright faces in order to characterize the intrinsic geometries shared by human faces. We combine these shape priors with simple appearance based constraints, e.g., texture, edge, and corner, into a quadratic optimization. Thereafter, the solution to facial point extraction can be found by the standard gradient descent. The inclusion of these shape priors renders the robustness to pose changes owing to their invariance to projective transformations. Extensive experiments on the Labeled Faces in the Wild, Labeled Face Parts in the Wild and Helen database, and cross-set faces with various changes demonstrate the effectiveness of the CN-based shape priors compared with the state of the art.
Xin Fan 0001, Zhongxuan Luo, Daiyun Luo
IEEE Trans. Image Process.3
2015 Towards Effective Bug Triage with Software Data Reduction Techniques
abstract
Software companies spend over 45 percent of cost in dealing with software bugs. An inevitable step of fixing bugs is bug triage, which aims to correctly assign a developer to a new bug. To decrease the time cost in manual work, text classification techniques are applied to conduct automatic bug triage. In this paper, we address the problem of data reduction for bug triage, i.e., how to reduce the scale and improve the quality of bug data. We combine instance selection with feature selection to simultaneously reduce data scale on the bug dimension and the word dimension. To determine the order of applying instance selection and feature selection, we extract attributes from historical bug data sets and build a predictive model for a new bug data set. We empirically investigate the performance of data reduction on totally 600,000 bug reports of two large open source projects, namely Eclipse and Mozilla. The results show that our data reduction can effectively reduce the data scale and improve the accuracy of bug triage. Ourwork provides an approach to leveraging techniques on data processing to form reduced and high-quality bug data in software development and maintenance.
Jifeng Xuan, He Jiang 0001, Zhilei Ren, Weiqin Zou, Zhongxuan Luo, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.6
2014 Real-time estimation of hand gestures based on manifold learning from monocular videos
Yi Wang 0037, Zhongxuan Luo, Xin Fan 0001, Yunzhen Wu
Multim. Tools Appl.2
2014 A new geometric descriptor for symbols with affine deformations
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Yu Liu 0012, He Guo 0001
Pattern Recognit. Lett.3
2014 New Insights Into Diversification of Hyper-Heuristics
abstract
There has been a growing research trend of applying hyper-heuristics for problem solving, due to their ability of balancing the intensification and the diversification with low level heuristics. Traditionally, the diversification mechanism is mostly realized by perturbing the incumbent solutions to escape from local optima. In this paper, we report our attempt toward providing a new diversification mechanism, which is based on the concept of instance perturbation. In contrast to existing approaches, the proposed mechanism achieves the diversification by perturbing the instance under solving, rather than the solutions. To tackle the challenge of incorporating instance perturbation into hyper-heuristics, we also design a new hyper-heuristic framework HIP-HOP (recursive acronym of HIP-HOP is an instance perturbation-based hyper-heuristic optimization procedure), which employs a grammar guided high level strategy to manipulate the low level heuristics. With the expressive power of the grammar, the constraints, such as the feasibility of the output solution could be easily satisfied. Numerical results and statistical tests over both the Ising spin glass problem and the p -median problem instances show that HIP-HOP is able to achieve promising performances. Furthermore, runtime distribution analysis reveals that, although being relatively slow at the beginning, HIP-HOP is able to achieve competitive solutions once given sufficient time.
Zhilei Ren, He Jiang 0001, Jifeng Xuan, Zhongxuan Luo
IEEE Trans. Cybern.5
2013 A shape descriptor based on new projective invariants
abstract
Great attention has been devoted to the development of shape descriptors that is the key to object recognition. Previous works have great success on either relatively simple shapes or limited transformations, e.g., translation, rotation and scaling. We propose a new projective invariant, named characteristic number (CN) that includes more points for complex shapes with rich inner structures. Moreover, we build a novel shape descriptor with CN values calculated on triangles that cover the convex hull of a shape. The matching based on the descriptor also runs fast since only one initial point for the triangular coverage needs to align based on its CN value prior to the matching. The performance of the proposed descriptor is validated by the experiments compared with the classical shape context (SC) and recently developed cross ratio spectrum (CRS) on 32 logos of television networks with a wide range of transformations (512 images in total).
Zhongxuan Luo, Daiyun Luo, Xin Fan 0001, Xinchen Zhou, Qi Jia 0001
ICIP1
2012 Haze filtering with aerial perspective
abstract
In this paper, we present haze filtering that is capable of editing the amount of haze in an image given a haze observation. Aerial perspective is taken into account to generate depth dependent haze. We re-formulate the transmission estimation, the key to haze removal or filtering so that users are able to change the amount of haze in an image by tuning maximum visibility. The guided filter is employed in order to efficiently refine the estimated transmission. Additionally, we develop color correction and sky compensation based on physical priors for quality improvements. Experimental results show that the proposed method is able to generate images with various degree of haze in a natural and efficient fashion. The results are also free of color distortion that typically occurs when shooting in fog weather.
Renjie Gao, Xin Fan 0001, Jielin Zhang, Zhongxuan Luo
ICIP4
2012 Face recognition using average invariant factor
abstract
The recent developed intrinsic discriminate analysis (IDA) demonstrates superior recognition rate compared with classical methods such as PCA and LDA. In this paper, we not only re-prove the core theorem of IDA from a new perspective, but also define the Average Invariant Factor (AIF) that generalizes IDA. Two new algorithms for face recognition are built upon the AIF by using SVD and QR decomposition. Moreover, this new formulation facilitates the kernel extensions for the recognition algorithms, which relax the linear assumption for IDA. The presented kernel based AIF algorithms also significantly lower down the computational expenses of the original IDA method. A series of experiments on YALE and ORL sets demonstrate higher performance in terms of recognition rate and efficiency compared with classical statistical analysis methods (e.g., PCA, KPCA and 2DPCA) and the IDA algorithm.
Zhongxuan Luo, Xin Fan 0001, Jielin Zhang
ICIP1
2012 Hyper-Heuristics with Low Level Parameter Adaptation
abstract
Recent years have witnessed the great success of hyper-heuristics applying to numerous real-world applications. Hyper-heuristics raise the generality of search methodologies by manipulating a set of low level heuristics (LLHs) to solve problems, and aim to automate the algorithm design process. However, those LLHs are usually parameterized, which may contradict the domain independent motivation of hyper-heuristics. In this paper, we show how to automatically maintain low level parameters (LLPs) using a hyper-heuristic with LLP adaptation (AD-HH), and exemplify the feasibility of AD-HH by adaptively maintaining the LLPs for two hyper-heuristic models. Furthermore, aiming at tackling the search space expansion due to the LLP adaptation, we apply a heuristic space reduction (SAR) mechanism to improve the AD-HH framework. The integration of the LLP adaptation and the SAR mechanism is able to explore the heuristic space more effectively and efficiently. To evaluate the performance of the proposed algorithms, we choose the p-median problem as a case study. The empirical results show that with the adaptation of the LLPs and the SAR mechanism, the proposed algorithms are able to achieve competitive results over the three heterogeneous classes of benchmark instances.
Zhilei Ren, He Jiang 0001, Jifeng Xuan, Zhongxuan Luo
Evol. Comput.4
2012 Solving the Large Scale Next Release Problem with a Backbone-Based Multilevel Algorithm
abstract
The Next Release Problem (NRP) aims to optimize customer profits and requirements selection for the software releases. The research on the NRP is restricted by the growing scale of requirements. In this paper, we propose a Backbone-based Multilevel Algorithm (BMA) to address the large scale NRP. In contrast to direct solving approaches, the BMA employs multilevel reductions to downgrade the problem scale and multilevel refinements to construct the final optimal set of customers. In both reductions and refinements, the backbone is built to fix the common part of the optimal customers. Since it is intractable to extract the backbone in practice, the approximate backbone is employed for the instance reduction while the soft backbone is proposed to augment the backbone application. In the experiments, to cope with the lack of open large requirements databases, we propose a method to extract instances from open bug repositories. Experimental results on 15 classic instances and 24 realistic instances demonstrate that the BMA can achieve better solutions on the large scale NRP instances than direct solving approaches. Our work provides a reduction approach for solving large scale problems in search-based requirements engineering.
Jifeng Xuan, He Jiang 0001, Zhilei Ren, Zhongxuan Luo
IEEE Trans. Software Eng.4
2012 An Accelerated-Limit-Crossing-Based Multilevel Algorithm for the p-Median Problem
abstract
In this paper, we investigate how to design an efficient heuristic algorithm under the guideline of the backbone and the fat, in the context of the p-median problem. Given a problem instance, the backbone variables are defined as the variables shared by all optimal solutions, and the fat variables are defined as the variables that are absent from every optimal solution. Identification of the backbone (fat) variables is essential for the heuristic algorithms exploiting such structures. Since the existing exact identification method, i.e., limit crossing (LC), is time consuming and sensitive to the upper bounds, it is hard to incorporate LC into heuristic algorithm design. In this paper, we develop the accelerated-LC (ALC)-based multilevel algorithm (ALCMA). In contrast to LC which repeatedly runs the time-consuming Lagrangian relaxation (LR) procedure, ALC is introduced in ALCMA such that LR is performed only once, and every backbone (fat) variable can be determined in O(1) time. Meanwhile, the upper bound sensitivity is eliminated by a dynamic pseudo upper bound mechanism. By combining ALC with the pseudo upper bound, ALCMA can efficiently find high-quality solutions within a series of reduced search spaces. Extensive empirical results demonstrate that ALCMA outperforms existing heuristic algorithms in terms of the average solution quality.
Zhilei Ren, He Jiang 0001, Jifeng Xuan, Zhongxuan Luo
IEEE Trans. Syst. Man Cybern. Part B4
2011 The nearest complex polynomial with a zero in a given complex domain
Zhongxuan Luo, Dongwoo Sheen
Theor. Comput. Sci.1
2010 Ant Based Hyper Heuristics with Space Reduction: A Case Study of the p-Median Problem
Zhilei Ren, He Jiang 0001, Jifeng Xuan, Zhongxuan Luo
PPSN (1)4
2010 Automatic Bug Triage using Semi-Supervised Text Classification
Jifeng Xuan, He Jiang 0001, Zhilei Ren, Jun Yan 0009, Zhongxuan Luo
SEKE5
2008 Layered deformation of solid model using conformal mapping
Zhongxuan Luo, Junxiao Xue
Comput. Graph.1
2007 3D Model Deformation via Conformal Mapping
abstract
In this paper, we introduce conformal mapping to 3D model deformation. The surface of the solid model is represented by base-patch and cross-section function. Accordingly, the shape of the 3D model can be modified by interactive means, such as changing the corresponding base-patch with conformal mapping and adjusting a cross-section function. There are two features of the proposed method. First, due to the properties of conformal mapping, the scheme satisfies the local demands on rigidity and still allows for considerable global deformation. Second, the deformation effect is predictable and the transformation function of the deformation can be expressed analytically. The proposed transformation of 3D deformation is continuous, not self- intersection, and topologically consistent. Some examples are presented in the end of the paper to illustrate the efficiency of our approach.
Zhongxuan Luo, Junxiao Xue
CAD/Graphics1