Lin Zhang 0014

dblp:37/1629-14 · DBLP profile ↗
← Back
120ranked-venue papers
32as first author
55since 2021 · last 2026
0000-0002-4360-5523ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 84 · 22 first-author · 44 since 2021Artificial intelligence and machine learning · 31 · 11 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 since 2021Computer networks · 10 · 10 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
abstract
Recent advances in generative AI have accelerated the production of ultra-high-resolution visual content. However, traditional image formats face significant limitations in efficient compression and real-time decoding, which restricts their applicability on end-user devices. Inspired by 3D Gaussian Splatting, 2D Gaussian image models have achieved notable progress in enhancing image representation efficiency and quality. Nevertheless, existing methods struggle to balance compression ratios and reconstruction fidelity in ultra-high-resolution scenarios. To address these challenges, we propose SmartSplat, a highly adaptive and feature-aware GS-based image compression framework that effectively supports arbitrary image resolutions and compression ratios. By leveraging image-aware features such as gradients and color variances, SmartSplat introduces a Gradient-Color Guided Variational Sampling strategy alongside an Exclusion-based Uniform Sampling scheme, significantly improving the non-overlapping coverage of Gaussian primitives in pixel space. Additionally, a Scale-Adaptive Gaussian Color Sampling method is proposed to enhance the initialization of Gaussian color attributes across scales. Through joint optimization of spatial layout, scale, and color initialization, SmartSplat can efficiently capture both local structures and global textures of images using a limited number of Gaussians, achieving superior reconstruction quality under high compression ratios. Extensive experiments on DIV8K and a newly created 16K dataset demonstrate that SmartSplat significantly outperforms state-of-the-art methods at comparable compression ratios and surpasses their compression limits, exhibiting strong scalability and practical applicability. This framework can effectively alleviate the storage and transmission burdens of ultra-high-resolution images, providing a robust foundation for future high-efficiency visual content processing.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
AAAI2
2026 Why Do Emotions Change? Appraisal-Guided Reasoning for Emotion-Cause Triplet Extraction in Conversations
abstract
Multimodal Emotion-Cause Triplet Extraction in Conversations (MECTEC) is fundamental for fine-grained affect understanding, yet it remains challenging in multi-turn, multispeaker settings.Existing methods often make locally plausible predictions but struggle to maintain conversation-level consistency under within-speaker emotion shifts and core events.To address this, we propose ECFlow, a unified framework that combines appraisal-guided structured generation with graph-structured reinforcement learning.ECFlow operationalizes cognitive appraisal theory into a controllable intermediate reasoning trace and constructs UMECS, a unified supervision dataset with cognitively grounded traces.It then lifts predicted and gold triplets into an Emotion-Cause Flow Graph and optimizes verifiable, structure-aware rewards for emotion-shift coherence and core-event consistency, together with task-oriented triplet rewards.Experiments on public MECTEC benchmarks show that ECFlow consistently outperforms strong baselines, achieving state-of-the-art triplet extraction and improved structure-aware metrics on emotion shifts and core events.
Ying Shen 0005, Lin Zhang 0014
ACL (1)5
2026 Decision-Invariant Sim-to-Real Vision-and-Language Navigation with Pseudo-Panoramic Observations
abstract
Following natural language instructions to complete navigation tasks is a crucial capability for real-world embodied robots. In vision-and-language navigation (VLN), agents typically assume access to complete, on-the-fly environmental observations and rely on them for decision making. However, most sim-to-real VLN approaches approximate the privileged complete panoramic sensing in simulation with monocular sensors, leading to significant semantic loss and incomplete perception. In this work, we present DIP2, a sim-to-real VLN framework that reduces the discrepancy between assumed and realizable observations by introducing a unified sensing and mapping representation shared across simulated and real-world domains. Specifically, dense pseudo-panoramic observations are synthesized from a self-assembled surround-view camera system to recover rich semantics comparable to simulation, while a structure-only metric map unifies simulated RGB-D inputs with real-world LiDAR observations. Furthermore, to promote sim-to-real decision invariance, DIP2 employs a learning-free local navigable waypoint prediction strategy via radial expansion, applied consistently in both domains. The predicted waypoints and dense semantic observations are integrated into a global topological map, allowing existing agents trained in simulation to be directly deployed in the real world. Extensive experiments in simulated and real environments demonstrate that DIP2 substantially improves sim-to-real navigation performance. Source code will be published at https://github.com/zheng19845/DIP2.
Yuanyu Zheng, Xumin Shen, Yunda Sun, Ying Shen 0005, Lin Zhang 0014
ICMR5
2026 ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
abstract
Large Language Models (LLMs) have shown strong potential as conversational agents. Yet, their effectiveness remains limited by deficiencies in robust long-term memory, particularly in complex, long-term web-based services such as online emotional support. However, existing long-term dialogue benchmarks primarily focus on static and explicit fact retrieval, failing to evaluate agents in critical scenarios where user information is dispersed, implicit, and continuously evolving. To address this gap, we introduce ES-MemEval, a comprehensive benchmark that systematically evaluates five core memory capabilities: information extraction, temporal reasoning, conflict detection, abstention, and user modeling, in long-term emotional support settings, covering question answering, summarization, and dialogue generation tasks. To support the benchmark, we also propose EvoEmo, a multi-session dataset for personalized long-term emotional support that captures fragmented, implicit user disclosures and evolving user states. Extensive experiments on open-source long-context, commercial, and retrieval-augmented (RAG) LLMs show that explicit long-term memory is essential for reducing hallucinations and enabling effective personalization. At the same time, RAG improves factual consistency but struggles with temporal dynamics and evolving user states. These findings highlight both the potential and limitations of current paradigms and motivate more robust integration of memory and retrieval for long-term personalized dialogue systems.
Jiaqi Lu 0004, Ying Shen 0005, Lin Zhang 0014
WWW4
2026 SETFusion: A semantic transformer for infrared and visible image fusion
Wei Tang 0018, Fazhi He, Lin Zhang 0014, Shengjie Zhao 0001
Pattern Recognit.3
2026 HiGS-Calib: A Hierarchical 3D Gaussian Splatting-Based Targetless Local-Consistent LiDAR-Camera Calibration Method
abstract
3D Gaussian Splatting (3DGS) has emerged as a powerful scene representation, offering geometrically dense and photometrically accurate modeling capabilities that present a promising new paradigm for accurate targetless sensor calibration. Current 3DGS-based LiDAR-camera calibration methods usually highly rely on the joint optimization of the Gaussian model and extrinsics, and thus suffer two critical limitations. On the one hand, the global optimization is usually sensitive to the accumulated localization error of LiDAR. On the other hand, the inaccurate extrinsics may cause oscillations during the joint optimization. Specifically, there is a fundamental dilemma: accurate extrinsic calibration requires accurate scene models, while constructing accurate models itself depends on accurate extrinsics. This dilemma frequently triggers oscillatory optimization trajectories, significantly increasing vulnerability to premature convergence at suboptimal states. To address these challenges, we propose HiGS-Calib, a novel 3DGS calibration pipeline featuring the integration of our proposed Local-Consistent Photometric-Geometric (LCPG) error model and the hierarchical architecture. The LCPG error leverages the spatial consistency within local windows to quantify pose misalignment using only geometric attributes of the 3DGS model, bypassing color-pose reliance constraints. Besides, diverging from joint optimization paradigms, HiGS-Calib implements coarse-to-fine iterative optimization, decoupling scene modeling from extrinsic refinement and thereby achieving stable and accurate calibration. Extensive evaluation demonstrates the significantly improved calibration accuracy and stability of our HiGS-Calib over other state-of-the-art methods. To make our results reproducible, the source code has been released at https://github.com/IRMVLab/HiGS-Calib.
Tianjun Zhang, Lin Zhang 0014, Hesheng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 MECI: A Multi-Model Motion Capture-Free Event Dataset Featuring Large-Scale Challenging Indoor Environments
abstract
Recently, several event-based datasets have emerged to foster the application of the new event camera to classic vision tasks like Simultaneous Localization and Mapping (SLAM). However, current indoor benchmark datasets depend on the expensive motion capture system to obtain ground-truth trajectories, restricting data acquisition to small object-centric scenes or single-room environments due to infrastructure costs and spatial limitations. Furthermore, these datasets lack sensor diversity, relying solely on a single event camera model that hinders practical cross-device generalization. To address the above limitations, we propose MECI, the first Multi-model Event dataset targeting Challenging Indoor environments, especially including large-scale scenes with complete trajectory ground-truth provided. Specifically, MECI includes 38 posed sequences of visual-inertial-event data from three different models of event cameras with varying resolutions and data frequencies. Apart from typical challenging factors such as illumination changes, motion blur, and dynamics, these sequences involve large-scale indoor scenes (across rooms and floors) with each room occupying approximately 60 m \({}^{2}\) and a maximum trajectory length of 343.8 m. We innovatively leverage ETS (Electronic Total Station) measurements and AprilTag markers to provide complete 6-DoF trajectories with low cost and minimal environmental modifications. Extensive experiments show that our MECI is an effective yet challenging benchmark dataset not only for visual SLAM, but also for event-image reconstruction. Our project page is https://cslinzhang.github.io/MECI_dataset/ .
Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.2
2026 CaneSpeaker: An LLM-Assisted Speaker for Generating Human-Like Navigation Instructions
abstract
Navigation instruction generation aims to address data scarcity in Vision-and-Language Navigation (VLN) by generating navigation instructions for unannotated routes from data sources like simulators or online data. However, existing methods usually suffer from high reliance on panoramic views, poor cross-task generalization ability, and limited availability of training data. To address these challenges, we propose a novel speaker, CaneSpeaker, to generate human-like instructions from front-facing images for a variety of VLN tasks. First, to mitigate the limited amount of speaker training data, we propose an Large Language Model (LLM)-based instruction augmentation method, LLM-IA, that utilizes an off-the-shelf LLM to create augmented instructions for training by distilling and reformulating existing instructions. This method allows us to collect an instruction-augmented dataset with human-level accuracy for speaker training, namely Rx2R. Second, to eliminate the dependency on panoramic views, we propose a novel Vision-Language Model (VLM)-based speaker architecture, VL-Sp. By leveraging the advanced reasoning capabilities of a pre-trained VLM, CaneSpeaker can effectively generate high-quality instructions directly from front-facing images without relying on panoramic views. Also, the prompt-based characteristic of the VLM allows us to devise a unified input representation to enable the processing of multiple VLN tasks, thus further addressing the problem of data scarcity by combining multiple datasets from different VLN tasks. Finally, we utilize CaneSpeaker to synthesize a large-scale augmented dataset, CANE, from unannotated routes in the Matterport3D Simulator. Comprehensive experiments demonstrate that CaneSpeaker generates precise instructions with diverse expressions across various VLN tasks, and the VLN agent trained on our datasets obviously outperforms its counterparts. The source codes and datasets are available at https://github.com/zheng19845/CaneSpeaker .
Yuanyu Zheng, Lin Zhang 0014, Yunda Sun, Ying Shen 0005, Shengjie Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Learning from Rendering: Realistic and Controllable Extreme Rainy Image Synthesis for Autonomous Driving Simulation
abstract
Autonomous driving simulators provide an effective and low-cost alternative for evaluating or enhancing visual perception models. However, the reliability of evaluation depends on the diversity and realism of the generated scenes. Extreme weather conditions, particularly extreme rainfalls, are rare and costly to capture in real-world settings. While simulated environments can help address this limitation, existing rainy image synthesizers often suffer from poor controllability over illumination and limited realism, which significantly undermines the effectiveness of the model evaluation. To that end, we propose a learning-from-rendering rainy image synthesizer, which combines the benefits of the realism of rendering-based methods and the controllability of learning-based methods. To validate the effectiveness and generalizability of our extreme rainy image synthesizer on the semantic segmentation task, a continuous set of pixel-accurately labeled extreme rainy images is necessary. By integrating the proposed synthesizer with the CARLA driving simulator, we develop CARLARain—an extreme rainy street scene simulator which can obtain paired rainy-clean images and labels under complex illumination conditions. Qualitative and quantitative experiments validate that CARLARain can effectively improve the accuracy of semantic segmentation models in extreme rainy scenes, with the models’ accuracy (mIoU) improved by \(5{-}8\%\) on the synthetic dataset and significantly enhanced in real extreme rainy scenarios under complex illuminations. Our source code and datasets are available at https://kb824999404.github.io/HRIG/ .
Kaibin Zhou, Kaifeng Huang 0001, Hao Deng 0002, Zelin Tao, Ziniu Liu, Lin Zhang 0014, Shengjie Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2025 Representing Sounds as Neural Amplitude Fields: A Benchmark of Coordinate-MLPs and a Fourier Kolmogorov-Arnold Framework
abstract
Although Coordinate-MLP-based implicit neural representations have excelled in representing radiance fields, 3D shapes, and images, their application to audio signals remains underexplored. To fill this gap, we investigate existing implicit neural representations, from which we extract 3 types of positional encoding and 16 commonly used activation functions. Through combinatorial design, we establish the first benchmark for Coordinate-MLPs in audio signal representations. Our benchmark reveals that Coordinate-MLPs require complex hyperparameter tuning and frequency-dependent initialization, limiting their robustness. To address these issues, we propose Fourier-ASR, a novel framework based on the Fourier series theorem and the Kolmogorov-Arnold representation theorem. Fourier-ASR introduces Fourier Kolmogorov-Arnold Networks (Fourier-KAN), which leverage periodicity and strong nonlinearity to represent audio signals, eliminating the need for additional positional encoding. Furthermore, a Frequency-adaptive Learning Strategy (FaLS) is proposed to enhance the convergence of Fourier-KAN by capturing high-frequency components and preventing overfitting of low-frequency signals. Extensive experiments conducted on natural speech and music datasets reveal that: (1) well-designed positional encoding and activation functions in Coordinate-MLPs can effectively improve audio representation quality; and (2) Fourier-ASR can robustly represent complex audio signals without extensive hyperparameter tuning. Looking ahead, the continuity and infinite resolution of implicit audio representations make our research highly promising for tasks such as audio compression, synthesis, and generation.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
AAAI2
2025 Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-Sources
abstract
Audio-visual navigation has received considerable attention in recent years. However, the majority of related investigations have focused on single sound-source scenarios. Studies in this field for multiple sound-source scenarios remain underexplored due to the limitations of two aspects. First, the existing audio-visual navigation dataset only has limited audio samples, making it difficult to simulate diverse multiple sound-source environments. Second, existing navigation frameworks are mainly designed for single sound-source scenarios, thus their performance is severely reduced in multiple sound-source scenarios. In this work, we make an attempt to fill in these two research gaps to some extent. First, we establish a large-scale BEnchmark Dataset for Audio-Vsual Navigation, namely BeDAViN. This dataset consists of 2,258 audio samples with a total duration of 10.8 hours, which is more than 33 times longer than the existing audio dataset employed in the audio-visual navigation task. Second, we propose a new Embodied Navigation framework for MUltiple Sound-Sources Scenarios called ENMuS3. There are mainly two essential components in ENMuS3, the sound event descriptor and the multi-scale scene memory transformer. The former component equips the agent with the ability to extract spatial and semantic features of the target sound-source among multiple sound-sources, while the latter provides the ability to track the target object effectively in noisy environments. Experimental results on our BeDAViN show that ENMuS3 strongly outperforms its counterparts with a significant improvement in success rates across diverse scenarios.
Zhanbo Shi, Lin Zhang 0014, Linfei Li, Ying Shen 0005
AAAI2
2025 WSGS: A Speech-Driven Zero-Shot System for 6D Robotic Arm Grasping
abstract
In the robotic vision industry, recent years have witnessed a growing interest in object detection and pose estimation for precise robotic arm grasping. In fact, various unfavorable factors, such as the size limitations of robotic grippers, the diversity and potential complexity of object shapes and poses, and the cluttered environment, make robotic arm grasping based on 6D object poses much harder than it seems. In this paper, to solve these issues to some extent, we proposed a speech-driven zero-shot system for robotic arm grasping, called WSGS (Whisper-SAM6D Grasping System). It enables speech-driven human-interactive grasping with the Franka Emika robotic arm by performing instance segmentation and pose estimation based on speech instructions. Specifically, WSGS accurately recognizes the 6D pose of an unknown object and adapts to its size to find the most suitable position for grasping. Comprehensive experiments on our real-world scenarios demonstrate that WSGS can produce high-accuracy instance segmentation and pose estimation results, achieving adaptive robotic arm grasping of unknown objects based on speech commands.
Yitong Ge, Lin Zhang 0014, Yang Chen 0037, Ying Shen 0005
ICME2
2025 GRE-SLAM: 6-DoF Pure Event-Based SLAM with Semi-Dense Depth Recovery Assisted Bundle Adjustment
abstract
Event cameras are innovative bioinspired vision sensors that output pixel-level brightness changes instead of standard intensity frames. Such cameras do not suffer from motion blur and cope well with scenes characterized by high dynamic range, which can benefit classic computer vision tasks such as pose estimation. However, currently developed event-based pose estimation methods either require extra data as inputs (such as IMU data or depths) or lack a global refinement step to alleviate accumulated drifts. To this end, we propose the first 6-DoF pure event-based SLAM system equipped with back-end global optimization, named GRE-SLAM (Globally Refined Event-based SLAM). For robustness and accuracy, first, 6-DoF motion compensation is introduced in the front-end to prepare sharp-edged event frames and a favorable initialization pose, mitigating unstable optimization during event registration brought by sparsity and noise of events. Second, a novel adaptive semi-dense depth recovery algorithm enriches front-end's sparse depths without additional sensors, helping establish long-term edge alignment constraints to support global BA in the back-end. Comprehensive experiments on real-world datasets demonstrate that our method can produce high-accuracy pose estimation results as well as recover a semi-dense depth map for each Image of Warped Events (IWE).
Yang Chen 0037, Lin Zhang 0014
ICMR2
2025 MaGo-I2P: Image-to-Point Cloud Registration with Mamba and Geometry Recovery
abstract
Estimating the relative poses between images and point clouds is a fundamental problem in multi-sensor fusion, with extensive applications in tasks such as robot localization and navigation. However, existing methods fall short in registration accuracy and efficiency due to the modality gaps and resource-consuming backbones. To address these issues, we propose the first Mamba-based I2P registration framework called MaGo-I2P. On the one hand, MaGo-I2P recovers the geometric structure of images through depth estimation, thereby constructing an implicit 3D representation of the image scene to alleviate the modality gap between images and point clouds, facilitating cross-modal feature extraction. On the other hand, unlike transformer-based backbones applied in existing methods, a Mamba-based backbone with linear time complexity is utilized in our MaGo-I2P. Such a backbone allows our method to possess both context-aware capability and fast inference speed. In addition, by adopting a coarse-to-fine matching strategy, MaGo-I2P eliminates outlier matches by progressively narrowing the matching region, establishing more accurate 2D-3D correspondences. Experiments on KITTI Odometry and Oxford Robotcar datasets suggest that our method achieves state-of-the-art registration accuracy while maintaining high-efficiency. Meanwhile, we also demonstrate the application potential of MaGo-I2P in LiDAR-camera calibration through qualitative experiments. The source code will be released at https://cslinzhang.github.io/MaGo-I2P.
Yunda Sun, Lin Zhang 0014
ICMR2
2025 Deep Unrolled Weighted Graph Laplacian Regularization for Depth Completion
Jin Zeng 0004, Qingpeng Zhu, Tongxuan Tian, Wenxiu Sun, Lin Zhang 0014, Shengjie Zhao 0001
Int. J. Comput. Vis.5
2025 Online indoor visual odometry with semantic assistance under implicit epipolar constraints
Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
Pattern Recognit.2
2025 All-Inclusive Image Enhancement for Degraded Images Exhibiting Low-Frequency Corruption
abstract
In this paper, a novel image enhancement method, called the all-inclusive image enhancement (AIIE), is proposed that can effectively enhance the degraded images for improving the visibility of image content. These imageries were acquired under various types of weather conditions such as haze, low-light, underwater, and sandstorm, etc. One commonality shared by this class of noise is that the resulted degradations on visual quality or visibility are caused by low-frequency interference. Existing image enhancement methods lack the ability to deal with all types of degradations from this class, while our proposed AIIE offers a unified treatment for them. To achieve this goal, a statistical property is obtained from the study of the discrete cosine transform (DCT) of 1,000 high- and 1000 low-quality images on their DCT domains. It shows that the normalized DCT coefficients (between 0 and 1) of high-quality images has about 95% fall in the interval [0, 0.2]; for low-quality images, almost all the coefficients are in the same interval. This fundamental property, called the DCT prior (DCT-P), is instrumental to the development of our AIIE algorithm proposed in this paper. Since the proposed DCT-P delineates the attributes of high- and low-quality images clearly, it becomes a highly effective ‘tool’ to convert low-quality images to its enhanced version. Extensive experimental results have clearly validated the superior performance of the AIIE conducted on different types of deteriorated images in terms of visual quality and efficiency as well as significant advantages on computational complexity, which is essential for real-time applications.
Mingye Ju, Chunming He, Can Ding 0002, Wenqi Ren, Lin Zhang 0014, Kai-Kuang Ma
IEEE Trans. Circuits Syst. Video Technol.5
2025 S²KAN-SLAM: Elastic Neural LiDAR SLAM With SDF Submaps and Kolmogorov-Arnold Networks
abstract
Traditional LiDAR SLAM approaches prioritize localization over mapping, yet high-precision dense maps are essential for numerous applications involving intelligent agents. Recent advancements have introduced methods leveraging neural fields to enhance mapping capabilities; however, these approaches still face several limitations. Firstly, concerning scene representation, they typically employ neural fields with high-dimensional features and multi-layer perceptron decoders utilizing non-continuous activation functions. This results in low learning efficiency and challenges in capturing high-frequency signals. Secondly, in terms of scene organization, these methods often treat the entire scene as a singular neural field, leading to inefficiencies, inflexibility, and difficulties in rectifying accumulated errors when mapping large-scale environments over extended periods. To tackle the first issue, we propose a lightweight continuous SDF regression approach by encoding the scene in single-valued embeddings and decoding SDF values from a Kolmogorov-Arnold Network. By minimizing discrepancies in measuring range, sampling distance, and decoded SDF values, we facilitate iterative frame-to-model tracking and bundle adjustment neural mapping. To mitigate the second challenge, we propose structuring the whole scene into multiple neural SDF submaps. By establishing node-node, node-submap, and loop closure constraints into a global pose graph, the system can create dense neural maps with global consistency across large-scale scenes. Experimental evaluations in both real-world and simulated settings indicate that our system achieves superior mapping completeness and accuracy, enhanced learning efficiency, reduced memory consumption, and greater flexibility compared to its counterparts.
Zhong Wang 0009, Lin Zhang 0014, Hesheng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 I-DACS: Always Maintaining Consistency Between Poses and the Field for Radiance Field Construction Without Pose Prior
abstract
The radiance field, emerging as a novel 3D scene representation, has found widespread application across diverse fields. Standard radiance field construction approaches rely on the ground-truth poses of key-frames, while building the field without pose prior remains a formidable challenge. Recent advancements have made strides in mitigating this challenge, albeit to a limited extent, by jointly optimizing poses and the radiance field. However, in these schemes, the consistency between the radiance field and poses is achieved completely by training. Once the poses of key-frames undergo changes, long-term training is required to readjust the field to fit them. To address such a limitation, we propose a new solution for radiance field construction without pose prior, namely I-DACS (Incremental radiance field construction with Direction-Aware Color Sampling). Diverging from most of the existing global optimization solutions, we choose to incrementally solve the poses and construct a radiance field within a sliding-window framework. The poses are unequivocally retrieved from the radiance field, devoid of any constraints and accompanying noise from other observation models, so as to achieve the consistency of poses to the field. Besides, in the radiance field, the color information is much higher-frequency and more time-consuming to learn compared with the density. To accelerate training, we isolate the color information to a distinct color field, and construct the color field based on an innovative direction-aware color sampling strategy, by which the color field can be derived directly from images without training. The color field obtained in this way is always consistent with the poses, and intricate details of training images can be retained to the utmost extent. Extensive experimental results evidently showcase both the remarkable training speed and the outstanding performance in rendering quality and localization accuracy achieved by I-DACS. To make our results reproducible, the source code has been released athttps://cslinzhang.github.io/I-DACS-MainPage/.
Tianjun Zhang, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2025 ATM-NeRF: Accelerating Training for NeRF Rendering on Mobile Devices via Geometric Regularization
abstract
Recently, an increasing number of researchers have been dedicated to transferring the impressive novel view synthesis capability of Neural Radiance Fields (NeRF) to resource-constrained mobile devices. One common solution is to pre-train NeRF and bake it into textured meshes which are well supported by mobile graphics hardware. However, the training process of existing methods often requires several hours even with multiple high-end NVIDIA V100 GPUs. The underlying reason is that these schemes mainly rely on photometric rendering loss, neglecting the geometric relationship between the pre-trained NeRF and the baked results. Standing on this point, we presentATM-NeRF(AcceleratingTraining forMobile rendering based onNeRF), which is the first to apply effective geometric regularization constraints during both the pre-training and the baking training stages for faster convergence. Specifically, in the initial NeRF pre-training stage, we enforce consistency of the multi-resolution density grids representing the scene geometry to mitigate the shape-radiance ambiguity problem to some extent, achieving a coarse mesh with smoothness. In the second stage, we utilize the positions and geometric features of 3D points projected from the pre-trained posed depths to provide geometric supervision for joint refinement of geometry and appearance of the coarse mesh. As a result, our ATM-NeRF achieves comparable rendering quality to MobileNeRF with a training speed that is about$30\times \sim 70\times$faster while maintaining finer structure details of the exported mesh.
Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
IEEE Trans. Multim.2
2025 Skeleton-Aware Graph-Based Adversarial Networks for Human Pose Estimation from Sparse IMUs
abstract
Recently, sparse-inertial human pose estimation (SI-HPE) with only a few IMUs has shown great potential in various fields. The most advanced work in this area achieved fairish results using only six IMUs. However, there are still two major issues that remain to be addressed. First, existing methods typically treat SI-HPE as a temporal sequential learning problem and often ignore the important spatial prior of skeletal topology. Second, there are far more synthetic data in their training data than real data, and the data distribution of synthetic data and real data is quite different, which makes it difficult for the model to be applied to more diverse real data. To address these issues, we propose “Graph-based Adversarial Inertial Poser (GAIP),” which tracks body movements using sparse data from six IMUs. To make full use of the spatial prior, we design a multi-stage pose regressor with graph convolution to explicitly learn the skeletal topology. A joint position loss is also introduced to implicitly mine spatial information. To enhance the generalization ability, we propose supervising the pose regression with an adversarial loss from a discriminator, bringing the ability of adversarial networks to learn implicit constraints into full play. Additionally, we construct a real dataset that includes hip support movements and a synthetic dataset containing various motion categories to enrich the diversity of inertial data for SI-HPE. Extensive experiments demonstrate that GAIP produces results with more precise limb movement amplitudes and relative joint positions, accompanied by smaller joint angle and position errors compared to state-of-the-art counterparts. The datasets and codes are publicly available at https://cslinzhang.github.io/GAIP/ .
Kaixin Chen 0003, Lin Zhang 0014, Zhong Wang 0009, Shengjie Zhao 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Towards a Robust Visual-Inertial-Surround-View SLAM System for Autonomous Indoor Parking
abstract
An autonomous parking system is a low-speed unmanned driving system applied in indoor parking environments. Real-time and high-precision vehicle localization and map construction of the environment are two core functional modules of the system. Camera and IMU (Inertial Measurement Unit) sensors provide complementary data to create a Visual-Inertial Simultaneous Localization and Mapping (VI-SLAM) system. However, existing SLAM systems face challenges in complex parking environments. Moreover, limitations inherent in VI-SLAM systems further compromise their perception accuracy, affecting both localization and optimization. This article addresses the shortcomings of current VI-SLAM systems by proposing the RVIS SLAM system. This robust semantic SLAM system integrates data from three sensors: a front-view camera, an IMU, and a surround-view system. To ensure localization accuracy, the system utilizes metric information from common semantic objects on the ground. These objects include parking-slots, speed bumps, and parking-slot numbers captured in surround-view images to build scale-aware constraints. These constraints refine the initial scale of the SLAM system, which is often compromised under low IMU excitation conditions. Additionally, in optimization, SLAM systems ideally assume that the front-end produces optimization graphs without data association outliers. However, in real-world indoor parking environments, sensor noise and vehicle vibrations make this assumption unrealistic. To mitigate the adverse effects of outliers in SLAM systems, this article proposes a robust surround-view semantic data association strategy. This strategy quantifies the uncertainty of surround-view semantic landmarks for the first time, ensuring reliable localization and mapping in challenging environments. Extensive experiments in typical indoor parking environments validate the effectiveness and efficiency of the proposed RVIS SLAM system.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Shengjie Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2024 IMU-Assisted Target-Free Extrinsic Calibration of Heterogeneous Lidars Based on Continuous-Time Optimization
abstract
Data fusion of heterogeneous LiDAR systems has gained significant attention due to its potential for providing wide-range sensing and high-density measurements for robots. However, existing LiDAR calibration methods primarily focus on homogeneous LiDAR systems and yield suboptimal outcomes when applied to heterogeneous setups. To this end, this paper proposes an IMU-Assisted Heterogeneous LiDAR extrinsics Calibration method, namely IA-HeLiC, which is a target-free method based on continuous-time optimization. Specifically, IA-HeLiC utilizes two types of errors, namely geometric constraint error and motion constraint error, and minimizes them within a B-spline-based continuous-time framework to achieve accurate extrinsic calibration. Using a parameter loopback mechanism, this optimization process is performed iteratively to further improve calibration accuracy. IA-HeLiC’s performance is corroborated through experiments using a ground-truth-known handheld device, by which multiple data sequences were collected in diverse real-world scenes. To make our results reproducible, the source code and the collected dataset have been released at https://cslinzhang.github.io/IA-HeLiC.
Zehao Yan, Lin Zhang 0014, Zhong Wang 0009, Shenjie Zhao
ICIP2
2024 GS3LAM: Gaussian Semantic Splatting SLAM
abstract
Recently, the multi-modal fusion of RGB, depth, and semantics has shown great potential in the domain of dense Simultaneous Localization and Mapping (SLAM), as known as dense semantic SLAM. Yet a prerequisite for generating consistent and continuous semantic maps is the availability of dense, efficient, and scalable scene representations. To date, existing semantic SLAM systems based on explicit scene representations (points/meshes/surfels) are limited by their resolutions and inabilities to predict unknown areas, thus failing to generate dense maps. Contrarily, a few implicit scene representations (Neural Radiance Fields) to deal with these problems rely on time-consuming ray tracing-based volume rendering technique, which cannot meet the real-time rendering requirements of SLAM. Fortunately, the Gaussian Splatting scene representation has recently emerged, which inherits the efficiency and scalability of point/surfel representations while smoothly represents geometric structures in a continuous manner, showing promise in addressing the aforementioned challenges. To this end, we propose GS3LAM, a Gaussian Semantic Splatting SLAM framework, which takes multimodal data as input and can render consistent, continuous dense semantic maps in real-time. To fuse multimodal data, GS3LAM models the scene as a Semantic Gaussian Field (SG-Field), and jointly optimizes camera poses and the field by establishing error constraints between observed and predicted data. Furthermore, a Depth-adaptive Scale Regularization (DSR) scheme is proposed to tackle the problem of misalignment between scale-invariant Gaussians and geometric surfaces within the SG-Field. To mitigate the forgetting phenomenon, we propose an effective Random Sampling-based Keyframe Mapping (RSKM) strategy, which exhibits notable superiority over local covisibility optimization strategies commonly utilized in 3DGS-based SLAM systems. Extensive experiments conducted on the benchmark datasets reveal that compared with state-of-the-art competitors, GS3 LAM demonstrates increased tracking robustness, superior real-time rendering quality, and enhanced semantic reconstruction precision. To make the results reproducible, the source code is available at https://github.com/lif314/GS3LAM.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
ACM Multimedia2
2024 Adversarial compact wrapping classifier learning for open set recognition
Lin Zhang 0014, Minghua Wan, Pu Huang 0004, Guowei Yang 0002
Inf. Sci.1
2024 MPEG: A Multi-Perspective Enhanced Graph Attention Network for Causal Emotion Entailment in Conversations
abstract
Emotion causes constitute a pivotal component in the comprehension of emotional conversations. Recently, a new task named Causal Emotion Entailment (CEE) has been proposed to identify the causal utterances for the target emotional utterance in a conversation. Although researchers have achieved some progress in solving this problem, they failed to adequately incorporate speaker characteristics and overlooked the effects of temporal relations in conversation structures. To fill such a research gap to some extent, we propose a novel causal emotion entailment framework, namely MPEG (Multi-Perspective Enhanced Graph attention network). The training of MPEG consists of three stages. Firstly, we utilize a speaker-aware pre-trained model and two attention mechanisms to obtain the utterance representations that incorporate local contexts as well as the speaker and emotional information. Then, these representations are fed into a graph attention network to model the conversation structures and emotional dynamics from both local and global perspectives. Finally, a fully-connected network is implemented to predict the relationships between emotional utterances and causal utterances. Experimental results show that MPEG achieves state-of-the-art performance. The source code is available athttps://github.com/slptongji/MPEG.
Ying Shen 0005, Xuri Chen, Lin Zhang 0014, Shengjie Zhao 0001
IEEE Trans. Affect. Comput.4
2024 TriKF: Triple-Perspective Knowledge Fusion Network for Empathetic Question Generation
abstract
Questioning is one of the essential tactics for demonstrating empathy in social dialogues. Effective questioning can guide individuals to express their experiences, feelings, and thoughts, aiming to establish emotional connections and deepen interpersonal understanding. However, how to generate empathetic questions in emotional support conversations remains an unresolved issue. To fill this research gap to some extent, we propose an empathetic question generation (QG) framework called triple-perspective knowledge fusion (TriKF), which incorporates external knowledge from the perspectives of events, cognition, and affection to comprehensively understand the dialogue context. Specifically, this framework acquires commonsense knowledge from these three perspectives and integrates them into the dialogue context to enrich the contextual information. To the best of our knowledge, this is the first method proposed for empathetic QG. Additionally, we construct an empathetic question dataset, namely EQ-EMAC. This dataset comprises 4213 dialogues with single user inputs and multiple empathetic question responses, which can be utilized to assess the effectiveness and generalization capability of empathetic QG models. Experimental results have demonstrated the effectiveness of TriKF on the task of empathetic QG compared with seven baseline models.
Ying Shen 0005, Xuri Chen, Lin Zhang 0014, Shengjie Zhao 0001
IEEE Trans. Comput. Soc. Syst.4
2024 Global Localization in Large-Scale Point Clouds via Roll-Pitch-Yaw Invariant Place Recognition and Low-Overlap Global Registration
abstract
For autonomous ground vehicles, global localization with 3D LiDAR is an indispensable part of tasks such as navigation. Usually, global localization using LiDAR is subdivided into two sub-problems, place recognition and global registration. For place recognition, the recent emerging schemes based on deep learning either rely on 3D convolution with high complexity or need to learn features from various forward perspectives. To mitigate this, we propose a model with roll-pitch-yaw invariance that represents point clouds as probabilistic voxels and generates occupancy grids from a bird’s-eye view, fulfilling robust place recognition by learning aggregated embeddings from a fixed perspective. For low-overlap global registration, the traditional handcraft feature-based methods are mostly limited to dense object-level point clouds, while the state-of-the-art learning-based approaches often rely on complex 3D convolution and additional feature association learning. To fill this gap to some extent, we propose to estimate the relative roll-pitch angles and vertical translation by fitting and aligning the ground plane of the point clouds and to determine the horizontal translations and yaw angle by matching their projected occupancy grids. Extensive experiments corroborate the superior recall and generalization ability of our place recognition model, as well as the advanced success rate and accuracy of our 3D registration approach. Especially in the recognition and registration of hard samples, our results far exceed those of our counterparts by large margins. To ensure full reproducibility, the relevant codes and data are made available online.
Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2024 Ct-LVI: A Framework Toward Continuous-Time Laser-Visual-Inertial Odometry and Mapping
abstract
Owing to the inherent complementarity among LiDAR, camera, and IMU, a growing effort has been paid to laser-visual-inertial SLAM recently. The existing approaches, however, are limited in two aspects. First, at the front-end, they usually employ a discrete-time representation that requires high-precision hardware/software synchronization and are based on geometric laser features, leading to low robustness and scalability. Second, at the backend, visual loop constraints suffer from scale ambiguity and the sparseness of the point cloud deteriorates the scan-to-scan loop detection. To solve these problems, for the front-end, we propose a continuous-time laser-visual-inertial odometry which formulates the carrier trajectory in continuous time, organizes point clouds in probabilistic submaps, and jointly optimizes the loss terms of laser anchors, visual reprojections, and IMU readings, achieving accurate pose estimation even with fast motion or in unstructured scenes where it is difficult to extract meaningful geometric features. At the backend, we propose building 5-DoF laser constraints by matching projected 2D submaps and 6-DoF visual constraints via laser-aided visual relocalization, ensuring mapping consistency in large-scale scenes. Results show that our framework achieves high-precision estimation and is more robust than its counterparts when the carrier works in large scenes or with fast motion. The relevant codes and data are open-sourced at https://cslinzhang.github.io/Ct-LVI/Ct-LVI.html.
Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2024 An Underwater Organism Image Dataset and a Lightweight Module Designed for Object Detection Networks
abstract
Long-term monitoring and recognition of underwater organism objects are of great significance in marine ecology, fisheries science and many other disciplines. Traditional techniques in this field, including manual fishing-based ones and sonar-based ones, are usually flawed. Specifically, the method based on manual fishing is time-consuming and unsuitable for scientific researches, while the sonar-based one, has the defects of low acoustic image accuracy and large echo errors. In recent years, the rapid development of deep learning and its excellent performance in computer vision tasks make vision-based solutions feasible. However, the researches in this area are still relatively insufficient in mainly two aspects. First, to our knowledge, there is still a lack of large-scale datasets of underwater organism images with accurate annotations. Second, in consideration of the limitation on hardware resources of underwater devices, an underwater organism detection algorithm that is both accurate and lightweight enough to be able to infer in real time is still lacking. As an attempt to fill in the aforementioned research gaps to some extent, we established the Multiple Kinds of Underwater Organisms (MKUO) dataset with accurate bounding box annotations of taxonomic information, which consists of 10,043 annotated images, covering eighty-four underwater organism categories. Based on our benchmark dataset, we evaluated a series of existing object detection algorithms to obtain their accuracy and complexity indicators as the baseline for future reference. In addition, we also propose a novel lightweight module, namely Sparse Ghost Module, designed especially for object detection networks. By substituting the standard convolution with our proposed one, the network complexity can be significantly reduced and the inference speed can be greatly improved without obvious detection accuracy loss. To make our results reproducible, the dataset and the source code are available online at https://cslinzhang.github.io/MKUO-and-Sparse-Ghost-Module/ .
Jiafeng Huang, Tianjun Zhang, Shengjie Zhao 0001, Lin Zhang 0014, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.4
2024 I2P Registration by Learning the Underlying Alignment Feature Space from Pixel-to-Point Similarities
abstract
Estimating the relative pose between a camera and a LiDAR holds paramount importance in facilitating complex task execution within multi-agent systems. Nonetheless, current methodologies encounter two primary limitations. First, amid the cross-modal feature extraction, they typically employ separate modal branches to extract cross-modal features from images and point clouds. This approach results in the feature spaces of images and point clouds being misaligned, thereby reducing the robustness of establishing correspondences. Second, due to the scale differences between images and point clouds, one-to-many pixel-point correspondences are inevitably encountered, which will mislead the pose optimization. To address these challenges, we propose a framework named I mage-to- P oint cloud registration by learning the underlying alignment feature space from P ixel-to- P oint SIM imilarities (I2P \({}_{\mathbf{ppsim}}\) ) . Central to \(\text{I2P}_{\text{ppsim}}\) is a Shared Feature Alignment Module (SFAM). It is designed under on a coarse-to-fine architecture and uses a weight-sharing network to construct an alignment feature space. Benefiting from SFAM, \(\text{I2P}_{\text{ppsim}}\) can effectively identify the co-view regions between images and point clouds and establish high-reliability 2D-3D correspondences. Moreover, to mitigate the one-to-many correspondence issue, we introduce a similarity maximization strategy termed point-max. This strategy effectively filters out outliers, thereby establishing accurate 2D-3D correspondences. To evaluate the efficacy of our framework, we conduct extensive experiments on KITTI Odometry and Oxford Robotcar. The results corroborate the effectiveness of our framework in improving image-to-point cloud registration. To make our results reproducible, the source codes have been released at https://cslinzhang.github.io/I2P
Yunda Sun, Lin Zhang 0014, Zhong Wang 0009, Yang Chen 0037, Shengjie Zhao 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Adaptively Hashing 3DLUTs for Lightweight Real-time Image Enhancement
abstract
Image enhancement is an essential and longstanding task in computer vision, in which the 3D lookup table (3DLUT) is widely used due to its powerful mapping capability and high time efficiency. However, the 3DLUT generally requires a large parameter amount since it caches the mapping results for all colors of the entire discrete color space with a 3D array. As a result, standard 3DLUT-based enhancement methods suffer from heavy memory footprints, limiting their practical applications. Based on the analyses of the inherent low grid utilization rate of 3DLUT, we propose HashLUT, an efficient hash form of the standard 3DLUT, and further build a lightweight real-time image enhancement network that adaptively learns HashLUTs and handles hash collisions end-to-end. Experiments on two benchmarks demonstrate that our model achieves comparable enhancement performance to the state-of-the-art methods with significantly fewer parameters. Source codes are available at https://github.com/Xian-Bei.
Lin Zhang 0014, Tianjun Zhang, Dongqing Wang
ICME2
2023 Extrinsic Self-Calibration of the Surround-View System: A Weakly Supervised Approach
abstract
An SVS usually consists of four wide-angle fisheye cameras mounted around the vehicle to sense the surrounding environment. From the images synchronously captured by all cameras, a top-down surround-view can be synthesized, on the premise that both intrinsics and extrinsics of the cameras have been calibrated. At present, the intrinsic calibration approach is relatively well-developed and can be pipelined, while the extrinsic calibration is still immature. On one hand, the existing manual calibration schemes are usually reliable, but need to be conducted by professionals in specific sites, which is undoubtedly cumbersome. On the other hand, the majority of the existing self- calibration schemes are based on low-level features and their stability and robustness are usually unsatisfactory. As far as we know, an effective extrinsic self-calibration scheme designed specially for the SVS is still lacking. To fill such a research gap to some extent, we propose a novel self-calibration scheme which follows a weakly supervised framework, namely WESNet (Weakly-supervised Extrinsic Self-calibration Network). The training of WESNet consists of two stages. First, we utilize the corners in a few calibration site images as the weak supervision to roughly optimize the network by minimizing the geometric loss. Then, after the convergency in the first stage, we additionally introduce a self-supervised photometric loss term that can be constructed by the photometric information from natural images for further fine-tuning. Besides, to support training, we totally collected 19,078 groups of synchronously captured fisheye images under various environmental conditions. To our knowledge, thus far this is the largest surround-view dataset containing original fisheye images. By means of learning prior knowledge from the training data, WESNet takes the original fisheye images synchronously collected as the input, and directly yields extrinsics end-to-end with little labor cost. Its efficiency and efficacy have been corroborated by extensive experiments conducted on our collected dataset. To make our results reproducible, source code and the collected dataset have been released at https://cslinzhang.github.io/WESNet/WESNet.html.
Yang Chen 0037, Lin Zhang 0014, Ying Shen 0005, Brian Nlong Zhao, Yicong Zhou
IEEE Trans. Multim.2
2023 D-LIOM: Tightly-Coupled Direct LiDAR-Inertial Odometry and Mapping
abstract
Simultaneous localization and mapping via LiDAR-Inertial fusion is a crucial technology in many automation-related applications. Recently, a number of approaches based on geometric features have evolved, yielding impressive results via tightly-coupled estimation. This sort of feature-based techniques, however, are inextricably linked to the scanning mechanism of the LiDAR, relying on stable feature detection, and thus are difficult to adapt to multi-LiDAR systems. A few “direct” solutions, on the other hand, register the raw point cloud with the built probability map, which is more computationally efficient and easy to be extended. But, the existing direct approaches are all loosely-coupled, lacking correction of the IMU biases, and thus only work well in 2D cases. To this end, we present D-LIOM, a tightly-coupled Direct LiDAR-Inertial Odometry and Mapping framework. In D-LIOM, a scan is directly registered to a probability submap, and the LiDAR odometry, the IMU pre-integration, and the gravity constraint are integrated to build a local factor graph in the submap's time window, allowing the system to perform real-time high-precision pose estimation. Furthermore, to eliminate accumulated errors in time, we detect loops and adjust the sparse pose graph based on mutual matching of projected 2D submaps, allowing D-LIOM to run stably in large-scale scenes. In addition, to improve its flexibility to varied sensor combinations, D-LIOM supports multi-LiDAR inputs and facilitates the initialization with a common 6-axis IMU. Extensive experiments demonstrate that D-LIOM largely outperforms the existing state-of-the-art counterparts in mapping effect and localization accuracy as well as with high time efficiency. Lastly, to ensure that our results are entirely reproducible, all necessary data and codes are made open-source available. One introduction video can also be found on the online website.
Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
IEEE Trans. Multim.2
2023 SLAM for Indoor Parking: A Comprehensive Benchmark Dataset and a Tightly Coupled Semantic Framework
abstract
For the task of autonomous indoor parking, various Visual-Inertial Simultaneous Localization And Mapping (SLAM) systems are expected to achieve comparable results with the benefit of complementary effects of visual cameras and the Inertial Measurement Units. To compare these competing SLAM systems, it is necessary to have publicly available datasets, offering an objective way to demonstrate the pros/cons of each SLAM system. However, the availability of such high-quality datasets is surprisingly limited due to the profound challenge of the groundtruth trajectory acquisition in the Global Positioning Satellite denied indoor parking environments. In this article, we establish BeVIS, a large-scale Be nchmark dataset with V isual (front-view), I nertial and S urround-view sensors for evaluating the performance of SLAM systems developed for autonomous indoor parking, which is the first of its kind where both the raw data and the groundtruth trajectories are available. In BeVIS, the groundtruth trajectories are obtained by tracking artificial landmarks scattered in the indoor parking environments, whose coordinates are recorded in a surveying manner with a high-precision Electronic Total Station. Moreover, the groundtruth trajectories are comprehensively evaluated in terms of two respects, the reprojection error and the pose volatility, respectively. Apart from BeVIS, we propose a novel tightly coupled semantic SLAM framework, namely VIS SLAM -2, leveraging V isual (front-view), I nertial, and S urround-view sensor modalities, specially for the task of autonomous indoor parking. It is the first work attempting to provide a general form to model various semantic objects on the ground. Experiments on BeVIS demonstrate the effectiveness of the proposed VIS SLAM -2. Our benchmark dataset BeVIS is publicly available at https://shaoxuan92.github.io/BeVIS .
Xuan Shao, Ying Shen 0005, Lin Zhang 0014, Shengjie Zhao 0001, Dandan Zhu 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.3
2022 An Automatical And Efficient Image Classification Based On Improved Genetic Programming
abstract
Image classification is a basic task in machine intelligence, but challenging due to high variations across images. Traditional methods use hand-crafted features to solve it, which require much domain knowledge. Genetic Programming (GP) can automatically solve problems without much knowledge about the structure and form of the solution. And GP is interpretable and needs less time to adjust the parameters compared with deep image classification methods. However, the existing GP-based image classification methods have some disadvantages, such as poor classification performance and long training time. This paper proposed a new image classification algorithm based on multilayer genetic programming with cache (MCGP). MCGP designs a new hierarchical individual program structure with a classification layer and uses a subtree cache strategy to reduce training time. The experiments show that MCGP can get better or competitive results compared with traditional methods, other GP methods, and convolutional neural network methods. In addition, the training speed of MCGP is much faster than other GP methods.
Fazhi He, Lin Zhang 0014
CSCWD4
2022 Self-augmented Unpaired Image Dehazing via Density and Depth Decomposition
abstract
To overcome the overfitting issue of dehazing models trained on synthetic hazy-clean image pairs, many recent methods attempted to improve models' generalization ability by training on unpaired data. Most of them simply formulate dehazing and rehazing cycles, yet ignore the physical properties of the real-world hazy environment, i.e. the haze varies with density and depth. In this paper, we propose a self-augmented image dehazing framework, termed D4 (Dehazing via Decomposing transmission map into Density and Depth) for haze generation and removal. Instead of merely estimating transmission maps or clean content, the proposed framework focuses on exploring scattering coefficient and depth information contained in hazy and clean images. With estimated scene depth, our method is capable of re-rendering hazy images with different thick-nesses which further benefits the training of the dehazing network. It is worth noting that the whole training process needs only unpaired hazy and clean images, yet succeeded in recovering the scattering coefficient, depth map and clean content from a single hazy image. Comprehensive experiments demonstrate our method outperforms state-of-the-art unpaired dehazing methods with much fewer parameters and FLOPs. Our code is available at https://github.com/YaN9-Y/D4.
Risheng Liu, Lin Zhang 0014, Xiaojie Guo 0001, Dacheng Tao
CVPR4
2022 Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated Learning
abstract
Federated Learning (FL) is an emerging distributed learning paradigm under privacy constraint. Data heterogeneity is one of the main challenges in FL, which results in slow convergence and degraded performance. Most existing approaches only tackle the heterogeneity challenge by restricting the local model update in client, ignoring the performance drop caused by direct global model aggregation. Instead, we propose a data-free knowledge distillation method to fine-tune the global model in the server (FedFTG), which relieves the issue of direct model aggregation. Concretely, FedFTG explores the input space of local models through a generator, and uses it to transfer the knowledge from local models to the global model. Besides, we propose a hard sample mining scheme to achieve effective knowledge distillation throughout the training. In addition, we develop customized label sampling and class-level ensemble to derive maximum utilization of knowledge, which implicitly mitigates the distribution discrepancy across clients. Extensive experiments show that our FedFTG significantly outperforms the state-of-the-art (SOTA) FL algorithms and can serve as a strong plugin for enhancing FedAvg, FedProx, FedDyn, and SCAFFOLD.
Lin Zhang 0014, Li Shen 0008, Liang Ding 0006, Dacheng Tao, Ling-Yu Duan
CVPR1
2022 Towards Controllable and Physical Interpretable Underwater Scene Simulation
abstract
The realistic simulation of underwater scenes has important significance for many researches related to underwater vision, such as underwater image restoration, underwater moving object monitoring, etc. To date, however, the existing underwater scene simulation pipelines are either too complicated due to the continuous spectra and camera parameters involved, or difficult to control since the empirically controlled distance based fog effect is usually used by them. In this paper, we try to fill in this research gap by proposing an Underwater Scene Simulation approach, namely USSim, which especially focuses on the influence of ocean water. In USSim, Jerlov water type and depth are regarded as main variables to control the simulation effects. In addition, the spectra of the incident light is decomposed into three primary components and their attenuations are modeled separately, and finally the simulated scene is generated via the hybrid underwater imaging model proposed by us. USSim greatly reduces the computational complexity and enables the fog effect to be controlled by variables with explicit physical meanings. The controllability, physical interpretability and simulation effects of our USSim under different conditions have been verified by extensive experiments. To make our results reproducible, the source code is made online available at https://cslinzhang.github.io/USSim/.
Kaixin Chen 0003, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
ICASSP2
2022 Chunkfusion: A Learning-Based RGB-D 3D Reconstruction Framework Via Chunk-Wise Integration
abstract
Recent years have witnessed a growing interest in online RGB-D 3D reconstruction. On the premise of ensuring the reconstruction accuracy with noisy depth scans, making the system scalable to various environments is still challenging. In this paper, we devote our efforts to try to fill in this research gap by proposing a scalable and robust RGB-D 3D reconstruction framework, namely Chunk-Fusion. In ChunkFusion, sparse voxel management is exploited to improve the scalability of online reconstruction. Besides, a chunk-wise TSDF (truncated signed distance function) fusion network is designed to perform a robust integration of the noisy depth measurements on the sparsely allocated voxel chunks. The proposed chunk-wise TSDF integration scheme can accurately restore surfaces with superior visual consistency from noisy depth maps and can guarantee the scalability of online reconstruction simultaneously, making our reconstruction framework widely applicable to scenes with various scales and depth scans with strong noises and outliers. The outstanding scalability and efficacy of our ChunkFusion have been corroborated by extensive experiments. To make our results reproducible, the source code is made online available at https://cslinzhang.github.io/ChunkFusion/.
Chaozheng Guo, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
ICASSP2
2022 Towards Underwater Image Restoration: A Physical-Accurate Pipeline and a Large Scale Full-Reference Benchmark
abstract
Underwater images always present low-quality features such as low contrast, blurred edges and color distortion, which brings great challenges to high-level underwater vision tasks. In this paper, a novel underwater image restoration method, namely MonoUIR (Monocular Underwater Image Restoration), is proposed, which is based on a more physical-accurate imaging model compared to existing schemes. And with the monocular depth estimation, MonoUIR has no dependence on extra ranging equipment or specific shooting operations. Experimental results demonstrate that MonoUIR overwhelmingly outperforms other physical model-based competitors. In addition, the Real-world Undersea Color Board (RUCB) dataset is established, providing the illconditioned underwater images collected in the East China Sea and the corresponding high-quality references. To our knowledge, this is the first full-reference underwater benchmark dataset collected entirely in a real-world marine environment, which will further support the fullreference evaluation of underwater image restoration approaches. The source code and the dataset are available at https://TongJiayan.github.io/MonoUIR-Homepage.
Jiayan Tong, Tianjun Zhang, Lin Zhang 0014
ICME3
2022 SiD-WaveFlow: A Low-Resource Vocoder Independent of Prior Knowledge
Ying Shen 0005, Dongqing Wang, Lin Zhang 0014
INTERSPEECH4
2022 LVI-ExC: A Target-free LiDAR-Visual-Inertial Extrinsic Calibration Framework
abstract
Recently, the multi-modal fusion with 3D LiDAR, camera, and IMU has shown great potential in applications of automation-related fields. Yet a prerequisite for a successful fusion is that the geometric relationships among the sensors are accurately determined, which is called an extrinsic calibration problem. To date, the existing target-based approaches to deal with this problem rely on sophisticated calibration objects (sites) and well-trained operators, which is time-consuming and inflexible in practical applications. Contrarily, a few target-free methods can overcome these shortcomings, while they only focus on the calibrations of two types of the sensors. Although it is possible to obtain LiDAR-visual-inertial extrinsics by chained calibrations, problems such as cumbersome operations, large cumulative errors, and weak geometric consistency still exist. To this end, we propose LVI-ExC, an integrated LiDAR-Visual-Inertial Extrinsic Calibration framework, which takes natural multi-modal data as input and yields sensor-to-sensor extrinsics end-to-end without any auxiliary object (site) or manual assistance. To fuse multi-modal data, we formulate the LiDAR-visual-inertial extrinsic calibration as a continuous-time simultaneous localization and mapping problem, in which the extrinsics, trajectories, time differences, and map points are jointly estimated by establishing sensor-to-sensor and sensor-to-trajectory constraints. Extensive experiments show that LVI-ExC can produce precise results. With LVI-ExC's outputs, the LiDAR-visual reprojection results and the reconstructed environment map are all highly consistent with the actual natural scenes, demonstrating LVI-ExC's outstanding performance. To ensure that our results are fully reproducible, all the relevant data and codes have been released publicly at https://cslinzhang.github.io/LVI-ExC/.
Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
ACM Multimedia2
2022 CLUT-Net: Learning Adaptively Compressed Representations of 3DLUTs for Lightweight Image Enhancement
abstract
Learning-based image enhancement has made great progress recently, among which the 3-Dimensional LookUp Table (3DLUT) based methods achieve a good balance between enhancement performance and time-efficiency. Generally, the more basis 3DLUTs are used in such methods, the more application scenarios could be covered, and thus the stronger enhancement capability could be achieved. However, more 3DLUTs would also lead to the rapid growth of the parameter amount, since a single 3DLUT has as many as D3 parameters where D is the table length. A large parameter amount not only hinders the practical application of the 3DLUT-based schemes but also gives rise to the training difficulty and does harm to the effectiveness of the basis 3DLUTs, leading to even worse performances with more utilized 3DLUTs. Through in-depth analysis of the inherent compressibility of 3DLUT, we propose an effective Compressed representation of 3-dimensional LookUp Table (CLUT) which maintains the powerful mapping capability of 3DLUT but with a significantly reduced parameter amount. Based on CLUT, we further construct a lightweight image enhancement network, namely CLUT-Net, in which image-adaptive and compression-adaptive CLUTs are learned in an end-to-end manner. Extensive experimental results on three benchmark datasets demonstrate that our proposed CLUT-Net outperforms the existing state-of-the-art image enhancement methods with orders of magnitude smaller parameter amounts. The source codes are available at https://github.com/Xian-Bei/CLUT-Net.
Tianjun Zhang, Lin Zhang 0014
ACM Multimedia4
2022 MOFISSLAM: A Multi-Object Semantic SLAM System With Front-View, Inertial, and Surround-View Sensors for Indoor Parking
abstract
The semantic SLAM (Simultaneous Localization And Mapping) system is a crucial module for autonomous indoor parking. Visual cameras (monocular/binocular) and IMU (Inertial Measurement Unit) constitute the basic configuration to build such a system. The performance of existing SLAM systems typically deteriorates in the presence of dynamically movable objects or objects with little texture. By contrast, semantic objects on the ground embody the most salient and stable features in the indoor parking environment. Due to their inabilities to perceive such features on the ground, existing SLAM systems are prone to tracking inconsistency during navigation. In this paper, we present MOFISSLAM, a novel tightly-coupled${M}$ulti-${O}$bject semantic SLAM system integrating${F}$ront-view,${I}$nertial, and${S}$urround-view sensors for autonomous indoor parking. The proposed system moves beyond existing semantic SLAM systems by complementing the sensor configuration with a surround-view system capturing images from a top-down viewpoint. In MOFISSLAM, apart from low-level visual features and inertial motion data, typical semantic objects (parking-slots, parking-slot IDs and speed bumps) detected in surround-views are also incorporated in optimization, forming robust surround-view constraints. Specifically, each surround-view feature imposes a surround-view constraint that can be split into a contact term and a registration term. The former pre-defines the position of each individual surround-view feature subject to whether it has semantic contact with other surround-view features. Three contact modes, defined ascomplementary,adjacentandcoincident, are identified to guarantee a unified form of all contact terms. The latter further constrains by registering each surround-view observation and its position in the world coordinate system. In parallel, to objectively evaluate SLAM studies for autonomous indoor parking, a large-scale dataset with groundtruth trajectories is collected, which is the first of its kind. Its groundtruth trajectories, commonly unavailable, are obtained by tracking artificial features scattered in the indoor parking environment, whose 3D coordinates are measured with an ETS (Electronic Total Station). The collected dataset has been made publicly available athttps://shaoxuan92.github.io/MOFIS.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2022 Progressive Motion Coherence for Remote Sensing Image Matching
abstract
In this article, we present a feature-based remote sensing (RS) image matching method termed progressive motion coherence (PMC). We formulate the matching problem into a mathematical model and derive a closed-form solution. The objective function is only based on two novel coherence constraints, namely, efficient neighborhood element coherence and relative order-aware motion coherence, and hence, it is general enough and can be applied to RS image matching with different image types and degradations. The efficient neighborhood element coherence uses the Jaccard distance to measure the dissimilarity of two neighborhoods, which are lists composed of$k$nearest neighbors of feature points. To prevent overpenalization on the outliers, we combine it with an exponential function, which is simple yet efficient. The relative order-aware motion coherence is an alternative to motion smoothness, which is based on the observation that the relative order of neighboring matches for inliers in a small region can be well preserved, while for outliers, the relative order changes greatly. The above two coherences are robust to large rotation changes and low ratio inliers. Extensive experiments on five RS image datasets compared with seven state of the arts demonstrate that our PMC is more efficient and robust than the competitors.
Yizhang Liu, Brian Nlong Zhao, Shengjie Zhao 0001, Lin Zhang 0014
IEEE Trans. Geosci. Remote. Sens.4
2022 CVIDS: A Collaborative Localization and Dense Mapping Framework for Multi-Agent Based Visual-Inertial SLAM
abstract
Nowadays, visual SLAM (Simultaneous Localization And Mapping) has become a hot research topic due to its low costs and wide application scopes. Traditional visual SLAM frameworks are usually designed for single-agent systems, completing both the localization and the mapping with sensors equipped on a single robot or a mobile device. However, the mobility and work capacity of the single agent are usually limited. In reality, robots or mobile devices sometimes may be deployed in the form of clusters, such as drone formations, wearable motion capture systems, and so on. As far as we know, existing SLAM systems designed for multi-agents are still sporadic, and most of them have non-negligible limitations in functions. Specifically, on one hand, most of the existing multi-agent SLAM systems can only extract some key features and build sparse maps. On the other hand, schemes that can reconstruct the environment densely cannot get rid of the dependence on depth sensors, such as RGBD cameras or LiDARs. Systems that can yield high-density maps just with monocular camera suites are temporarily lacking. As an attempt to fill in the research gap to some extent, we design a novel collaborative SLAM system, namely CVIDS (Collaborative Visual-Inertial Dense SLAM), which follows a centralized and loosely coupled framework and can be integrated with any existing Visual-Inertial Odometry (VIO) to accomplish the co-localization and the dense reconstruction. Integrating our proposed robust loop closure detection module and two-stage pose-graph optimization pipeline, the co-localization module of CVIDS can estimate the poses of different agents in a unified coordinate system efficiently from the packed images and local poses sent by the client-ends of different agents. Besides, our motion-based dense mapping module can effectively recover the 3D structures of selected keyframes and then fuse their depth information to the global map for reconstruction. The superior performance of CVIDS is corroborated by both quantitative and qualitative experimental results. To make our results reproducible, the source code has been released at https://cslinzhang.github.io/CVIDS.
Tianjun Zhang, Lin Zhang 0014, Yang Chen 0037, Yicong Zhou
IEEE Trans. Image Process.2
2022 Intrinsic Performance Influence-based Participant Contribution Estimation for Horizontal Federated Learning
abstract
The rapid development of modern artificial intelligence technique is mainly attributed to sufficient and high-quality data. However, in the data collection, personal privacy is at risk of being leaked. This issue can be addressed by federated learning, which is proposed to achieve efficient model training among multiple data providers without direct data access and aggregation. To encourage more parties owning high-quality data to participate in the federated learning, it is important to evaluate and reward the participant contribution in a reasonable, robust, and efficient manner. To achieve this goal, we propose a novel contribution estimation method: Intrinsic Performance Influence-based Contribution Estimation (IPICE). In particular, the class-level intrinsic performance influence is adopted as the contribution estimation criteria in IPICE, and a neural network is employed to exploit the non-linear relationship between the performance change and estimated contribution. Extensive experiments are conducted on various datasets, and the results demonstrate that IPICE is more accurate and stable than the counterpart in various data distribution settings. The computational complexity is significantly reduced in our IPICE, especially when a new party joins the federation. IPICE assigns small contributions to bad/garbage data and thus prevent them from participating and deteriorating the learning ecosystem.
Lin Zhang 0014, Lixin Fan, Yong Luo 0002, Ling-Yu Duan
ACM Trans. Intell. Syst. Technol.1
2022 Online Correction of Camera Poses for the Surround-view System: A Sparse Direct Approach
abstract
The surround-view module is an indispensable component of a modern advanced driving assistance system. By calibrating the intrinsics and extrinsics of the surround-view cameras accurately, a top-down surround-view can be generated from raw fisheye images. However, poses of these cameras sometimes may change. At present, how to correct poses of cameras in a surround-view system online without re-calibration is still an open issue. To settle this problem, we introduce the sparse direct framework and propose a novel optimization scheme of a cascade structure. This scheme is actually composed of two levels of optimization and two corresponding photometric error based models are proposed. The model for the first-level optimization is called the ground model, as its photometric errors are measured on the ground plane. For the second level of the optimization, it’s based on the so-called ground-camera model, in which photometric errors are computed on the imaging planes. With these models, the pose correction task is formulated as a nonlinear least-squares problem to minimize photometric errors in overlapping regions of adjacent bird’s-eye-view images. With a cascade structure of these two levels of optimization, an appropriate balance between the speed and the accuracy can be achieved. Experiments show that our method can effectively eliminate the misalignment caused by cameras’ moderate pose changes in the surround-view system. Source code and test cases are available online at https://cslinzhang.github.io/CamPoseCorrection/ .
Tianjun Zhang, Hao Deng 0002, Lin Zhang 0014, Shengjie Zhao 0001, Xiao Liu 0030, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Federated Learning for Non-IID Data via Unified Feature Learning and Optimization Objective Alignment
abstract
Federated Learning (FL) aims to establish a shared model across decentralized clients under the privacy-preserving constraint. Despite certain success, it is still challenging for FL to deal with non-IID (non-independent and identical distribution) client data, which is a general scenario in real-world FL tasks. It has been demonstrated that the performance of FL will be reduced greatly under the non-IID scenario, since the discrepant data distributions will induce optimization inconsistency and feature divergence issues. Besides, naively minimizing an aggregate loss function in this scenario may have negative impacts on some clients and thus deteriorate their personal model performance. To address these issues, we propose a Unified Feature learning and Optimization objectives alignment method (FedUFO) for non-IID FL. In particular, an adversary module is proposed to reduce the divergence on feature representation among different clients, and two consensus losses are proposed to reduce the inconsistency on optimization objectives from two perspectives. Extensive experiments demonstrate that our FedUFO can outperform the state-of-the-art approaches, including the competitive one data-sharing method. Besides, FedUFO can enable more reasonable and balanced model performance among different clients.
Lin Zhang 0014, Yong Luo 0002, Bo Du 0001, Ling-Yu Duan
ICCV1
2021 ROECS: A Robust Semi-direct Pipeline Towards Online Extrinsics Correction of the Surround-view System
abstract
Generally, a surround-view system (SVS), which is an indispensable component of advanced driving assistant systems (ADAS), consists of four to six wide-angle fisheye cameras. As long as both intrinsics and extrinsics of all cameras have been calibrated, a top-down surround-view with the real scale can be synthesized at runtime from fisheye images captured by these cameras. However, when the vehicle is driving on the road, relative poses between cameras in the SVS may change from the initial calibrated states due to bumps or collisions. In case that extrinsics' representations are not adjusted accordingly, on the surround-view, obvious geometric misalignment will appear. Currently, the researches on correcting the extrinsics of the SVS in an online manner are quite sporadic, and a mature and robust pipeline is still lacking. As an attempt to fill this research gap to some extent, in this work, we present a novel extrinsics correction pipeline designed specially for the SVS, namely ROECS (Robust Online Extrinsics Correction of the Surround-view system). Specifically, a "refined bi-camera error" model is firstly designed. Then, by minimizing the overall "bi-camera error" within a sparse and semi-direct framework, the SVS's extrinsics can be iteratively optimized and become accurate eventually. Besides, an innovative three-step pixel selection strategy is also proposed. The superior robustness and the generalization capability of ROECS are validated by both quantitative and qualitative experimental results. To make the results reproducible, the collected data and the source code have been released at https://cslinzhang.github.io/ROECS/.
Tianjun Zhang, Brian Nlong Zhao, Ying Shen 0005, Xuan Shao, Lin Zhang 0014, Yicong Zhou
ACM Multimedia5
2021 Simulation of Atmospheric Visibility Impairment
abstract
Changes in aerosol composition and its proportions can cause changes in atmospheric visibility. Vision systems deployed outdoors must take into account the negative effects brought by visibility impairment. In order to develop vision algorithms that can adapt to low atmospheric visibility conditions, a large-scale dataset containing pairs of clear images and their visibility-impaired versions (along with other annotations if necessary) is usually indispensable. However, it is almost impossible to collect large amounts of such image pairs in a real physical environment. A natural and reasonable solution is to use virtual simulation technologies, which is also the focus of this paper. In this paper, we first deeply analyze the limitations and irrationalities of the existing work specializing on simulation of atmospheric visibility impairment. We point out that many simulation schemes actually even violate the assumptions of the Koschmieder's law. Second, more importantly, based on a thorough investigation of the relevant studies in the field of atmospheric science, we present simulation strategies for five most commonly encountered visibility impairment phenomena, including mist, fog, natural haze, smog, and Asian dust. Our work establishes a direct link between the fields of atmospheric science and computer vision. In addition, as a byproduct, with the proposed simulation schemes, a large-scale synthetic dataset is established, comprising 40,000 clear source images and their 800,000 visibility-impaired versions. To make our work reproducible, source codes and the dataset have been released at https://cslinzhang.github.io/AVID/.
Lin Zhang 0014, Shiyu Zhao 0001, Yicong Zhou
IEEE Trans. Image Process.1
2021 RefineDNet: A Weakly Supervised Refinement Framework for Single Image Dehazing
abstract
Haze-free images are the prerequisites of many vision systems and algorithms, and thus single image dehazing is of paramount importance in computer vision. In this field, prior-based methods have achieved initial success. However, they often introduce annoying artifacts to outputs because their priors can hardly fit all situations. By contrast, learning-based methods can generate more natural results. Nonetheless, due to the lack of paired foggy and clear outdoor images of the same scenes as training samples, their haze removal abilities are limited. In this work, we attempt to merge the merits of prior-based and learning-based approaches by dividing the dehazing task into two sub-tasks, i.e., visibility restoration and realness improvement. Specifically, we propose a two-stage weakly supervised dehazing framework, RefineDNet. In the first stage, RefineDNet adopts the dark channel prior to restore visibility. Then, in the second stage, it refines preliminary dehazing results of the first stage to improve realness via adversarial learning with unpaired foggy and clear images. To get more qualified results, we also propose an effective perceptual fusion strategy to blend different dehazing outputs. Extensive experiments corroborate that RefineDNet with the perceptual fusion has an outstanding haze removal capability and can also produce visually pleasing results. Even implemented with basic backbone networks, RefineDNet can outperform supervised dehazing approaches as well as other state-of-the-art methods on indoor and outdoor datasets. To make our results reproducible, relevant code and data are available at https://github.com/xiaofeng94/RefineDNet-for-dehazing.
Shiyu Zhao 0001, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
IEEE Trans. Image Process.2
2021 Learning Compact Multifeature Codes for Palmprint Recognition From a Single Training Image per Palm
abstract
In this article, we propose a multifeature learning method to jointly learn compact multifeature codes (LCMFCs) for palmprint recognition with a single training sample per palm. Unlike most existing hand-crafted methods that extract single-type features from raw pixels, we first form the multi-type data vectors such as the direction-data, and texture-data to completely sample the multiple information of a palmprint image. Then, we learn the discriminative multifeatures from multi-type data vectors by maximizing the inter-palm distance, and minimizing the energy loss between the learned codes, and the original data. Moreover, our LCMFC method adaptively learns the optimal weights of multi-type features to jointly learn the compact multifeature codes. Finally, we cluster the nonoverlapping blockwise histograms of the compact multifeature codes into a feature vector for palmprint representation. Extensive experimental results on six benchmark palmprint databases are presented to show the effectiveness of the proposed method.
Lunke Fei, Bob Zhang 0001, Lin Zhang 0014, Wei Jia 0001, Jie Wen 0001, Jigang Wu
IEEE Trans. Multim.3
2021 Pedestrian-Aware Panoramic Video Stitching Based on a Structured Camera Array
abstract
The panorama stitching system is an indispensable module in surveillance or space exploration. Such a system enables the viewer to understand the surroundings instantly by aligning the surrounding images on a plane and fusing them naturally. The bottleneck of existing systems mainly lies in alignment and naturalness of the transition of adjacent images. When facing dynamic foregrounds, they may produce outputs with misaligned semantic objects, which is evident and sensitive to human perception. We solve three key issues in the existing workflow that can affect its efficiency and the quality of the obtained panoramic video and present Pedestrian360, a panoramic video system based on a structured camera array (a spatial surround-view camera system). First, to get a geometrically aligned 360○ view in the horizontal direction, we build a unified multi-camera coordinate system via a novel refinement approach that jointly optimizes camera poses. Second, to eliminate the brightness and color difference of images taken by different cameras, we design a photometric alignment approach by introducing a bias to the baseline linear adjustment model and solving it with two-step least-squares. Third, considering that the human visual system is more sensitive to high-level semantic objects, such as pedestrians and vehicles, we integrate the results of instance segmentation into the framework of dynamic programming in the seam-cutting step. To our knowledge, we are the first to introduce instance segmentation to the seam-cutting problem, which can ensure the integrity of the salient objects in a panorama. Specifically, in our surveillance oriented system, we choose the most significant target, pedestrians, as the seam avoidance target, and this accounts for the name Pedestrian360 . To validate the effectiveness and efficiency of Pedestrian360, a large-scale dataset composed of videos with pedestrians in five scenes is established. The test results on this dataset demonstrate the superiority of Pedestrian360 compared to its competitors. Experimental results show that Pedestrian360 can stitch videos at a speed of 12 to 26 fps, which depends on the number of objects in the shooting scene and their frequencies of movements. To make our reported results reproducible, the relevant code and collected data are publicly available at https://cslinzhang.github.io/Pedestrian360-Homepage/ .
Lin Zhang 0014, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.2
2020 A Study Of Parking-Slot Detection With The Aid Of Pixel-Level Domain Adaptation
abstract
The self-parking system is an important component of self-driving vehicles. Such a system needs to detect and locate the parking-slots from surround-view images, and then guide the vehicle to the designated parking-slot. In the real world, the appearances and environmental conditions of parking-slots can be rich and varied. Thus, to train the parking-slot detection model, it is necessary to collect and label a huge quantity of surround-view images covering as many real cases as possible. Such a process is cumbersome and costly, and will be repeated whenever encountering an unseen parking condition that is quite different from the ones covered by existing training set. To this end, in this paper we propose an extensible pipeline, namely FakePS, to assist parking-slot detection model training by making use of synthetic data. Specifically, with FakePS, we can first build various simulated parking scenes and collect labeled surround-view images automatically. Besides, we resort to pixel-level domain adaptation strategies to enhance the realism of the synthetic images using unlabeled real images while preserving their label information. The efficacy of FakePS has been corroborated by experimental results.
Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou
ICME2
2020 Oecs: Towards Online Extrinsics Correction For The Surround-View System
abstract
A typical surround-view system consists of four fisheye cameras. By performing an offline calibration that determines both the intrinsics and extrinsics of the system, surround-view images can be synthesized at runtime. However, poses of calibrated cameras sometimes may change. In such a case, if cameras' extrinsics are not updated accordingly, observable geometric misalignment will appear in surround-views. Most existing solutions to this problem resort to re-calibration, which is quite cumbersome. Thus, how to correct cameras' extrinsics in an online manner without using re-calibration is still an open issue. In this paper, we attempt to propose a novel solution to this problem and the proposed solution is referred to as “Online Extrinsics Correction for the Surround-view system OECS for short. We first design a Bi-Camera error model, measuring the photometric discrepancy between two corresponding pixels on images captured by two adjacent cameras. Then, by minimizing the system's overall BiCamera error, cameras' extrinsics can be optimized and the optimization is conducted within a sparse direct framework. The efficacy and efficiency of OECS are validated by experiments. Data and source code used in this work are publicly available at https://z619850002.github.io/OECage/.
Tianjun Zhang, Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou
ICME2
2020 Zero-Shot Restoration of Underexposed Images via Robust Retinex Decomposition
abstract
Underexposed images often suffer from serious quality degradation such as poor visibility and latent noise in the dark. Most previous methods for underexposed images restoration ignore the noise and amplify it during stretching contrast. We predict the noise explicitly to achieve the goal of denoising while restoring the underexposed image. Specifically, a novel three-branch convolution neural network, namely RRDNet (short for Robust Retinex Decomposition Network), is proposed to decompose the input image into three components, illumination, reflectance and noise. As an image-specific network, RRDNet doesn't need any prior image examples or prior training. Instead, the weights of RRDNet will be updated by a zero-shot scheme of iteratively minimizing a specially designed loss function. Such a loss function is devised to evaluate the current decomposition of the test image and guide noise estimation. Experiments demonstrate that RRDNet can achieve robust correction with overall naturalness and pleasing visual quality. To make the results reproducible, the source code has been made publicly available at https://aaaaangel.github.io/RRDNet-Homepage.
Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou
ICME2
2020 A Tightly-coupled Semantic SLAM System with Visual, Inertial and Surround-view Sensors for Autonomous Indoor Parking
abstract
The semantic SLAM (simultaneous localization and mapping) system is an indispensable module for autonomous indoor parking. Monocular and binocular visual cameras constitute the basic configuration to build such a system. Features used in existing SLAM systems are often dynamically movable, blurred and repetitively textured. By contrast, semantic features on the ground are more stable and consistent in the indoor parking environment. Due to their inabilities to perceive salient features on the ground, existing SLAM systems are prone to tracking loss during navigation. Therefore, a surround-view camera system capturing images from a top-down viewpoint is necessarily called for. To this end, this paper proposes a novel tightly-coupled semantic SLAM system by integrating Visual, Inertial, and Surround-view sensors, VIS SLAM for short, for autonomous indoor parking. In VIS SLAM, apart from low-level visual features and IMU (inertial measurement unit) motion data, parking-slots in surround-view images are also detected and geometrically associated, forming semantic constraints. Specifically, each parking-slot can impose a surround-view constraint that can be split into an adjacency term and a registration term. The former pre-defines the position of each individual parking-slot subject to whether it has an adjacent neighbor. The latter further constrains by registering between each observed parking-slot and its position in the world coordinate system. To validate the effectiveness and efficiency of VIS SLAM, a large-scale dataset composed of synchronous multi-sensor data collected from typical indoor parking sites is established, which is the first of its kind. The collected dataset has been made publicly available at https://cslinzhang.github.io/VISSLAM/.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Hongyu Li 0001, Yicong Zhou
ACM Multimedia2
2020 Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and Baselines
abstract
On benchmark images, modern dehazing methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. Thus, it is imperative to adopt quantitative evaluation on a vast number of hazy images. However, existing quantitative evaluation schemes are not convincing due to a lack of appropriate datasets and poor correlations between metrics and human perceptions. In this work, we attempt to address these issues, and we make two contributions. First, we establish two benchmark datasets, i.e., the BEnchmark Dataset for Dehazing Evaluation (BeDDE) and the EXtension of the BeDDE (exBeDDE), which had been lacking for a long period of time. The BeDDE is used to evaluate dehazing methods via full reference image quality assessment (FR-IQA) metrics. It provides hazy images, clear references, haze level labels, and manually labeled masks that indicate the regions of interest (ROIs) in image pairs. The exBeDDE is used to assess the performance of dehazing evaluation metrics. It provides extra dehazed images and subjective scores from people. To the best of our knowledge, the BeDDE is the first dehazing dataset whose image pairs were collected in natural outdoor scenes without any simulation. Second, we provide a new insight that dehazing involves two separate aspects, i.e., visibility restoration and realness restoration, which should be evaluated independently; thus, to characterize them, we establish two criteria, i.e., the visibility index (VI) and the realness index (RI), respectively. The effectiveness of the criteria is verified through extensive experiments. Furthermore, 14 representative dehazing methods are evaluated as baselines using our criteria on BeDDE. Our datasets and relevant code are available at https://github.com/xiaofeng94/BeDDE-for-defogging.
Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001
IEEE Trans. Image Process.2
2019 Seamless 3D Surround View with a Novel Burger Model
abstract
In recent years, the 3D surround view (3D-SV) system has become a hot research topic in the field of Advanced Driver Assistance Systems (ADAS). It can be used to form a stereoscopic view of the surrounding 3D environment by using 4 car-mounted surround cameras, and users can switch the viewpoints for virtual observation conveniently. However, there are still many problems in how to stitch calibrated images to the panoramic view and how to project the panorama to the surround view. In this paper, we introduce the graph cut algorithm and multi-band blending to the panorama stitching phase. In addition, we design a new hamburger-shaped 3D geometric model to be the carrier of the panorama for texture mapping. Our 3D-SV system can make drivers have an immersive visual experience. Experimental results show that the 3D-SV generated by our method is less distorted and looks more natural than the other competitors.
Lin Zhang 0014, Ying Shen 0005, Shengjie Zhao 0001
ICIP1
2019 DMPR-PS: A Novel Approach for Parking-Slot Detection Using Directional Marking-Point Regression
abstract
The self-parking system plays an important role in autonomous driving, and one of its critical issues is parking-slot detection. Previous studies in this field are mostly based on off-the-shelf models designed for universal purposes, which have various limitations in solving specific problems. In this paper, we propose a parking-slot detection method using directional marking-point regression, namely DMPR-PS. Instead of utilizing multiple off-the-shelf models, DMPR-PS uses a novel CNN-based model specially designed for directional marking-point regression. Given a surround-view image I, the model predicts position, shape and orientation of each marking-point on I. From marking-points, parking-slots on I could be easily inferred using geometric rules. DMPR-PS outperforms state-of-the-art competitors on the benchmark dataset with a precision rate of 99.42% and a recall rate of 99.37%, while achieving a real-time detection speed of 12ms per frame on Nvidia Titan Xp. To make the results reproducible, the source code is available at https://github.com/Teoge/DMPR-PS.
Lin Zhang 0014, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang
ICME2
2019 Revisit Surround-view Camera System Calibration
abstract
The surround-view system is an essential component of an advanced driver assistance system especially when the vehicle runs in tight parking space or on a narrow road. To ensure successful maneuvering, a panoramic bird's-eye image with no blind spots is necessarily called for. Hence, a typical surround-view system consists of several cameras mounted around the vehicle capturing images from a top-down viewpoint, and an accurate extrinsic calibration for such system is prerequisite for providing a seamless surround-view image. To achieve this goal, this paper presents a novel extrinsic calibration pipeline which is both easy-to-use and reliable to operate on multiple cameras. Instead of taking the vehicle to a fixed position in a specific calibration site, a single chessboard is the only demand. We adopt a novel refinement procedure that jointly optimizes camera poses in a closed-loop manner. The effectiveness and efficiency of the proposed pipeline to calibrate a surround-view camera system has been corroborated by experiments.
Xuan Shao, Xiao Liu 0030, Lin Zhang 0014, Shengjie Zhao 0001, Ying Shen 0005, Yukai Yang
ICME3
2019 From Market to Dish: Multi-ingredient Image Recognition for Personalized Recipe Recommendation
abstract
Recognition of food ingredients enables applications on recipe recommendation for developing a healthier eating habit. Existing ingredients recognition methods largely rely on ideal images captured in a controlled environment, while ingredients are usually displayed unorderly in a complex environment in the market. We propose the multi-ingredient recognition problem in the market and develop a Spatial Regularization Network (SRN) based method to solve it by using a newly collected multiple vegetable image dataset captured in the market. We further use the recognition result to develop a recipe recommendation system to satisfy the daily nutrition requirements and individual preference of each user. Experiments show that our multi-ingredient recognition outperforms previous methods over 14% in mAP and recommendation model shows an improvement of over 23% in HR@10.
Lin Zhang 0014, Jianbo Zhao 0002, Si Li 0001, Boxin Shi, Ling-Yu Duan
ICME1
2019 Pay By Showing Your Palm: A Study of Palmprint Verification on Mobile Platforms
abstract
With the fast development of smart mobile devices, mobile phones have gradually become an indispensable part of people's lives. Many biometric technologies based on mobile platforms have also developed rapidly, such as face verification and fingerprint recognition. However, the great potential of palmprint has been neglected. In this paper, we conducted a thorough study of palmprint verification on mobile devices for the first time. Firstly, we established an annotated, palmprint dataset named MPD, which was collected by multi-brands phones in two different sessions. As the largest dataset in this field, MPD contains 16,000 palm images from 200 subjects. Secondly, we built a DCNN-based palmprint verification system named DeepMPV for mobile platforms. The efficiency and performance of our system have been corroborated on our collected dataset. The labelled dataset and the source code are publicly available at https://cslinzhang.github.io/deepmpv/.
Lin Zhang 0014, Xiao Liu 0030, Shengjie Zhao 0001, Ying Shen 0005, Yukai Yang
ICME2
2019 Evaluation of Defogging: A Real-World Benchmark Dataset, A New Criterion and Baselines
abstract
Modern defogging methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. On the other hand, existing quantitative evaluation methods are also not convincing due to a lack of proper datasets. In this work, we attempt to address these issues and establish a long-term lacking benchmark dataset, namely BeDDE (BEnchmark Dataset for Defogging Evaluation), for evaluating the performance of defogging algorithms. To our knowledge, BeDDE is the first real-world dataset comprising foggy images with their registered clear counterparts. Using BeDDE, we set up a new criterion for evaluating defogging methods where VSI, a full reference image quality assessment metric, is calculated and averaged on registered ROIs of all image pairs. The evaluation results of the proposed criterion correlate well with human judgements. 10 state-of-the-art defogging methods are evaluated as baselines on BeDDE. BeDDE is available online.
Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang
ICME2
2019 Online Camera Pose Optimization for the Surround-view System
abstract
Surround-view system is an important information medium for drivers to monitor the driving environment. A typical surround-view system consists of four to six fish-eye cameras arranged around the vehicle. From these camera inputs, a top-down image of the ground around the vehicle, namely the surround-view image can be generated with well calibrated camera poses. Although existing surround-view system solutions can estimate camera poses accurately in off-line environment, how to correct the camera poses' change in online environment is still an open issue. In this paper, we propose a camera pose optimization method for surround-view system in online environment. Our method consists of two models: Ground Model and Ground-Camera Model, both of which correct the camera poses by minimizing photometric errors between ground projections of adjacent cameras. Experiments show that our method can effectively correct the geometric misalignment of the surround-view image caused by camera poses' change. Since our method is highly automated with low requirement of calibration site and manual operation, it has a wide range of applications and is convenient for the end-users. To make the results reproducible, the source code is publicly available at https://cslinzhang.github.io/CamPoseOpt/.
Xiao Liu 0030, Lin Zhang 0014, Ying Shen 0005, Shaoming Zhang, Shengjie Zhao 0001
ACM Multimedia2
2019 Zero-Shot Restoration of Back-lit Images Using Deep Internal Learning
abstract
How to restore back-lit images still remains a challenging task. State-of-the-art methods in this field are based on supervised learning and thus they are usually restricted to specific training data. In this paper, we propose a "zero-shot" scheme for back-lit image restoration, which exploits the power of deep learning, but does not rely on any prior image examples or prior training. Specifically, we train a small image-specific CNN, namely ExCNet (short for Exposure Correction Network) at test time, to estimate the "S-curve" that best fits the test back-lit image. Once the S-curve is estimated, the test image can be then restored straightforwardly. ExCNet can adapt itself to different settings per image. This makes our approach widely applicable to different shooting scenes and kinds of back-lighting conditions. Statistical studies performed on 1512 real back-lit images demonstrate that our approach can outperform the competitors by a large margin. To the best of our knowledge, our scheme is the first unsupervised CNN-based back-lit image restoration method. To make the results reproducible, the source code is available at https://cslinzhang.github.io/ExCNet/.
Lin Zhang 0014, Lijun Zhang 0005, Xiao Liu 0030, Ying Shen 0005, Shaoming Zhang, Shengjie Zhao 0001
ACM Multimedia1
2018 A CNN-Based Depth Estimation Approach with Multi-scale Sub-pixel Convolutions and a Smoothness Constraint
Shiyu Zhao 0001, Lin Zhang 0014, Ying Shen 0005, Yongning Zhu
ACCV (2)2
2018 Image Exposure Assessment: A Benchmark and a Deep Convolutional Neural Networks Based Model
abstract
In the camera equipment manufacturing industry, the exposure calibration is one of the basic steps for manufacturers to consider before launching their products to the market. To this end, a method that can objectively and automatically assess the exposure levels of images taken by the camera is highly desired. However, few studies have been conducted in this area. In this paper, we attempt to solve this issue to some extent and our contributions are twofold. Firstly, in order to facilitate the study of image exposure assessment, an Image Exposure Database$(IE_{ps}D)$is established. In this database, there are 15, 582 images with various exposure levels, and for each image there is an associated subjective exposure score which could reflect its perceptual exposure level. Secondly, we propose a novel highly accurate DCNN-based model, namely$IE_{ps}M$(Image Exposure Metric), to predict the exposure level of a given image.
Lijun Zhang 0005, Lin Zhang 0014, Xiao Liu 0030, Ying Shen 0005, Dongqing Wang
ICME2
2018 Vision-Based Parking-Slot Detection: A DCNN-Based Approach and a Large-Scale Benchmark Dataset
abstract
In the automobile industry, recent years have witnessed a growing interest in developing self-parking systems. For such systems, how to accurately and efficiently detect and localize the parking-slots defined by regular line segments near the vehicle is a key and still unresolved issue. In fact, kinds of unfavorable factors, such as the diversity of ground materials, changes in illumination conditions, and unpredictable shadows caused by nearby trees, make the vision-based parking-slot detection much harder than it looks. In this paper, we attempt to solve this issue to some extent and our contributions are twofold. First, we propose a novel DCNN (Deep Convolutional Neural Networks) based parking-slot detection approach, namely DeepPS, which takes the surround-view image as the input. There are two key steps in DeepPS, identifying all the marking-points on the input image and classifying local image patterns formed by pairs of markingpoints. We formulate both of them as learning problems, which can be solved naturally by modern DCNN models. Second, to facilitate the study of vision-based parking-slot detection, a largescale labeled dataset is established. This dataset is the largest in this field, comprising 12,165 surround-view images collected from typical indoor and outdoor parking sites. For each image, the marking-points and parking-slots are carefully labeled. The efficacy and efficiency of DeepPS have been corroborated on our collected dataset. To make our results fully reproducible, all the relevant source codes and the dataset have been made publicly available at https://cslinzhang.github.io/deepps/.
Lin Zhang 0014, Xiyuan Li
IEEE Trans. Image Process.1
2017 The Hasp Motif: A New Type of RNA Tertiary Interactions
Ying Shen 0005, Lin Zhang 0014
ICIC (2)2
2017 Vision-based parking-slot detection: A benchmark and a learning-based approach
abstract
Recent years have witnessed a growing interest in developing automatic parking systems in the field of intelligent vehicle. However, how to effectively and efficiently locating parking-slots using a vision-based system is still an unresolved issue. In this paper, we attempt to fill this research gap to some extent and our contributions are twofold. Firstly, to facilitate the study of vision-based parking-slot detection, a large-scale parking-slot image database is established. For each image in this database, the marking-points and parking-slots are carefully labelled. Such a database can serve as a benchmark to design and validate parking-slot detection algorithms. Secondly, a learning based parking-slot detection approach is proposed. With this approach, given a test image, the marking-points will be detected at first and then the valid parking-slots can be inferred. Its efficacy and efficiency have been corroborated on our database. The labeled database and the source codes are publicly available at http://sse.tongji.edu.cn/linzhang/ps/index.htm.
Linshen Li, Lin Zhang 0014, Xiyuan Li, Xiao Liu 0030, Ying Shen 0005
ICME2
2017 Towards Simulating Foggy and Hazy Images and Evaluating Their Authenticity
Ning Zhang 0007, Lin Zhang 0014, Zaixi Cheng
ICONIP (3)2
2017 Illumination Quality Assessment for Face Images: A Benchmark and a Convolutional Neural Networks Based Model
Lijun Zhang 0005, Lin Zhang 0014, Lida Li
ICONIP (3)2
2017 Image set classification based on synthetic examples and reverse training
Lin Zhang 0014, Qingjun Liang, Ying Shen 0005, Meng Yang 0001, Feng Liu 0013
Neurocomputing1
2017 Towards contactless palmprint recognition: A novel device, a new benchmark, and a collaborative representation based identification approach
Lin Zhang 0014, Lida Li, Anqi Yang, Ying Shen 0005, Meng Yang 0001
Pattern Recognit.1
2016 Multi-dictionary Based Collaborative Representation for 3D Biometrics
Anqi Yang, Lin Zhang 0014, Lida Li, Hongyu Li 0001
ICIC (1)2
2016 3D Ear Identification Using Block-Wise Statistics-Based Features and LC-KSVD
abstract
Biometrics authentication has been corroborated to be an effective method for recognizing a person's identity with high confidence. In this field, the use of three-dimensional (3D) ear shape is a recent trend. As a biometric identifier, the ear has several inherent merits. However, although a great deal of efforts have been devoted, there is still large room for improvement in developing a highly effective and efficient 3D ear identification approach. In this paper, we attempt to fill this gap to some extent by proposing a novel 3D ear classification scheme that makes use of the label consistent K-SVD (LC-KSVD) framework. As an effective supervised dictionary learning algorithm, LC-KSVD learns a single compact discriminative dictionary for sparse coding and a multi-class linear classifier simultaneously. To use the LC-KSVD framework, one key issue is how to extract feature vectors from 3D ear scans. To this end, we propose a blockwise statistics-based feature extraction scheme. Specifically, we divide a 3D ear region of interest into uniform blocks and extract a histogram of surface types from each block; histograms from all blocks are then concatenated to form the desired feature vector. Feature vectors extracted in this way are highly discriminative and are robust to mere misalignment between samples. Experiments demonstrate that our approach can achieve better recognition accuracy than the other state-of-the-art methods. More importantly, its computational complexity is extremely low, making it quite suitable for the large-scale identification applications. MATLAB source codes are publicly online available at http://sse.tongji.edu.cn/linzhang/LCKSVDEar/LCKSVDEar. htm.
Lin Zhang 0014, Lida Li, Hongyu Li 0001, Meng Yang 0001
IEEE Trans. Multim.1
2015 Palmprint Recognition Based on Image Sets
Qingjun Liang, Lin Zhang 0014, Hongyu Li 0001
ICIC (1)2
2015 Image Set Classification Based on Synthetic Examples and Reverse Training
Qingjun Liang, Lin Zhang 0014, Hongyu Li 0001
ICIC (3)2
2015 The λ-Turn: A New Structural Motif in Ribosomal RNA
Huizhu Ren, Ying Shen 0005, Lin Zhang 0014
ICIC (2)3
2015 3D ear identification using LC-KSVD and local histograms of surface types
abstract
In this paper, we propose a novel 3D ear classification scheme, making use of the label consistent K-SVD (LC-KSVD) framework. As an effective supervised dictionary learning algorithm, LC-KSVD learns a compact discriminative dictionary for sparse coding and a multi-class linear classifier simultaneously. To use LC-KSVD, one key issue is how to extract feature vectors from 3D ear scans. To this end, we propose a block-wise statistics based scheme. Specifically, we divide a 3D ear ROI into blocks and extract a histogram of surface types from each block; histograms from all blocks are concatenated to form the desired feature vector. Feature vectors extracted in this way are highly discriminative and are robust to mere misalignment. Experimental results demonstrate that the proposed approach can achieve much better recognition accuracy than the other state-of-the-art methods. More importantly, its computational complexity is extremely low at the classification stage.
Lida Li, Lin Zhang 0014, Hongyu Li 0001
ICME2
2015 3D Palmprint Identification Using Block-Wise Features and Collaborative Representation
abstract
Developing 3D palmprint recognition systems has recently begun to draw attention of researchers. Compared with its 2D counterpart, 3D palmprint has several unique merits. However, most of the existing 3D palmprint matching methods are designed for one-to-one verification and they are not efficient to cope with the one-to-many identification case. In this paper, we fill this gap by proposing a collaborative representation (CR) based framework with l1-norm or l2-norm regularizations for 3D palmprint identification. The effects of different regularization terms have been evaluated in experiments. To use the CR-based classification framework, one key issue is how to extract feature vectors. To this end, we propose a block-wise statistics based feature extraction scheme. We divide a 3D palmprint ROI into uniform blocks and extract a histogram of surface types from each block; histograms from all blocks are then concatenated to form a feature vector. Such feature vectors are highly discriminative and are robust to mere misalignment. Experiments demonstrate that the proposed CR-based framework with an l2-norm regularization term can achieve much better recognition accuracy than the other methods. More importantly, its computational complexity is extremely low, making it quite suitable for the large-scale identification application. Source codes are available at http://sse.tongji.edu.cn/linzhang/cr3dpalm/cr3dpalm.htm.
Lin Zhang 0014, Ying Shen 0005, Hongyu Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 A Feature-Enriched Completely Blind Image Quality Evaluator
abstract
Existing blind image quality assessment (BIQA) methods are mostly opinion-aware. They learn regression models from training images with associated human subjective scores to predict the perceptual quality of test images. Such opinion-aware methods, however, require a large amount of training samples with associated human subjective scores and of a variety of distortion types. The BIQA models learned by opinion-aware methods often have weak generalization capability, hereby limiting their usability in practice. By comparison, opinion-unaware methods do not need human subjective scores for training, and thus have greater potential for good generalization capability. Unfortunately, thus far no opinion-unaware BIQA method has shown consistently better quality prediction accuracy than the opinion-aware methods. Here, we aim to develop an opinion-unaware BIQA method that can compete with, and perhaps outperform, the existing opinion-aware methods. By integrating the features of natural image statistics derived from multiple cues, we learn a multivariate Gaussian model of image patches from a collection of pristine natural images. Using the learned multivariate Gaussian model, a Bhattacharyya-like distance is used to measure the quality of each image patch, and then an overall quality score is obtained by average pooling. The proposed BIQA method does not need any distorted sample images nor subjective quality scores for training, yet extensive experiments demonstrate its superior quality-prediction performance to the state-of-the-art opinion-aware BIQA methods. The MATLAB source code of our algorithm is publicly available at www.comp.polyu.edu.hk/~cslzhang/IQA/ILNIQE/ILNIQE.htm.
Lin Zhang 0014, Lei Zhang 0006, Alan C. Bovik
IEEE Trans. Image Process.1
2014 Integrating Visual Saliency Information into Objective Quality Assessment of Tone-Mapped Images
Xueyabo Liu, Lin Zhang 0014, Hongyu Li 0001
ICIC (1)2
2014 Learning quality-aware filters for no-reference image quality assessment
abstract
With the rapid development of the usage of digital imaging and communication technologies, there appears to be a great demand for fast and practical approaches for image quality assessment (IQA) algorithms that can match human judgements. In this paper, we propose a novel general-purpose no-reference IQA (NR-IQA) framework by means of learning quality-aware filters (QAF). Using these filters for image encoding, we can obtain effective image representations for quality estimation. Additionally, random forest is used to learn the mapping from feature space to human subjective scores. Extensive experiments conducted on LIVE and CSIQ databases demonstrate that the proposed NR-IQA metric QAF can achieve better prediction performance than all the other state-of-the-art NR-IQA approaches in terms of both prediction accuracy and generalization capabilities.
Zhongyi Gu, Lin Zhang 0014, Hongyu Li 0001
ICME2
2014 3DMKDSRC: A novel approach for 3D face recognition
abstract
Recent years have witnessed a growing interest in developing methods for 3D face recognition. However, 3D scans often suffer from the problems of missing parts, large facial expressions, and occlusions. In this paper, we propose a novel general approach to deal with the 3D face recognition problem by making use of multiple keypoint descriptors (MKD) and the sparse representation-based classifier (SRC). We call the proposed method 3DMKDSRC for short. Specifically, with 3DMKDSRC, each 3D face scan is represented as a set of descriptor vectors extracted from keypoints by meshSIFT. Descriptor vectors of gallery samples form the gallery dictionary. Given a probe 3D face scan, its descriptors are extracted at first and then its identity can be determined by using a multitask SRC. The effectiveness of 3DMKDSRC has been corroborated by extensive experiments.
Lin Zhang 0014, Zhixuan Ding, Hongyu Li 0001
ICME1
2014 Integration of multiple orientation and texture information for finger-knuckle-print verification
Guangwei Gao, Jian Yang 0003, Jianjun Qian, Lin Zhang 0014
Neurocomputing4
2014 VSI: A Visual Saliency-Induced Index for Perceptual Image Quality Assessment
abstract
Perceptual image quality assessment (IQA) aims to use computational models to measure the image quality in consistent with subjective evaluations. Visual saliency (VS) has been widely studied by psychologists, neurobiologists, and computer scientists during the last decade to investigate, which areas of an image will attract the most attention of the human visual system. Intuitively, VS is closely related to IQA in that suprathreshold distortions can largely affect VS maps of images. With this consideration, we propose a simple but very effective full reference IQA method using VS. In our proposed IQA model, the role of VS is twofold. First, VS is used as a feature when computing the local quality map of the distorted image. Second, when pooling the quality score, VS is employed as a weighting function to reflect the importance of a local region. The proposed IQA index is called visual saliency-based index (VSI). Several prominent computational VS models have been investigated in the context of IQA and the best one is chosen for VSI. Extensive experiments performed on four large-scale benchmark databases demonstrate that the proposed IQA index VSI works better in terms of the prediction accuracy than all state-of-the-art IQA indices we can find while maintaining a moderate computational complexity. The MATLAB source code of VSI and the evaluation results are publicly available online at http://sse.tongji.edu.cn/linzhang/IQA/VSI/VSI.htm.
Lin Zhang 0014, Ying Shen 0005, Hongyu Li 0001
IEEE Trans. Image Process.1
2013 Efficient 3D Reconstruction for Urban Scenes
Weichao Fu, Lin Zhang 0014, Hongyu Li 0001, Xinfeng Zhang 0004
ICIC (1)2
2013 Real-Time Visual Tracking Based on an Appearance Model and a Motion Mode
Guizi Li, Lin Zhang 0014, Hongyu Li 0001
ICIC (2)2
2013 Improving Classification Accuracy Using Gene Ontology Information
Ying Shen 0005, Lin Zhang 0014
ICIC (3)2
2013 SDSP: A novel saliency detection method by combining simple priors
abstract
Salient regions detection from images is an important and fundamental research problem in neuroscience and psychology and it serves as an indispensible step for numerous machine vision tasks. In this paper, we propose a novel conceptually simple salient region detection method, namely SDSP, by combining three simple priors. At first, the behavior that the human visual system detects salient objects in a visual scene can be well modeled by band-pass filtering. Secondly, people are more likely to pay their attention on the center of an image. Thirdly, warm colors are more attractive to people than cold colors are. Extensive experiments conducted on the benchmark dataset indicate that SDSP could outperform the other state-of-the-art algorithms by yielding higher saliency prediction accuracy. Moreover, SDSP has a quite low computational complexity, rendering it an outstanding candidate for time critical applications. The Matlab source code of SDSP and the evaluation results have been made online available at http://sse.tongji.edu.cn/linzhang/va/SDSP/SDSP.htm.
Lin Zhang 0014, Zhongyi Gu, Hongyu Li 0001
ICIP1
2013 A novel 3D ear identification approach based on sparse representation
abstract
Recently, ear shape has attracted tremendous interests in biometric research due to its richness of feature and ease of acquisition. In this paper, we present a novel 3D ear identification approach based on the sparse representation framework. To this end, at first, we propose a template-based ear detection method. By utilizing such a method, the extracted ear regions are represented in a common standard coordinate system determined by the template, which facilitates the following feature extraction and classification. For each 3D ear, a feature vector can be generated as its representation. With respect to the ear identification, we resort to the l1-minimization based sparse representation. Experiments conducted on a benchmark dataset corroborate the effectiveness and efficacy of the proposed approach. The associated Matlab source code and the evaluation results have been made online available at http://sse.tongji.edu.cn/linzhang/ear/srcear/srcear.htm.
Zhixuan Ding, Lin Zhang 0014, Hongyu Li 0001
ICIP2
2013 Learning a blind image quality index based on visual saliency guided sampling and Gabor filtering
abstract
The goal of no-reference image quality assessment (NR-IQA) is to estimate the quality of an image consistent with the human perception of the image automatically without any prior of the reference image. In this paper, we present a simple yet efficient and effective approach to learn a blind Image Quality index based on Visual saliency guided sampling and Gabor filtering, namely IQVG. Given an image, we at first randomly sample a sufficient number of image patches guided by the image's visual saliency map and convolve each patch with Gabor filters to get a bag of features. Then, the image is represented by using a histogram to encode the bag of features. Support vector regression (SVR) is used to learn the mapping from feature space to image quality. Extensive experiments conducted on the LIVE IQA database demonstrate the overall superiority of our IQVG over the other state-of-the-art NR-IQA algorithms evaluated. The Matlab source code of IQVG and the evaluation results are available online at http://sse.tongji.edu.cn/linzhang/IQA/IQVG/IQVG.htm.
Zhongyi Gu, Lin Zhang 0014, Hongyu Li 0001
ICIP2
2013 SR-LLA: A novel spectral reconstruction method based on locally linear approximation
abstract
Compared with tristimulus, spectrum contains much more information of a color, which can be used in many fields, such as disease diagnosis and material recognition. In order to get an accurate and stable reconstruction of spectral data from a tristimulus input, a method based on locally linear approximation is proposed in this paper, namely SR-LLA. To test the performance of SR-LLA, we conduct experiments on three Munsell databases and present a comprehensive analysis of its accuracy and stability. We also compare the performance of SR-LLA with the other two spectral reconstruction methods based on BP neural network and PCA, respectively. Experimental results indicate that SR-LLA could outperform other competitors in terms of both accuracy and stability for spectral reconstruction.
Hongyu Li 0001, Zhujing Wu, Lin Zhang 0014, Jussi Parkkinen
ICIP3
2013 Is local dominant orientation necessary for the classification of rotation invariant texture?
Zhenhua Guo 0001, Qin Li 0001, Lin Zhang 0014, Jane You, David Zhang 0001, Wenhuang Liu
Neurocomputing3
2013 Reconstruction Based Finger-Knuckle-Print Verification With Score Level Adaptive Binary Fusion
abstract
Recently, a new biometrics identifier, namely finger knuckle print (FKP), has been proposed for personal authentication with very interesting results. One of the advantages of FKP verification lies in its user friendliness in data collection. However, the user flexibility in positioning fingers also leads to a certain degree of pose variations in the collected query FKP images. The widely used Gabor filtering based competitive coding scheme is sensitive to such variations, resulting in many false rejections. We propose to alleviate this problem by reconstructing the query sample with a dictionary learned from the template samples in the gallery set. The reconstructed FKP image can reduce much the enlarged matching distance caused by finger pose variations; however, both the intra-class and inter-class distances will be reduced. We then propose a score level adaptive binary fusion rule to adaptively fuse the matching distances before and after reconstruction, aiming to reduce the false rejections without increasing much the false acceptances. Experimental results on the benchmark PolyU FKP database show that the proposed method significantly improves the FKP verification accuracy.
Guangwei Gao, Lei Zhang 0006, Jian Yang 0003, Lin Zhang 0014, David Zhang 0001
IEEE Trans. Image Process.4
2012 SR-SIM: A fast and high performance IQA index based on spectral residual
abstract
Automatic image quality assessment (IQA) attempts to use computational models to measure the image quality in consistency with subjective ratings. In the past decades, dozens of IQA models have been proposed. Though some of them can predict subjective image quality accurately, their computational costs are usually very high. To meet real-time requirements, in this paper, we propose a novel fast and effective IQA index, namely spectral residual based similarity (SR-SIM), based on a specific visual saliency model, spectral residual visual saliency. SR-SIM is designed based on the hypothesis that an image's visual saliency map is closely related to its perceived quality. Extensive experiments conducted on three large-scale IQA datasets indicate that SR-SIM could achieve better prediction performance than the other state-of-the-art IQA indices evaluated. Moreover, SR-SIM can have a quite low computational complexity. The Matlab source code of SR-SIM and the evaluation results are available online at http://sse.tongji.edu.cn/linzhang/IQA/SR-SIM/SR-SIM.htm.
Lin Zhang 0014, Hongyu Li 0001
ICIP1
2012 Binary Gabor pattern: An efficient and robust descriptor for texture classification
abstract
In this paper, we present a simple yet efficient and effective multi-resolution approach to gray-scale and rotation invariant texture classification. Given a texture image, we at first convolve it with J Gabor filters sharing the same parameters except the parameter of orientation. Then by binarizing the obtained responses, we can get J bits at each location. Then, each location can be assigned a unique integer, namely “rotation invariant binary Gabor pattern (BGPri)”, formed from J bits associated with it using some rule. The classification is based on the image's histogram of its BGPris at multiple scales. Using BGPri, there is no need for a pre-training step to learn a texton dictionary, as required in methods based on clustering such as MR8. Extensive experiments conducted on the CUReT database demonstrate the overall superiority of BGPriover the other state-of-the-art texture representation methods evaluated. The Matlab source codes are publicly available at http://sse.tongji.edu.cn/linzhang/IQA/BGP/BGP.htm.
Lin Zhang 0014, Hongyu Li 0001
ICIP1
2012 A comprehensive evaluation of full reference image quality assessment algorithms
abstract
Recent years have witnessed a growing interest in developing objective image quality assessment (IQA) algorithms that can measure the image quality consistently with subjective evaluations. For the full reference (FR) IQA problem, great progress has been made in the past decade. On the other hand, several new large scale image datasets have been released for evaluating FR IQA methods in recent years. Meanwhile, no work has been reported to evaluate and compare the performance of state-of-the-art and representative FR IQA methods on all the available datasets. In this paper, we aim to fulfill this task by reporting the performance of eleven selected FR IQA algorithms on all the seven public IQA image datasets. Our evaluation results and the associated discussions will be very helpful for relevant researchers to have a clearer understanding about the status of modern FR IQA indices. Evaluation results presented in this paper are also online available at http://sse.tongji.edu.cn/linzhang/IQA/IQA.htm.
Lin Zhang 0014, Lei Zhang 0006, Xuanqin Mou, David Zhang 0001
ICIP1
2012 Manifold Analysis of Spectral Munsell Colors
Hongyu Li 0001, Chen Lin 0001, Junyu Niu, Lin Zhang 0014, Jussi Parkkinen
ICONIP (1)4
2012 Entropy Based Image Semantic Cycle for Image Classification
Hongyu Li 0001, Junyu Niu, Lin Zhang 0014
ICONIP (5)3
2012 Spatio-temporal LTSA and Its Application to Motion Decomposition
Hongyu Li 0001, Junyu Niu, Lin Zhang 0014
ICONIP (5)3
2012 Local tangent space based manifold entropy for image retrieval
Hongyu Li 0001, Junyu Niu, Lin Zhang 0014
ICPR4
2012 Encoding local image patterns using Riesz transforms: With applications to palmprint and finger-knuckle-print recognition
Lin Zhang 0014, Hongyu Li 0001
Image Vis. Comput.1
2012 Phase congruency induced local features for finger-knuckle-print recognition
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Zhenhua Guo 0001
Pattern Recognit.1
2012 Fragile Bits in Palmprint Recognition
abstract
Recent years have witnessed a growing interest in developing automatic palmprint recognition methods. Among them, coding-based ones, representing the texture of a palmprint using a binary code, are most prevalent and successful. We find that not all bits in a code map generated by a specific coding scheme are equally consistent. A bit is deemed fragile if its value changes across code maps created from different images of the same palmprint. In this paper, we first analyze the fragile bits phenomenon in a state-of-the-art palmprint coding scheme, namely, binary orientation co-occurrence vector (BOCV). Then, based on our analysis, we extend BOCV to E-BOCV by incorporating fragile bits information in appropriate ways. Experiments conducted on the benchmark dataset demonstrate that E-BOCV can achieve the highest verification accuracy among all the state-of-the-art palmprint verification methods evaluated. To our knowledge, this is the first work investigating the fragile bits of coding-based palmprint recognition approaches.
Lin Zhang 0014, Hongyu Li 0001, Junyu Niu
IEEE Signal Process. Lett.1
2011 Texture Image Classification Using Complex Texton
Zhenhua Guo 0001, Qin Li 0001, Lin Zhang 0014, Jane You, Wenhuang Liu
ICIC (2)3
2011 Ensemble of local and global information for finger-knuckle-print recognition
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Hailong Zhu
Pattern Recognit.1
2011 FSIM: A Feature Similarity Index for Image Quality Assessment
abstract
Image quality assessment (IQA) aims to use computational models to measure the image quality consistently with subjective evaluations. The well-known structural similarity index brings IQA from pixel- to structure-based stage. In this paper, a novel feature similarity (FSIM) index for full reference IQA is proposed based on the fact that human visual system (HVS) understands an image mainly according to its low-level features. Specifically, the phase congruency (PC), which is a dimensionless measure of the significance of a local structure, is used as the primary feature in FSIM. Considering that PC is contrast invariant while the contrast information does affect HVS' perception of image quality, the image gradient magnitude (GM) is employed as the secondary feature in FSIM. PC and GM play complementary roles in characterizing the image local quality. After obtaining the local quality map, we use PC again as a weighting function to derive a single quality score. Extensive experiments performed on six benchmark IQA databases demonstrate that FSIM can achieve much higher consistency with the subjective evaluations than state-of-the-art IQA metrics.
Lin Zhang 0014, Lei Zhang 0006, Xuanqin Mou, David Zhang 0001
IEEE Trans. Image Process.1
2010 Monogenic-LBP: A new approach for rotation invariant texture classification
abstract
Analysis of two-dimensional textures has many potential applications in computer vision. In this paper, we investigate the problem of rotation invariant texture classification, and propose a novel texture feature extractor, namely Monogenic-LBP (M-LBP). M-LBP integrates the traditional Local Binary Pattern (LBP) operator with the other two rotation invariant measures: the local phase and the local surface type computed by the 1st-order and 2nd-order Riesz transforms, respectively. The classification is based on the image's histogram of M-LBP responses. Extensive experiments conducted on the CUReT database demonstrate the overall superiority of M-LBP over the other state-of-the-art methods evaluated.
Lin Zhang 0014, Lei Zhang 0006, Zhenhua Guo 0001, David Zhang 0001
ICIP1
2010 RFSIM: A feature based image quality assessment metric using Riesz transforms
abstract
Image quality assessment (IQA) aims to provide computational models to measure the image quality in a perceptually consistent manner. In this paper, a novel feature based IQA model, namely Riesz-transform based Feature SIMilarity metric (RFSIM), is proposed based on the fact that the human vision system (HVS) perceives an image mainly according to its low-level features. The 1st-order and 2nd-order Riesz transform coefficients of the image are taken as image features, while a feature mask is defined as the edge locations of the image. The similarity index between the reference and distorted images is measured by comparing the two feature maps at key locations marked by the feature mask. Extensive experiments on the comprehensive TID2008 database indicate that the proposed RFSIM metric is more consistent with the subjective evaluation than all the other competing methods evaluated.
Lin Zhang 0014, Lei Zhang 0006, Xuanqin Mou
ICIP1
2010 Incremental Nyström Low-Rank Decomposition for Dynamic Learning
abstract
Eigen-decomposition is a key step in spectral clustering and some kernel methods. The Nyström method is often used to speed up kernel matrix decomposition. However, it cannot effectively update eigenvectors of matrices when datasets dynamically increase with time. In this paper, we propose an incremental Nyström method for dynamic learning. Experimental results demonstrate the feasibility and effectiveness of the proposed method.
Lin Zhang 0014, Hongyu Li 0001
ICMLA1
2010 Monogenic Binary Pattern (MBP): A Novel Feature Extraction and Representation Model for Face Recognition
abstract
A novel feature extraction method, namely monogenic binary pattern (MBP), is proposed in this paper based on the theory of monogenic signal analysis, and the histogram of MBP (HMBP) is subsequently presented for robust face representation and recognition. MBP consists of two parts: one is monogenic magnitude encoded via uniform LBP, and the other is monogenic orientation encoded as quadrant-bit codes. The HMBP is established by concatenating the histograms of MBP of all sub-regions. Compared with the well-known and powerful Gabor filtering based LBP schemes, one clear advantage of HMBP is its lower time and space complexity because monogenic signal analysis needs fewer convolutions and generates more compact feature vectors. The experimental results on the AR and FERET face databases validate that the proposed MBP algorithm has better performance than or comparable performance with state-of-the-art local feature based methods but with significantly lower time and space complexity.
Meng Yang 0001, Lei Zhang 0006, Lin Zhang 0014, David Zhang 0001
ICPR3
2010 Online finger-knuckle-print verification for personal authentication
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Hailong Zhu
Pattern Recognit.1
2009 A Multi-scale Bilateral Structure Tensor Based Corner Detector
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001
ACCV (2)1
2009 Finger-Knuckle-Print Verification Based on Band-Limited Phase-Only Correlation
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001
CAIP1
2009 Finger-knuckle-print: A new biometric identifier
abstract
This paper presents a new biometric identifier, namely finger-knuckle-print (FKP), for personal identity authentication. First a specific data acquisition device is constructed to capture the FKP images, and then an efficient FKP recognition algorithm is presented to process the acquired data. The local convex direction map of the FKP image is extracted, based on which a coordinate system is defined to align the images and a region of interest (ROI) is cropped for feature extraction. A competitive coding scheme, which uses 2D Gabor filters to extract the image local orientation information, is employed to extract and represent the FKP features. When matching, the angular distance is used to measure the similarity between two competitive code maps. An FKP database was established to examine the performance of the proposed system, and the experimental results demonstrated the efficiency and effectiveness of this new biometric characteristic.
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001
ICIP1