EDBT 2026 Demo / reviewers in the wild / expert
Lei Ma 0008
dblp:20/6534-8
· DBLP profile ↗
48ranked-venue papers
3as first author
44since 2021 · last 2026
0000-0001-6024-3854ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 20 · 20 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Splats in Splats: Robust and Effective 3D Steganography Towards Gaussian Splattingabstract3D Gaussian splatting (3DGS) has demonstrated impressive 3D reconstruction performance with explicit scene representations. Given the widespread application of 3DGS in 3D reconstruction and generation tasks, there is an urgent need to protect the copyright of 3DGS assets. However, existing copyright protection techniques for 3DGS overlook the usability of 3D assets, posing challenges for practical deployment. Here we describe splats in splats, the first 3DGS steganography framework that embeds 3D content in 3DGS itself without modifying any attributes. To achieve this, we take a deep insight into spherical harmonics (SH) and devise an importance-graded SH coefficient encryption strategy to embed the hidden SH coefficients. Furthermore, we employ a convolutional autoencoder to establish a mapping between the original Gaussian primitives' opacity and the hidden Gaussian primitives' opacity. Extensive experiments indicate that our method significantly outperforms existing 3D steganography techniques, with 5.31% higher scene fidelity and 3x faster rendering speed, while ensuring security, robustness, and user experience. Yijia Guo, Wenkai Huang 0003, Gaolei Li, Hang Zhang 0010, Liwen Hu 0002, Jianhua Li 0001, Tiejun Huang 0001, Lei Ma 0008 |
AAAI | 9 |
| 2026 | Can Protective Watermarking Safeguard the Copyright of 3D Gaussian Splatting?abstract3D Gaussian Splatting (3DGS) has emerged as a powerful representation for 3D scenes, widely adopted due to its exceptional efficiency and high-fidelity visual quality. Given the significant value of 3DGS assets, recent works have introduced specialized watermarking schemes to ensure copyright protection and ownership verification. However, can existing 3D Gaussian watermarking approaches genuinely guarantee robust protection of the 3D assets? In this paper, for the first time, we systematically explore and validate possible vulnerabilities of 3DGS watermarking frameworks. We demonstrate that conventional watermark removal techniques designed for 2D images do not effectively generalize to the 3DGS scenario due to the specialized rendering pipeline and unique attributes of each gaussian primitives. Motivated by this insight, we propose GSPure, the first watermark purification framework specifically for 3DGS watermarking representations. By analyzing view-dependent rendering contributions and exploiting geometrically accurate feature clustering, GSPure precisely isolates and effectively removes watermark-related Gaussian primitives while preserving scene integrity. Extensive experiments demonstrate that our GSPure achieves the best watermark purification performance, reducing watermark PSNR by up to 16.34dB while minimizing degradation to original scene fidelity with less than 1dB PSNR loss. Moreover, it consistently outperforms existing methods in both effectiveness and generalization. Wenkai Huang 0003, Yijia Guo, Gaolei Li, Lei Ma 0008, Hang Zhang 0010, Liwen Hu 0002, Jiazheng Wang 0001, Jianhua Li 0001, Tiejun Huang 0001 |
AAAI | 4 |
| 2026 | Learn to Enhance Sparse Spike StreamsabstractHigh-speed vision tasks have long been a challenge in computer vision. Recently, the spike camera has shown great potential in these tasks due to its high temporal resolution. Unlike traditional cameras, it emits asynchronous spike signals to capture visual information. However, under low-light conditions, spike signals become highly sparse, and the sparse spike stream severely hinders the effectiveness of existing spike-based methods in high-speed scenarios. To address this challenge, we introduce SS2DS, the first deep learning framework that enhances sparse spike streams into dense spike streams. SS2DS first estimates the spike firing frequency within sparse streams. Subsequently, the spike firing frequency is enhanced by a neural network. Finally, SS2DS decodes the enhanced spike stream from the enhanced spike firing frequency sequence. SS2DS can adjust the temporal distribution of sparse spike streams and improve the performance degradation of existing methods in low-light and high-speed scenarios. To evaluate sparse spike stream enhancement, we construct both synthetic and real sparse spike stream datasets. By comparing the reconstruction results, enhanced spike streams achieve an average improvement of +0.78 MA, -18.42 BRISQUE, and -1.42 NIQE over sparse spike streams. Moreover, the enhanced spike streams also benefit other spike-based vision tasks, such as 3D reconstruction (+1.325 dB PSNR, +0.005 SSIM, and -0.01 LPIPS) and superresolution (+0.63 MA, -13.67 BRISQUE, and -1.28 NIQE). Liwen Hu 0002, Yijia Guo, Mianzhi Liu, Shengbo Chen, Lei Ma 0008, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | SpikeGS: Reconstruct 3D Scene Captured by a Fast-Moving Bio-Inspired Cameraabstract3D Gaussian Splatting (3DGS) has been proven to exhibit exceptional performance in reconstructing 3D scenes. However, the effectiveness of 3DGS heavily relies on sharp images, and fulfilling this requirement presents challenges in real-world scenarios particularly when utilizing fast-moving cameras. This limitation severely constrains the practical application of 3DGS and may compromise the feasibility of real-time reconstruction. To mitigate these challenges, we proposed Spike Gaussian Splatting (SpikeGS), the first framework that integrates the Bayer-pattern spike streams into the 3DGS pipeline to reconstruct 3D scenes captured by a fast-moving high temporal color spike camera in one second. With accumulation rasterization, interval supervision, and a special designed pipeline, SpikeGS realizes continuous spatiotemporal perception while extracts detailed structure and texture from Bayer-pattern spike stream which is unstable and lacks details. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of SpikeGS compared with existing spike-based and deblur 3D scene reconstruction methods. Yijia Guo, Liwen Hu 0002, Yuanxi Bai, Jiawei Yao, Lei Ma 0008, Tiejun Huang 0001 |
AAAI | 5 |
| 2025 | MindCustomer: Multi-Context Image Generation Blended with Brain SignalabstractAdvancements in generative models have promoted text- and image-based multi-context image generation. Brain signals, offering a direct representation of user intent, present new opportunities for image customization. However, it faces challenges in brain interpretation, cross-modal context fusion and retention. In this paper, we present MindCustomer to explore the blending of visual brain signals in multi-context image generation. We first design shared neural data augmentation for stable cross-subject brain embedding by introducing the Image-Brain Translator (IBT) to generate brain responses from visual images. Then, we propose an effective cross-modal information fusion pipeline that mask-freely adapts distinct semantics from image and brain contexts within a diffusion model. It resolves semantic conflicts for context preservation and enables harmonious context integration. During the fusion pipeline, we further utilize the IBT to transfer image context to the brain representation to mitigate the cross-modal disparity. MindCustomer enables cross-subject generation, delivering unified, high-quality, and natural image outputs. Moreover, it exhibits strong generalization for new subjects via few-shot learning, indicating the potential for practical application. As the first work for multi-context blending with brain signal, MindCustomer lays a foundational exploration and inspiration for future brain-controlled generative technologies. Muzhou Yu, Shuyun Lin, Lei Ma 0008, Kaisheng Ma |
ICML | 3 |
| 2025 | Robust Indoor Person Re-Identification With Multimodal TrainingabstractExisting person re-identification (ReID) methods mainly rely on images and videos to match persons across cameras, yet visual data captured by cameras are vulnerable to environmental interferences (e.g. illumination and occlusion) or personal appearance changes, leading to performance degradation under such scenes. Meanwhile, the popularization of Wi-Fi networks has allowed probe requests to be captured for mobile sensing applications such as crowd counting and trajectory estimation. However, the MAC address randomization technique adopted by modern devices breaks the association of probe requests and adversely affects the functionality of these applications. In this paper, we propose MaRPA, the first multimodal training approach that incorporates both videos and Wi-Fi probe requests to simultaneously promote tasks of probe requests association and person ReID. MaRPA first distinguishes among pairwise probe request frames through a contrastive learning model. It then matches video and probe request sequences by exploring their similarities from the position and the vision aspects. Matched videos and probe requests provide complementary information and generate more robust features for both tasks. To evaluate MaRPA, we contribute a new dataset containing synchronous videos and probe requests data for probe requests association and person ReID. Experimental results demonstrate the effectiveness of our approach. For probe requests association, it achieves > 85% discrimination accuracy and > 0.90 V-measure score; for person ReID, it achieves 75.8% mAP and 90.6% Rank-1, improving state-of-the-art video-based ReID methods by over 40% Can Su, Xinlei Xue, Lei Ma 0008, Wei Yan 0007, Kaigui Bian |
IEEE Internet Things J. | 3 |
| 2024 | E/I Balanced Adaptive Sequential Neural Posterior Estimation for Inferring the Connection Weights in Mouse V1 ModelabstractEffectively utilizing biological firing rate data to estimate the numerous connection weights in the mouse primary visual cortex (V1) model from the Allen Institute is a challenging task. The existing iterative grid-search algorithm cannot enable the mouse V1 model to better fit the biological firing rate data. To tackle this issue, we propose an excitation-inhibition balanced adaptive sequential neural posterior estimation (E/I balanced ASNPE) approach to accurately infer the connection weights of the mouse V1 model, allowing the neurons’ firing rates to converge to the given biological data. This method fully leverages the structural information of the mouse V1 model, reducing the dimensionality of the weight parameters to be optimized. Initially, sampling is performed in the prior distribution based on the proposed non-dominated sorting adaptive genetic algorithm (NSAGA). This algorithm optimizes the sorting, crossover and mutation processes based on the fitness scores of the current samples and updates the proposal distribution based on these samples, increasing the likelihood of identifying high posterior probability regions in the prior distribution. To avoid bad simulations, we also explore the E/I balance in each layer of the mouse V1 model, adding biological constraints during weight inference with Automatic Posterior Transformation (APT). Experimental results confirm that the proposed E/I balanced ASNPE method significantly outperforms the baseline in all five firing rate fitness scores in the mouse V1 model. This study is pioneering in applying non-dominated sorting genetic algorithms combined with sequential neural posterior estimation to optimize connection weights in large-scale complex biological models. Luntian Mou, Peize Li, Lei Ma 0008, Tiejun Huang 0001 |
BIBM | 4 |
| 2024 | Correspondence-Free Non-Rigid Point Set Registration Using Unsupervised Clustering AnalysisabstractThis paper presents a novel non-rigid point set registration method that is inspired by unsupervised clustering analysis. Unlike previous approaches that treat the source and target point sets as separate entities, we develop a holistic framework where they are formulated as clustering centroids and clustering members, separately. We then adopt Tikhonov regularization with an$\ell_{1}$-induced Laplacian kernel instead of the commonly used Gaussian kernel to ensure smooth and more robust displacement fields. Our formulation delivers closed-form solutions, theoretical guarantees, independence from dimensions, and the ability to handle large deformations. Subsequently, we introduce a clustering-improved Nyström method to effectively reduce the computational complexity and storage of the Gram matrix to linear, while providing a rigorous bound for the low-rank approximation. Our method achieves high accuracy results across various scenarios and surpasses competitors by a significant margin, particularly on shapes with sub-stantial deformations. Additionally, we demonstrate the versatility of our method in challenging tasks such as shape transfer and medical registration. [Code release] Mingyang Zhao 0001, Jingen Jiang 0001, Lei Ma 0008, Shi-Qing Xin, Gaofeng Meng, Dong-Ming Yan 0001 |
CVPR | 3 |
| 2024 | Learning to Robustly Reconstruct Dynamic Scenes from Low-Light Spike Streams
Liwen Hu 0002, Ziluo Ding, Mianzhi Liu, Lei Ma 0008, Tiejun Huang 0001 |
ECCV (17) | 4 |
| 2024 | Spike-NeRF: Neural Radiance Field Based On Spike CameraabstractAs a neuromorphic sensor with high temporal resolution, spike cameras offer notable advantages over traditional cameras in high-speed vision applications such as high-speed optical estimation, depth estimation, and object tracking. Inspired by the success of the spike camera, we proposed Spike-NeRF, the first Neural Radiance Field derived from spike data, to achieve 3D reconstruction and novel viewpoint synthesis for high-speed scenes. Instead of the multi-view images at the same as time of NeRF, the inputs of Spike-NeRF are continuous spike streams captured by a moving spike camera in a very short time. To reconstruct a correct and stable 3D scene from high-frequency but unstable spike data, we devised spike masks along with a distinctive loss function. We evaluate our method qualitatively and quantitatively on several challenging synthetic scenes generated using Blender with the spike camera simulator. Our results demonstrate that Spike-NeRF produces more visually appealing results than the existing methods and the baseline we proposed in high-speed scenes. Our code is available at https://github.com/yijiaguo02/SpikeNerf Yijia Guo, Yuanxi Bai, Liwen Hu 0002, Mianzhi Liu, Lei Ma 0008, Tiejun Huang 0001 |
ICME | 6 |
| 2024 | SCSim: A Realistic Spike Cameras SimulatorabstractSpike cameras, with their exceptional temporal resolution, are revolutionizing high-speed visual applications. Large-scale synthetic datasets have significantly accelerated the development of these cameras, particularly in reconstruction and optical flow. However, current synthetic datasets for spike cameras lack sophistication. Addressing this gap, we introduce SCSim, a novel and more realistic spike camera simulator with a comprehensive noise model. SCSim is adept at autonomously generating driving scenarios and synthesizing corresponding spike streams. To enhance the fidelity of these streams, we’ve developed a comprehensive noise model tailored to the unique circuitry of spike cameras. Our evaluations demonstrate that SCSim outperforms existing simulation methods in generating authentic spike streams. Crucially, SCSim simplifies the creation of datasets, thereby greatly advancing spike-based visual tasks like reconstruction. Our project refers to https://github.com/Acnext/SCSim. Liwen Hu 0002, Lei Ma 0008, Yijia Guo, Tiejun Huang 0001 |
ICME | 2 |
| 2024 | ShapeMamba-EM: Fine-Tuning Foundation Model with Local Shape Descriptors and Mamba Blocks for 3D EM Image Segmentation
Ruohua Shi, Qiufan Pang, Lei Ma 0008, Ling-Yu Duan, Tiejun Huang 0001, Tingting Jiang 0001 |
MICCAI (12) | 3 |
| 2024 | PRTGS: Precomputed Radiance Transfer of Gaussian Splats for Real-Time High-Quality RelightingabstractWe proposed Precomputed Radiance Transfer of Gaussian Splats (PRTGS), a real-time high-quality relighting method for Gaussian splats in low-frequency lighting environments that captures soft shadows and interreflections by precomputing 3D Gaussian splats' radiance transfer. Existing studies have demonstrated that 3D Gaussian splatting (3DGS) outperforms neural fields in efficiency for dynamic lighting scenarios. However, the current relighting method based on 3DGS is still struggling to compute high-quality shadow and indirect illumination in real time for dynamic light, leading to unrealistic rendering results. We solve this problem by precomputing the expensive transport simulations required for complex transfer functions like shadowing, the resulting transfer functions are represented as dense sets of vectors or matrices for every Gaussian splat. We introduce distinct precomputing methods tailored for training and rendering stages, along with unique ray tracing and indirect lighting precomputation techniques for 3D Gaussian splats to accelerate training speed and compute accurate indirect lighting related to environment light. Experimental analyses demonstrate that our approach achieves state-of-the-art visual quality while maintaining competitive training times and importantly allows high-quality real-time (30+ fps) relighting for dynamic light and relatively complex scenes at 1080p resolution. Yijia Guo, Yuanxi Bai, Liwen Hu 0002, Mianzhi Liu, Yu Cai 0008, Tiejun Huang 0001, Lei Ma 0008 |
ACM Multimedia | 8 |
| 2024 | Learning from Pattern Completion: Self-supervised Controllable GenerationabstractThe human brain exhibits a strong ability to spontaneously associate different visual attributes of the same or similar visual scene, such as associating sketches and graffiti with real-world visual objects, usually without supervising information. In contrast, in the field of artificial intelligence, controllable generation methods like ControlNet heavily rely on annotated training datasets such as depth maps, semantic segmentation maps, and poses, which limits the method’s scalability. Inspired by the neural mechanisms that may contribute to the brain’s associative power, specifically the cortical modularization and hippocampal pattern completion, here we propose a self-supervised controllable generation (SCG) framework. Firstly, we introduce an equivariance constraint to promote inter-module independence and intra-module correlation in a modular autoencoder network, thereby achieving functional specialization. Subsequently, based on these specialized modules, we employ a self-supervised pattern completion approach for controllable generation training. Experimental results demonstrate that the proposed modular autoencoder effectively achieves functional specialization, including the modular processing of color, brightness, and edge detection, and exhibits brain-like features including orientation selectivity, color antagonism, and center-surround receptive fields. Through self-supervised training, associative generation capabilities spontaneously emerge in SCG, demonstrating excellent zero-shot generalization ability to various tasks such as superresolution, dehaze and associative or conditional generation on painting, sketches, and ancient graffiti. Compared to the previous representative method ControlNet, our proposed approach not only demonstrates superior robustness in more challenging high-noise scenarios but also possesses more promising scalability potential due to its self-supervised manner. Codes are released on Github and Gitee. Guofan Fan, Jinying Gao, Lei Ma 0008, Tiejun Huang 0001 |
NeurIPS | 4 |
| 2024 | Retrospective for the Dynamic Sensorium Competition for predicting large-scale mouse primary visual cortex activity from videosabstractUnderstanding how biological visual systems process information is challenging because of the nonlinear relationship between visual input and neuronal responses. Artificial neural networks allow computational neuroscientists to create predictive models that connect biological and machine vision.Machine learning has benefited tremendously from benchmarks that compare different models on the same task under standardized conditions. However, there was no standardized benchmark to identify state-of-the-art dynamic models of the mouse visual system.To address this gap, we established the SENSORIUM 2023 Benchmark Competition with dynamic input, featuring a new large-scale dataset from the primary visual cortex of ten mice. This dataset includes responses from 78,853 neurons to 2 hours of dynamic stimuli per neuron, together with behavioral measurements such as running speed, pupil dilation, and eye movements.The competition ranked models in two tracks based on predictive performance for neuronal responses on a held-out test set: one focusing on predicting in-domain natural stimuli and another on out-of-distribution (OOD) stimuli to assess model generalization.As part of the NeurIPS 2023 Competition Track, we received more than 160 model submissions from 22 teams. Several new architectures for predictive models were proposed, and the winning teams improved the previous state-of-the-art model by 50\%. Access to the dataset as well as the benchmarking infrastructure will remain online at www.sensorium-competition.net. Polina Turishcheva, Paul G. Fahey, Michaela Vystrcilová, Laura Hansel, Rachel Froebe, Kayla Ponder, Yongrong Qiu, Konstantin Willeke, Mohammad Bashiri, Ruslan Baikulov, Yu Zhu 0008, Lei Ma 0008, Tiejun Huang 0001, Bryan Li, Wolf De Wulf, Nina Kudryashova, Matthias H. Hennig, Nathalie Rochefort, Arno Onken, Eric Y. Wang, Zhiwei Ding, Andreas S. Tolias, Fabian H. Sinz, Alexander S. Ecker |
NeurIPS | 12 |
| 2024 | Correlation-aware Encoder-Decoder with Adapters for SVBRDF AcquisitionabstractFig. 1.By modeling the correlation among input images with an encoder, together with an adapter-equipped decoder, our network achieves high-quality SVBRDF recovery on both isotropic and anisotropic (with roughness encoded in red and green channels) materials.Here we show re-rendered views for four materials under environment illumination.(Please use Adobe Acrobat and click the renderings to see the animation.)Capturing materials from the real world avoids laborious manual material authoring.However, recovering high-fidelity Spatially Varying Bidirectional Reflectance Distribution Function (SVBRDF) maps from a few captured images is challenging due to its ill-posed nature.Existing approaches have made extensive efforts to alleviate this ambiguity issue by leveraging generative models with latent space optimization or extracting features with variant encoder-decoders.Albeit the rendered images at input views can match input images, the problematic decomposition among maps leads to significant differences when rendered under novel views/lighting.We observe that for human eyes, besides individual images, the correlation (or the highlights variation) among input images also serves as an important hint to recognize * Contribute equally. Hanxiao Sun, Lei Ma 0008, Jian Yang 0003, Beibei Wang 0002 |
SIGGRAPH Asia | 3 |
| 2024 | Light Flickering Guided Reflection Removal
Yuchen Hong, Yakun Chang, Jinxiu Liang, Lei Ma 0008, Tiejun Huang 0001, Boxin Shi |
Int. J. Comput. Vis. | 4 |
| 2024 | An improved hierarchical deep reinforcement learning algorithm for multi-intelligent vehicle lane change
Hongbo Gao 0001, Chengbo Wang 0001, Lin Zhou 0012, Yafei Wang 0002, Lei Ma 0008, Bo Cheng 0003, Zhenyu Wu 0007, Yuansheng Li |
Neurocomputing | 7 |
| 2024 | Embedded prompt tuning: Towards enhanced calibration of pretrained models for medical images
Wenqiang Zu, Shenghao Xie 0002, Lei Ma 0008 |
Medical Image Anal. | 5 |
| 2024 | A Bayesian Approach Toward Robust Multidimensional Ellipsoid-Specific FittingabstractThis work presents a novel and effective method for fitting multidimensional ellipsoids (i.e., ellipsoids embedded in [Formula: see text]) to scattered data in the contamination of noise and outliers. Unlike conventional algebraic or geometric fitting paradigms that assume each measurement point is a noisy version of its nearest point on the ellipsoid, we approach the problem as a Bayesian parameter estimate process and maximize the posterior probability of a certain ellipsoidal solution given the data. We establish a more robust correlation between these points based on the predictive distribution within the Bayesian framework, i.e., considering each model point as a potential source for generating each measurement. Concretely, we incorporate a uniform prior distribution to constrain the search for primitive parameters within an ellipsoidal domain, ensuring ellipsoid-specific results regardless of inputs. We then establish the connection between measurement point and model data via Bayes' rule to enhance the method's robustness against noise. Due to independent of spatial dimensions, the proposed method not only delivers high-quality fittings to challenging elongated ellipsoids but also generalizes well to multidimensional spaces. To address outlier disturbances, often overlooked by previous approaches, we further introduce a uniform distribution on top of the predictive distribution to significantly enhance the algorithm's robustness against outliers. Thanks to the uniform prior, our maximum a posterior probability coincides with a more tractable maximum likelihood estimation problem, which is subsequently solved by a numerically stable Expectation Maximization (EM) framework. Moreover, we introduce an ε-accelerated technique to expedite the convergence of EM considerably. We also investigate the relationship between our algorithm and conventional least-squares-based ones, during which we theoretically prove our method's superior robustness. To the best of our knowledge, this is the first comprehensive method capable of performing multidimensional ellipsoid-specific fitting within the Bayesian optimization paradigm under diverse disturbances. We evaluate it across lower and higher dimensional spaces in the presence of heavy noise, outliers, and substantial variations in axis ratios. Also, we apply it to a wide range of practical applications such as microscopy cell counting, 3D reconstruction, geometric shape approximation, and magnetometer calibration tasks. In all these test contexts, our method consistently delivers flexible, robust, ellipsoid-specific performance, and achieves the state-of-the-art results. Mingyang Zhao 0001, Xiaohong Jia 0001, Lei Ma 0008, Yuke Shi, Jingen Jiang 0001, Qizhai Li, Dong-Ming Yan 0001, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Coherent chord computation and cross ratio for accurate ellipse detection
Mingyang Zhao 0001, Xiaohong Jia 0001, Lei Ma 0008, Liming Hu, Dong-Ming Yan 0001 |
Pattern Recognit. | 3 |
| 2024 | Spike Camera Image Reconstruction Using Deep Spiking Neural NetworksabstractSpike camera is a bio-inspired sensor with ultra-high temporal resolution and low energy consumption. It captures visual signals using an “integrate-and-fire" mechanism and outputs a continuous stream of binary spikes. Reconstructing image sequence from spikes streams is critical for spike camera. Several reconstruction methods have been proposed in recent years. However, the computational cost of these methods is relatively high. Inspired by the fact that spiking neural networks (SNNs) are energy efficient and support time-series signal processing inherently, we propose a lightweight SNN for spike camera image reconstruction (abbreviated to SSIR). Experimental results show that SSIR achieves comparable performance with the state-of-the-art (SOTA) methods at much lower computation and energy cost. Rui Zhao 0010, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Shuyuan Zhu, Lei Ma 0008, Tiejun Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Efficient Video Transformers via Spatial-temporal Token Merging for Action RecognitionabstractTransformer has exhibited promising performance in various video recognition tasks but brings a huge computational cost in modeling spatial-temporal cues. This work aims to boost the efficiency of existing video transformers for action recognition through eliminating redundancies in their tokens and efficiently learning motion cues of moving objects. We propose a lightweight and plug-and-play module, namely Spatial-temporal Token Merger (STTM), to merge the tokens belonging to the same object into a more compact representation. STTM first adaptively identifies crucial object clues underlying the video as meta tokens. Similarity scores between input tokens and meta tokens are hence computed and used to guide the fusion of similar tokens in both spatial and temporal domains, respectively. To compensate for motion cues lost in the merging procedure, we compute the linear aggregation of spatial-temporal positions of tokens as motion features. STTM hence outputs a compact set of tokens fusing both appearance and motion features of moving objects. This procedure substantially decreases the number of tokens that need to be processed by each Transformer block and boosts the efficiency. As a general module, STTM can be applied to different layers of various video Transformers. Extensive experiments on the action recognition datasets Kinectics-400 and SSv2 demonstrate its promising performance. For example, it reduces the computation complexity of ViT by 38% while maintaining a similar performance on Kinectics-400. It also brings 1.7% gains of top-1 accuracy on SSv2 under the same computational cost. Zhanzhou Feng, Lei Ma 0008, Shiliang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Accurate Registration of Cross-Modality Geometry via Consistent ClusteringabstractThe registration of unitary-modality geometric data has been successfully explored over past decades. However, existing approaches typically struggle to handle cross-modality data due to the intrinsic difference between different models. To address this problem, in this article, we formulate the cross-modality registration problem as a consistent clustering process. First, we study the structure similarity between different modalities based on an adaptive fuzzy shape clustering, from which a coarse alignment is successfully operated. Then, we optimize the result using fuzzy clustering consistently, in which the source and target models are formulated as clustering memberships and centroids, respectively. This optimization casts new insight into point set registration, and substantially improves the robustness against outliers. Additionally, we investigate the effect of fuzzier in fuzzy clustering on the cross-modality registration problem, from which we theoretically prove that the classical Iterative Closest Point (ICP) algorithm is a special case of our newly defined objective function. Comprehensive experiments and analysis are conducted on both synthetic and real-world cross-modality datasets. Qualitative and quantitative results demonstrate that our method outperforms state-of-the-art approaches with higher accuracy and robustness. Our code is publicly available at https://github.com/zikai1/CrossModReg. Mingyang Zhao 0001, Xiaoshui Huang, Jingen Jiang 0001, Luntian Mou, Dong-Ming Yan 0001, Lei Ma 0008 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing BiasabstractThe scarcity of data presents a critical obstacle to the efficacy of medical vision-language pre-training (VLP). A potential solution lies in the combination of datasets from various language communities.
Nevertheless, the main challenge stems from the complexity of integrating diverse syntax and semantics, language-specific medical terminology, and culture-specific implicit knowledge. Therefore, one crucial aspect to consider is the presence of community bias caused by different languages.
This paper presents a novel framework named Unifying Cross-Lingual Medical Vision-Language Pre-Training (\textbf{Med-UniC}), designed to integrate multi-modal medical data from the two most prevalent languages, English and Spanish.
Specifically, we propose \textbf{C}ross-lingual \textbf{T}ext Alignment \textbf{R}egularization (\textbf{CTR}) to explicitly unify cross-lingual semantic representations of medical reports originating from diverse language communities.
\textbf{CTR} is optimized through latent language disentanglement, rendering our optimization objective to not depend on negative samples, thereby significantly mitigating the bias from determining positive-negative sample pairs within analogous medical reports. Furthermore, it ensures that the cross-lingual representation is not biased toward any specific language community.
\textbf{Med-UniC} reaches superior performance across 5 medical image tasks and 10 datasets encompassing over 30 diseases, offering a versatile framework for unifying multi-modal medical data within diverse linguistic communities.
The experimental outcomes highlight the presence of community bias in cross-lingual VLP. Reducing this bias enhances the performance not only in vision-language tasks but also in uni-modal visual tasks. Zhongwei Wan, Che Liu 0002, Mi Zhang 0002, Jie Fu 0001, Benyou Wang, Sibo Cheng, Lei Ma 0008, César Quilodrán Casas, Rossella Arcucci |
NeurIPS | 7 |
| 2023 | Unsupervised Optical Flow Estimation with Dynamic Timing Representation for Spike CameraabstractEfficiently selecting an appropriate spike stream data length to extract precise information is the key to the spike vision tasks. To address this issue, we propose a dynamic timing representation for spike streams. Based on multi-layers architecture, it applies dilated convolutions on temporal dimension to extract features on multi-temporal scales with few parameters. And we design layer attention to dynamically fuse these features. Moreover, we propose an unsupervised learning method for optical flow estimation in a spike-based manner to break the dependence on labeled data. In addition, to verify the robustness, we also build a spike-based synthetic validation dataset for extreme scenarios in autonomous driving, denoted as SSES dataset. It consists of various corner cases. Experiments show that our method can predict optical flow from spike streams in different high-speed scenes, including real scenes. For instance, our method achieves $15\%$ and $19\%$ error reduction on PHM dataset compared to the best spike-based work, SCFlow, in $\Delta t=10$ and $\Delta t=20$ respectively, using the same settings as in previous works. The source code and dataset are available at \href{https://github.com/Bosserhead/USFlow}{https://github.com/Bosserhead/USFlow}. Lujie Xia, Ziluo Ding, Rui Zhao 0010, Jiyuan Zhang 0005, Lei Ma 0008, Zhaofei Yu, Tiejun Huang 0001, Ruiqin Xiong |
NeurIPS | 5 |
| 2023 | PS-Net: human perception-guided segmentation network for EM cell membraneabstractMOTIVATION: Cell membrane segmentation in electron microscopy (EM) images is a crucial step in EM image processing. However, while popular approaches have achieved performance comparable to that of humans on low-resolution EM datasets, they have shown limited success when applied to high-resolution EM datasets. The human visual system, on the other hand, displays consistently excellent performance on both low and high resolutions. To better understand this limitation, we conducted eye movement and perceptual consistency experiments. Our data showed that human observers are more sensitive to the structure of the membrane while tolerating misalignment, contrary to commonly used evaluation criteria. Additionally, our results indicated that the human visual system processes images in both global-local and coarse-to-fine manners. RESULTS: Based on these observations, we propose a computational framework for membrane segmentation that incorporates these characteristics of human perception. This framework includes a novel evaluation metric, the perceptual Hausdorff distance (PHD), and an end-to-end network called the PHD-guided segmentation network (PS-Net) that is trained using adaptively tuned PHD loss functions and a multiscale architecture. Our subjective experiments showed that the PHD metric is more consistent with human perception than other criteria, and our proposed PS-Net outperformed state-of-the-art methods on both low- and high-resolution EM image datasets as well as other natural image datasets. AVAILABILITY AND IMPLEMENTATION: The code and dataset can be found at https://github.com/EmmaSRH/PS-Net. Ruohua Shi, Keyan Bi, Lei Ma 0008, Fang Fang 0003, Ling-Yu Duan, Tingting Jiang 0001, Tiejun Huang 0001 |
Bioinform. | 4 |
| 2023 | Hybrid High Dynamic Range Imaging fusing Neuromorphic and Conventional ImagesabstractReconstruction of high dynamic range image from a single low dynamic range image captured by a conventional RGB camera, which suffers from over- or under-exposure, is an ill-posed problem. In contrast, recent neuromorphic cameras like event camera and spike camera can record high dynamic range scenes in the form of intensity maps, but with much lower spatial resolution and no color information. In this article, we propose a hybrid imaging system (denoted as NeurImg) that captures and fuses the visual information from a neuromorphic camera and ordinary images from an RGB camera to reconstruct high-quality high dynamic range images and videos. The proposed NeurImg-HDR+ network consists of specially designed modules, which bridges the domain gaps on resolution, dynamic range, and color representation between two types of sensors and images to reconstruct high-resolution, high dynamic range images and videos. We capture a test dataset of hybrid signals on various HDR scenes using the hybrid camera, and analyze the advantages of the proposed fusing strategy by comparing it to state-of-the-art inverse tone mapping methods and merging two low dynamic range images approaches. Quantitative and qualitative experiments on both synthetic data and real-world scenarios demonstrate the effectiveness of the proposed hybrid high dynamic range imaging system. Code and dataset can be found at: https://github.com/hjynwa/NeurImg-HDR. Jin Han 0001, Yixin Yang 0008, Peiqi Duan 0002, Chu Zhou, Lei Ma 0008, Chao Xu 0006, Tiejun Huang 0001, Imari Sato, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Driver Emotion Recognition With a Hybrid Attentional Multimodal Fusion FrameworkabstractNegative emotions may induce dangerous driving behaviors leading to extremely serious traffic accidents. Therefore, it is necessary to establish a system that can automatically recognize driver emotions so that some actions can be taken to avoid traffic accidents. Existing studies on driver emotion recognition have mainly used facial data and physiological data. However, there are fewer studies on multimodal data with contextual characteristics of driving. In addition, fully fusing multimodal data in the feature fusion layer to improve the performance of emotion recognition is still a challenge. To this end, we propose to recognize driver emotion using a novel multimodal fusion framework based on convolutional long-short term memory network (ConvLSTM), and hybrid attention mechanism to fuse non-invasive multimodal data of eye, vehicle, and environment. In order to verify the effectiveness of the proposed method, extensive experiments have been carried out on a dataset collected using an advanced driving simulator. The experimental results demonstrate the effectiveness of the proposed method. Finally, a preliminary exploration on the correlation between driver emotion and stress is performed. Luntian Mou, Yiyuan Zhao, Bahareh Nakisa, Mohammad Naim Rastgoo, Lei Ma 0008, Tiejun Huang 0001, Ramesh Jain 0001, Wen Gao 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | AMSA: Adaptive Multimodal Learning for Sentiment AnalysisabstractEfficient recognition of emotions has attracted extensive research interest, which makes new applications in many fields possible, such as human-computer interaction, disease diagnosis, service robots, and so forth. Although existing work on sentiment analysis relying on sensors or unimodal methods performs well for simple contexts like business recommendation and facial expression recognition, it does far below expectations for complex scenes, such as sarcasm, disdain, and metaphors. In this article, we propose a novel two-stage multimodal learning framework, called AMSA, to adaptively learn correlation and complementarity between modalities for dynamic fusion, achieving more stable and precise sentiment analysis results. Specifically, a multiscale attention model with a slice positioning scheme is proposed to get stable quintuplets of sentiment in images, texts, and speeches in the first stage. Then a Transformer-based self-adaptive network is proposed to assign weights flexibly for multimodal fusion in the second stage and update the parameters of the loss function through compensation iteration. To quickly locate key areas for efficient affective computing, a patch-based selection scheme is proposed to iteratively remove redundant information through a novel loss function before fusion. Extensive experiments have been conducted on both machine weakly labeled and manually annotated datasets of self-made Video-SA, CMU-MOSEI, and CMU-MOSI. The results demonstrate the superiority of our approach through comparison with baselines. Luntian Mou, Lei Ma 0008, Tiejun Huang 0001, Wen Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | On-line Single Machine Scheduling with Release Dates and Submodular Rejection Penalties
Yaoyu Zhu, Weidong Li 0002, Lei Ma 0008 |
AAIM | 4 |
| 2022 | Effect of arsenic trioxide on human ventricular myocytes: a model studyabstractArsenic trioxide $(As2\mathrm{O}_{3}$), an antileukemia drug, has been used to treat acute promyelocytic leukemia (APL) for more than fifty years, and its therapeutic effect has been elucidated at the molecular level. However, several side effects were observed in APL patients administrated with $As2\mathrm{O}_{3}$, such as long QT (LQT) syndrome, torsade de pointes tachycardia, and even sudden cardiac death. This means that the clinically relevant dosage may induce severe cardiotoxicity. Accordingly, it is essential to determine the underlying mechanisms of arrhythmia induced by $\mathrm{As}2\mathrm{O}_{3}$. Some biological experiments indicated that $\mathrm{As}2\mathrm{O}_{3}$ can impair the human ether-à-go-gorelated gene (hERG), thus inhibiting rapid delayed rectifier potassium current $(I_{Kr})$ and prolonging action potential duration (APD), which was regarded as the reason for LQT syndrome. However, previous experiments did not illuminate the deep mechanisms of $\mathrm{As}2\mathrm{O}_{3}$-induced side effects, which is important in clinical treatment. In addition, the experimental data were restricted to animal studies, so human cellular data were lacking. In this study, we investigated $\mathrm{As}2\mathrm{O}_{3}$-related cardiotoxicity through a human ventricular model study. Based on the current experimental data, the effects of $\mathrm{As}2\mathrm{O}_{3}$ on ventricular myocytes (VMs) were predicted at various $\mathrm{As}2\mathrm{O}_{3}$ concentrations. In addition, the potential hazard of $\mathrm{As}2\mathrm{O}_{3}$ was simulated and illustrated under different stimulation protocols. Moreover, electrocardiograms (ECGs) were estimated in heterogeneous ventricular cables, by which the clinical phenomenon was verified and explained. Based on the present modeling study, deep reasons for arrhythmia caused by $\mathrm{As}2\mathrm{O}_{3}$ were uncovered. $\mathrm{As}2\mathrm{O}_{3}$ not only led to a prolonged APD but also alternated action potentials and exacerbated heterogeneity among VMs. Moreover, the degree of arrhythmia risk was susceptible to $\mathrm{As}2\mathrm{O}_{3}$ dosage. These new findings may provide targets for attenuating $\mathrm{As}2\mathrm{O}_{3}$ toxicity and may help to improve the APL therapeutic regimen. Yacong Li, Jun Liu 0080, Runlan Wan, Lei Ma 0008, Henggui Zhang |
BIBM | 4 |
| 2022 | Effect of cell coupling between pacemaker cells on the biological pacemaker in cardiac tissue modelabstractBiological pacemaker is a therapy for cardiac rhythm disease, which can be transformed from ventricular myocytes (VMs) by overexpressing HCN gene which codes the expression of hyperpolarization-activated current (${\mathrm {I}}_{\mathrm{f}}$) and knocking off Kir2.1 gene which codes inward-rectifier potassium current (${\mathrm {I}}_{\mathrm{K1}}$). Our previous study built a biological pacemaker single cell model and clarified the underlying mechanisms of how gene expressing levels influence the pacemaking activity of single pacemaker cell. But the pacemaking ability of pacemaker tissue has not been researched systematically. And what factors may have effects on pacemaker’s synchronization and spontaneous beating propagation are not clear. Biological research indicated that both sinoatrial node and pacemaker cells has less expression of connexin than unexcitable cardiac cells, which provides a possibility that improve pacemaking ability of pacemaker by decreasing its cell coupling. Another possible factor is the number of pacemaker cells. According to the common sense, increasing cell number can promote pacemaking behaviours, but overmuch pacemaker cells is unreasonable in clinic. As a result, the balance between pacemaker number and cell coupling is important when applying biological pacemaker. In this study, we constructed a two-dimensional cardiac tissue model with the description of electrophysiology to illustrate the relationship between gap junction and cell number. Based on this model, we modified the cell coupling between pacemaker cells by adjusting the diffusion coefficient of tissue with different pacemaker number. In different condition, the synchronization, pacemaking cycle length and electrical signal propagation were evaluated. It can be concluded that weakening cell coupling among pacemaker cells can lift the efficiency of bio-pacemaker therapy. This study may contribute to produce effective pacemaker in clinic. Yacong Li, Lei Ma 0008, Qince Li, Henggui Zhang, Kuanquan Wang |
BIBM | 2 |
| 2022 | Optical Flow Estimation for Spiking CameraabstractAs a bio-inspired sensor with high temporal resolution, the spiking camera has an enormous potential in real applications, especially for motion estimation in high-speed scenes. However, frame-based and event-based methods are not well suited to spike streams from the spiking camera due to the different data modalities. To this end, we present, SCFlow, a tailored deep learning pipeline to estimate optical flow in high-speed scenes from spike streams. Importantly, a novel input representation is introduced which can adaptively remove the motion blur in spike streams according to the prior motion. Further, for training SCFlow, we synthesize two sets of optical flow data for the spiking camera, SPIkingly Flying Things and Photo-realistic Highspeed Motion, denoted as SPIFT and PHM respectively, corresponding to random high-speed and well-designed scenes. Experimental results show that the SCFlow can predict optical flow from spike streams in different high-speed scenes. Moreover, SCFlow shows promising generalization on real spike streams. Codes and datasets refer to https://github.com/Acnext/Optical-Flow-For-Spiking-Camera. Liwen Hu 0002, Rui Zhao 0010, Ziluo Ding, Lei Ma 0008, Boxin Shi, Ruiqin Xiong, Tiejun Huang 0001 |
CVPR | 4 |
| 2022 | Modeling The Detection Capability Of High-Speed Spiking CamerasabstractThe novel working principle enables spiking cameras to capture high-speed moving objects. However, the applications of spiking cameras can be affected by many factors, such as brightness intensity, detectable distance, and the maximum speed of moving targets. Improper settings such as weak ambient brightness and too short object-camera distance, will lead to failure in the application of such cameras. To address the issue, this paper proposes a modeling algorithm that studies the detection capability of spiking cameras. The algorithm deduces the maximum detectable speed of spiking cameras corresponding to different scenario settings (e.g., brightness intensity, camera lens, and object-camera distance) based on the basic technical parameters of cameras (e.g., pixel size, spatial and temporal resolution). Thereby, the proper camera settings for various applications can be determined. Extensive experiments verify the effectiveness of the modeling algorithm. To our best knowledge, it is the first work to investigate the detection capability of spiking cameras. Junwei Zhao 0003, Zhaofei Yu, Lei Ma 0008, Ziluo Ding, Shiliang Zhang, Yonghong Tian 0001, Tiejun Huang 0001 |
ICASSP | 3 |
| 2022 | SpikingSIM: A Bio-Inspired Spiking SimulatorabstractLarge-scale neuromorphic dataset is costly to construct and difficult to annotate because of the unique high-speed asynchronous imaging principle of bio-inspired cameras. Lacking of large-scale annotated neuromorphic datasets has significantly hindered the applications of bio-inspired cameras in deep neural networks. Synthesizing neuromorphic data from annotated RGB images can be considered to alleviate this challenge. This paper proposes a simulator to generate simulated spiking data from images recorded by frame cameras. To minimize the deviations between synthetic data and real data, the proposed simulator named SpikingSIM considers the sensing principle of spiking cameras, and generates high-quality simulated spiking data, e.g., the noises in real data are also simulated. Experimental results show that, our simulator generates more realistic spiking data than existing methods. We hence train deep neural networks with synthesized spiking data. Experiments show that, the net- work trained by our simulated data generalizes well on real spiking data. The source code of SpikingSIM is available at http://github.com/Evin-X/SpikingSIM. Junwei Zhao 0003, Shiliang Zhang, Lei Ma 0008, Zhaofei Yu, Tiejun Huang 0001 |
ISCAS | 3 |
| 2022 | Structure inference of networked system with the synergy of deep residual network and fully connected layer network
Keke Huang, Wenfeng Deng, Zhaofei Yu, Lei Ma 0008 |
Neural Networks | 5 |
| 2022 | GraphReg: Dynamical Point Cloud Registration With Geometry-Aware Graph Signal ProcessingabstractThis study presents a high-accuracy, efficient, and physically induced method for 3D point cloud registration, which is the core of many important 3D vision problems. In contrast to existing physics-based methods that merely consider spatial point information and ignore surface geometry, we explore geometry aware rigid-body dynamics to regulate the particle (point) motion, which results in more precise and robust registration. Our proposed method consists of four major modules. First, we leverage the graph signal processing (GSP) framework to define a new signature, i.e., point response intensity for each point, by which we succeed in describing the local surface variation, resampling keypoints, and distinguishing different particles. Then, to address the shortcomings of current physics-based approaches that are sensitive to outliers, we accommodate the defined point response intensity to median absolute deviation (MAD) in robust statistics and adopt the X84 principle for adaptive outlier depression, ensuring a robust and stable registration. Subsequently, we propose a novel geometric invariant under rigid transformations to incorporate higher-order features of point clouds, which is further embedded for force modeling to guide the correspondence between pairwise scans credibly. Finally, we introduce an adaptive simulated annealing (ASA) method to search for the global optimum and substantially accelerate the registration process. We perform comprehensive experiments to evaluate the proposed method on various datasets captured from range scanners to LiDAR. Results demonstrate that our proposed method outperforms representative state-of-the-art approaches in terms of accuracy and is more suitable for registering large-scale point clouds. Furthermore, it is considerably faster and more robust than most competitors. Our implementation is publicly available at https://github.com/zikai1/GraphReg. Mingyang Zhao 0001, Lei Ma 0008, Xiaohong Jia 0001, Dong-Ming Yan 0001, Tiejun Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Constant-Cost Spatio-Angular Prefiltering of Glinty Appearance Using Tensor DecompositionabstractThe detailed glinty appearance from complex surface microstructures enhances the level of realism but is both - and time-consuming to render, especially when viewed from far away (large spatial coverage) and/or illuminated by area lights (large angular coverage). In this article, we formulate the glinty appearance rendering process as a spatio-angular range query problem of the Normal Distribution Functions (NDFs), and introduce an efficient spatio-angular prefiltering solution to it. We start by exhaustively precomputing all possible NDFs with differently sized positional coverages. Then we compress the precomputed data using tensor rank decomposition, which enables accurate and fast angular range queries. With our spatio-angular prefiltering scheme, we are able to solve both the storage and performance issues at the same time, leading to efficient rendering of glinty appearance with both constant storage and constant performance, regardless of the range of spatio-angular queries. Finally, we demonstrate that our method easily applies to practical rendering applications that were traditionally considered difficult. For example, efficient bidirectional reflection distribution function evaluation accurate NDF importance sampling, fast global illumination between glinty objects, high-frequency preserving rendering with environment lighting, and tile-based synthesis of glinty appearance. Hong Deng, Yang Liu 0288, Beibei Wang 0002, Jian Yang 0003, Lei Ma 0008, Nicolas Holzschuch, Lingqi Yan 0001 |
ACM Trans. Graph. | 5 |
| 2022 | Parallel Computation of 3D Clipped Voronoi DiagramsabstractComputing the Voronoi diagram of a given set of points in a restricted domain (e.g., inside a 2D polygon, on a 3D surface, or within a volume) has many applications. Although existing algorithms can compute 2D and surface Voronoi diagrams in parallel on graphics hardware, computing clipped Voronoi diagrams within volumes remains a challenge. This article proposes an efficient GPU algorithm to tackle this problem. A preprocessing step discretizes the input volume into a tetrahedral mesh. Then, unlike existing approaches which use the bisecting planes of the Voronoi cells to clip the tetrahedra, we use the four planes of each tetrahedron to clip the Voronoi cells. This strategy drastically simplifies the computation, and as a result, it outperforms state-of-the-art CPU methods up to an order of magnitude. Lei Ma 0008, Jianwei Guo 0003, Dong-Ming Yan 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Robust Ellipsoid-specific Fitting via Expectation Maximization
Mingyang Zhao 0001, Xiaohong Jia 0001, Lei Ma 0008, Xinlin Qiu, Xin Jiang 0008, Dong-Ming Yan 0001 |
BMVC | 3 |
| 2021 | A Deep Learning Method for 2D Image Stippling
Zhongmin Xue, Beibei Wang 0002, Lei Ma 0008 |
CGI | 3 |
| 2021 | DCNet: Dual-Task Cycle Network for End-to-End Image DehazingabstractSingle image dehazing is an important technology in the field of computer vision. In this paper, we propose an image dehazing via dual learning strategy, named dual-task cycle network (DCNet). The core of DCNet is a dual learning framework, which consists of two tasks: the dehazing task and the haze generation task. The dehazing task completes the image dehazing, while the haze generation task achieves the restoration from the dehazed image to the haze image and can form a cycle to provide additional supervision. Our method uses the duality between each task as a constraint to learn and train two tasks jointly, so that the effects of the dehazing model can be improved. Since the haze generation process does not depend on clear images, the DCNet can satisfy the requirements for limited supervision. Extensive experiments demonstrate that our DCNet performs favorably on haze removal. Yu Zhou 0066, Ping Li 0016, Xiaoyu Chi, Lei Ma 0008, Bin Sheng 0001 |
ICME | 5 |
| 2021 | Milliseconds Color StipplingabstractStippling is a popular and fascinating sketching art in stylized illustrations. Various digital stippling techniques have been proposed to reduce tedious manual work. In this paper, we present a novel method to create high-quality color stippling from an input image in milliseconds. The key idea is to obtain stipples with predetermined incremental 2D sample sequences, which algorithms generate with sequential incrementality and distributional uniformity features. Two typical sequences are employed in our work: one is constructed from incremental Voronoi sets, and the other is from Poisson disk distributions. A threshold-based algorithm is then applied to determine stipple appearance and guarantee result quality. We extend color stippling with multitone level and radius adjustment to achieve improved visual quality. Detailed comparisons of the two sequences are conducted to explore further the strengths and weaknesses of the proposed method. For more information, please visit https://gitlab.com/maleiwhat/milliseconds-color-stippling. Lei Ma 0008, Yanyun Chen |
ACM Multimedia | 1 |
| 2020 | Simulation of multi-solvent stains on textile
Lei Ma 0008, Yanyun Chen, Guangzheng Fei, Bin Sheng 0001, Enhua Wu |
Vis. Comput. | 2 |
| 2019 | Layered leaf texturing using structure-guided model
Yinling Qian, Hanqiu Sun, Lei Ma 0008, Yanyun Chen, Qiong Wang 0001, Pheng-Ann Heng |
Graph. Model. | 4 |
| 2018 | Instant Stippling on 3D ScenesabstractAbstract In this paper, we present a novel real‐time approach to generate high‐quality stippling on 3D scenes. The proposed method is built on a precomputed 2D sample sequence called incremental Voronoi set with blue‐noise properties. A rejection sampling scheme is then applied to achieve tone reproduction, by thresholding the sample indices proportional to the inverse target tonal value to produce a suitable stipple density. Our approach is suitable for stippling large‐scale or even dynamic scenes because the thresholding of individual stipples is trivially parallelizable. In addition, the static nature of the underlying sequence benefits the frame‐to‐frame coherence of the stippling. Finally, we propose an extension that supports stipples of varying sizes and tonal values, leading to smoother spatial and temporal transitions. Experimental results reveal that the temporal coherence and real‐time performance of our approach are superior to those of previous approaches. Lei Ma 0008, Jianwei Guo 0003, Dong-Ming Yan 0001, Hanqiu Sun, Yanyun Chen |
Comput. Graph. Forum | 1 |
| 2018 | Incremental Voronoi sets for instant stippling
Lei Ma 0008, Yanyun Chen, Yinling Qian, Hanqiu Sun |
Vis. Comput. | 1 |