VLDB 2026 Research / reviewers in the wild / expert
Zhongpai Gao
dblp:149/4942
· DBLP profile ↗
49ranked-venue papers
18as first author
33since 2021 · last 2026
0000-0003-4344-4501ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 11 first-author · 22 since 2021Artificial intelligence and machine learning · 22 · 9 first-author · 22 since 2021Systems, architecture and hardware · 5 · 2 first-authorComputer networks · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Mesh Representation Learning With Spectral Dictionary EmbeddingabstractLearning mesh representation is important for many 3D tasks. Conventional convolution for regular data (i.e., images) cannot directly be applied to meshes since each vertex's neighbors are unordered. Previous methods use isotropic filters or predefined local coordinate systems or learning weighting matrices for each template vertex to overcome the irregularity. Learning weighting matrices to resample the vertex's neighbors into an implicit canonical order is the most effective way to capture the local structure of each vertex. However, learning weighting matrices for each vertex increases the model size linearly with the vertex number. Thus, large parameters are required for high-resolution 3D shapes, which is not favorable for many applications. In this paper, we learn spectral dictionary (i.e., bases) for the weighting matrices such that the model size is independent of the resolution of 3D shapes. The coefficients of the weighting matrix bases are learned from the spectral features of the template and its hierarchical levels in a weight-sharing manner. Furthermore, we introduce an adaptive sampling method that learns the hierarchical mapping matrices directly to improve the performance without increasing the model size at the inference stage. Comprehensive experiments demonstrate that our model produces state-of-the-art results with a much smaller model size. Zhongpai Gao, Junchi Yan, Tianyu Luan, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | FreeQR: Free Lunch for Aesthetic QR Codes Emerging From the Latent Space in Diffusion ModelsabstractIn the modern digital age, Quick Response (QR) codes serve as a critical interface for bridging the physical and virtual worlds, widely utilized in multimedia applications. However, traditional binary QR codes often lack the visual appeal desired in contexts. Aesthetic QR codes address this limitation by enabling the customization of QR code patterns to enhance visual attractiveness while retaining compatibility with standard QR decoders. Previous works have explored the use of diffusion models for generating such codes but often require extensive training of ControlNets and face challenges in maintaining scannability. To address these issues, we present FreeQR, a streamlined and effective approach that enables the stable generation of QR code images with diffusion models. Our methodology involves the strategic fusion between the specific channel in the latent space of the denoising process with the noised latent representations of the QR blueprint image at corresponding timesteps. This ensures that the generated images adhere to the brightness distribution required for effective scanning while achieving a balance between aesthetics and functionality. Additionally, we introduce gradient guidance based on scanning errors directly in the latent space, enabling the generation of scannable QR codes in seconds without additional model parameters. Experimental results demonstrate that FreeQR significantly enhances the aesthetics and scannability of QR codes compared to existing methods, making it a lightweight and efficient solution for multimedia applications. Yiwei Yang 0007, Jun Jia, Zheyuan Liu 0011, Zhongpai Gao, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 4 |
| 2025 | Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal GroundingabstractTemporal awareness is essential for video large language models (LLMs) to understand and reason about events within long videos, enabling applications like dense video captioning and temporal video grounding in a unified system. However, the scarcity of long videos with detailed captions and precise temporal annotations limits their temporal awareness. In this paper, we propose Seq2Time, a data-oriented training paradigm that leverages sequences of images and short video clips to enhance temporal awareness in long videos. By converting sequence positions into temporal annotations, we transform large-scale image and clip captioning datasets into sequences that mimic the temporal structure of long videos, enabling self-supervised training with abundant time-sensitive data. To enable sequence-to-time knowledge transfer, we introduce a novel time representation that unifies positional information across image sequences, clip sequences, and long videos. Experiments demonstrate the effectiveness of our method, achieving a 27.6% improvement in F1 score and 44.8% in CIDEr on the YouCook2 benchmark and a 14.7% increase in recall on the Charades-STA benchmark compared to the baseline. Project available at: https://seq2time.github.io/ Andong Deng, Zhongpai Gao, Anwesa Choudhuri, Benjamin Planche, Meng Zheng 0002, Bin Wang 0068, Terrence Chen, Chen Chen 0001, Ziyan Wu 0001 |
CVPR | 2 |
| 2025 | CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single ImageabstractReconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on the assumption that the human subject is in an occlusion-free environment. Thus, when encountering in-the-wild occluded images, these algorithms produce multiview inconsistent and fragmented reconstructions. Additionally, most algorithms for monocular 3D human reconstruction leverage geometric priors such as SMPL annotations for training and inference, which are extremely challenging to acquire in real-world applications. To address these limitations, we propose CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-ConsistEncy from a Single Image, a novel pipeline designed to reconstruct occlusion-resilient 3D humans with multiview consistency from a single occluded image, without requiring either ground-truth geometric prior annotations or 3D supervision. Specifically, CHROME leverages a multiview diffusion model to first synthesize occlusion-free human images from the occluded input, compatible with off-the-shelf pose control to explicitly enforce cross-view consistency during synthesis. A 3D reconstruction model is then trained to predict a set of 3D Gaussians conditioned on both the occluded input and synthesized views, aligning cross-view details to produce a cohesive and accurate 3D representation. CHROME achieves significant improvements in terms of both novel view synthesis (upto 3 db PSNR) and geometric reconstruction under challenging conditions. Arindam Dutta, Meng Zheng 0002, Zhongpai Gao, Benjamin Planche, Anwesa Choudhuri, Terrence Chen, Amit K. Roy-Chowdhury, Ziyan Wu 0001 |
ICCV | 3 |
| 2025 | 7DGS: Unified Spatial-Temporal-Angular Gaussian SplattingabstractReal-time rendering of dynamic scenes with view-dependent effects remains a fundamental challenge in computer graphics. While recent advances in Gaussian Splatting have shown promising results separately handling dynamic scenes (4DGS) and view-dependent effects (6DGS), no existing method unifies these capabilities while maintaining real-time performance. We present 7D Gaussian Splatting (7DGS), a unified framework representing scene elements as seven-dimensional Gaussians spanning position (3D), time (1D), and viewing direction (3D). Our key contribution is an efficient conditional slicing mechanism that transforms 7D Gaussians into view- and time-conditioned 3D Gaussians, maintaining compatibility with existing 3D Gaussian Splatting pipelines while enabling joint optimization. Experiments demonstrate that 7DGS outperforms prior methods by up to 7.36 dB in PSNR while achieving real-time rendering (401 FPS) on challenging dynamic scenes with complex view-dependent effects. The project page is: https://gaozhongpai.github.io/7dgs/. Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Anwesa Choudhuri, Terrence Chen, Ziyan Wu 0001 |
ICCV | 1 |
| 2025 | Order-aware Interactive SegmentationabstractInteractive segmentation aims to accurately segment target objects with minimal user interactions. However, current methods often fail to accurately separate target objects from the background, due to a limited understanding of order, the relative depth between objects in a scene. To address this issue, we propose OIS: order-aware interactive segmentation, where we explicitly encode the relative depth between objects into order maps. We introduce a novel order-aware attention, where the order maps seamlessly guide the user interactions (in the form of clicks) to attend to the image features. We further present an object-aware attention module to incorporate a strong object-level understanding to better differentiate objects with similar order. Our approach allows both dense and sparse integration of user clicks, enhancing both accuracy and efficiency as compared to prior works. Experimental results demonstrate that OIS achieves state-of-the-art performance, improving mIoU after one click by 7.61 on the HQSeg44K dataset and 1.32 on the DAVIS dataset as compared to the previous state-of-the-art SegNext, while also doubling inference speed compared to current leading methods. Bin Wang 0068, Anwesa Choudhuri, Meng Zheng 0002, Zhongpai Gao, Benjamin Planche, Andong Deng, Qin Liu 0008, Terrence Chen, Ulas Bagci, Ziyan Wu 0001 |
ICLR | 4 |
| 2025 | 6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric RenderingabstractNovel view synthesis has advanced significantly with the development of neural radiance fields (NeRF) and 3D Gaussian splatting (3DGS). However, achieving high quality without compromising real-time rendering remains challenging, particularly for physically-based rendering using ray/path tracing with view-dependent effects. Recently, N-dimensional Gaussians (N-DG) introduced a 6D spatial-angular representation to better incorporate view-dependent effects, but the Gaussian representation and control scheme are sub-optimal. In this paper, we revisit 6D Gaussians and introduce 6D Gaussian Splatting (6DGS), which enhances color and opacity representations and leverages the additional directional information in the 6D space for optimized Gaussian control. Our approach is fully compatible with the 3DGS framework and significantly improves real-time radiance field rendering by better modeling view-dependent effects and fine details. Experiments demonstrate that 6DGS significantly outperforms 3DGS and N-DG, achieving up to a 15.73 dB improvement in PSNR with a reduction of 66.5\% Gaussian points compared to 3DGS. The project page is: https://gaozhongpai.github.io/6dgs/. Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Anwesa Choudhuri, Terrence Chen, Ziyan Wu 0001 |
ICLR | 1 |
| 2025 | 3D Vision-Language Gaussian SplattingabstractRecent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches have naively embedded semantic representations into 3D reconstruction methods without striking a balance between visual and language modalities, which leads to unsatisfying semantic rasterization of translucent or reflective objects, as well as over-fitting on color modality. To alleviate these limitations, we propose a solution that adequately handles the distinct visual and semantic modalities, i.e., a 3D vision-language Gaussian splatting model for scene understanding, to put emphasis on the representation learning of language modality. We propose a novel cross-modal rasterizer, using modality fusion along with a smoothed semantic indicator for enhancing semantic rasterization. We also employ a camera-view blending technique to improve semantic consistency between existing and synthesized views, thereby effectively mitigating over-fitting. Extensive experiments demonstrate that our method achieves state-of-the-art performance in open-vocabulary semantic segmentation, surpassing existing methods by a significant margin. Qucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng 0002, Anwesa Choudhuri, Terrence Chen, Chen Chen 0001, Ziyan Wu 0001 |
ICLR | 3 |
| 2025 | PolypSegTrack: Unified Foundation Model for Colonoscopy Video Analysis
Anwesa Choudhuri, Zhongpai Gao, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Ziyan Wu 0001 |
MICCAI (11) | 2 |
| 2025 | Automated Patient Positioning with Learned 3D Hand Gestures
Zhongpai Gao, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Ziyan Wu 0001 |
WACV | 1 |
| 2024 | Disguise without Disruption: Utility-Preserving Face De-identificationabstractWith the rise of cameras and smart sensors, humanity generates an exponential amount of data. This valuable information, including underrepresented cases like AI in medical settings, can fuel new deep-learning tools. However, data scientists must prioritize ensuring privacy for individuals in these untapped datasets, especially for images or videos with faces, which are prime targets for identification methods. Proposed solutions to de-identify such images often compromise non-identifying facial attributes relevant to downstream tasks. In this paper, we introduce Disguise, a novel algorithm that seamlessly de-identifies facial images while ensuring the usability of the modified data. Unlike previous approaches, our solution is firmly grounded in the domains of differential privacy and ensemble-learning research. Our method involves extracting and substituting depicted identities with synthetic ones, generated using variational mechanisms to maximize obfuscation and non-invertibility. Additionally, we leverage supervision from a mixture-of-experts to disentangle and preserve other utility attributes. We extensively evaluate our method using multiple datasets, demonstrating a higher de-identification rate and superior consistency compared to prior approaches in various downstream tasks. Zikui Cai, Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Terrence Chen, Muhammad Salman Asif, Ziyan Wu 0001 |
AAAI | 2 |
| 2024 | Implicit Modeling of Non-rigid Objects with Cross-Category SignalsabstractDeep implicit functions (DIFs) have emerged as a potent and articulate means of representing 3D shapes. However, methods modeling object categories or non-rigid entities have mainly focused on single-object scenarios. In this work, we propose MODIF, a multi-object deep implicit function that jointly learns the deformation fields and instance-specific latent codes for multiple objects at once. Our emphasis is on non-rigid, non-interpenetrating entities such as organs. To effectively capture the interrelation between these entities and ensure precise, collision-free representations, our approach facilitates signaling between category-specific fields to adequately rectify shapes. We also introduce novel inter-object supervision: an attraction-repulsion loss is formulated to refine contact regions between objects. Our approach is demonstrated on various medical benchmarks, involving modeling different groups of intricate anatomical entities. Experimental results illustrate that our model can proficiently learn the shape representation of each organ and their relations to others, to the point that shapes missing from unseen instances can be consistently recovered by our method. Finally, MODIF can also propagate semantic information throughout the population via accurate point correspondences. Yuchun Liu, Benjamin Planche, Meng Zheng 0002, Zhongpai Gao, Pierre Sibut-Bourde, Fan Yang 0035, Terrence Chen, Ziyan Wu 0001 |
AAAI | 4 |
| 2024 | DaReNeRF: Direction-aware Representation for Dynamic ScenesabstractAddressing the intricate challenge of modeling and re-rendering dynamic scenes, most recent approaches have sought to simplify these complexities using plane-based explicit representations, overcoming the slow training time issues associated with methods like Neural Radiance Fields (NeRF) and implicit representations. However, the straight-forward decomposition of 4D dynamic scenes into multiple 2D plane-based representations proves insufficient for re-rendering high-fidelity scenes with complex motions. In response, we present a novel direction-aware representation (DaRe) approach that captures scene dynamics from six different directions. This learned representation under-goes an inverse dual-tree complex wavelet transformation (DTCWT) to recover plane-based information. DaReNeRF computes features for each space-time point by fusing vectors from these recovered planes. Combining DaReNeRF with a tiny MLP for color regression and leveraging volume rendering in training yield state-of-the-art performance in novel view synthesis for complex dynamic scenes. Notably, to address redundancy introduced by the six real and six imag-inary direction-aware wavelet coefficients, we introduce a trainable masking approach, mitigating storage issues without significant performance decline. Moreover, DaReNeRF maintains a 2 × reduction in training time compared to prior art while delivering superior performance. Ange Lou, Benjamin Planche, Zhongpai Gao, Tianyu Luan, Hao Ding 0021, Terrence Chen, Jack H. Noble, Ziyan Wu 0001 |
CVPR | 3 |
| 2024 | Divide and Fuse: Body Part Mesh Recovery from Partially Visible Human Images
Tianyu Luan, Zhongpai Gao, Luyuan Xie, Hao Ding 0021, Benjamin Planche, Meng Zheng 0002, Ange Lou, Terrence Chen, Junsong Yuan 0001, Ziyan Wu 0001 |
ECCV (24) | 2 |
| 2024 | PBADet: A One-Stage Anchor-Free Approach for Part-Body AssociationabstractThe detection of human parts (e.g., hands, face) and their correct association with individuals is an essential task, e.g., for ubiquitous human-machine interfaces and action recognition. Traditional methods often employ multi-stage processes, rely on cumbersome anchor-based systems, or do not scale well to larger part sets. This paper presents PBADet, a novel one-stage, anchor-free approach for part-body association detection. Building upon the anchor-free object representation across multi-scale feature maps, we introduce a singular part-to-body center offset that effectively encapsulates the relationship between parts and their parent bodies. Our design is inherently versatile and capable of managing multiple parts-to-body associations without compromising on detection accuracy or robustness. Comprehensive experiments on various datasets underscore the efficacy of our approach, which not only outperforms existing state-of-the-art techniques but also offers a more streamlined and efficient solution to the part-body association challenge. Zhongpai Gao, Huayi Zhou 0001, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Ziyan Wu 0001 |
ICLR | 1 |
| 2024 | DiffStega: Towards Universal Training-Free Coverless Image Steganography with Diffusion Models
Yiwei Yang 0007, Zheyuan Liu 0011, Jun Jia, Zhongpai Gao, Wei Sun 0029, Xiaohong Liu 0001, Guangtao Zhai |
IJCAI | 4 |
| 2024 | Few-Shot 3D Volumetric Segmentation with Multi-surrogate Fusion
Meng Zheng 0002, Benjamin Planche, Zhongpai Gao, Terrence Chen, Richard J. Radke, Ziyan Wu 0001 |
MICCAI (9) | 3 |
| 2024 | DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic Volume RenderingabstractDigitally reconstructed radiographs (DRRs) are simulated 2D X-ray images generated from 3D CT volumes, widely used in preoperative settings but limited in intraoperative applications due to computational bottlenecks. Physics-based Monte Carlo simulations provide accurate representations but are extremely computationally intensity. Analytical DRR renderers are much more efficient, but at the price of ignoring anisotropic X-ray image formation phenomena such as Compton scattering. We propose a novel approach that balances realistic physics-inspired X-ray simulation with efficient, differentiable DRR generation using 3D Gaussian splatting (3DGS). Our direction-disentangled 3DGS (DDGS) method decomposes the radiosity contribution into isotropic and direction-dependent components, able to approximate complex anisotropic interactions without complex runtime simulations. Additionally, we adapt the 3DGS initialization to account for tomography data properties, enhancing accuracy and efficiency. Our method outperforms state-of-the-art techniques in image accuracy and inference speed, demonstrating its potential for intraoperative applications and inverse problems like pose registration. Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Terrence Chen, Ziyan Wu 0001 |
NeurIPS | 1 |
| 2024 | Synergetic Assessment of Quality and Aesthetic: Approach and Comprehensive Benchmark DatasetabstractQuantifications of image quality and aesthetic have been regarded as two independent fields in computer vision. Generally, image quality assessment aims at measuring image distortions and image aesthetic is judged by commonly established photography rules. However, either measuring image quality or aesthetic alone is not sufficient to qualitatively rank images. Therefore, this paper puts forward the synergetic assessment of quality and aesthetic to help understand the subjective human preferences of digital pictures more comprehensively. Specifically, considering that the images of existing benchmark datasets are only labeled with single attribute, we first establish a new dataset which contains 9042 real-world images with the corresponding human rated pair-wise quality-aesthetic scores. Previously, these images are only labeled with aesthetic score, and we evaluate the subjective quality score of them, so that it can make up the lack of image dataset with double attributes. Moreover, since the existing methods are mostly designed for individual attribute prediction. We then propose a two-stream learning network to assess both quality and aesthetic of images in parallel. This network follows the top-down perception mechanism which learns from both fined grained details and holistic image layout simultaneously. Furthermore, we introduce a Channel-Diversity loss, which can be deployed in grouped convolution operation, and can constrain channels to be mutually exclusive across the spatial dimensions. To some extent, this contributes to spotlight different local discriminative regions with a finer granularity. Finally, experiments demonstrate that our method outperforms the state-of-the-art methods on our established benchmark dataset and other benchmark datasets in terms of image quality and aesthetic assessment. We hope this paper could serve as a potent reference and be useful for future research on the study of image ranking. Both the benchmark dataset and the code will be publicly available to facilitate further research. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Zhongpai Gao, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Hidden Barcode in Sub-Images with Invisible Locating MarkerabstractThe prevalence of the Internet of Things (IoT) has led to the widespread adoption of 2D barcodes as a means of offline-to-online communication. Whereas, 2D barcodes are not ideal for publicity materials, due to their space-consuming nature. Recent works have proposed 2D image barcodes that contain invisible codes or hyperlinks to transmit hidden information from offline to online. However, these methods undermine the purpose of the codes being invisible, due to the the requirement of markers to locate them. The conference version of this work has presents a novel imperceptible information embedding framework for display or print-camera scenarios, which includes not only hiding and recvoery but also locating and correcting. With the assistance of learned invisible markers, hidden codes can be rendered truly imperceptible. A highly effective multi-stage training scheme is proposed to achieve high visual fidelity and retrieval resiliency, wherein information is concealed in a sub-region rather than the entire image. However, our conference version does not address the optimal sub-region for hiding, which is crucial when dealing with local region concealment problems. In this paper extension, we consider human perceptual characteristics and introduce an optimal hiding region recommendation algorithm that comprehensively incorporates Just Noticeable Difference (JND) and visual saliency factors into consideration. Extensive experiments demonstrate superior visual quality and robustness compared to state-of-the-art methods. With the assistance of our proposed hiding region recommendation algorithm, concealed information becomes even less visible than the results of our conference version without compromising robustness. Jun Jia, Zhongpai Gao, Yiwei Yang 0007, Wei Sun 0029, Dandan Zhu 0001, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Learning Continuous Mesh Representation with Spherical Implicit SurfaceabstractAs the most common representation for 3D shapes, mesh is often stored discretely with arrays of vertices and faces. However, 3D shapes in the real world are presented continuously. In this paper, we propose to learn a continuous representation for meshes with fixed topology, a common and practical setting in many faces-, hand-, and body-related applications. First, we split the template into multiple closed manifold genus-0 meshes so that each genus-0 mesh can be parameterized onto the unit sphere. Then we learn spherical implicit surface (SIS), which takes a spherical coordinate and a global feature or a set of local features around the coordinate as inputs, predicting the vertex corresponding to the coordinate as an output. Since the spherical coordinates are continuous, SIS can depict a mesh in an arbitrary resolution. SIS representation builds a bridge between discrete and continuous representation in 3D shapes. Specifically, we train SIS networks in a self-supervised manner for two tasks: a reconstruction task and a super-resolution task. Experiments show that our SIS representation is comparable with state-of-the-art methods that are specifically designed for meshes with a fixed resolution and significantly outperforms methods that work in arbitrary resolutions. Zhongpai Gao |
FG | 1 |
| 2023 | Computed Tomography and 3-D Face Scan Fusion for IoT-Based Diagnostic SolutionsabstractIn clinical diagnosis, multimodal medical image fusion is meaningful and necessary, for the reason that some diseases need to be diagnosed in combination with the situation of different tissues of patients. Spiral computed tomography (CT) realizes the high precision and smooth reconstruction of bone tissue, while it can not represent the color and texture information in soft tissue reconstruction with high accuracy. The face scan precisely records the color and shape of the maxillofacial region. The diagnosis of some diseases (like cavernous hemangioma and jaw deformity caused by idiopathic condylar resorption) needs to combine the information of maxillofacial soft tissue and bone, so it is of great significance to fuse spiral CT and face scan images. In this article, a novel intelligent Internet of Things scene is proposed: a multimodal medical images acquisition and fusion system, and by combining CT machine and face scan equipment, the CT and face scan of patients can be synchronously collected. Deep point neural networks are used to extract feature points and a threshold iterative closest point algorithm performing registration with deep feature points and contributed region segmentation is applied. Finally, high-precision fused modal data is output at the mobile terminal to facilitate diagnosis and analysis and improve the efficiency of doctor–patient communication. Quantitative experiments show promising results, and clinical experiments prove that our method enables patients and doctors to better understand the state of an illness and improves the efficiency of doctor–patient communication. Zhiyuan Qu, Hongyi Jing, Guo Bai, Zhongpai Gao, Leilei Yu, Guangtao Zhai, Chi Yang |
IEEE Internet Things J. | 4 |
| 2023 | RIVIE: Robust Inherent Video Information EmbeddingabstractImagine an interesting situation when watching a movie, we can scan the screen using our smartphones to get some extra information about this movie such as the cast, the release date, the movie's homepage, etc. Our prospect is a world where each video contains invisible information that can be delivered to us through mobile devices with cameras. This paper proposes the first deep learning-based information hiding method for videos to achieve information transmission from screens to cameras. Compared with hiding information in single images, the methods for videos need to maintain visual quality in both spatial and temporal domains. Furthermore, the training of video models builds on a large video dataset, which needs much more computational resources than training models for images. To reduce the computational complexity, we propose to simulate data on-the-fly to generate simulated sequences from single images. Then, we use the simulated data to train a spatio-temporal generator that hides information in videos while maintaining visual quality. During training, a temporal loss function based on the simulated data is exploited to ensure the temporal consistency of generated videos. After embedding, we use a decoder to recover the hidden information. To simulate the imaging pipeline from screens to cameras in the real world, we insert a distortion network between the generator and decoder. The distortion network is based on differentiable 3D rendering to cover possible distortions introduced in the procedure of camera imaging. Experimental results show that the hidden information in videos can be extracted by cameras without impacting the visual quality. Our work can be applied to many fields, such as advertisement, entertainment, and education. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Menghan Hu, Guangtao Zhai |
IEEE Trans. Multim. | 2 |
| 2023 | Robust Mesh Representation Learning via Efficient Local Structure-Aware Anisotropic ConvolutionabstractMesh is a type of data structure commonly used for 3-D shapes. Representation learning for 3-D meshes is essential in many computer vision and graphics applications. The recent success of convolutional neural networks (CNNs) for structured data (e.g., images) suggests the value of adapting insights from CNN for 3-D shapes. However, 3-D shape data are irregular since each node's neighbors are unordered. Various graph neural networks for 3-D shapes have been developed with isotropic filters or predefined local coordinate systems to overcome the node inconsistency on graphs. However, isotropic filters or predefined local coordinate systems limit the representation power. In this article, we propose a local structure-aware anisotropic convolutional operation (LSA-Conv) that learns adaptive weighting matrices for each template's node according to its neighboring structure and performs shared anisotropic filters. In fact, the learnable weighting matrix is similar to the attention matrix in the random synthesizer-a new Transformer model for natural language processing (NLP). Since the learnable weighting matrices require large amounts of parameters for high-resolution 3-D shapes, we introduce a matrix factorization technique to notably reduce the parameter size, denoted as LSA-small. Furthermore, a residual connection with a linear transformation is introduced to improve the performance of our LSA-Conv. Comprehensive experiments demonstrate that our model produces significant improvement in 3-D shape reconstruction compared to state-of-the-art methods. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Juyong Zhang, Xiaokang Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Learning Invisible Markers for Hidden Codes in Offline-to-online PhotographyabstractQR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible codes/hyperlinks that can convey hidden information from offline to online. However, they require markers to locate invisible codes, which fails the purpose of invisible codes to be visible because of the markers. This paper proposes a novel invisible information hiding architecture for display/print-camera scenarios, consisting of hiding, locating, correcting, and recovery, where invisible markers are learned to make hidden codes truly invisible. We hide information in a sub-image rather than the entire image and include a localization module in the end-to-end framework. To achieve both high visual quality and high recovering robustness, an effective multi-stage training strategy is proposed. The experimental results show that the proposed method outperforms the state-of-the-art information hiding methods in both visual quality and robustness. In addition, the automatic localization of hidden codes significantly reduces the time of manually correcting geometric distortions for photos, which is a revolutionary innovation for information hiding in mobile applications. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
CVPR | 2 |
| 2022 | RIHOOP: Robust Invisible Hyperlinks in Offline and Online PhotographsabstractIn the era of multimedia and Internet, the quick response (QR) code helps people obtain information from offline to online quickly. However, the QR code is often limited in many scenarios because of its random and dull appearance. Therefore, this article proposes a novel approach to embed hyperlinks into common images, making the hyperlinks invisible for human eyes but detectable for mobile devices equipped with a camera. Our approach is an end-to-end neural network with an encoder to hide messages and a decoder to extract messages. To maintain the hidden message resilient to cameras, we build a distortion network between the encoder and the decoder to augment the encoded images. The distortion network uses differentiable 3-D rendering operations, which can simulate the distortion introduced by camera imaging in both printing and display scenarios. To maintain the visual attraction of the image with hyperlinks, a loss function conforming to the human visual system (HVS) is used to supervise the training of the encoder. Experimental results show that the proposed approach outperforms the previous work on both robustness and quality. Based on the proposed approach, many applications become possible, for example, "image hyperlinks" for advertisement on TV, website, or poster, and "invisible watermark" for copyright protection on digital resources or product packagings. Jun Jia, Zhongpai Gao, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Learning Local Neighboring Structure for Robust 3D Shape RepresentationabstractMesh is a powerful data structure for 3D shapes. Representation learning for 3D meshes is important in many computer vision and graphics applications. The recent success of convolutional neural networks (CNNs) for structured data (e.g., images) suggests the value of adapting insight from CNN for 3D shapes. However, 3D shape data are irregular since each node's neighbors are unordered. Various graph neural networks for 3D shapes have been developed with isotropic filters or predefined local coordinate systems to overcome the node inconsistency on graphs. However, isotropic filters or predefined local coordinate systems limit the representation power. In this paper, we propose a local structure-aware anisotropic convolutional operation (LSA-Conv) that learns adaptive weighting matrices for each node according to the local neighboring structure and performs shared anisotropic filters. In fact, the learnable weighting matrix is similar to the attention matrix in random synthesizer -- a new Transformer model for natural language processing (NLP). Comprehensive experiments demonstrate that our model produces significant improvement in 3D shape reconstruction compared to state-of-the-art methods. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Juyong Zhang, Yiyan Yang, Xiaokang Yang 0001 |
AAAI | 1 |
| 2021 | Looking here or there? Gaze Following in 360-Degree ImagesabstractGaze following, i.e., detecting the gaze target of a human subject, in 2D images has become an active topic in computer vision. However, it usually suffers from the out of frame issue due to the limited field-of-view (FoV) of 2D images. In this paper, we introduce a novel task, gaze following in 360-degree images which provide an omnidirectional FoV and can alleviate the out of frame issue. We collect the first dataset, "GazeFollow360"1, for this task, containing around 10,000 360-degree images with complex gaze behaviors under various scenes. Existing 2D gaze following methods suffer from performance degradation in 360degree images since they may use the assumption that a gaze target is in the 2D gaze sight line. However, this assumption is no longer true for long-distance gaze behaviors in 360-degree images, due to the distortion brought by sphere-to-plane projection. To address this challenge, we propose a 3D sight line guided dual-pathway framework, to detect the gaze target within a local region (here) and from a distant region (there), parallelly. Specifically, the local region is obtained as a 2D cone-shaped field along the 2D projection of the sight line starting at the human subject’s head position, and the distant region is obtained by searching along the sight line in 3D sphere space. Finally, the location of the gaze target is determined by fusing the estimations from both the local region and the distant region. Experimental results show that our method achieves significant improvements over previous 2D gaze following methods on our GazeFollow360 dataset. Wei Shen 0002, Zhongpai Gao, Yucheng Zhu, Guangtao Zhai, Guodong Guo |
ICCV | 3 |
| 2021 | Learning Spectral Dictionary for Local Representation of MeshabstractFor meshes, sharing the topology of a template is a common and practical setting in face-, hand-, and body-related applications. Meshes are irregular since each vertex's neighbors are unordered and their orientations are inconsistent with other vertices. Previous methods use isotropic filters or predefined local coordinate systems or learning weighting matrices for each vertex of the template to overcome the irregularity. Learning weighting matrices for each vertex to soft-permute the vertex's neighbors into an implicit canonical order is an effective way to capture the local structure of each vertex. However, learning weighting matrices for each vertex increases the parameter size linearly with the number of vertices and large amounts of parameters are required for high-resolution 3D shapes. In this paper, we learn spectral dictionary (i.e., bases) for the weighting matrices such that the parameter size is independent of the resolution of 3D shapes. The coefficients of the weighting matrix bases for each vertex are learned from the spectral features of the template's vertex and its neighbors in a weight-sharing manner. Comprehensive experiments demonstrate that our model produces state-of-the-art results with a much smaller model size. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Xiaokang Yang 0001 |
IJCAI | 1 |
| 2021 | Mask Region-oriented Diabetic Retinopathies Detection in Ophthalmic Medical Images via Non-local AttentionabstractAccurate lesions detection on retinopathy images is crucial for the diagnosis of diabetes. However, it is always hampered by various characteristics of lesions such as shape, color, texture, and similarities. Most advanced algorithms still cannot automatically detect common lesions, e.g. exudate, hemorrhage, and cotton-wool spots, being used for comprehensive analysis of disease state. To this end, we present a multi-functional detection model for diabetic retinopathies and further analyze disease mechanisms overall. Specifically, this paper attempts to implement a multi-lesion detector via modified Mask region-oriented CNN, which can be used for the aforementioned retinopathies. Meanwhile, a non-local attention module is introduced to the detector as a spatial attention mechanism for handling the global information missing problem. In addition, to boost the effectiveness of the detector, the dilated operation is implemented for dataset preprocessing. Improvement is achieved both algorithmically and architecturally, via investigating thoroughly the most probable lesion category with a novel ensemble learning framework. Extensive experiments on standard datasets for three different tasks evidence the superior performance of the proposed method over state-of-the-art methods. Hongtian Zhao, Haoyue Peng, Zhongpai Gao, Shibao Zheng |
IJCNN | 3 |
| 2021 | LRS-Net: invisible QR Code embedding, detection, and restorationabstractQR code is a powerful tool to bridge the offline and online worlds. It has been widely used because it can store a large amount of information in a small space. However, the black-and-white style of QR codes is not attractive to the human eyes when embedded in videos, which greatly affects the viewing experience. Invisible QR code has proposed based on temporal psycho-visual modulation (TPVM) to embed invisible hyperlinks in shopping websites, copyright watermarks in movies, etc. However, existing embedding and detection methods are not robust enough. In this paper, we adopt a novel embedding method to greatly improve the visual quality of the embedded video. Furthermore, we build a new dataset of invisible QR codes named 'IQRCodes' to train deep neural networks. At last, we propose localization, refinement, and segmentation neural netowrks (LRS-Net) to efficiently detect and restore invisible QR codes that are captured by mobile phones. Yiyan Yang, Zhongpai Gao, Guangtao Zhai |
VCIP | 2 |
| 2021 | SalGFCN: Graph Based Fully Convolutional Network for Panoramic Saliency PredictionabstractThe saliency prediction of panoramic images is dramatically affected by the distortion caused by non-Euclidean geometry characteristic. Traditional CNN based saliency pre-diction algorithms for 2D images are no longer suitable for 360-degree images. Intuitively, we propose a graph based fully convolutional network for saliency prediction of 360-degree images, which can reasonably map panoramic pixels to spherical graph data structures for representation. The saliency prediction network is based on residual U-Net architecture, with dilated graph convolutions and attention mechanism in the bottleneck. Furthermore, we design a fully convolutional layer for graph pooling and unpooling operations in spherical graph space to retain node-to-node features. Experimental results show that our proposed method outperforms other state-of-the-art saliency models on the large-scale dataset. Yiwei Yang 0007, Yucheng Zhu, Zhongpai Gao, Guangtao Zhai |
VCIP | 3 |
| 2021 | Psycho-visual modulation based information display: introduction and survey
Zhongpai Gao, Jia Wang 0004, Guangtao Zhai |
Frontiers Comput. Sci. | 2 |
| 2020 | An Improved Algorithm for Real-Time Dual-View DisplayabstractDual-view display based on spatial psychovisual modulation (SPVM) aims to present two different views on a single screen. Users with special glasses see a personal view. Users without special glasses see a shared view which can be completely irrelevant to the personal view. This dual-view display technology can be used for information security, QR hiding, education, etc. In this paper, we propose an improved algorithm called pseudo-inv algorithm, using the pseudo-inverse method to solve the optimization problem. Moreover, a new Gaussian integration window and a new method to adjust the luminance range, are implemented for a better personal view. The experiments prove that the new algorithm results in a better image quality of both the shared view and the personal view. The time cost of the new algorithm is much less than that of the previous algorithms. Zhongpai Gao, Guangtao Zhai, Zhaodi Wang 0002, Yucheng Zhu |
ISCAS | 2 |
| 2019 | EMBDN: An Efficient Multiclass Barcode Detection Network for Complicated EnvironmentsabstractThis article presents a novel method for efficient barcodes detection in real and complicated environments using a convolutional neural network (CNN)-based model. The method is developed as a preprocess-module of existing decoders to enhance decoding rates. Our method is trained as an end-to-end model to determine accurate locations of four barcode vertexes. Our method consists of four modules: 1) base net module; 2) region proposals generator; 3) classification and regression module; and 4) distortion removal module. The feature of barcodes extracted from the base net is fed to the next module. Region proposals are generated and selected as region of interest (ROI). Then the ROI are forward propagated to the classification and regression module to determine the positions and shapes of the barcodes. Finally, the distortion removal module is used to remove the geometric distortion according to regression parameters acquired from the previous step. The accurate position and distorted barcodes shape can be determined and corrected by our method. We validate our method on a challenging large-scale dataset in experiments. Compared with the previous methods, our method provides an end-to-end solution to determine accurate locations of barcode vertexes, which shows an excellent performance on detection accuracy. In addition, our method can enhance decoding rate through distortion removal. Jun Jia, Guangtao Zhai, Zhongpai Gao, Zehao Zhu, Xiongkuo Min, Xiaokang Yang 0001, Guodong Guo |
IEEE Internet Things J. | 4 |
| 2016 | Live demonstration: Screen piracy protection using saturation laser attack and TPVMabstractThis live demonstration presents a new screen piracy protection technology by using “saturation laser attack” and temporal psychovisual modulation(TPVM). “Saturation laser attack” is using lasers to emit beams of colored light into the camera, so that the camera is saturated and instantly blind. TPVM is a new information display technology using the interplay of signal processing, optoelectronics and psychophysics and provides a solution for information security. Using the TPVM technique, we are able to show the disguise information on the screen while authorized viewers can truly see the secret information. An automatic target recognition algorithm is used to locate the position of the target mobile phone, then the device sends out lasers to the target phone to defeat unauthorized recording screens. By combining laser attack and TPVM, the whole system can prevent disclosure of confidential and personal information through unauthorized recording screens effectively. Feng Lan, Guangtao Zhai, Zhongpai Gao, Xiaokang Yang 0001 |
ISCAS | 3 |
| 2016 | Quality assessment for dual-view display systemabstractSpatial psychovisual modulation (SPVM) is a new information display technology, which aims to generate multiple visual percepts for different viewers on a single display simultaneously. After the proposal of SPVM, lots of efforts have been made and several applications (i.e., dual-view display system) have been implemented based on this technology. The dual-view display (DVD) system is considered as an effective digital image hiding system based on SPVM theory, but little work has been dedicated to the perceptual quality assessment of DVD system. Up to now, there is no clear and standard method to evaluate the performance of the dual-view display system. It is important for the viewers to see a clear and non-aliasing image when they are front of the screen. Therefore, in this paper, we will build a DVD database and carry out a subjective experiment to evaluate the performance of the DVD system, and then we investigate and analyze the performance of prevailing no-reference (NR) image quality metrics on the particular DVD system. We have a sufficient belief that this paper can supply the guideline for the performance on the DVD system and serve as a good testing bed for future research of SPVM technology. Yuanchun Chen, Guangtao Zhai, Ke Gu 0001, Jia Wang 0004, Zhongpai Gao, Yucheng Zhu |
VCIP | 6 |
| 2016 | Factorization Algorithms for Temporal Psychovisual Modulation DisplayabstractTemporal psychovisual modulation (TPVM) is a new information display technology which aims to generate multiple visual percepts for different viewers on a single display simultaneously. In a TPVM system, the viewers wearing different active liquid crystal (LC) glasses with varying transparency levels can see different images (called personal views). The viewers without LC glasses can also see a semantically meaningful image (called shared view). The display frames and weights for the LC glasses in the TPVM system can be computed through nonnegative matrix factorization (NMF) with three additional constrains: the values of images and modulation weights should have upper bound (i.e., limited luminance of the display and transparency level of the LC); the shared view without using viewing devices should be considered (i.e., the sum of all basis images should be a meaningful image); and the sparsity of modulation weights should be considered due to the material property of LC. In this paper, we proposed to solve the constrained NMF problem by a modified version of hierarchical alternating least squares (HALS) algorithms. Through experiments, we analyze the choice of parameters in the setup of TPVM system. This work serves as a guideline for practical implementation of TPVM display system. Zhongpai Gao, Guangtao Zhai, Jiantao Zhou 0001 |
IEEE Trans. Multim. | 1 |
| 2015 | Adapting hierarchical ALS algorithms for temporal psychovisual modulationabstractTemporal psychovisual modulation (TPVM) is a new information display technology, which aims to generate multiple visual percepts for different viewers on a single display simultaneously. In TPVM system, the viewers with different active liquid crystal (LC) glasses (i.e., different modulation weights) which are synchronized with the display can see different images (called personal views). TPVM can be implemented by nonnegative matrix factorization (NMF) with three additional constrains: the values of images and modulation weights should have upper bound; a special view (called shared view) without using viewing devices should be considered (i.e., the sum of all basis images should be a meaningful image); the sparsity of modulation weights should be considered because of the material property of LC. In this paper, we solve the constrained NMF problem by the modified hierarchical alternating least squares (HALS) algorithms. Through experiments, we analyse the influence of different parameters of TPVM to provide a guideline for parameter selection. This paper will provide an algorithmic guidance for the applications of TPVM. Zhongpai Gao, Guangtao Zhai, Xiao Gu 0001, Jiantao Zhou 0001 |
ISCAS | 1 |
| 2015 | The Invisible QR CodeabstractQR (Quick Response) Codes are widely used as a convenient unidirectional communication channel to convey information, such as emails, hyperlinks, or phone numbers, from publicity materials to mobile devices. But the QR Code is not visually appealing and takes up valuable space of publicity materials. In this paper, we propose a new method to embed QR Code on digital screen via temporal psychovisual modulation (TPVM). By exploiting the difference between human eyes and semiconductor imaging sensors in temporal convolution of optical signals, we make QR Code perceptually transparent to human but detectable for mobile devices. Based on the idea of invisible QR Code, many applications can be implemented, e.g., "physical hyperlink" for something interesting on TV or digital signage , "invisible watermark" for anti-piracy in theater. A prototype system introduced in this paper serves as a proof-of-concept of the invisible QR Code and can be improved in future works. Zhongpai Gao, Guangtao Zhai, Chunjia Hu |
ACM Multimedia | 1 |
| 2015 | Visible Light Communication via Temporal Psycho-Visual ModulationabstractIn this paper we propose a new paradigm for visible light communication (VLC) using the emerging display technology of Temporal Psycho-Visual Modulation (TPVM) that exploits the interaction between human visual system and modern electro-optical display devices. Unlike traditional VLC, no specifically designed light emitter and receiver are required. In the proposed system, light projector is used as the information source and digital cameras act as information decoder. The emitted light is designed in a specific way such that it can carry meaningful information (or simply works as an illumination source) for human eyes while other message can be decoded by the digital camera due to the fundamental difference in the imaging mechanism of the human eye and digital devices. We further describe two applications of this new type of VLC in ubiquitous augmented reality and illegal camcorder-recording prevention with extensive experimental results. Chunjia Hu, Guangtao Zhai, Zhongpai Gao |
ACM Multimedia | 3 |
| 2014 | Dual-view medical image visualization based on spatial-temporal psychovisual modulationabstractMedical imaging technologies such as magnetic resonance imaging (MRI) and computerized tomographic (CT) are used to diagnose a wide range of medical diseases. Medical images are generated by detecting density differences between different tissues in the body. Multiple medical image visualization is of critical importance to diagnosis. This paper introduces a dual-view medical image visualization prototype based on spatial-temporal psychovisual modulation (STPVM). Temporal psychovisual modulation (TPVM) enables a single display to generate multiple visual content for different viewers. Spatial psychovisual modulation (SPVM) extends the idea of TPVM to spatial domain. STPVM combines TPVM and SPVM by exploiting both temporal and spatial redundancy of modern displays. Based on STPVM technology, one display can present even more images simultaneously. In this demo, two kinds of medical images e.g. T1 and T2 weighted MRI images, are presented simultaneously. Physicians can switch between either image by just moving the eye fixations. Since T1 and T2 are shown simultaneously and are aligned on the screen, it is more convenient for the physicians to get different information of the same spot from the T1 and T2 images. The developed demo is useful for physicians during surgery navigation and effectively reduces the burden of mental transfer. Zhongpai Gao, Guangtao Zhai, Chunjia Hu, Xiongkuo Min |
ICIP | 1 |
| 2014 | Information security display system based on Spatial Psychovisual ModulationabstractPrivacy protection is of increasing importance in this era of information explosion. This paper introduces an information security display system based on the idea of Spatial Psycho-visual Modulation (SPVM). With the rapid advance of modern manufacturing techniques, display devices now support very high pixel density (e.g. the retina display of Apple). Meanwhile the human visual system (HVS) cannot distinguish image signals with spatial frequency above a threshold, as predicted by the contrast sensitivity function (CSF). Therefore, it is now possible for us to devise a type of information security display using the mismatch between resolutions of modern display devices and the HVS. Given the desired visual stimuli for both the bystanders and the authorized users, we propose a method to design display signals accordingly. We select polarization as a way to effectively differentiate the bystanders and the authorized users. Operationally, applying complementary polarization to light emitted from different spatial sections of the display could in itself be a challenge. Fortunately, the development of stereoscopic display technologies has made available a lot of polarization based spatial multiplexing type of display devices. Hardware of the information security display system is based on a polarization based stereoscopic screen made by LG. Software of the information security display system is written in C++ with SDKs of DirectX and etc. Kinect is also included into our system to enhance the experience of human-computer interaction. Extended experimental results will be given in this paper to justify the effectiveness and robustness of the system. The developed system serves both as a proof-of-concept of the SPVM method, as well as a test bed for future research of SPVM based display technology. Chunjia Hu, Guangtao Zhai, Zhongpai Gao, Xiongkuo Min |
ICME | 3 |
| 2014 | Influence of compression artifacts on visual attentionabstractVisual attention is an important function of the human visual system (HVS). In the long term research of visual attention, various computational models have been proposed with encouraging results. However, most of those work were conducted on images with ideal visual quality. In practice, outputs of most visual communication systems contain different levels of artifacts, e.g. noise, blurring, blockiness and etc. Therefore, it is interesting to investigate the impacts of artifacts on visual attention. In this paper, we question into the problem of how the widely encountered JPEG compression artifacts affect visual attention. We designed eye-tracking experiments on images with different levels of compression and viewing time and quantitatively compared the recorded eye movement data. We found that compression level does have impacts on visual attention, and yet this influence can be negligible for low levels of compression. For high levels of compression, the visual artifacts alter visual attention in a systematic way. Dependence of the influence on viewing duration was also analyzed and it was observed that too short or too long viewing time reduces the impact of compression artifacts on visual attention. Xiongkuo Min, Guangtao Zhai, Zhongpai Gao, Chunjia Hu |
ICME | 3 |
| 2014 | Information security display system based on temporal psychovisual modulationabstractThis paper introduces an information security display system using temporal psychovisual modulation (TPVM). TPVM was proposed as a new information display technology using the interplay of signal processing, optoelectronics and psychophysics. Since the human visual system cannot detect quick temporal changes above the flicker fusion frequency (about 60 Hz) and yet modern display technologies offer much higher refresh rates, there is a chance for a single display to simultaneously serve different contents to multiple observers. A TPVM display broadcasts a set of images called atom frames at a high speed, and those atom frames are then weighted by liquid crystal (LC) shutter based viewing devices that are synchronized with the display before entering the human visual system and fusing into the desired visual stimuli. And through different viewing devices, people can see different information. In this work, we develop a TPVM based information security display prototype. There are two kinds of viewers, those authorized viewers with the viewing devices who can see the secret information and those unauthorized viewers (bystanders) without the viewing devices who only see mask/disguise images. The prototype is built on a 120 Hz LCD screen with synchronized LC shutter glasses that were originally developed for stereoscopic display. The system is written in C++ language with SDKs of Nvidia 3D Vision, DirectX, CEGUI, MuPDF and etc. We also added human-computer interaction support of the system using Kinect. The information security display system developed in this work serves as a proof-of-concept of the TPVM paradigm, as well as a testbed for future research of TPVM technology. Zhongpai Gao, Guangtao Zhai, Xiongkuo Min |
ISCAS | 1 |
| 2014 | Visual attention data for image quality assessment databasesabstractImages usually contain areas that particularly attract people's attention and visual attention is an important feature of human visual system (HVS). Visual attention had been shown to be effective in improving performance of existing image quality assessment (IQA) metrics. However, with the quick advancement of IQA research, the booming of open IQA databases calls for associated comprehensive and accurate visual attention dataset. Despite of the large number of existing computational attention/saliency models, the most accurate measure of human attention is still human based. In this research, we first conduct extensive eye tracking experiments for all the pristine images from the seven widely used IQA databases (LIVE, TID2008, CSIQ, Toyama, LIVE Multiply Distortion, IVC and A57 databases). Then we propose a gaze-duration adaptive weighting approach to generate saliency maps from the eye tracking data. When applied on the IQA databases, experimental results suggest that accuracy of benchmark quality metrics, e.g. PSNR and SSIM can be systematically improved, outperforming existing saliency datasets. Both the eye tracking data and the saliency maps in this research will be made publicly available at gvsp.sjtu.edu.cn. Xiongkuo Min, Guangtao Zhai, Zhongpai Gao, Ke Gu 0001 |
ISCAS | 3 |
| 2014 | Demo: DLP based anti-piracy display systemabstractCamcorder piracy has great impact on the movie industry. Although there are many methods to prevent recording in theatre, no recognized technology satisfies the need of defeating camcorder piracy as well as having no effect on the audience. To realize anti-piracy, we uses a new paradigm of information display technology, called temporal psychovisual modulation (TPVM). TPVM exploits the difference in image formation mechanisms of human eyes and imaging sensors. Based on this difference, we build a prototype system on the platform of DLP® LightCrafter 4500™ which features high speed pattern display. The display system serves as a proof-of-concept of anti-piracy system. Zhongpai Gao, Guangtao Zhai, Xiaolin Wu 0001, Xiongkuo Min, Chunjia Hu |
VCIP | 1 |
| 2014 | DLP based anti-piracy display systemabstractCamcorder piracy has great impact on the movie industry. Although there are many methods to prevent recording in theatre, no recognized technology satisfies the need of defeating camcorder piracy as well as having no effect on the audience. This paper presents a new projector display technique to defeat camcorder piracy in the theatre using a new paradigm of information display technology, called temporal psychovisual modulation (TPVM). TPVM exploits the difference in image formation mechanisms of human eyes and imaging sensors. The images formed in human vision is continuous integration of the light field while discrete sampling is used in digital video acquisition which has "blackout" period in each sampling cycle. Based on this difference, we can decompose a movie into a set of display frames and broadcast them out at high speed so that the audience can not notice any disturbance, while the video frames captured by camcorder will contain highly objectionable artifacts. The proposed prototype system built on the platform of DLP® LightCrafter 4500™ serves as a proof-of-concept of anti-piracy system. Zhongpai Gao, Guangtao Zhai, Xiaolin Wu 0001, Xiongkuo Min, Cheng Zhi |
VCIP | 1 |
| 2014 | Information security display via uncrowded windowabstractWith the booming of visual media, people pay more and more attention to privacy protection in public environments. Most existing research on information security such as cryptography and steganography is mainly concerned about transmission and yet little has been done to prevent the information displayed on screens from reaching eyes of the bystanders. This "security of the last foot (SOLF)" problem, if left without being taken care of, will inevitably lead to the total failure of a trustable information communication system. To deal with the SOLF problem, for the application of text-reading, we proposed an eye tracking based solution using the newly revealed concept of uncrowded window from vision research. The theory of uncrowded window suggests that human vision can only effectively recognize objects inside a small window. Object features outside the window may still be detectable but the feature detection results cannot be efficiently combined properly and therefore those objects will not be recognizable. We use eye-tracker to locate fixation points of the authorized reader in real time, and only the area inside the uncrowded window displays the private information we want to protect. A number of dummy windows with fake messages are displayed around the real uncrowded window as diversions. And without the precise knowledge about the fixations of the authorized reader, the chance for bystanders to capture the private message from those surrounding area and the dummy windows is very low. Meanwhile, since the authorized reader can only read within the uncrowded window, detrimental impact of those dummy windows is almost negligible. The proposed prototype system was written in C++ with SDKs of Direct3D, Tobii Gaze SDK, CEGUI, MuPDF, OpenCV and etc. Extended demonstration of the system will be provided to show that the proposed method is an effective solution to SOLF problem of information communication and display. Zhongpai Gao, Guangtao Zhai, Jiantao Zhou 0001, Xiongkuo Min, Chunjia Hu |
VCIP | 1 |