Ying Zhang 0047

dblp:13/6769-47 · DBLP profile ↗
← Back
45ranked-venue papers
15as first author
26since 2021 · last 2026
0000-0001-6411-4486ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 10 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 ArtEEGAttention: an advanced deep learning approach for art brain decoding
Shuming Hu, Shu Zhang 0006, Ying Zhang 0047, Zhu Wang 0001, Bin Guo 0001, Zhiwen Yu 0001
Frontiers Comput. Sci.3
2025 CollageNoter: Real-Time and Adaptive Collage Layout Design for Screenshot-Based E-Note-Taking
abstract
To enhance the processing of complex multi-modal documents (e.g. e-books, long web pages, etc.), it is an efficient way for users to take digital screenshots of key parts and reorganize them into a new collage E-Note. Existing methods for assisting collage layout design primarily employ a semantic relevance-first strategy, with arranging related contents together. Though capable, it can not ensure the visual readability of screenshots and may conflict with human natural reading patterns. In this paper, we introduce CollageNoter for real-time collage layout design that adapts to various devices (e.g. laptop, tablet, phone, etc.), offering users with visually and cognitively well-organized screenshot-based E-Notes. Specifically, we construct a novel two-stage pipeline for collage design, including 1) readability-first layout generation and 2) cognitive-driven layout adjustment. In addition, to achieve real-time response and adaptive model training, we propose a cascade transformer-based layout generator named CollageFormer and a size-aware collage layout builder for automatic dataset construction. Extensive experimental results have confirmed the effectiveness of our CollageNoter.
Qiuyun Zhang, Bin Guo 0001, Lina Yao 0001, Xiaotian Qiao, Ying Zhang 0047, Zhiwen Yu 0001
AAAI5
2025 HEMPF: Efficient High-Resolution Multimodal Inference on Heterogeneous Edges via Perception-Driven Segmentation and Dynamic Offloading
Ying Zhang 0047, Bin Guo 0001, Zhiwen Yu 0001
ICA3PP (6)2
2025 CompMTL: Layer-Wise Competitive Multi-Task Learning
abstract
It is challenging to simultaneously address multiple related tasks using a unified multi-task model and consistently balance conflicts across these tasks. The conflicts arise because each task competes to update the shared module in a manner that can better align with its own requirements. To address the conflicts, existing multi-task learning methods primarily balance task losses or gradients of the shared module, but frequently overlook the differences in layer-wise conflicts, which can enable a more fine-grained conflict-averse. For details, each task establishes its competitive influence over a particular layer by assigning varying degrees of importance to it. Attenuating the gradients of tasks with relatively lower importance during updates in specific layers may not only balance gradient conflicts but also facilitate the progression of other tasks. Based on this, we proposed Layer-wise Competitive Multi-task Learning, wherein multiple tasks compete for gradient update weights within the shared modules of a multi-task model. This approach aims to achieve layer-specific gradient balance by considering each task’s relative importance at different layers within the shared module. Tasks of relatively lower importance for a specific layer will receive smaller gradient updates, thereby facilitating faster convergence for more important tasks.
Tiancong Cheng, Ying Zhang 0047, Rajiv Ratn Shah, Roger Zimmermann, Zhiwen Yu 0001, Bin Guo 0001
ICASSP2
2025 JointDistill: Adaptive Multi-Task Distillation for Joint Depth Estimation and Scene Segmentation
abstract
Depth estimation and scene segmentation are two important tasks in intelligent transportation systems. A joint modeling of these two tasks will reduce the requirement for both the storage and training efforts. This work explores how the multi-task distillation could be used to improve such unified modeling. While existing solutions transfer multiple teachers’ knowledge in a static way, we propose a self-adaptive distillation method that can dynamically adjust the knowledge amount from each teacher according to the student’s current learning ability. Furthermore, as multiple teachers exist, the student’s gradient update direction in the distillation is more prone to be erroneous where knowledge forgetting may occur. To avoid this, we propose a knowledge trajectory to record the most essential information that a model has learnt in the past, based on which a trajectory-based distillation loss is designed to guide the student to follow the learning curve similarly in a cost-effective way. We evaluate our method on multiple benchmarking datasets including Cityscapes and NYU-v2. Compared to the state-of-the-art solutions, our method achieves a clearly improvement. The code is provided in the supplementary materials.
Tiancong Cheng, Ying Zhang 0047, Yuxuan Liang 0002, Roger Zimmermann, Zhiwen Yu 0001, Bin Guo 0001
ICME2
2025 Tree-of-AdEditor: Heuristic Tree Reasoning for Automated Video Advertisement Editing with Large Language Model
abstract
Video advertising has become a popular marketing strategy on e-commerce platforms, requiring high-level semantic reasoning like selling point discovery, narrative organization. Previous rule-based methods struggle with these complex tasks, and learning-based approaches demand large datasets and high training costs. Recently, Large Language Models have opened incredible opportunities for advancing intelligent video advertisement editing. However, Input-output (IO) prompting and Chain-of-Thought (CoT) struggle to adapt to the nonlinear thinking hierarchy of video editing, where editors iteratively select shots or revert them to explore potential editing solutions. While Tree-of-Thought (ToT) offers a conceptual structure that mirrors this hierarchy, it falls short in aligning with effective video advertising strategies and lacks robust fact-checking mechanisms. To address these, we propose a novel framework, Tree-of-AdEditor (ToAE), which constructs a reasoning tree to mimic human editors, and incorporates domain-specific theories and heuristic fact-checking to identify optimal editing solutions. Specifically, motivated by effective advertisement principles, we develop a "local-global" mechanism to guide LLM in both the shot level and sequence level decision-making. We introduce a visual incoherence pruning module to provide external heuristic fact-checking, ensuring visual attractiveness and reducing computation costs. Quantitative experiments and expert evaluation demonstrate the superiority of our method compared to baselines.
Bin Guo 0001, Ying Zhang 0047, Shijie Wang 0002, Zhiwen Yu 0001, Qing Li 0001
IJCAI4
2025 Cascade context-oriented spatio-temporal attention network for efficient and fine-grained video-grounded dialogues
Hao Wang 0182, Bin Guo 0001, Qiuyun Zhang, Yasan Ding, Ying Zhang 0047, Zhiwen Yu 0001
Frontiers Comput. Sci.6
2025 Design Graph Guided Element Importance-Aware Layout Generation With Multimodality Cascade Transformer
abstract
Graphic designs are pervasive in our daily lives and widely used to communicate information hierarchically to humans. To achieve this, the layout plays an essential role in guiding readers to understand the importance of different elements and comprehend the content. To deal with the rapidly increasing demands of graphic designs, recent studies attempt to automatically generate layouts based on category information and spatial relations, often resulting in layouts with poor communication quality. In this article, we make the first attempt to explore element importance-aware layout generation under the guidance of a novel design graph, which attracts readers’ attention to a layout by formulating aesthetic relations implicitly involved in graphic designs between element pairs. The core of our approach is a learning-based framework with a new multimodality cascade transformer (MCT) in a coarse-to-fine manner. A hierarchical multimodality fusion (HMF) mechanism and two new losses are introduced to guide the training process progressively. We further collect a new fine-grained advertisement poster layout dataset containing more than 30 K layouts labeled with 91 element labels. Both qualitative and quantitative experiments demonstrate the effectiveness of our approach against existing works. We also conduct user studies and cognitive experiments to evaluate the direct adaptability and attractiveness of generated layouts.
Qiuyun Zhang, Bin Guo 0001, Lina Yao 0001, Xiaotian Qiao, Hao Wang 0182, Ying Zhang 0047, Zhiwen Yu 0001
IEEE Trans. Hum. Mach. Syst.6
2025 Cinematographic-Aware Coherent Shot Assembly for How-To Vlog Generation
abstract
The how-to vlog has gained popularity as a medium for learning practical skills, such as cooking and handcrafting. The sequencing process of these videos should meticulously integrate storytelling and cinematographic elements to ensure audience comprehension and engagement, posing a significant challenge for novice creators. Pioneer efforts have adhered to predefined editing rules or assembled clips based on textual logic, limiting their applicability across diverse educational scenarios and causing visual discontinuities. In contrast, we aim to extract rich cinematographic patterns from well-edited professional how-to vlogs to enhance shot assembly. To this end, we identify two crucial aspects influencing educational value: narrative continuity and scale transition. We model the shot assembly process as a task of selecting the next shot and devise a cinematographic-aware contrastive model to learn representations that distinguish between the good next shots against the bad ones. This method incorporates a two-stream cinematographic-aware encoding module for explicit factor encoding and a situation-adaptive attention-based integration module to accommodate varying assembly scenarios. Quantitative results from our novel professional user-generated vlogs dataset (proVlog-HowTo) clearly demonstrate the proposed method’s effectiveness. User study results further indicate its superiority in generating videos with narrative continuity and smooth transitions.
Bin Guo 0001, Ying Zhang 0047, Qianru Wang, Zhiwen Yu 0001, Qing Li 0001
IEEE Trans. Hum. Mach. Syst.3
2025 Enabling Harmonious Human-Machine Interaction with Visual-Context Augmented Dialogue System: A Review
abstract
The intelligent dialogue system, aiming at communicating with humans harmoniously with natural language, is brilliant for promoting the advancement of human-machine interaction in the era of artificial intelligence. With the gradually complex human-computer interaction requirements, it is difficult for traditional text-based dialogue system to meet the demands for more vivid and convenient interaction. Consequently, Visual-Context Augmented Dialogue (VAD) System, which has the potential to communicate with humans by perceiving and understanding multimodal information (i.e., visual context in images or videos, textual dialogue history), has become a predominant research paradigm. Benefiting from the consistency and complementarity between visual and textual context, VAD possesses the potential to generate engaging and context-aware responses. To depict the development of VAD, we first characterize the concept model of VAD and then present its generic system architecture to illustrate the system workflow, followed by a summary of multimodal fusion techniques. Subsequently, several research challenges and representative works are investigated, followed by the summary of authoritative benchmarks and real-world application of VAD. We conclude this article by putting forward some open issues and promising research trends for VAD, e.g., the cognitive mechanisms of human-machine dialogue under cross-modal dialogue context, mobile and lightweight deployment of VAD.
Hao Wang 0182, Bin Guo 0001, Yating Zeng, Yasan Ding, Ying Zhang 0047, Lina Yao 0001, Zhiwen Yu 0001
ACM Trans. Inf. Syst.6
2024 Traj2Former: A Local Context-aware Snapshot and Sequential Dual Fusion Transformer for Trajectory Classification
abstract
The wide use of mobile devices has led to a proliferated creation of extensive trajectory data, rendering trajectory classification increasingly vital and challenging for downstream applications. Existing deep learning methods offer powerful feature extraction capabilities to detect nuanced variances in trajectory classification tasks. However, their effectiveness remains compromised by the following two unsolved challenges. First, identifying the distribution of nearby trajectories based on noisy and sparse GPS coordinates poses a significant challenge, providing critical contextual features to the classification. Second, though efforts have been made to incorporate a shape feature by rendering trajectories into images, they fail to model the local correspondence between GPS points and image pixels. To address these issues, we propose a novel model termed Traj2Former to spotlight the spatial distribution of the adjacent trajectory points (i.e., contextual snapshot) and enhance the snapshot fusion between the trajectory data and the corresponding spatial contexts. We propose a new GPS rendering method to generate contextual snapshots, but it can be applied from a trajectory database to a digital map. Moreover, to capture diverse temporal patterns, we conduct a multi-scale sequential fusion by compressing the trajectory data with differing rates. Extensive experiments have been conducted to verify the superiority of the Traj2Former model.
Yichen Zhang 0002, Yifang Yin, Sheng Zhang 0023, Ying Zhang 0047, Rajiv Ratn Shah, Roger Zimmermann, Guoqing Xiao 0001
ACM Multimedia5
2024 UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation
abstract
Urbanization challenges underscore the necessity for effective satellite image-text retrieval methods to swiftly access specific information enriched with geographic semantics for urban applications. However, existing methods often overlook significant domain gaps across diverse urban landscapes, primarily focusing on enhancing retrieval performance within single domains. To tackle this issue, we present UrbanCross, a new framework for cross-domain satellite image-text retrieval. UrbanCross leverages a high-quality, cross-domain dataset enriched with extensive geo-tags from three countries to highlight domain diversity. It employs the Large Multimodal Model (LMM) for textual refinement and the Segment Anything Model (SAM) for visual augmentation, achieving fine-grained alignment of images, segments and texts, yielding a 10% improvement in retrieval performance. Additionally, UrbanCross incorporates an adaptive curriculum-based source sampler and a weighted adversarial cross-domain fine-tuning module, progressively enhancing adaptability across various domains. Extensive experiments confirm UrbanCross's superior efficiency in retrieval and adaptation to new urban environments, demonstrating an average performance increase of 15% over its version without domain adaptation mechanisms, effectively bridging the domain gap. Our code and dataset are publicly accessible at https://github.com/siruzhong/UrbanCross.
Siru Zhong, Xixuan Hao, Ying Zhang 0047, Yangqiu Song, Yuxuan Liang 0002
ACM Multimedia4
2024 Memory-Enhanced Emotional Support Conversations with Motivation-Driven Strategy Inference
Hao Wang 0182, Bin Guo 0001, Yasan Ding, Qiuyun Zhang, Ying Zhang 0047, Zhiwen Yu 0001
ECML/PKDD (5)6
2024 Limits of predictability in top-N recommendation
En Xu, Zhiwen Yu 0001, Ying Zhang 0047, Bin Guo 0001, Lina Yao 0001
Inf. Process. Manag.4
2023 A Multi-Teacher Assisted Knowledge Distillation Approach for Enhanced Face Image Authentication
abstract
Recent deep-learning-based face recognition systems have achieved significant success. However, most existing face recognition systems are vulnerable to spoofing attacks where a copy of the face image is used to deceive the authentication. A number of solutions are developed to overcome this problem by building a separate face anti-spoofing model, which however brings in additional storage and computation requirements. Since both recognition and face anti-spoofing tasks stem from the analysis of the same face image, this paper explores a unified approach to reduce the original dual-model redundancy. To this end, we introduce a compressed multi-task model to simultaneously perform both tasks in a lightweight manner, which has the potential to benefit lightweight IoT applications. Concretely, we regard the original two single-task deep models as teacher networks and propose a novel multi-teacher-assisted knowledge distillation method to guide our lightweight multi-task model to achieve satisfying performance on both tasks. Additionally, to reduce the large gap between the deep teachers and the light student, a comprehensive feature alignment is further integrated by distilling multi-layer features. Extensive experiments are carried out on two benchmark datasets, where we achieve the task accuracy of 93% meanwhile reducing the model size by 97% and reducing the inference time by 56% compared to the original dual-model.
Tiancong Cheng, Ying Zhang 0047, Yifang Yin, Roger Zimmermann, Zhiwen Yu 0001, Bin Guo 0001
ICMR2
2023 FaceLivePlus: A Unified System for Face Liveness Detection and Face Verification
abstract
Face verification is a trending way to verify someone’s identity in broad applications. But such systems are vulnerable to face spoofing attacks via, for example, a fraudulent copy of a photo, making it necessary to include face liveness detection as an additional safeguard. Among most existing studies, the face liveness detection is realized in a separate machine learning model in addition to the model for face verification. Such a two-model configuration may face challenges when deployed onto platforms with limited computation power and storage (e.g. mobile phone, IoT devices), especially considering each model may have millions of parameters. Inspired by the fact that humans can verify a person’s identity and liveness at a single glance from a face, we develop a novel system, named FaceLivePlus, to learn a single and universal face descriptor for the two tasks (face verification and liveness detection) so that the computational workload and storage space can be halved. To achieve this, we formulate the underlying relationship between the two tasks, and seamlessly embed this relationship in a distance ranking deep model. The model directly works on features rather than classification labels, which makes the system well generalized on unseen data. Extensive experiments show that our average half total error rate (HTER) has at least 15% and 8% improvement from the state-of-the-arts on two benchmark datasets. We anticipate this approach could become a new direction for face authentication.
Ying Zhang 0047, Lilei Zheng, Vrizlynn L. L. Thing, Roger Zimmermann, Bin Guo 0001, Zhiwen Yu 0001
ICMR1
2023 ALDA: An Adaptive Layout Design Assistant for Diverse Posters throughout the Design Process
abstract
Layout generation is important in the field of graphic design and has attracted intensive research attention recently. To further prompt human-computer interactions, we construct the ALDA to assist users throughout the design process, which achieves adaptive and diverse content expansions upon only one element for beginners and generates high-quality posters subsequently. Specifically, to obtain diverse contents, we propose aesthetic-aware design graphs (AGs) for effective poster representation and propose a self-constrained blending strategy upon related examples. In addition, we build a novel layout generator to better arrange elements conditioned on our AGs. Finally, we implement ALDA as an online tool with a set of controllable factors to enhance its practicality.
Qiuyun Zhang, Bin Guo 0001, Lina Yao 0001, Han Wang 0005, Ying Zhang 0047, Zhiwen Yu 0001
ACM Multimedia5
2023 Prototypical Cross-domain Knowledge Transfer for Cervical Dysplasia Visual Inspection
abstract
Early detection of dysplasia of the cervix is critical for cervical cancer treatment. However, automatic cervical dysplasia diagnosis via visual inspection, which is more appropriate in low-resource settings, remains a challenging problem. Though promising results have been obtained by recent deep learning models, their performance is significantly hindered by the limited scale of the available cervix datasets. Distinct from previous methods that learn from a single dataset, we propose to leverage cross-domain cervical images that were collected in different but related clinical studies to improve the model's performance on the targeted cervix dataset. To robustly learn the transferable information across datasets, we propose a novel prototype-based knowledge filtering method to estimate the transferability of cross-domain samples. We further optimize the shared feature space by aligning the cross-domain image representations simultaneously on domain level with early alignment and class level with supervised contrastive learning, which endows model training and knowledge transfer with stronger robustness. The empirical results on three real-world benchmark cervical image datasets show that our proposed method outperforms the state-of-the-art cervical dysplasia visual inspection by an absolute improvement of 4.7% in top-1 accuracy, 7.0% in precision, 1.4% in recall, 4.6% in F1 score, and 0.05 in ROC-AUC.
Yichen Zhang 0002, Yifang Yin, Ying Zhang 0047, Zhenguang Liu, Zheng Wang 0007, Roger Zimmermann
ACM Multimedia3
2022 Parallel Edge-Image Learning for Image Inpainting
abstract
The primary goal of image inpainting is to fix holes in a damaged image with natural contents. A key challenge is that a damaged image contains complex structures in differ-ent ways, with each consisting of its configuration of edges and spatial dependencies. As a result, filled images often converge to unnatural and implausible results. Currently, the edge-image inpainting methods adopt two stages to recover edges and images successively, which suffer from feature in-consistency and error accumulation. This leads us to present a parallel edge-image learning framework that explicitly char-acterizes these internal configurations in a single stage. The framework introduces a dual parallel network-based decoder to generate the image and the edges concurrently, leading to feature consistency at the semantic level. Also, a new cross-fire mechanism aims to exchange edge-image information in the decoder, avoiding error accumulation. Empirical evaluations on benchmark datasets suggest that our approach out-performs the state-of-the-art methods on image inpainting.
Junfeng Hu 0001, Chengxin Wang, Ying Zhang 0047, Li Liu 0001, Yifang Yin, Roger Zimmermann
ICME3
2022 Mix-Up Self-Supervised Learning for Contrast-Agnostic Applications
abstract
Contrastive self-supervised learning has attracted significant research attention recently. It learns effective visual represen-tations from unlabeled data by embedding augmented views of the same image close to each other while pushing away embeddings of different images. Despite its great success on ImageNet classification, COCO object detection, etc., its performance degrades on contrast-agnostic applications, e.g., medical image classification, where all images are visually similar to each other. This creates difficulties in optimizing the embedding space as the distance between images is rather small. To solve this issue, we present the first mix-up self-supervised learning framework for contrast-agnostic applications. We address the low variance across images based on cross-domain mix-up and build the pretext task based on two synergistic objectives: image reconstruction and transparency prediction. Experimental results on two benchmark datasets validate the effectiveness of our method, where an improve-ment of 2.5% ~ 7.4% in top-1 accuracy was obtained compared to existing self-supervised learning methods.
Yichen Zhang 0002, Yifang Yin, Ying Zhang 0047, Roger Zimmermann
ICME3
2022 Evaluation of a new dataset for visual detection of cervical precancerous lesions
Ying Zhang 0047, Yonit Zall, Ronen Nissim, Satyam, Roger Zimmermann
Expert Syst. Appl.1
2022 GPS2Vec: Pre-Trained Semantic Embeddings for Worldwide GPS Coordinates
abstract
GPS coordinates are fine-grained location indicators that are difficult to be effectively utilized by classifiers in geo-aware applications. Previous GPS encoding methods concentrate on generating hand-crafted features for small areas of interest. However, many real world applications require a machine learning model, analogous to the pre-trained ImageNet model for images, that can efficiently generate semantically-enriched features for planet-scale GPS coordinates. To address this issue, we propose a novel two-level grid-based framework, termed GPS2Vec, which is able to extract geo-aware features in real-time for locations worldwide. The Earth’s surface is first discretized by the Universal Transverse Mercator (UTM) coordinate system. Each UTM zone is then considered as a local area of interest that is further divided into fine-grained cells to perform the initial GPS encoding. We train a neural network in each UTM zone to learn the semantic embeddings from the initial GPS encoding. The training labels can be automatically derived from large-scale geotagged documents such as tweets, check-ins, and images that are available from social sharing platforms. We conducted comprehensive experiments on three geo-aware applications, namely place semantic annotation, geotagged image classification, and next location prediction. Experimental results demonstrate the effectiveness of our approach, as prediction accuracy improves significantly based on a simple multi-feature early fusion strategy with deep neural networks, including both CNNs and RNNs.
Yifang Yin, Ying Zhang 0047, Zhenguang Liu, Sheng Wang 0011, Rajiv Ratn Shah, Roger Zimmermann
IEEE Trans. Multim.2
2021 A Spatial Regulated Patch-Wise Approach for Cervical Dysplasia Diagnosis
abstract
Cervical dysplasia diagnosis via visual investigation is a challenging problem. Recent approaches use deep learning techniques to extract features and require the downsampling of high-resolution cervical screening images to smaller sizes for training. Such a reduction may result in the loss of visual details that appear weakly and locally within a cervical image. To overcome this challenge, our work divides an image into patches and then represents it from patch features. We aggregate patch patterns into an image feature in a weighted manner by considering the patch--image label relation. The weights are visualized as a heatmap to explain where the diagnosis results come from. We further introduce a spatial regulator to guide the classifier to focus on the cervix region and to adjust the weight distribution, without requiring any manual annotations of the cervix region. A novel iterative algorithm is designed to refine the regulator, which is able to capture the variations in cervix center locations and shapes. Experiments on an 18-year real-world dataset indicate a minimal of 3.47%, 4.59%, 8.54% improvements over the state-of-the-art in accuracy, F1, and recall measures, respectively.
Ying Zhang 0047, Yifang Yin, Zhenguang Liu, Roger Zimmermann
AAAI1
2021 Enhanced Audio Tagging via Multi- to Single-Modal Teacher-Student Mutual Learning
abstract
Recognizing ongoing events based on acoustic clues has been a critical yet challenging problem that has attracted significant research attention in recent years. Joint audio-visual analysis can improve the event detection accuracy but may not always be feasible as under many circumstances only audio recordings are available in real-world scenarios. To solve the challenges, we present a novel visual-assisted teacher-student mutual learning framework for robust sound event detection from audio recordings. Our model adopts a multi-modal teacher network based on both acoustic and visual clues, and a single-modal student network based on acoustic clues only. Conventional teacher-student learning performs unsatisfactorily for knowledge transfer from a multi-modality network to a single-modality network. We thus present a mutual learning framework by introducing a single-modal transfer loss and a cross-modal transfer loss to collaboratively learn the audio-visual correlations between the two networks. Our proposed solution takes the advantages of joint audio-visual analysis in training while maximizing the feasibility of the model in use cases. Our extensive experiments on the DCASE17 and the DCASE18 sound event detection datasets show that our proposed method outperforms the state-of-the-art audio tagging approaches.
Yifang Yin, Harsh Shrivastava 0002, Ying Zhang 0047, Zhenguang Liu, Rajiv Ratn Shah, Roger Zimmermann
AAAI3
2021 Multimodal Fusion of Satellite Images and Crowdsourced GPS Traces for Robust Road Attribute Detection
abstract
Automatic inference of missing road attributes (e.g., road type and speed limit) for enriching digital maps has attracted significant research attention in recent years. A number of machine learning based approaches have been proposed to detect road attributes from GPS traces, dash-cam videos, or satellite images. However, existing solutions mostly focus on a single modality without modeling the correlations among multiple data sources. To bridge the gap, we present a multimodal road attribute detection method, which improves the robustness by performing pixel-level fusion of crowdsourced GPS traces and satellite images. A GPS trace is usually given by a sequence of location, bearing, and speed. To align it with satellite imagery in the spatial domain, we render GPS traces into a sequence of multi-channel images that simultaneously capture the global distribution of the GPS points, the local distribution of vehicles' moving directions and speeds, and their temporal changes over time, at each pixel. Unlike previous GPS based road feature extraction methods, our proposed GPS rendering does not require map matching in the data preprocessing step. Moreover, our multimodal solution addresses single-modal challenges such as occlusions in satellite images and data sparsity in GPS traces by learning the pixel-wise correspondences among different data sources. Extensive experiments have been conducted on two real-world datasets in Singapore and Jakarta. Compared with previous work, our method is able to improve the detection accuracy on road attributes by a large margin.
Yifang Yin, An Tran, Ying Zhang 0047, Wenmiao Hu, Guanfeng Wang, Jagannadan Varadarajan, Roger Zimmermann, See-Kiong Ng
SIGSPATIAL/GIS3
2021 Learning Multi-context Aware Location Representations from Large-scale Geotagged Images
abstract
With the ubiquity of sensor-equipped smartphones, it is common to have multimedia documents uploaded to the Internet that have GPS coordinates associated with them. Utilizing such geotags as an additional feature is intuitively appealing for improving the performance of location-aware applications. However, raw GPS coordinates are fine-grained location indicators without any semantic information. Existing methods on geotag semantic encoding mostly extract hand-crafted, application-specific location representations that heavily depend on large-scale supplementary data and thus cannot perform efficiently on mobile devices. In this paper, we present a machine learning based approach, termed GPS2Vec+, which learns rich location representations by capitalizing on the world-wide geotagged images. Once trained, the model has no dependence on the auxiliary data anymore so it encodes geotags highly efficiently by inference. We extract visual and semantic knowledge from image content and user-generated tags, and transfer the information into locations by using geotagged images as a bridge. To adapt to different application domains, we further present an attention-based fusion framework that estimates the importance of the learnt location representations under different contexts for effective feature fusion. Our location representations yield significant performance improvements over the state-of-the-art geotag encoding methods on image classification and venue annotation.
Yifang Yin, Ying Zhang 0047, Zhenguang Liu, Yuxuan Liang 0002, Sheng Wang 0011, Rajiv Ratn Shah, Roger Zimmermann
ACM Multimedia2
2019 GPS2Vec: Towards Generating Worldwide GPS Embeddings
abstract
GPS coordinates are fine-grained location indicators that are difficult to be effectively utilized by classifiers in geo-aware applications. Previous GPS embedding methods are mostly tailored for specific problems that are taken place within areas of interest. When it comes to the scale of the entire planet, existing approaches always suffer from extensive computational cost and significant information loss. To solve these issues, we present a novel two-level grid based framework to learn semantic embeddings for geo-coordinates worldwide. The Earth's surface is first discretized by the Universal Transverse Mercator (UTM) coordinate system. Each UTM zone is next processed as a local area of interest that is further divided into fine-grained cells to perform the initial GPS encoding. We train a neural network in each UTM zone to learn the semantic embeddings from the initial GPS encoding. The training labels can be automatically derived from large-scale geotagged documents such as tweets, check-ins, and images that are available from social sharing platforms. We evaluate the effectiveness of our proposed GPS embeddings in geotagged image classification. Improved classification results have been obtained based on a simple early feature fusion technique.
Yifang Yin, Zhenguang Liu, Ying Zhang 0047, Sheng Wang 0011, Rajiv Ratn Shah, Roger Zimmermann
SIGSPATIAL/GIS3
2019 Importance of Truncation Activation in Pre-Processing for Spatial and Jpeg Image Steganalysis
abstract
In recent years, research on the task of image steganalysis have increasingly tapped on the power of deep learning algorithms to build more complex and deeper CNN to improve detection performance of embedded secrets in both grayscale and/or JPEG images. This paper will present an empirical study on the effectiveness of having a truncation activation function at the pre-processing phase of a CNN-based stegan-alyzer. Specifically, with commonly-used high-pass filters and the truncation activation, a domain-specific steganalyzer can now be expanded for multi-domain steganalysis, i.e., for both spatial and JPEG domain steganalysis. In our experiments, we investigated two state-of-the-art CNN-based steganalyzers, namely the Yedroudj-Net and Dense-Net which were originally built solely for spatial and JPEG steganalysis respectively. The truncation activation in pre-processing has shown to improve detection rate and accelerate the training phase.
Yew Yi Lu, Zhong Liang Ou Yang, Lilei Zheng, Ying Zhang 0047
ICIP4
2019 Coverless Image Steganography Framework with Increased Payload Capacity
abstract
This paper proposes a coverless image steganog- raphy framework with increased payload hiding capacity for confidential information transmission. This is superior to existing coverless schemes which require transmitting a stream of stego images in sequence due to the issue of small hiding capacity per stego image. In addition, by directly linking the secret with existing or generated images, coverless schemes are proven to be more secure as compared to traditional image steganography which embeds confidential message by modifying pixel values of a given image (known as cover). Specifically, we build a dictionary- based encoding scheme using images of the MNIST handwritten digit dataset to generate the stego image. In this dataset, thousands of images carry the same digit but have variance in their shapes and pixel depth. Thus it serves as a good codebook to represent information. The proposed coverless steganography scheme holds necessary resistance to image processing operations like JPEG compression, image binarization, image scaling and image noise, i.e., the hidden message can be successfully decoded from the processed stego images.
Ying Zhang 0047, Lilei Zheng, Yew Yi Lu, Vrizlynn L. L. Thing, Roger Zimmermann
ISM1
2019 A survey on image tampering and its detection in real-world photos
Lilei Zheng, Ying Zhang 0047, Vrizlynn L. L. Thing
J. Vis. Commun. Image Represent.2
2018 Towards Building a Remote Anti-spoofing Face Authentication System
abstract
The ability to offer remote access to services or data through various platforms has become a public expectation of many applications and systems nowadays. Recently, we can see the growth in trend where wide range of service providers offer the option to use biometric as a form of authentication, replacing conventional password system. While we gradually migrate towards such authentication method, it is important not to overlook the vulnerabilities of such systems to spoof attack. In fact, spoof attack must be prevented as a mandatory prerequisite in all biometric systems. In this paper, we present a systematic approach for face authentication that incorporates state-of-the-art liveness detection and face verification algorithms to safeguard a system against such attack. We explore and examine the feasibility of the application of such approach on generic devices and systems without incurring hardware dependencies or requiring extensive user-cooperation.
Chien Eao Lee, Lilei Zheng, Ying Zhang 0047, Vrizlynn L. L. Thing, Ying Yu Chu
TENCON3
2018 Face Spoofing Video Detection Using Spatio-Temporal Statistical Binary Pattern
abstract
Face has been used as a popular biometric trait to identify a person. However, attacking such face recognition system is not challenging today by using for example a fake face photo or video in front of a camera. In this paper, we present a novel feature, namely, SBP-TOP, to effectively identify such spoofings. SBP-TOP presents the texture information for the face region from both spatial and temporal perspectives. We have tested the proposed feature on two well-known face spoofing datasets: new Michigan State University mobile face spoofing database (MSU MFSD) and CASIA Face Anti-Spoofing Database (CASIA). The results indicate an accuracy over 95% on both datasets and there is an improvement over the state-of-the-art feature by around 10% and 3.2%, respectively.
Ying Zhang 0047, Rohit Kumar Dubey, Guang Hua 0001, Vrizlynn L. L. Thing
TENCON1
2018 A semi-feature learning approach for tampered region localization across multi-format images
Ying Zhang 0047, Vrizlynn L. L. Thing
Multim. Tools Appl.1
2017 Efficient Cache-Supported Path Planning on Roads (Extended Abstract)
abstract
Owing to the wide availability of the global positioning system (GPS) and digital mapping of roads, road network navigation services have become a basic application on many mobile devices. Path planning, a fundamental function of road network navigation services, finds a route between the specified start location and destination. The efficiency of this path planning function is critical for mobile users on roads due to various dynamic scenarios, such as a sudden change in driving direction, unexpected traffic conditions, lost or unstable GPS signals, and so on. In these scenarios, the path planning service needs to be delivered in a timely fashion. In this paper, we propose a system, namely, Path Planning by Caching (PPC), to answer a new path planning query in real time by efficiently caching and reusing historical queried-paths. Unlike the conventional cachebased path planning systems, where a queried-path in cache is used only when it matches perfectly with the new query, PPC leverages the partially matched queries to answer part(s) of the new query. Comprehensive experimentation on a real road network database shows that our system outperforms the state of-the-art path planning techniques by reducing 32% of the computation latency on average.
Ying Zhang 0047, Yu-Ling Hsueh, Wang-Chien Lee, Yi-Hao Jhang
ICDE1
2016 Geometric discriminative features for aerial image retrieval in social media
Yingjie Xia, Ying Zhang 0047
Multim. Syst.4
2016 Audio Authentication by Exploring the Absolute-Error-Map of ENF Signals
abstract
Recently, the electric network frequency (ENF), a natural signature embedded in many audio recordings, has been utilized as a criterion to examine the authenticity of audio recordings. ENF-based audio authentication system involves extraction of the ENF signal from a questioned audio recording, and matching it with the reference signal stored in an ENF database. This establishes a popular application of audio timestamp verification. In this paper, we explore another important application, i.e., ENF-based audio tampering detection, which has received less research attention. Specifically, we introduce the absolute-error-map (AEM) between the ENF signals obtained from the testing audio recording and the database. The AEM serves as an ensemble of the raw data associated with the ENF matching process. Through intensive analysis of the AEM, we propose two algorithms to jointly deal with timestamp verification and tampering detection, including insertion, deletion, and splicing attacks, respectively. The first algorithm is based on exhaustive point search and measurement, while the second algorithm leverages the image erosion technique to achieve fast detection of tampering type and tampered region, thus the second algorithm sacrifices some accuracy for speed. The authentication mechanism is that the system first determines if the testing data have been tampered with, and then outputs the timestamp information if no tampering is detected. Otherwise, it outputs the tampering type and tampered region. We demonstrate the effectiveness of the proposed solution via both synthetic and practical examples from our practically deployed audio authentication system.
Guang Hua 0001, Ying Zhang 0047, Jonathan Goh, Vrizlynn L. L. Thing
IEEE Trans. Inf. Forensics Secur.2
2016 Efficient Cache-Supported Path Planning on Roads
abstract
Owing to the wide availability of the global positioning system (GPS) and digital mapping of roads, road network navigation services have become a basic application on many mobile devices. Path planning, a fundamental function of road network navigation services, finds a route between the specified start location and destination. The efficiency of this path planning function is critical for mobile users on roads due to various dynamic scenarios, such as a sudden change in driving direction, unexpected traffic conditions, lost or unstable GPS signals, and so on. In these scenarios, the path planning service needs to be delivered in atimelyfashion. In this paper, we propose a system, namely,Path Planning by Caching (PPC), to answer a new path planning query in real time by efficiently caching and reusing historical queried-paths. Unlike the conventional cache-based path planning systems, where a queried-path in cache is used only when it matches perfectly with the new query, PPC leverages the partially matched queries to answer part(s) of the new query. As a result, the server only needs to compute the unmatched path segments, thus significantly reducing the overall system workload. Comprehensive experimentation on a real road network database shows that our system outperforms the state-of-the-art path planning techniques by reducing 32 percent of the computation latency on average.
Ying Zhang 0047, Yu-Ling Hsueh, Wang-Chien Lee, Yi-Hao Jhang
IEEE Trans. Knowl. Data Eng.1
2016 Efficient Summarization From Multiple Georeferenced User-Generated Videos
abstract
The rapid developments in camera technology and mobile devices bring a flourish of user-generated videos with rich geographic metadata, which can be great information resources for prospective tourists to preview a place of interest. In a video retrieval system, a simple query will return many videos. To provide users with a convenient way to explore videos, we generate summarization frommultiplegeoreferenced videos, which is composed of segments from different input videos but complementing each other to preserve the regions of interest (ROIs) among the original inputs. Different from conventional ROI detection techniques which extract objects from a single video with data-intensive computing, the proposed strategy solely leverages the geographic metadata among multiple videos (we term the ROI as Geographic-ROI/GROI). A Gaussian-based model is proposed to formulate the capture intention distribution in geo-space for each video-frame. By selecting such characteristics from keyframes in each video, we successfully detect the GROIs in a few milliseconds with average error distances within a few meters. Based on this, we select representative video segments capturing popular GROIs and compose them in a coherent manner as a final summarization. To speed up the processing, we represent videos according to their geographic characteristics and actively select their representative pieces (geo-keys). The highest quality videos among the local scope (key- neighborhood) for geo-keys are added to the summary with an appropriate travel route determined by their spatial consistency. Experiments indicate our summarization approach out-performs the state-of-the-arts by preserving the original GROIs with better accuracy and within less time.
Ying Zhang 0047, Roger Zimmermann
IEEE Trans. Multim.1
2015 Learning a Probabilistic Topology Discovering Model for Scene Categorization
abstract
A recent advance in scene categorization prefers a topological based modeling to capture the existence and relationships among different scene components. To that effect, local features are typically used to handle photographing variances such as occlusions and clutters. However, in many cases, the local features alone cannot well capture the scene semantics since they are extracted from tiny regions (e.g., 4×4 patches) within an image. In this paper, we mine a discriminative topology and a low-redundant topology from the local descriptors under a probabilistic perspective, which are further integrated into a boosting framework for scene categorization. In particular, by decomposing a scene image into basic components, a graphlet model is used to describe their spatial interactions. Accordingly, scene categorization is formulated as an intergraphlet matching problem. The above procedure is further accelerated by introducing a probabilistic based representative topology selection scheme that makes the pairwise graphlet comparison trackable despite their exponentially increasing volumes. The selected graphlets are highly discriminative and independent, characterizing the topological characteristics of scene images. A weak learner is subsequently trained for each topology, which are boosted together to jointly describe the scene image. In our experiment, the visualized graphlets demonstrate that the mined topological patterns are representative to scene categories, and our proposed method beats state-of-the-art models on five popular scene data sets.
Rongrong Ji, Yingjie Xia, Ying Zhang 0047, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2014 Points of Interest Detection from Multiple Sensor-Rich Videos in Geo-Space
abstract
Recently, the popularity of user generated videos has highlighted efficient video indexing and browsing as an urgent problem. Points of interest (POI) detection is a technique to address this issue by establishing the implicit relationship among different media resources. The majority of existing studies detect POI by visual similarity, leveraging computer vision techniques. However, these methods suffer from high computational complexity when processing large-scale video sets and are challenging due to the sparse visual correlation among different consumer videos. The advent of geo-referenced videos provides an opportunity to detect POIs in an efficient manner. In this work, we first propose a probability model to formulate the capture intention distribution for video frames. Second, we detect the POIs by considering the contributions from multiple videos. Evaluations demonstrate that our algorithm successfully detects POIs with a reduced error distance from several dozens to a few meters.
Ying Zhang 0047, Roger Zimmermann, David A. Shamma
ACM Multimedia1
2014 Aesthetics-Guided Summarization from Multiple User Generated Videos
abstract
In recent years, with the rapid development of camera technology and portable devices, we have witnessed a flourish of user generated videos, which are gradually reshaping the traditional professional video oriented media market. The volume of user generated videos in repositories is increasing at a rapid rate. In today's video retrieval systems, a simple query will return many videos which seriously increase the viewing burden. To manage these video retrievals and provide viewers with an efficient way to browse, we introduce a system to automatically generate a summarization from multiple user generated videos and present their salience to viewers in an enjoyable manner. Among multiple consumer videos, we find their qualities to be highly diverse due to various factors such as a photographer's experience or environmental conditions at the time of capture. Such quality inspires us to include a video quality evaluation component into the video summarization since videos with poor qualities can seriously degrade the viewing experience. We first propose a probabilistic model to evaluate the aesthetic quality of each user generated video. This model compares the rich aesthetics information from several well-known photo databases with generic unlabeled consumer videos, under a human perception component indicating the correlation between a video and its constituting frames. Subjective studies were carried out with the results indicating that our method is reliable. Then a novel graph-based formulation is proposed for the multi-video summarization task. Desirable summarization criteria is incorporated as the graph attributes and the problem is solved through a dynamic programming framework. Comparisons with several state-of-the-art methods demonstrate that our algorithm performs better than other methods in generating a skimming video in preserving the essential scenes from the original multiple input videos, with smooth transitions among consecutive segments and appealing aesthetics overall.
Ying Zhang 0047, Roger Zimmermann
ACM Trans. Multim. Comput. Commun. Appl.1
2013 Camera Shooting Location Recommendations for Landmarks in Geo-space
abstract
Taking photos of landmarks is a favorite and popular way for travellers to keep memories of places they have visited. Community-contributed photo collections, such as on Flickr, provide us an opportunity to gain a more in-depth understanding of a landmark's visual appeal. While much current research is focusing on recommending which representative photos should be selected from such pervasive photo sources, our work aims to find out where a visitor can capture his or her own, beautiful and personal photo of a queried landmark. We believe that this aspect of helping users to take memorable photos has not been well studied. We propose a method to recommend a list of shooting locations that have the utmost potential to capture appealing photos for a landmark of interest. A Gaussian Mixture Model based clustering approach is applied to the camera locations from an existing photo repository, generating a set of regions each of which covers an area with sufficient semantics, e.g., a route section. The scores and ranks among these camera locations are evaluated through multiple criteria, including their potential for better visual aesthetics, overall social attractiveness, popularity, etc. Additionally, we investigate the temporal characteristics of these locations by considering the spatio-temporal space. A number of different recommendations are generated from these results, such as the best camera positions at different times throughout a single day, or the best visiting time in the same spatial area. Subjective evaluation studies have been conducted, which indicate that our work can generate promising results.
Ying Zhang 0047, Roger Zimmermann
MASCOTS1
2013 Dynamic Multi-video Summarization of Sensor-Rich Videos in Geo-Space
Ying Zhang 0047, Roger Zimmermann
MMM (1)1
2012 DVS: a dynamic multi-video summarization system of sensor-rich videos in geo-space
abstract
It is now very easy to produce user generated videos (UGV) due to progress in camera and recording technologies on mobile devices, such as smartphones. Additionally, the ubiquitous, built-in sensors in digital devices can greatly enrich these videos with sensor descriptions, for example geo-spatial properties. A repository of such sensor-rich videos is a valuable source of information for prospective tourists when they plan to visit a city and would like to get a preview of its main attractions. Inspired by this, we have built an interactive geo-video search system. On a given map, a user specifies a start point and a destination and the system dynamically retrieves a video summarization along the path between the two points. Moreover, the user can interactively update the query during the video playback by dragging the route on the map. The main features of our technique are, first, that it is fully automatic and leverages sensor meta-data information which is acquired in conjunction with videos. Second, the system dynamically adapts to query updates in real-time. Third, a concise but comprehensive summarization from multiple user generated videos is presented for any queried route.
Ying Zhang 0047, Roger Zimmermann
ACM Multimedia1
2012 Multi-video summary and skim generation of sensor-rich videos in geo-space
abstract
User-generated videos have become increasingly popular in recent years. Due to advances in camera technology it is now very easy and convenient to record videos with mobile devices, such as smartphones. Here we consider an application where users collect and share a large set of videos that are related to a geographic area, say a city. Such a repository can be a great source of information for prospective tourists when they plan to visit a city and would like to get a preview of its main areas. The challenge that we address is how to automatically create a preview video summary from a large set of source videos.
Ying Zhang 0047, Guanfeng Wang, Beomjoo Seo, Roger Zimmermann
MMSys1