Zihang Lin

dblp:251/3444 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Style Nursing with Spatial and Semantic Guidance for Zero-Shot Traffic Scene Style Transfer
abstract
Recent advances in text-to-image diffusion models have shown an outstanding ability in zero-shot style transfer. However, existing methods often struggle to balance preserving the semantic content of the input image and faithfully transferring the target style in line with the edit prompt. Especially when applied to complex traffic scenes with diverse objects, layouts, and stylistic variations, current diffusion models tend to exhibit Style Neglection, i.e., failing to generate the required style in the prompt. To address this issue, we propose Style Nursing, which directs the model to focus on style subject tokens in the text prompt and excites their corresponding visual activations. Moreover, we introduce Spatial and Semantic Guidance to guide the preservation of content after editing, which utilizes spatial features from the DDIM sampling process together with attention maps from the semantic reconstruction. To evaluate the performance of zero-shot style transfer methods in traffic scenes, we present STREET-6K, a new benchmark dataset comprising 6000 images showcasing diverse traffic scenes and style transfer variations, accompanied by comprehensive annotations and evaluation metrics. Our approach beats state-of-the-art image translation methods in comprehensive quantitative metrics and human evaluations on traffic scene image synthesis while seamlessly generalizing to various other types of images without training or fine-tuning. Further experiments on detection and segmentation tasks show that fine-tuning perception models on our synthesized images improves Recall and mean Intersection over Union (mIoU) by over 10% and 3% respectively in rarely-seen traffic scenes.
Zihang Lin, Yuehu Liu, Chi Zhang 0020
AAAI2
2025 T-Stars-Poster: A Framework for Product-Centric Advertising Image Design
Hongyu Chen 0005, Zihang Lin, Bo Xiao 0006, Tiezheng Ge, Bo Zheng 0007
CIKM6
2025 PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering
abstract
Product posters, which integrate subject, scene, and text, are crucial promotional tools for attracting customers. Creating such posters using modern image generation methods is valuable, while the main challenge lies in accurately rendering text, especially for complex writing systems like Chinese, which contains over 10,000 individual characters. In this work, we identify the key to precise text rendering as constructing a character-discriminative visual feature as a control signal. Based on this insight, we propose a robust character-wise representation as control and we develop TextRenderNet, which achieves a high text rendering accuracy of over 90%. Another challenge in poster generation is maintaining the fidelity of user-specific products. We address this by introducing SceneGenNet, an inpainting-based model, and propose subject fidelity feedback learning to further enhance fidelity. Based on TextRenderNet and SceneGenNet, we present PosterMaker, an end-to-end generation framework. To optimize PosterMaker efficiently, we implement a two-stage training strategy that decouples text rendering and background generation learning. Experimental results show that PosterMaker outperforms existing baselines by a remarkable margin, which demonstrates its effectiveness.
Yifan Gao 0011, Zihang Lin, Chuanbin Liu 0001, Tiezheng Ge, Bo Zheng 0007, Hongtao Xie 0001
CVPR2
2024 LearnedFTL: A Learning-Based Page-Level FTL for Reducing Double Reads in Flash-Based SSDs
abstract
We present LearnedFTL, a new on-demand pagelevel flash translation layer (FTL) design, which employs learned indexes to improve the address translation efficiency of flashbased SSDs. The first of its kind, it reduces the number of double reads induced by address translation in random read accesses. LearnedFTL proposes three key techniques: an in-place-update linear model to build learned indexes efficiently, a virtual PPN representation to obtain contiguous PPNs for sorted LPNs, and a group-based allocation and model training via GC/rewrite strategy to reduce the training overhead. By tightly integrating the aforementioned key techniques, LearnedFTL considerably speeds up address translation while reducing the number of flash read accesses caused by the address translation. Our extensive experiments on a FEMU-based prototype show that LearnedFTL can reduce up to 55.5% address translation-induced double reads. As a result, LearnedFTL reduces the P99 tail latency by 2.9× ∼ 12.2× with an average of 5.5× and 8.2× compared to the state-of-the-art TPFTL and LeaFTL schemes, respectively.
Shengzhe Wang 0001, Zihang Lin, Suzhen Wu, Hong Jiang 0001, Jie Zhang 0048, Bo Mao 0003
HPCA2
2024 SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
Chaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi, Wei-Yi Pei, Yexin Wang, Ying Shan, Wei-Shi Zheng 0001, Jianfang Hu
ACM Multimedia2
2024 FSDedup: Feature-Aware and Selective Deduplication for Improving Performance of Encrypted Non-Volatile Main Memory
abstract
Enhancing the endurance, performance, and energy efficiency of encrypted Non-Volatile Main Memory (NVMM) can be achieved by minimizing written data through inline deduplication. However, existing approaches applying inline deduplication to encrypted NVMM suffer from substantial performance degradation due to high computing, memory footprint, and index-lookup overhead to generate, store, and query the cryptographic hash (fingerprint). In the preliminary ESD [ 14 ], we proposed the Error Correcting Code (ECC) assisted selective deduplication scheme, utilizing the ECC information as a fingerprint to identify similar data effectively and then leveraging the selective deduplication technique to eliminate a large amount of redundant data with high reference counts. In this article, we proposed FSDedup. Compared with ESD, FSDedup could leverage the prefetch cache to reduce the read overhead during similarity comparison and utilize the cache refresh mechanism to identify further and eliminate more redundant data. Extensive experimental evaluations demonstrate that FSDedup can enhance the performance of the NVMM system further than the ESD. Experimental results show that FSDedup can improve both write and read speed by up to 1.8×, enhance Instructions Per Cycle by up to 1.5×, and reduce energy consumption by up to 2.0×, compared to ESD.
Chunfeng Du, Zihang Lin, Suzhen Wu, Yifei Chen 0011, Jiapeng Wu, Shengzhe Wang 0001, Weichun Wang 0002, Bo Mao 0003
ACM Trans. Storage2
2023 Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding
abstract
Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the language description, which requires effective joint modeling of spatiotemporal visuallinguistic dependencies. In this work, we propose a novel framework in which a static vision-language stream and a dynamic vision-language stream are developed to collaboratively reason the target tube. The static stream performs cross-modal understanding in a single frame and learns to attend to the target object spatially according to intraframe visual cues like object appearances. The dynamic stream models visual-linguistic dependencies across multiple consecutive frames to capture dynamic cues like motions. We further design a novel cross-stream collaborative block between the two streams, which enables the static and dynamic streams to transfer useful and complementary information from each other to achieve collaborative reasoning. Experimental results show the effectiveness of the collaboration of the two streams and our overall frame-work achieves new state-of-the-art performance on both HCSTVG and VidSTG datasets.
Zihang Lin, Chaolei Tan, Jianfang Hu, Zhi Jin 0002, Tiancai Ye, Wei-Shi Zheng 0001
CVPR1
2023 Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding
abstract
Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex semantic relations between visual and textual modalities. Previous methods focus on modeling the contextual information between the video and text from a single-level perspective (i.e., the sentence level), ignoring rich visual-textual correspondence relations at different semantic levels, e.g., the video-word and video-paragraph correspondence. To this end, we propose a novel Hierarchical Semantic Correspondence Network (HSCNet), which explores multi-level visual-textual correspondence by learning hierarchical semantic alignment and utilizes dense supervision by grounding diverse levels of queries. Specifically, we develop a hierarchical encoder that encodes the multi-modal inputs into semantics-aligned representations at different levels. To exploit the hierarchical semantic correspondence learned in the encoder for multi-level supervision, we further design a hierarchical decoder that progressively performs finer grounding for lower-level queries conditioned on higher-level semantics. Extensive experiments demonstrate the effectiveness of HSCNet and our method significantly outstrips the state-of-the-arts on two challenging benchmarks, i.e., ActivityNet-Captions and TACoS.
Chaolei Tan, Zihang Lin, Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai
CVPR2
2022 APANet: Auto-Path Aggregation for Future Instance Segmentation Prediction
abstract
Despite the remarkable progress achieved in conventional instance segmentation, the problem of predicting instance segmentation results for unobserved future frames remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting features of future frames. However, these methods always treat features of multiple levels (e.g., coarse-to-fine pyramid features) independently and do not exploit them collaboratively, which results in inaccurate prediction for future frames; and moreover, such a weakness can partially hinder self-adaption of a future segmentation prediction model for different input samples. To solve this problem, we propose an adaptive aggregation approach called Auto-Path Aggregation Network (APANet), where the spatio-temporal contextual information obtained in the features of each individual level is selectively aggregated using the developed "auto-path". The "auto-path" connects each pair of features extracted at different pyramid levels for task-specific hierarchical contextual information aggregation, which enables selective and adaptive aggregation of pyramid features in accordance with different videos/frames. Our APANet can be further optimized jointly with the Mask R-CNN head as a feature decoder and a Feature Pyramid Network (FPN) feature encoder, forming a joint learning system for future instance segmentation prediction. We experimentally show that the proposed method can achieve state-of-the-art performance on three video-based instance segmentation benchmarks for future instance segmentation prediction.
Jianfang Hu, Jiangxin Sun, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Structural Hole Theory in Social Network Analysis: A Review
abstract
Social networks now connect billions of people around the world, where individuals occupying different positions often represent different social roles and show different characteristics in their behaviors. The structural hole (SH) theory demonstrates that users occupying the bridging positions between different communities have advantages since they control the key information diffusion paths. Users of this type, known as SH spanners, are important when it comes to assimilating social network structures and user behaviors. In this article, we review the use of SHs theory in social network analysis, where SH spanners take advantage of both information and control benefits. We investigate the existing algorithms of SH spanner detection and classify them into information flow-based algorithms and network centrality-based algorithms. For practitioners, we further illustrate the applications of SH theory in various practical scenarios, including enterprise settings, information diffusion in social networks, software development, mobile applications, and machine learning (ML)-based social prediction. Our review provides a comprehensive discussion on the foundation, detection, and practical applications of SHs. The insights can facilitate researchers and service providers to better apply the theory and derive value-added tools with advanced ML techniques. To inspire follow-up research, we identify potential research trends in this area, especially on the dynamics of networks.
Zihang Lin, Qingyuan Gong, Yang Chen 0001, Atte Oksanen, Aaron Yi Ding
IEEE Trans. Comput. Soc. Syst.1
2022 Density-Aware Haze Image Synthesis by Self-Supervised Content-Style Disentanglement
abstract
The key procedure of haze image synthesis with adversarial training lies in the disentanglement of the feature involved only in haze synthesis, i.e.,the style feature, from the feature representing the invariant semantic content, i.e.,the content feature. Previous methods introduced a binary classifier to constrain the domain membership from being distinguished through the learned content feature during the training stage, thereby the style information is separated from the content feature. However, we find that these methods cannot achieve complete content-style disentanglement. The entanglement of the flawed style feature with content information inevitably leads to the inferior rendering of haze images. To address this issue, we propose a self-supervised style regression model with stochastic linear interpolation that can suppress the content information in the style feature. Ablative experiments demonstrate the disentangling completeness and its superiority in density-aware haze image synthesis. Moreover, the synthesized haze data are applied to test the generalization ability of vehicle detectors. Further study on the relation between haze density and detection performance shows that haze has an obvious impact on the generalization ability of vehicle detectors and that the degree of performance degradation is linearly correlated to the haze density, which in turn validates the effectiveness of the proposed method.
Chi Zhang 0020, Zihang Lin, Liheng Xu, Zongliang Li, Wei Tang 0016, Yuehu Liu, Gaofeng Meng, Le Wang 0003, Li Li 0013
IEEE Trans. Circuits Syst. Video Technol.2
2021 Predictive Feature Learning for Future Segmentation Prediction
abstract
Future segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation models. However, these segmentation features are learned to be local discriminative (with rich details) and are always of high resolution/dimension. Hence, the complicated spatiotemporal variations of these features are difficult to predict, which motivates us to learn a more predictive representation. In this work, we develop a novel framework called Predictive Feature Autoencoder. In the proposed framework, we construct an autoencoder which serves as a bridge between the segmentation features and the predictor. In the latent feature learned by the autoencoder, global structures are enhanced and local details are suppressed so that it is more predictive. In order to reduce the risk of vanishing the suppressed details during recurrent feature prediction, we further introduce a reconstruction constraint in the prediction module. Extensive experiments show the effectiveness of the proposed approach and our method outperforms state-of-the-arts by a considerable margin.
Zihang Lin, Jiangxin Sun, Jianfang Hu, Qi-Zhi Yu, Jian-Huang Lai, Wei-Shi Zheng 0001
ICCV1
2021 Action-guided 3D Human Motion Prediction
abstract
The ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human motion is closely related to the action being performed, we propose to explore action context to guide motion prediction. Specifically, we construct an action-specific memory bank to store representative motion dynamics for each action category, and design a query-read process to retrieve some motion dynamics from the memory bank. The retrieved dynamics are consistent with the action depicted in the observed video frames and serve as a strong prior knowledge to guide motion prediction. We further formulate an action constraint loss to ensure the global semantic consistency of the predicted motion. Extensive experiments demonstrate the effectiveness of the proposed approach, and we achieve state-of-the-art performance on 3D human motion prediction.
Jiangxin Sun, Zihang Lin, Xintong Han, Jianfang Hu, Jia Xu 0011, Wei-Shi Zheng 0001
NeurIPS2
2019 Predicting Future Instance Segmentation with Contextual Pyramid ConvLSTMs
abstract
Despite the remarkable progress in instance segmentation, the problem of predicting future instance segmentation remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting pyramid features to represent unobserved future frames. However, they mainly predict features for each pyramid level independently, and ignore the underlying structural relationship between features of different levels.
Jiangxin Sun, Jiafeng Xie, Jianfang Hu, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001
ACM Multimedia4