Zhilong Zhou

dblp:228/1352 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
3since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Segmentation and scene understanding · 50% Generative modeling · 37% Deep learning architectures and training · 13%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 87% Computational photography and imaging · 13%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
edge detection
0.822020
Bottom-up and Top-down: Bidirectional Additive Net for Edge Detection · IJCAI 2020
Cumulative Nets for Edge Detection · ACM Multimedia 2018
Machine learning › Generative modeling › diffusion model › image restoration
image inpainting
0.612022
CreaGAN: An Automatic Creative Generation Framework for Display Advertising · ACM Multimedia 2022
Visual content generation and editing
virtual try-on
0.612022
A High-resolution Image-based Virtual Try-on System in Taobao E-commerce Scenario · ACM Multimedia 2022
Web and social media mining
e-commerce
0.212022
A High-resolution Image-based Virtual Try-on System in Taobao E-commerce Scenario · ACM Multimedia 2022
Machine learning › Deep learning architectures and training › convolutional neural network
atrous convolution
0.112018
Cumulative Nets for Edge Detection · ACM Multimedia 2018
Machine learning › Deep learning architectures and training
convolutional neural network
0.112018
Cumulative Nets for Edge Detection · ACM Multimedia 2018

Methods — techniques the papers use, named apart from their topics

inpainting model · 1.1image generation · 1.1GAN · 1.1convolutional neural network · 0.4attention mechanism · 0.4cumulative residual attention · 0.3atrous convolution · 0.3CNN · 0.3ASPP · 0.3
YearPublicationVenuePosition
2023 Deep Task-specific Bottom Representation Network for Multi-Task Recommendation
abstract
Neural-based multi-task learning (MTL) has gained significant improvement, and it has been successfully applied to recommendation system (RS). Recent deep MTL methods for RS (e.g. MMoE, PLE) focus on designing soft gating-based parameter-sharing networks that implicitly learn a generalized representation for each task. However, MTL methods may suffer from performance degeneration when dealing with conflicting tasks, as negative transfer effects can occur on the task-shared bottom representation. This can result in a reduced capacity for MTL methods to capture task-specific characteristics, ultimately impeding their effectiveness and hindering the ability to generalize well on all tasks. In this paper, we focus on the bottom representation learning of MTL in RS and propose the Deep Task-specific Bottom Representation Network (DTRN) to alleviate the negative transfer problem. DTRN obtains task-specific bottom representation explicitly by making each task have its own representation learning network in the bottom representation modeling stage. Specifically, it extracts the user's interests from multiple types of behavior sequences for each task through the parameter-efficient hypernetwork. To further obtain the dedicated representation for each task, DTRN refines the representation of each feature by employing a SENet-like network for each task. The two proposed modules can achieve the purpose of getting task-specific bottom representation to relieve tasks' mutual interference. Moreover, the proposed DTRN is flexible to combine with existing MTL methods. Experiments on one public dataset and one industrial dataset demonstrate the effectiveness of the proposed DTRN.
Qi Liu 0003, Zhilong Zhou, Gangwei Jiang, Tiezheng Ge, Defu Lian
CIKM2
2022 CreaGAN: An Automatic Creative Generation Framework for Display Advertising
abstract
Creatives are an effective form of delivering product information on the E-commerce platform. Designing an exquisite creative is a time-consuming but crucial task for sellers. In order to accelerate the process, we propose an automatic creative generation framework, named CreaGAN, to make the design procedure easier, faster, and more accurate. Given a well-designed creative for one product, our method can generalize to other product materials by utilizing existing design elements (e.g., background material). The framework consists of two major parts: aesthetics-aware placement and creative inpainting model. The placement model aims to generate plausible locations for new products by considering aesthetic principles. And the inpainting model focus on filling the mismatched regions through contextual information. We conduct experiments on both the public dataset and the real-world creative dataset. Quantitative and qualitative results demonstrate that our method outperforms current state-of-the-art methods and obtains more reasonable and aesthetic visualization results.
Shiyao Wang 0001, Qi Liu 0003, Yicheng Zhong, Zhilong Zhou, Tiezheng Ge, Defu Lian, Yuning Jiang 0001
ACM Multimedia4
2022 A High-resolution Image-based Virtual Try-on System in Taobao E-commerce Scenario
abstract
On an e-commerce platform, virtual try-on not only improves consumers' shopping experience but also attracts more consumers. However, the existing public virtual try-on dataset can't be generalized to Taobao fashion items because of low resolution, racial differences, and small dataset size. In this work, we build a large-scale and high-resolution virtual try-on dataset for Taobao e-commerce. Based on the dataset, we propose a virtual try-on system to generate high-resolution and attractive virtual try-on images without time-consuming image preprocessing. The system has covered more than 10000 fashion items and produced millions of virtual try-on results in Taobao e-commerce scenario.
Zhilong Zhou, Shiyao Wang 0001, Tiezheng Ge, Yuning Jiang 0001
ACM Multimedia1
2020 Bottom-up and Top-down: Bidirectional Additive Net for Edge Detection
abstract
Image edge detection is considered as a cornerstone task in computer vision. Due to the nature of hierarchical representations learned in CNN, it is intuitive to design side networks utilizing the richer convolutional features to improve the edge detection. However, there is no consensus way to integrate the hierarchical information. In this paper, we propose an effective and end-to-end framework, named Bidirectional Additive Net (BAN), for image edge detection. In the proposed framework, we focus on two main problems: 1) how to design a universal network for incorporating hierarchical information sufficiently; and 2) how to achieve effective information flow between different stages and gradually improve the edge map stage by stage. To tackle these problems, we design a consecutive bottom-up and top-down architecture, where a bottom-up branch can gradually remove detailed or sharp boundaries to enable accurate edge detection and a top-down branch offers a chance of error-correcting by revisiting the low-level features that contain rich textual and spatial information. And attended additive module (AAM) is designed to cumulatively refine edges by selecting pivotal features in each stage. Experimental results show that our proposed methods can improve the edge detection performance to new records and achieve state-of-the-art results on two public benchmarks: BSDS500 and NYUDv2.
Lianli Gao, Zhilong Zhou, Heng Tao Shen, Jingkuan Song
IJCAI2
2019 Residual attention-based LSTM for video captioning
Zhilong Zhou, Lijiang Chen, Lianli Gao
World Wide Web2
2018 Cumulative Nets for Edge Detection
abstract
Lots of recent progress have been made by using Convolutional Neural Networks (CNN) for edge detection. Due to the nature of hierarchical representations learned in CNN, it is intuitive to design side networks utilizing the richer convolutional features to improve the edge detection. However, different side networks are isolated, and the final results are usually weighted sum of the side outputs with uneven qualities. To tackle these issues, we propose a Cumulative Network (C-Net), which learns the side network cumulatively based on current visual features and low-level side outputs, to gradually remove detailed or sharp boundaries to enable high-resolution and accurate edge detection. Therefore, the lower-level edge information is cumulatively inherited while the superfluous details are progressively abandoned. In fact, recursively Learningwhere to remove superfluous details from the current edge map with the supervision of a higher-level visual feature is challenging. Furthermore, we employ atrous convolution (AC) and atrous convolution pyramid pooling (ASPP) to robustly detect object boundaries at multiple scales and aspect ratios. Also, cumulatively refining edges using high-level visual information and lower-lever edge maps is achieved by our designed cumulative residual attention (CRA) block. Experimental results show that our C-Net sets new records for edge detection on both two benchmark datasets: BSDS500 (i.e., .819 ODS, .835 OIS and .862 AP) and NYUDV2 (i.e., .762 ODS, .781 OIS, .797 AP). C-Net has great potential to be applied to other deep learning based applications, e.g., image classification and segmentation.
Jingkuan Song, Zhilong Zhou, Lianli Gao, Xing Xu 0001, Heng Tao Shen
ACM Multimedia2