Ziling Huang

dblp:150/6487 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
9since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 23% Segmentation and scene understanding · 20% Face, body and person analysis · 12%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.812024
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024
Machine learning › Efficient and distributed learning › distributed training
recommendation model training
0.812024
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024
Computer vision › Vision and language
visual grounding
0.812024
LoA-Trans: Enhancing Visual Grounding by Location-Aware Transformers · ECCV (7) 2024
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.812024
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024
Machine learning › Deep learning architectures and training › feedforward neural network
cascaded network
0.712023
Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment Network · EMNLP 2023
Computer vision › Segmentation and scene understanding › image segmentation
mask prediction
0.712023
Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment Network · EMNLP 2023
Computer vision › Segmentation and scene understanding
referring image segmentation
0.712023
Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment Network · EMNLP 2023
Natural language and speech › Information extraction and text analysis › relation extraction › social network extraction
social relation extraction
0.512021
FL-MSRE: A Few-Shot Learning based Approach to Multimodal Social Relation Extraction · AAAI 2021
Machine learning › Generative modeling › diffusion model › image restoration
image dehazing
0.412020
HardGAN: A Haze-Aware Representation Distillation GAN for Single Image Dehazing · ECCV (6) 2020
Machine learning › Transfer learning and domain adaptation › knowledge transfer › representation transfer
representation distillation
0.412020
HardGAN: A Haze-Aware Representation Distillation GAN for Single Image Dehazing · ECCV (6) 2020
Computer vision › Face, body and person analysis › person re-identification
group re-identification
0.412019
DoT-GNN: Domain-Transferred Graph Neural Network for Group Re-identification · ACM Multimedia 2019
Computer vision › Face, body and person analysis
person re-identification
0.412019
DoT-GNN: Domain-Transferred Graph Neural Network for Group Re-identification · ACM Multimedia 2019
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.112021
FL-MSRE: A Few-Shot Learning based Approach to Multimodal Social Relation Extraction · AAAI 2021

Methods — techniques the papers use, named apart from their topics

resource-performance models · 1.5heuristic resource allocation · 1.5transformer · 0.8location-aware attention · 0.8progressive alignment network · 0.7contextual embedding learning · 0.7few-shot learning · 0.5BERT · 0.5knowledge distillation · 0.4generative adversarial network · 0.4
YearPublicationVenuePosition
2027 Lightweight speech enhancement guided target speech extraction in noisy scenarios
Ziling Huang, Junnan Wu, Lichun Fan, Haixin Guan, Yanhua Long
Comput. Speech Lang.1
2026 DLRover-LM: LLM Pre-Training Framework With Thousands of Accelerators in AntGroup
Ziling Huang, Zhengmao Ye, Qingsong Cai, Zelong Huang, Bo Sang, Jian Sha, Tingfeng Lan, Hui Lu 0001, Yuanchun Zhou, MingJie Tang
ICDE1
2025 SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation
abstract
Personalized speech enhancement (PSE) methods typically rely on pre-trained speaker verification models or self-designed speaker encoders to extract target speaker clues, guiding the PSE model in isolating the desired speech. However, these approaches suffer from significant model complexity and often underutilize enrollment speaker information, limiting the potential performance of the PSE model. To address these limitations, we propose a novel Speaker Encoder-Free PSE network, termed SEF-PNet, which fully exploits the information present in both the enrollment speech and noisy mixtures. SEF-PNet incorporates two key innovations: Interactive Speaker Adaptation (ISA) and Local-Global Context Aggregation (LCA). ISA dynamically modulates the interactions between enrollment and noisy signals to enhance the speaker adaptation, while LCA employs advanced channel attention within the PSE encoder to effectively integrate local and global contextual information, thus improving feature learning. Experiments on the Libri2Mix dataset demonstrate that SEF-PNet significantly outperforms baseline models, achieving state-of-the-art PSE performance. Our source code is available at https://github.com/isHuangZiling/SEF-PNet.
Ziling Huang, Haixin Guan, Yanhua Long
ICASSP1
2025 A survey on action and event recognition with detectors for video data
abstract
In this paper, the development of action and event detectors over the past three decades is summarized. The detectors are divided into 2D detectors, 3D detectors and deep learning detectors according to whether they contain spatial information and whether they use deep learning. This paper briefly introduces the typical detectors of the different types mentioned above, and explains the advantages, disadvantages and characteristics respectively, and compares them. Comparing traditional feature detection methods with ones based on deep learning, we found that the method of first detecting microscopic details such as point, line, surface angle, etc., and then performing action and event recognition is no longer the mainstream of current research. Due to the strong generalization ability, end-to-end action and event recognition methods based on deep learning perform better than traditional methods. Finally, this paper proposes three research directions for action recognition and event recognition based on feature detectors.
Changyu Liu, Zhankai Gao, Jiawei Pei, Ziling Huang, Yicong He
Web Intell.5
2024 LoA-Trans: Enhancing Visual Grounding by Location-Aware Transformers
Ziling Huang, Shin'ichi Satoh 0001
ECCV (7)1
2024 DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud
abstract
Deep learning recommendation models (DLRM) rely on large embedding tables to manage categorical sparse features. Expanding such embedding tables can significantly enhance model performance, but at the cost of increased GPU/CPU/memory usage. Meanwhile, tech companies have built extensive cloud-based services to accelerate training DLRM models at scale. In this paper, we conduct a deep investigation of the DLRM training platforms at AntGroup and reveal two critical challenges: low resource utilization due to suboptimal configurations by users and the tendency to encounter abnormalities due to an unstable cloud environment. To overcome them, we introduce DLRover, an elastic training framework for DLRMs designed to increase resource utilization and handle the instability of a cloud environment. DLRover develops a resource-performance model by considering the unique characteristics of DLRMs and a three-stage heuristic strategy to automatically allocate and dynamically adjust resources for DLRM training jobs for higher resource utilization. Further, DLRover develops multiple mechanisms to ensure efficient and reliable execution of DLRM training jobs. Our extensive evaluation shows that DLRover reduces job completion times by 31%, increases the job completion rate by 6%, enhances CPU usage by 15%, and improves memory utilization by 20%, compared to state-of-the-art resource scheduling frameworks. DLRover has been widely deployed at AntGroup and processes thousands of DLRM training jobs on a daily basis. DLRover is open-sourced and has been adopted by 10+ companies.
Qinlong Wang, Tingfeng Lan, Yinghao Tang, Bo Sang, Ziling Huang, Yiheng Du, Jian Sha, Hui Lu 0001, Yuanchun Zhou, Ke Zhang 0048, MingJie Tang
Proc. VLDB Endow.5
2023 Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment Network
abstract
Referring image segmentation is a task that aims to predict pixel-wise masks corresponding to objects in an image described by natural language expressions.Previous methods for referring image segmentation employ a cascade framework to break down complex problems into multiple stages.However, its limitations are also apparent: existing methods within the cascade framework may encounter challenges in both maintaining a strong focus on the most relevant information during specific stages of the referring image segmentation process and rectifying errors propagated from early stages, which can ultimately result in sub-optimal performance.To address these limitations, we propose the Joint Mask Contextual Embedding Learning Network (JMCELN).JMCELN is designed to enhance the Cascade Framework by incorporating a Learnable Contextual Embedding and a Progressive Alignment Network (PAN).The Learnable Contextual Embedding module dynamically stores and utilizes reasoning information based on the current mask prediction results, enabling the network to adaptively capture and refine pertinent information for improved mask prediction accuracy.Furthermore, the Progressive Alignment Network (PAN) is introduced as an integral part of JMCELN.PAN leverages the output from the previous layer as a filter for the current output, effectively reducing inconsistencies between predictions from different stages.By iteratively aligning the predictions, PAN guides the Learnable Contextual Embedding to incorporate more discriminative information for reasoning, leading to enhanced prediction quality and a reduction in error propagation.With these methods, we achieved state-of-the-art results on three commonly used benchmarks, especially in more intricate datasets.
Ziling Huang, Shin'ichi Satoh 0001
EMNLP1
2021 FL-MSRE: A Few-Shot Learning based Approach to Multimodal Social Relation Extraction
abstract
Social relation extraction (SRE for short), which aims to infer the social relation between two people in daily life, has been demonstrated to be of great value in reality. Existing methods for SRE consider extracting social relation only from unimodal information such as text or image, ignoring the high coupling of multimodal information. Moreover, previous studies overlook the serious unbalance distribution on social relations. To address these issues, this paper proposes FL-MSRE, a few-shot learning based approach to extracting social relations from both texts and face images. Considering the lack of multimodal social relation datasets, this paper also presents three multimodal datasets annotated from four classical masterpieces and corresponding TV series. Inspired by the success of BERT, we propose a strong BERT based baseline to extract social relation from text only. FL-MSRE is empirically shown to outperform the baseline significantly. This demonstrates that using face images benefits text-based SRE. Further experiments also show that using two faces from different images achieves similar performance as from the same image. This means that FL-MSRE is suitable for a wide range of SRE applications where the faces of two people can only be collected from different images.
Hai Wan, Manrong Zhang, Jianfeng Du, Ziling Huang, Jeff Z. Pan
AAAI4
2021 DotSCN: Group Re-Identification via Domain-Transferred Single and Couple Representation Learning
abstract
Group re-identification (G-ReID) is an important yet less-studied task. Its challenges not only lie in appearance changes of individuals, but also involve group layout and membership changes. To address these issues, the key task of G-ReID is to learn group representations robust to such changes. Nevertheless, unlike ReID tasks, there still lacks comprehensive publicly available G-ReID datasets, making it difficult to learn effective representations using deep learning models. In this article, we propose a Domain-Transferred Single and Couple Representation Learning Network (DotSCN). Its merits are two aspects: 1) Owing to the lack of labelled training samples for G-ReID, existing G-ReID methods mainly rely on unsatisfactory hand-crafted features. To gain the power of deep learning models in representation learning, we first treat a group as a collection of multiple individuals and propose transferring the representation of individuals learned from an existing labeled ReID dataset to a target G-ReID domain without a suitable training dataset. 2) Taking into account the neighborhood relationship in a group, we further propose learning a novel couple representation between two group members, that achieves better discriminative power in G-ReID tasks. In addition, we propose a weight learning method to adaptively fuse the domain-transferred individual and couple representations based on an L-shape prior. Extensive experimental results demonstrate the effectiveness of our approach that significantly outperforms state-of-the-art methods by 11.7% CMC-1 on the Road Group dataset and by 39.0% CMC-1 on the DukeMCMT dataset.
Ziling Huang, Zheng Wang 0007, Chung-Chi Tsai, Shin'ichi Satoh 0001, Chia-Wen Lin
IEEE Trans. Circuits Syst. Video Technol.1
2020 HardGAN: A Haze-Aware Representation Distillation GAN for Single Image Dehazing
Qili Deng, Ziling Huang, Chung-Chi Tsai, Chia-Wen Lin
ECCV (6)2
2019 Consistency Constrained Reconstruction of Depth Maps from Epipolar Plane Images
abstract
In this paper, we propose a method of reconstructing the depth map of a set of multiview images from the epipolar plane images (EPIs) of multiview Images. Our method involves two steps: finding support points and estimating depth. First, we propose to include a consistency term and a smoothness term in the objective function for edge point detection, where the consistency term is used to identify edge points and the smoothness term is applied to mitigate false edge detection due to light density variations caused by viewpoint changes. Then, based on the detected edge points, a depth map can be estimated by solving a energy minimization problem, in which a line uniformness term and a matching error term are introduced to ensure the line traces estimated from EPIs for depth estimation match the colors of edge points well. The depths of non-edge points are then estimated by introducing an additional prior term. In order to speed up our algorithm, the depth estimation problem is aggregated by a winner-take-all strategy. Experiments show that our method outperforms the state-of-the-art schemes in reconstructing depth map with fine details.
Ziling Huang, Chia-Wen Lin, Hao-Chiang Shao, Xiangsheng Huang
ICASSP1
2019 DoT-GNN: Domain-Transferred Graph Neural Network for Group Re-identification
abstract
Most person re-identification (ReID) approaches focus on retrieving a person-of-interest from a database of collected individual images. In addition to the individual ReID task, matching a group of persons across different camera views also plays an important role in surveillance applications. This kind of Group Re-identification (GReID) task is very challenging since we face the obstacles not only from the appearance changes of individuals, but also from the group layout and membership changes. In order to obtain robust representation for the group image, we design a Domain-Transferred Graph Neural Network (DoT-GNN) method. The merits are three aspects: 1) Transferred Style. Due to the lack of training samples, we transfer the labeled ReID dataset to the G-ReID dataset style, and feed the transferred samples to the deep learning model. Taking the superiority of deep learning models, we achieve a discriminative individual feature model. 2) Graph Generation. We treat a group as a graph, where each node denotes the individual feature and each edge represents the relation of a couple of individuals. We propose a graph generation strategy to create sufficient graph samples. 3) Graph Neural Network. Employing the generated graph samples, we train the GNN so as to acquire graph features which are robust to large graph variations. The key to the success of DoT-GNN is that the transferred graph addresses the challenge of the appearance change, while the graph representation in GNN overcomes the challenge of the layout and membership change. Extensive experimental results demonstrate the effectiveness of our approach, outperforming the state-of-the-art method by 1.8% CMC-1 on Road Group dataset and 6.0% CMC-1 on DukeMCMT dataset respectively.
Ziling Huang, Zheng Wang 0007, Wei Hu 0003, Chia-Wen Lin, Shin'ichi Satoh 0001
ACM Multimedia1
2018 Joint Pairwise Learning and Image Clustering Based on a Siamese CNN
abstract
How to use a deep convolutional neural network (CNN) to efficiently and effectively learn representations of a large unlabeled set of images and group them into clusters remains a challenging problem. To address this problem, we propose a Siamese clustering CNN (SC-CNN) to iteratively learn discriminative representations for image clustering. Based on the proposed SC-CNN, we employ a mini-batch-based joint pairwise representation learning and clustering scheme to make the computation and storage cost efficient for large-scale image clustering on a personal computer with a commercial GPU graphic card. On top of SC-CNN, the proposed pairwise learning scheme effectively learns discriminative representations by appropriately selecting same-cluster and different-cluster image pairs from the results of each clustering iteration. Experimental results demonstrate that the proposed method outperforms start-of-the-art clustering schemes in clustering accuracy on public image sets.
Weng-Tai Su, Chih-Chung Hsu, Ziling Huang, Chia-Wen Lin, Gene Cheung
ICIP3
2016 A semi-global matching method for large-scale light field images
abstract
Semi-Global Matching (SGM) is a robust method in traditional stereo matching. It maintains precise boundary with low computational cost. However, directly applying SGM to light field stereo matching degrades the results greatly due to the sparsity of support points. In this letter, we proposes a novel stereo matching approach for large-scale light field images. We observe that adding weak edges to support points efficiently stabilizes the depth propagation. Based on this observation, we apply a cross detector to obtain support points, and then we propagate the depth of support points to homogeneous region. By solving a semi-global energy minimization problem, the depth information can be well estimated from epipolar plane images. Besides, we introduce a new strategy to deal with occlusion. We iteratively sample the pixels under current disparity hypothesis and the consistency scores are aggregated by a weighted winner-take-all strategy. Our method allows for significant reduce of the disparity search space, the time is halved and the depth is more robust at the occurrence of occlusion. For every pixel, the calculation is based on a single EPI and locally independent. Implementation on GPU shows that our method can achieve state-of-art results with less computational cost.
Xiangsheng Huang, Ziling Huang, Weili Ding
ICASSP2
2014 Propeller: A Scalable Real-Time File-Search Service in Distributed Systems
abstract
File-search service is a valuable facility to accelerate many analytics applications, because it can drastically reduce the scale of the input data. The main challenge facing the design of large-scale and accurate file-search services is how to support real-time indexing in an efficient and scalable way. To address this challenge, we propose a distributed file-search service, called Propeller, which utilizes a special file-access pattern, called access-causality, to partition file-indices in order to expose substantial access locality and parallelism to accelerate the file-indexing process. The extensive evaluations of Propeller show that it is real-time in file-indexing operations, accurate in file-search results, and scalable in large datasets. It achieves significantly better file-indexing and file-search performance (up to 250x) than a centralized solution (MySQL) and much higher accuracy and substantially lower query latency (up to 22x than a state-of-the-art desktop search engine (Spotlight).
Lei Xu 0038, Hong Jiang 0001, Lei Tian 0001, Ziling Huang
ICDCS4