VLDB 2026 Research / reviewers in the wild / expert
Siyu Huang
dblp:146/9031
· DBLP profile ↗
56ranked-venue papers
15as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 7 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 11 first-author · 18 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KScaNN: Scalable Approximate Nearest Neighbor Search on KunpengabstractApproximate Nearest Neighbor Search (ANNS) is a cornerstone algorithm for information retrieval, recommendation systems, and machine learning applications. While x86-based architectures have historically dominated this domain, the increasing adoption of ARM-based servers in industry presents a critical need for ANNS solutions optimized on ARM architectures. A naive port of existing x86 ANNS algorithms to ARM platforms results in a substantial performance deficit, failing to leverage the unique capabilities of the underlying hardware. To address this challenge, we introduce KScaNN, a novel ANNS algorithm co-designed for the Kunpeng 920 ARM architecture. KScaNN embodies a holistic approach that synergizes sophisticated, data aware algorithmic refinements with carefully-designed hardware specific optimizations. Its core contributions include: 1) novel algorithmic techniques, including a hybrid intra-cluster search strategy and an improved PQ residual calculation method, which optimize the search process at a higher level; 2) an ML-driven adaptive search module that provides adaptive, per-query tuning of search parameters, eliminating the inefficiencies of static configurations; and 3) highly-optimized SIMD kernels for ARM that maximize hardware utilization for the critical distance computation workloads. The experimental results demonstrate that KScaNN not only closes the performance gap but establishes a new standard, achieving up to a 1.63x speedup over the fastest x86-based solution. This work provides a definitive blueprint for achieving leadership-class performance for vector search on modern ARM architectures and underscores Oleg Senkevich, Siyang Xu, Tianyi Jiang, Alexander Radionov, Jan Tabaszewski, Dmitriy Malyshev, Daihao Xue, Licheng Yu, Weidi Zeng, Xin Yao 0008, Siyu Huang, Gleb Neshchetkin, Qiuling Pan, Yaoyao Fu |
ICDE | 13 |
| 2026 | Toward Inherently Robust VLMs Against Visual Perception AttacksabstractAutonomous vehicles rely on deep neural networks (DNNs) for traffic sign recognition, lane centering, and vehicle detection, yet these models are vulnerable to attacks that induce misclassification and threaten safety. Existing defenses (e.g., adversarial training) often fail to generalize and degrade clean accuracy. We introduce Vehicle Vision-Language Models (V2LMs), fine-tuned vision-language models specialized for autonomous vehicle perception, and show that they are inherently more robust to unseen attacks without adversarial training, maintaining substantially higher adversarial accuracy than conventional DNNs. We study two deployments: Solo (task-specific V2LMs) and Tandem (a single V2LM for all three tasks). Under attacks, DNNs drop 33-74%, whereas V2LMs decline by under 8% on average. Tandem achieves comparable robustness to Solo while being more memory-efficient. We also explore integrating V2LMs in parallel with existing perception stacks to enhance resilience. Our results suggest V2LMs are a promising path toward secure, robust AV perception. Pedram MohajerAnsari, Amir Salarpour, Michael Kühr, Siyu Huang, Mohammad Hamad, Habeeb Olufowobi, Sebastian Steinhorst, Mert D. Pesé |
IV | 4 |
| 2026 | New kiwifruit quality grading methods based on image multi-feature fusion and ResTNet model
Honggang Miao, Xiyuan Lu, Yuhan Lou, Jianxiang Ma, Siyu Huang, Zhengwei Ren, Yunliu Zeng |
Expert Syst. Appl. | 9 |
| 2025 | SoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow RemovalabstractRecent advancements in deep learning have yielded promising results for the image shadow removal task. However, most existing methods rely on binary pre-generated shadow masks. The binary nature of such masks could potentially lead to artifacts near the boundary between shadow and non-shadow areas. In view of this, inspired by the physical model of shadow formation, we introduce novel soft shadow masks specifically designed for shadow removal. To achieve such soft masks, we propose a SoftShadow framework by leveraging the prior knowledge of pretrained SAM and integrating physical constraints. Specifically, we jointly tune the SAM and the subsequent shadow removal network using penumbra formation constraint loss, mask reconstruction loss, and shadow removal loss. This framework enables accurate predictions of penumbra (partially shaded) and umbra (fully shaded) areas while simultaneously facilitating end-to-end shadow removal. Through extensive experiments on popular datasets, we found that our Soft-Shadow framework, which generates soft masks, can better restore boundary artifacts, achieve state-of-the-art performance, and demonstrate superior generalizability. Xinrui Wang 0004, Lanqing Guo, Siyu Huang, Bihan Wen |
CVPR | 4 |
| 2025 | HetKEN: Heterogeneous Kernel-Enhanced Representation Network
Siyu Huang |
ICIC (16) | 1 |
| 2025 | Latent Radiance Fields with 3D-aware 2D RepresentationsabstractLatent 3D reconstruction has shown great promise in empowering 3D semantic understanding and 3D generation by distilling 2D features into the 3D space. However, existing approaches struggle with the domain gap between 2D feature space and 3D representations, resulting in degraded rendering performance. To address this challenge, we propose a novel framework that integrates 3D awareness into the 2D latent space. The framework consists of three stages: (1) a correspondence-aware autoencoding method that enhances the 3D consistency of 2D latent representations, (2) a latent radiance field (LRF) that lifts these 3D-aware 2D representations into 3D space, and (3) a VAE-Radiance Field (VAE-RF) alignment strategy that improves image decoding from the rendered 2D representations. Extensive experiments demonstrate that our method outperforms the state-of-the-art latent 3D reconstruction approaches in terms of synthesis performance and cross-dataset generalizability across diverse indoor and outdoor scenes. To our knowledge, this is the first work showing the radiance field representations constructed from 2D latent representations can yield photorealistic 3D reconstruction performance. Chaoyi Zhou, Siyu Huang |
ICLR | 4 |
| 2025 | A Low-Rank Defense Method for Adversarial Attack on Diffusion ModelsabstractRecently, adversarial attacks for diffusion models as well as their fine-tuning process have been developed rapidly. To prevent the abuse of these attack algorithms from affecting the practical application of diffusion models, it is critical to develop corresponding defensive strategies. In this work, we propose an efficient defensive strategy, named Low-Rank Defense (LoRD), to defend the adversarial attack on Latent Diffusion Models (LDMs). LoRD introduces the merging idea and a balance parameter, combined with the low-rank adaptation (LoRA) modules, to detect and defend the adversarial samples. Based on LoRD, we build up a defense pipeline that applies the learned LoRD modules to help diffusion models defend against attack algorithms. Our method ensures that the LDM fine-tuned on both adversarial and clean samples can still generate high-quality images. To demonstrate the effectiveness of our approach, we conduct extensive experiments on facial and landscape images, and our method shows significantly better defense performance compared to the baseline methods. Jiaxuan Zhu, Siyu Huang |
ICME | 2 |
| 2025 | AutoAL: Automated Active Learning with Differentiable Query Strategy SearchabstractAs deep learning continues to evolve, the need for data efficiency becomes increasingly important. Considering labeling large datasets is both time-consuming and expensive, active learning (AL) provides a promising solution to this challenge by iteratively selecting the most informative subsets of examples to train deep neural networks, thereby reducing the labeling cost. However, the effectiveness of different AL algorithms can vary significantly across data scenarios, and determining which AL algorithm best fits a given task remains a challenging problem. This work presents the first differentiable AL strategy search method, named AutoAL, which is designed on top of existing AL sampling strategies. AutoAL consists of two neural nets, named SearchNet and FitNet, which are optimized concurrently under a differentiable bi-level optimization framework. For any given task, SearchNet and FitNet are iteratively co-optimized using the labeled data, learning how well a set of candidate AL algorithms perform on that task. With the optimal AL strategies identified, SearchNet selects a small subset from the unlabeled pool for querying their annotations, enabling efficient training of the task model. Experimental results demonstrate that AutoAL consistently achieves superior accuracy compared to all candidate AL algorithms and other selective AL approaches, showcasing its potential for adapting and integrating multiple existing AL methods across diverse tasks and domains. Xueying Zhan, Siyu Huang |
ICML | 3 |
| 2025 | A Hierarchical Compilation Method for Programmable Analog-to-Digital Converter ArraysabstractThis paper introduces a hierarchical compilation approach for multi-channel reconfigurable analog-to-digital converter (ADC) systems, motivated by the need for highly flexible and scalable solutions in programmable converter arrays (PCAs). Unlike existing methods that mainly rely on manual circuitlevel adjustments with high complexity, limited scalability, and low flexibility, this method provides a structured and scalable hierarchical mapping scheme. This facilitates flexible and efficient configuration by integrating software and hardware design, making it highly suitable for automation and future expansion. It paves the way for the automatic synthesis, optimization and future expansion of PCAs. Zhishuai Zhang, Siyu Huang, Yi Zhong 0002, Nan Sun 0001, Lu Jie 0001 |
ISCAS | 2 |
| 2025 | Few-Shot Generalized Category Discovery With Retrieval-Guided Decision Boundary Enhancement
Yunhan Ren, Feng Luo 0001, Siyu Huang |
ICMR | 3 |
| 2025 | Diffusion Network Inference for Cross-layer CascadesabstractA cascade over a network refers to the diffusion process where behavior changes occurring in one part of an interconnected population lead to a series of sequential changes throughout the entire population. In recent years, there has been a surge in interest and efforts to understand and model cascade mechanisms since they motivate many significant research topics across different disciplines. The propagation structure of cascades is governed by underlying diffusion networks that are often hidden. Inferring diffusion networks thus enables interventions in cascading process to maximize information propagation and provides insights into the Granger causality of interaction mechanisms among individuals. In this project, we propose a novel double network mixture model for inferring latent diffusion network in presence of strong cascade heterogeneity. The new model represents cascade pathways as a distributional mixture over diffusion networks that capture different cascading patterns at the population level. We develop a data-driven optimization method to infer diffusion networks using only visible temporal cascade records, avoiding the need to model complex and heterogeneous individual states. Both statistical and computational guarantees are established for the proposed method. We apply the proposed model to analyze research topic cascades in social sciences across U.S. universities and uncover the latent research topic diffusion network among top U.S. social science programs. Siyu Huang, Yubai Yuan, Abdul Basit Adeel |
NeurIPS | 1 |
| 2025 | Bézier Splatting for Fast and Differentiable Vector Graphics RenderingabstractDifferentiable vector graphics (VGs) are widely used in image vectorization and vector synthesis, while existing representations are costly to optimize and struggle to achieve high-quality rendering results for high-resolution images. This work introduces a new differentiable VG representation, dubbed Bézier Splatting, that enables fast yet high-fidelity VG rasterization. Bézier Splatting samples 2D Gaussians along Bézier curves, which naturally provide positional gradients at object boundaries. Thanks to the efficient splatting-based differentiable rasterizer, Bézier Splatting achieves 30× and 150× faster per forward and backward rasterization step for open curves compared to DiffVG. Additionally, we introduce an adaptive pruning and densification strategy that dynamically adjusts the spatial distribution of curves to escape local minima, further improving VG quality. Furthermore, our new VG representation supports conversion to standard XML-based SVG format, enhancing interoperability with existing VG tools and pipelines. Experimental results show that Bézier Splatting significantly outperforms existing methods with better visual fidelity and significant optimization speedup. Chaoyi Zhou, Nanxuan Zhao, Siyu Huang |
NeurIPS | 4 |
| 2025 | Graphormer-Based Bayesian Network Conditional Normalizing Flow for Multivariate Time Series Anomaly Detection in Communication NetworksabstractHigh-dimensional time series data are becoming more widespread in many domains, including large-scale wireless networks for communication. However, because of its high dimensionality, label scarcity, and complicated temporal connections, anomaly detection in such data is difficult. This work proposes a Bayesian network conditional normalizing flow model for multivariate time series anomaly detection, called Graphormer-based Bayesian Network Conditional Normalizing Flow (GBNCNF), based on a graph Transformer (Graphormer) to convert the spatial and temporal dependencies of high-dimensional time series into simple evaluable conditional densities. It models the causal links between numerous time series using a Bayesian network, and it obtains representations of the interdependencies between different time series by combining LSTM modules with Graphormer modules. These representations are introduced as conditional information into the normalizing flow for density estimation, and data corresponding to low density are judged as anomalies. Experiments are conducted on two real datasets and show that our method detects anomalies more accurately than baseline methods, accurately captures the correlations between sensors, and allows users to infer the root causes of detected anomalies. Zeyu Tan, Shiwen He, Hang Zhan, Yongming Huang 0001, Siyu Huang |
WCNC | 5 |
| 2025 | An Integrated Tradeoff Design of Reduced-Order Diagnostic Observer for Event-Triggered Fault DetectionabstractReduced-order diagnostic observer (RODO) is more desirable in the monitoring of electrical systems due to the advantages of low computation consumption and flexible structure, while its tradeoff design remains a longstanding unsettled issue in model-based fault diagnosis research. In this paper, an integrated design scheme of RODO is developed for event-triggered systems, making, for the first time, a direct tradeoff between false alarm rate (FAR) and fault detection rate (FDR). Firstly, by considering the effects of disturbance, fault and event-triggered transmission error on the residual, along with the principle of residual evaluation, the event-triggered FAR and FDR are defined and their relationship is established. Then, by introducing the so-called quasi-Luenberger constraints, the Luenberger equality challenges in the optimization of performance tradeoff are removed and an iterative algorithm is developed for the tradeoff design of full-order or high-order diagnostic observer. Based on this, the deviations between the quasi-Luenberger and Luenberger equalities are further eliminated with the assistance of another iterative algorithm to achieve the tradeoff of RODO. Finally, a DC microgrid simulation is adopted to demonstrate the effectiveness of the proposed design scheme as well as its advantages over other schemes. Aibing Qiu, Siyu Huang, Juping Gu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Programmable Analog-to-Digital Converter Array Supporting Architecture Restructuring and Mode ConcurrencyabstractThis work presents a new analog-to-digital converter (ADC) architecture named programmable converter array (PCA) for multi-standard signal acquisition. Unlike prior reconfigurable ADCs that are mainly configured at the circuit level, PCA is highly flexible at the architecture level and can process multiple input signals simultaneously for mode concurrency. The elemental units in the converter array are Conversion Blocks (CBs) based on successive approximation register (SAR) ADCs. Multiple CBs can interleave or run synergically through a bus system to form sophisticated architectures. Fabricated in 28nm CMOS, the prototype converter array can be configured as over 16 modes, with an SNDR range from 30dB to 82dB and an aggregate bandwidth from sub-MHz to 1000MHz. The prototype achieves a peak Schreier figure of merit (FoMs) of 176dB while maintaining FoMs over 165dB in most configurations, and occupies only 0.1mm2 of silicon area. Zhishuai Zhang, Mingtao Zhan, Zijie Gao, Siyu Huang, Yunsong Tao, Xiyu He, Chitian Yuan, Yi Zhong 0002, Nan Sun 0001, Lu Jie 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
Lanqing Guo, Yingqing He, Haoxin Chen, Menghan Xia, Xiaodong Cun, Yufei Wang 0006, Siyu Huang, Yong Zhang 0034, Xintao Wang 0002, Qifeng Chen 0001, Ying Shan, Bihan Wen |
ECCV (36) | 7 |
| 2024 | Fundus2Video: Cross-Modal Angiography Video Generation from Static Fundus Photography with Clinical Knowledge Guidance
Weiyi Zhang 0004, Siyu Huang, Jiancheng Yang, ZongYuan Ge, Yingfeng Zheng, Danli Shi, Mingguang He |
MICCAI (1) | 2 |
| 2024 | 3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion PriorsabstractNovel-view synthesis aims to generate novel views of a scene from multiple input
images or videos, and recent advancements like 3D Gaussian splatting (3DGS)
have achieved notable success in producing photorealistic renderings with efficient
pipelines. However, generating high-quality novel views under challenging settings,
such as sparse input views, remains difficult due to insufficient information in
under-sampled areas, often resulting in noticeable artifacts. This paper presents
3DGS-Enhancer, a novel pipeline for enhancing the representation quality of
3DGS representations. We leverage 2D video diffusion priors to address the
challenging 3D view consistency problem, reformulating it as achieving temporal
consistency within a video generation process. 3DGS-Enhancer restores view-
consistent latent features of rendered novel views and integrates them with the
input views through a spatial-temporal decoder. The enhanced views are then
used to fine-tune the initial 3DGS model, significantly improving its rendering
performance. Extensive experiments on large-scale datasets of unbounded scenes
demonstrate that 3DGS-Enhancer yields superior reconstruction performance and
high-fidelity rendering results compared to state-of-the-art methods. The project
webpage is https://xiliu8006.github.io/3DGS-Enhancer-project. Chaoyi Zhou, Siyu Huang |
NeurIPS | 3 |
| 2024 | HyperXRC: Hybrid In-Person + Remote Extended Reality Classroom - A Design StudyabstractThis paper investigates HyperXRC, a hybrid classroom design that accommodates both local and remote students. The instructor wears an extended reality (XR) headset that shows the local classroom and the local students, as well as remote students modeled with video sprites. The remote students are displayed either on virtual banners hanging off the classroom ceiling, or on virtual billboards placed in empty classroom seats. Thereby, the remote students are integrated into the field of view of the instructor, who remains aware of the remote students while teaching. A controlled user study with two experiments evaluated the HyperXRC design from the instructor and from the local students perspective. In the first experiment (N = 15) participants served as instructors to a hybrid classroom of 14 local and 15 remote students. Participants were more likely to detect hand-raising and head-on-desk remote student actions in the HyperXRC conditions (59%) than in a conventional videoconferencing condition (36%). This advantage did not come at the cost of decreasing the detection rate of local student actions. Furthermore, instructor participants preferred the HyperXRC to the videoconferencing approach. In the second experiment (N = 16) participants served as local students. The participants preferred the lecture when the instructor used videoconferencing to the one when the instructor used HyperXRC, wearing the XR headset. Siyu Huang, Voicu Popescu |
VR | 1 |
| 2024 | Toward Robust Image Denoising via Flow-Based Joint Image and Noise ModelabstractOne of the fundamental challenges in image restoration is denoising, where the objective is to estimate the clean image from its noisy measurements. Existing denoising approaches generally focus on exploiting effective natural image priors to remove the noise. However, the utilization and analysis of the noise model are often ignored, although the noise model can provide complementary information to the denoising algorithms. As a result, they are very sensitive to different noise distributions. To tackle this issue and hence towards a robust image denoiser in practice, in this paper, we propose a novel Flow-based joint Image and NOise model (FINO) that distinctly decouples the image and noise in the latent space and losslessly reconstructs them via a series of invertible transformations. We further present a variable swapping strategy to align structural information in images and a noise correlation matrix to constrain the noise based on spatially minimized correlation information. Experimental results demonstrate FINO’s capacity to remove both synthetic additive white Gaussian noise (AWGN) and real noise. Furthermore, the generalization of FINO to the removal of spatially variant noise and noise with inaccurate estimation surpasses that of the popular and state-of-the-art methods by large margins. Lanqing Guo, Siyu Huang, Haosen Liu 0001, Bihan Wen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Temporal Output Discrepancy for Loss Estimation-Based Active LearningabstractWhile deep learning succeeds in a wide range of tasks, it highly depends on the massive collection of annotated data which is expensive and time-consuming. To lower the cost of data annotation, active learning has been proposed to interactively query an oracle to annotate a small proportion of informative samples in an unlabeled dataset. Inspired by the fact that the samples with higher loss are usually more informative to the model than the samples with lower loss, in this article we present a novel deep active learning approach that queries the oracle for data annotation when the unlabeled sample is believed to incorporate high loss. The core of our approach is a measurement temporal output discrepancy (TOD) that estimates the sample loss by evaluating the discrepancy of outputs given by models at different optimization steps. Our theoretical investigation shows that TOD lower-bounds the accumulated sample loss thus it can be used to select informative unlabeled samples. On basis of TOD, we further develop an effective unlabeled data sampling strategy as well as an unsupervised learning criterion for active learning. Due to the simplicity of TOD, our methods are efficient, flexible, and task-agnostic. Extensive experimental results demonstrate that our approach achieves superior performances than the state-of-the-art active learning methods on image classification and semantic segmentation tasks. In addition, we show that TOD can be utilized to select the best model of potentially the highest testing accuracy from a pool of candidate models. Siyu Huang, Tianyang Wang 0004, Haoyi Xiong, Bihan Wen, Jun Huan, Dejing Dou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ShadowFormer: Global Context Helps Shadow RemovalabstractRecent deep learning methods have achieved promising results in image shadow removal. However, most of the existing approaches focus on working locally within shadow and non-shadow regions, resulting in severe artifacts around the shadow boundaries as well as inconsistent illumination between shadow and non-shadow regions. It is still challenging for the deep shadow removal model to exploit the global contextual correlation between shadow and non-shadow regions. In this work, we first propose a Retinex-based shadow model, from which we derive a novel transformer-based network, dubbed ShandowFormer, to exploit non-shadow regions to help shadow region restoration. A multi-scale channel attention framework is employed to hierarchically capture the global information. Based on that, we propose a Shadow-Interaction Module (SIM) with Shadow-Interaction Attention (SIA) in the bottleneck stage to effectively model the context correlation between shadow and non-shadow regions. We conduct extensive experiments on three popular public datasets, including ISTD, ISTD+, and SRD, to evaluate the proposed method. Our method achieves state-of-the-art performance by using up to 150X fewer model parameters. Lanqing Guo, Siyu Huang, Ding Liu 0001, Hao Cheng 0016, Bihan Wen |
AAAI | 2 |
| 2023 | ShadowDiffusion: When Degradation Prior Meets Diffusion Model for Shadow RemovalabstractRecent deep learning methods have achieved promising results in image shadow removal. However, their restored images still suffer from unsatisfactory boundary artifacts, due to the lack of degradation prior embedding and the deficiency in modeling capacity. Our work addresses these issues by proposing a unified diffusion framework that integrates both the image and degradation priors for highly effective shadow removal. In detail, we first propose a shadow degradation model, which inspires us to build a novel unrolling diffusion model, dubbed ShandowDiffusion. It remarkably improves the model's capacity in shadow removal via progressively refining the desired output with both degradation prior and diffusive generative prior, which by nature can serve as a new strong baseline for image restoration. Furthermore, ShadowDiffusion progressively refines the estimated shadow mask as an auxiliary task of the diffusion generator, which leads to more accurate and robust shadow-free image generation. We conduct extensive experiments on three popular public datasets, including ISTD, ISTD+, and SRD, to validate our method's effectiveness. Compared to the state-of-the-art methods, our model achieves a significant improvement in terms of PSNR, increasing from 31.69dB to 34. 73dB over SRD dataset.11https://github.com/GuoLanqing/ShadowDiffusion Lanqing Guo, Chong Wang 0011, Wenhan Yang, Siyu Huang, Yufei Wang 0006, Hanspeter Pfister, Bihan Wen |
CVPR | 4 |
| 2023 | QuantArt: Quantizing Image Style Transfer Towards High Visual FidelityabstractThe mechanism of existing style transfer algorithms is by minimizing a hybrid loss function to push the generated image toward high similarities in both content and style. However, this type of approach cannot guarantee visual fidelity, i.e., the generated artworks should be indistinguishable from real ones. In this paper, we devise a new style transfer framework called QunntArt for high visual-fidelity stylization. QuantArt pushes the latent representation of the generated artwork toward the centroids of the real artwork distribution with vector quantization. By fusing the quantized and continuous latent representations, QuantArt allows flexible control over the generated artworks in terms of content preservation, style similarity, and visual fidelity. Experiments on various style transfer settings show that our QuantArt framework achieves significantly higher visual fidelity compared with the existing style transfer methods. Siyu Huang, Jie An 0002, Donglai Wei 0001, Jiebo Luo 0001, Hanspeter Pfister |
CVPR | 1 |
| 2023 | ContRE: A Complementary Measure for Robustness Evaluation of Deep Networks via Contrastive ExamplesabstractTraining images with data transformations, e.g., crops, shifts, rotations and color distortions, have been suggested as contrastive examples to evaluate the robustness of deep neural networks against data noises [1]. In this work, we propose a practical framework ContRE (which is the meaning of “against” in French) that uses Contrastive examples for DNN Robustness Estimation. Specifically, ContRE follows the assumption in [2], [3] that robust DNN models with good generalization performance are capable of extracting a consistent set of features and making consistent predictions from the same image under varying data transformations. Incorporating with a set of randomized strategies for well-designed data transformations over the training set, ContREadopts classification errors and Fisher ratios on the generated contrastive examples to assess and analyze the robustness of DNN models, which correlates to the models’ generalization performance. To show the effectiveness and efficiency of ContRE, extensive experiments have been done using various DNN models, e.g., ResNet, VGGNet, DenseNet, EfficientNet, etc., on three open source benchmark datasets, i.e., CIFAR-10, CIFAR-100, and ImageNet, with thorough ablation studies and applicability analyses. Our experiment results confirm that ❨1❩ behaviors of deep models on contrastive examples are strongly correlated to what on the testing set, and ❨2❩ the robustness that ContRE calculates is a robust measure of generalization performance complementing to the testing set in various settings. Codes is to be publicly available. Xuhong Li 0002, Xuanyu Wu, Linghe Kong, Xiao Zhang 0001, Siyu Huang, Dejing Dou, Haoyi Xiong |
ICDM | 5 |
| 2023 | Cross-model consensus of explanations and beyond for image classification models: an empirical study
Xuhong Li 0002, Haoyi Xiong, Siyu Huang, Shilei Ji, Dejing Dou |
Mach. Learn. | 3 |
| 2023 | Domain-Scalable Unpaired Image Translation via Latent Space AnchoringabstractUnpaired image-to-image translation (UNIT) aims to map images between two visual domains without paired training data. However, given a UNIT model trained on certain domains, it is difficult for current methods to incorporate new domains because they often need to train the full model on both existing and new domains. To address this problem, we propose a new domain-scalable UNIT method, termed as latent space anchoring, which can be efficiently extended to new visual domains and does not need to fine-tune encoders and decoders of existing domains. Our method anchors images of different domains to the same latent space of frozen GANs by learning lightweight encoder and regressor models to reconstruct single-domain images. In the inference phase, the learned encoders and decoders of different domains can be arbitrarily combined to translate images between any two domains without fine-tuning. Experiments on various datasets show that the proposed method achieves superior performance on both standard and domain-scalable UNIT tasks in comparison with the state-of-the-art methods. Siyu Huang, Jie An 0002, Donglai Wei 0001, Zudi Lin, Jiebo Luo 0001, Hanspeter Pfister |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | 3D Domain Adaptive Instance Segmentation via Cyclic Segmentation GANsabstract3D instance segmentation for unlabeled imaging modalities is a challenging but essential task as collecting expert annotation can be expensive and time-consuming. Existing works segment a new modality by either deploying pre-trained models optimized on diverse training data or sequentially conducting image translation and segmentation with two relatively independent networks. In this work, we propose a novel Cyclic Segmentation Generative Adversarial Network (CySGAN) that conducts image translation and instance segmentation simultaneously using a unified network with weight sharing. Since the image translation layer can be removed at inference time, our proposed model does not introduce additional computational cost upon a standard segmentation model. For optimizing CySGAN, besides the CycleGAN losses for image translation and supervised losses for the annotated source domain, we also utilize self-supervised and segmentation-based adversarial objectives to enhance the model performance by leveraging unlabeled target domain images. We benchmark our approach on the task of 3D neuronal nuclei segmentation with annotated electron microscopy (EM) images and unlabeled expansion microscopy (ExM) data. The proposed CySGAN outperforms pre-trained generalist models, feature-level domain adaptation models, and the baselines that conduct image translation and segmentation sequentially. Our implementation and the newly collected, densely annotated ExM zebrafish brain nuclei dataset, named NucExM, are publicly available at https://connectomics-bazaar.github.io/proj/CySGAN/index.html. Leander Lauenburg, Zudi Lin, Ruihan Zhang 0003, Márcia dos Santos, Siyu Huang, Ignacio Arganda-Carreras, Edward S. Boyden, Hanspeter Pfister, Donglai Wei 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Boosting Active Learning via Improving Test PerformanceabstractCentral to active learning (AL) is what data should be selected for annotation. Existing works attempt to select highly uncertain or informative data for annotation. Nevertheless, it remains unclear how selected data impacts the test performance of the task model used in AL. In this work, we explore such an impact by theoretically proving that selecting unlabeled data of higher gradient norm leads to a lower upper-bound of test loss, resulting in a better test performance. However, due to the lack of label information, directly computing gradient norm for unlabeled data is infeasible. To address this challenge, we propose two schemes, namely expected-gradnorm and entropy-gradnorm. The former computes the gradient norm by constructing an expected empirical loss while the latter constructs an unsupervised loss with entropy. Furthermore, we integrate the two schemes in a universal AL framework. We evaluate our method on classical image classification and semantic segmentation tasks. To demonstrate its competency in domain applications and its robustness to noise, we also validate our method on a cellular imaging analysis task, namely cryo-Electron Tomography subtomogram classification. Results demonstrate that our method achieves superior performance against the state of the art. We refer readers to https://arxiv.org/pdf/2112.05683.pdf for the full version of this paper which includes the appendix and source code link. Tianyang Wang 0004, Xingjian Li 0002, Pengkun Yang, Guosheng Hu, Siyu Huang, Cheng-Zhong Xu 0001, Min Xu 0009 |
AAAI | 6 |
| 2022 | BM-NAS: Bilevel Multimodal Neural Architecture SearchabstractDeep neural networks (DNNs) have shown superior performances on various multimodal learning problems. However, it often requires huge efforts to adapt DNNs to individual multimodal tasks by manually engineering unimodal features and designing multimodal feature fusion strategies. This paper proposes Bilevel Multimodal Neural Architecture Search (BM-NAS) framework, which makes the architecture of multimodal fusion models fully searchable via a bilevel searching scheme. At the upper level, BM-NAS selects the inter/intra-modal feature pairs from the pretrained unimodal backbones. At the lower level, BM-NAS learns the fusion strategy for each feature pair, which is a combination of predefined primitive operations. The primitive operations are elaborately designed and they can be flexibly combined to accommodate various effective feature fusion modules such as multi-head attention (Transformer) and Attention on Attention (AoA). Experimental results on three multimodal tasks demonstrate the effectiveness and efficiency of the proposed BM-NAS framework. BM-NAS achieves competitive performances with much less search time and fewer model parameters in comparison with the existing generalized multimodal NAS methods. Our code is available at https://github.com/Somedaywilldo/BM-NAS. Yihang Yin, Siyu Huang |
AAAI | 2 |
| 2022 | AutoGCL: Automated Graph Contrastive Learning via Learnable View GeneratorsabstractContrastive learning has been widely applied to graph representation learning, where the view generators play a vital role in generating effective contrastive samples. Most of the existing contrastive learning methods employ pre-defined view generation methods, e.g., node drop or edge perturbation, which usually cannot adapt to input data or preserve the original semantic structures well. To address this issue, we propose a novel framework named Automated Graph Contrastive Learning (AutoGCL) in this paper. Specifically, AutoGCL employs a set of learnable graph view generators orchestrated by an auto augmentation strategy, where every graph view generator learns a probability distribution of graphs conditioned by the input. While the graph view generators in AutoGCL preserve the most representative structures of the original graph in generation of every contrastive sample, the auto augmentation learns policies to introduce adequate augmentation variances in the whole contrastive learning procedure. Furthermore, AutoGCL adopts a joint training strategy to train the learnable view generators, the graph encoder, and the classifier in an end-to-end manner, resulting in topological heterogeneity yet semantic similarity in the generation of contrastive samples. Extensive experiments on semi-supervised learning, unsupervised learning, and transfer learning demonstrate the superiority of our AutoGCL framework over the state-of-the-arts in graph contrastive learning. In addition, the visualization results further confirm that the learnable view generators can deliver more compact and semantically meaningful contrastive samples compared against the existing view generation methods. Our code is available at https://github.com/Somedaywilldo/AutoGCL. Yihang Yin, Qingzhong Wang, Siyu Huang, Haoyi Xiong |
AAAI | 3 |
| 2022 | Combining Dynamic Mode Decomposition and Difference-in-Differences in an Analysis of At-Risk YouthabstractWe analyze the impact of the Los Angeles Mayor’s Office of Gang Reduction Youth Development (GRYD) prevention programming using quasi-experimental data. We model the evolution of questionnaire scores and apply Dynamic Mode Decomposition (DMD) to describe the asymptotic behavior of the dynamical system. The analysis indicates that risk decreased for youth who enrolled in GRYD prevention services, while it increased or remained the same for those who were in the control group. We augment these observations using a difference-in-differences (DID) model, showing that the decrease in risk can be attributed to enrolment in prevention services. We draw a connection between DMD and DID using both mathematical analysis and empirical evidence from the questionnaire data. Combining DMD and DID with factor analysis, we investigate the effectiveness of prevention services with respect to different attitudinal domains. We conclude that gang prevention is most effective in impacting attitudes towards negative peer obedience and least effective in impacting attitudes towards violence for self defense. Our analytical approach can be extended to other types of repeated questionnaires. Marc Andrew Choi, Siyu Huang, Hengyuan Qi, Marco Scialanga, Emerson McMullen, Axel Sanchez Moreno, Yifei Lou, Andrea L. Bertozzi, P. Jeffrey Brantingham |
IEEE Big Data | 2 |
| 2022 | Parameter-Free Style Projection for Arbitrary Image Style TransferabstractArbitrary image style transfer is a challenging task which aims to stylize a content image conditioned on arbitrary style images. In this task the feature-level content-style transformation plays a vital role for proper fusion of features. Existing feature transformation algorithms often suffer from loss of content or style details, non-natural stroke patterns, and unstable training. To mitigate these issues, this paper proposes a new feature-level style transformation technique, named Style Projection, for parameter-free, fast, and effective content-style transformation. This paper further presents a real-time feed-forward model to leverage Style Projection for arbitrary image style transfer, which includes a regularization term for matching the semantics between input contents and stylized outputs. Extensive qualitative analysis, quantitative evaluation, and user study have demonstrated the effectiveness and efficiency of the proposed methods. Siyu Huang, Haoyi Xiong, Tianyang Wang 0004, Bihan Wen, Qingzhong Wang, Jun Huan, Dejing Dou |
ICASSP | 1 |
| 2022 | MUSCLE: Multi-task Self-supervised Continual Learning to Pre-train Deep Models for X-Ray Images of Multiple Body Parts
Weibin Liao, Haoyi Xiong, Qingzhong Wang, Yan Mo, Xuhong Li 0002, Yi Liu 0040, Siyu Huang, Dejing Dou |
MICCAI (8) | 8 |
| 2022 | A Unified Framework for Bidirectional Prototype Learning From Contaminated Faces Across Heterogeneous DomainsabstractExisting heterogeneous face synthesis (HFS) methods focus on performing accurate image-to-image translation across domains, while they cannot effectively remove the nuisance facial variations such as poses, expressions or occlusions. To address such challenges, this paper studies a new practical heterogeneous prototype learning (HPL) problem. To be specific, given a face image contaminated by facial variations from a source domain, HPL aims to reconstruct the variation-free prototype in a specified target domain. To tackle HPL, we propose a unified and end-to-end framework named bidirectional heterogeneous prototype learning (BHPL). As a bidirectional learning framework, BHPL is able to simultaneously reconstruct the heterogeneous prototypes acrosssource-to-targetas well astarget-to-sourcedomains. Furthermore, BHPL is capable of learning the identity prototype features for the contaminated face images from both source and target domains in order to perform robust heterogeneous face recognition. BHPL consists of an encoder-decoder structural generator and two dual-task discriminators, which play an adversarial game such that the generator learns the identity prototype feature and generates the cross-domain identity-preserved prototype for each input face image from both domains, and the discriminators accurately predict face identity and distinguish real versus fake prototypes. Empirically studies on multiple heterogeneous face datasets containing facial variations demonstrate the effectiveness of BHPL. Binghui Wang, Siyu Huang, Yiu-Ming Cheung, Bihan Wen |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | ArtFlow: Unbiased Image Style Transfer via Reversible Neural FlowsabstractUniversal style transfer retains styles from reference images in content images. While existing methods have achieved state-of-the-art style transfer performance, they are not aware of the content leak phenomenon that the image content may corrupt after several rounds of stylization process. In this paper, we propose ArtFlow to prevent content leak during universal style transfer. ArtFlow consists of reversible neural flows and an unbiased feature transfer module. It supports both forward and backward inferences and operates in a projection-transfer-reversion scheme. The forward inference projects input images into deep features, while the backward inference remaps deep features back to input images in a lossless and unbiased way. Extensive experiments demonstrate that ArtFlow achieves comparable performance to state-of-the-art style transfer methods while avoiding content leak. Jie An 0002, Siyu Huang, Yibing Song, Dejing Dou, Wei Liu 0005, Jiebo Luo 0001 |
CVPR | 2 |
| 2021 | Semi-Supervised Active Learning with Temporal Output DiscrepancyabstractWhile deep learning succeeds in a wide range of tasks, it highly depends on the massive collection of annotated data which is expensive and time-consuming. To lower the cost of data annotation, active learning has been proposed to interactively query an oracle to annotate a small proportion of informative samples in an unlabeled dataset. Inspired by the fact that the samples with higher loss are usually more informative to the model than the samples with lower loss, in this paper we present a novel deep active learning approach that queries the oracle for data annotation when the unlabeled sample is believed to incorporate high loss. The core of our approach is a measurement Temporal Output Discrepancy (TOD) that estimates the sample loss by evaluating the discrepancy of outputs given by models at different optimization steps. Our theoretical investigation shows that TOD lower-bounds the accumulated sample loss thus it can be used to select informative unlabeled samples. On basis of TOD, we further develop an effective unlabeled data sampling strategy as well as an unsupervised learning criterion that enhances model performance by incorporating the unlabeled data. Due to the simplicity of TOD, our active learning approach is efficient, flexible, and task-agnostic. Extensive experimental results demonstrate that our approach achieves superior performances than the state-of-the-art active learning methods on image classification and semantic segmentation tasks. Siyu Huang, Tianyang Wang 0004, Haoyi Xiong, Jun Huan, Dejing Dou |
ICCV | 1 |
| 2021 | Contour Primitive of Interest Extraction Network Based on One-Shot Learning for Object-Agnostic Vision MeasurementabstractImage contour based vision measurement is widely applied in robot manipulation and industrial automation. It is appealing to realize object-agnostic vision system, which can be conveniently reused for various types of objects. We propose the contour primitive of interest extraction network (CPieNet) based on the one-shot learning framework. First, CPieNet is featured by that its contour primitive of interest (CPI) output, a designated regular contour part lying on a specified object, provides the essential geometric information for vision measurement. Second, CPieNet has the one-shot learning ability, utilizing a support sample to assist the perception of the novel object. To realize lower-cost training, we generate support-query sample pairs from unpaired online public images, which cover a wide range of object categories. To obtain single-pixel wide contour for precise measurement, the Gabor-filters based non-maximum suppression is designed to thin the raw contour. For the novel CPI extraction task, we built the Object Contour Primitives dataset using online public images, and the Robotic Object Contour Measurement dataset using a camera mounted on a robot. The effectiveness of the proposed methods is validated by a series of experiments. Fangbo Qin, Siyu Huang, De Xu |
ICRA | 3 |
| 2021 | ReLLIE: Deep Reinforcement Learning for Customized Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) is a pervasive yet challenging problem, since: 1) low-light measurements may vary due to different imaging conditions in practice; 2) images can be enlightened subjectively according to diverse preference by each individual. To tackle these two challenges, this paper presents a novel deep reinforcement learning based method, dubbed ReLLIE, for customized low-light enhancement. ReLLIE models LLIE as a markov decision process, i.e., estimating the pixel-wise image-specific curves sequentially and recurrently. Given the reward computed from a set of carefully crafted non-reference loss functions, a lightweight network is proposed to estimate the curves for enlightening of a low-light image input. As ReLLIE learns a policy instead of one-one image translation, it can handle various low-light measurements and provide customized enhanced outputs by flexibly applying the policy different times. Furthermore, ReLLIE can enhance real-world images with hybrid corruptions, i.e., noise, by using a plug-and-play denoiser easily. Extensive experiments on various benchmarks demonstrate the advantages of ReLLIE, comparing to the state-of-the-art methods. (Code is available: https://github.com/GuoLanqing/ReLLIE.) Rongkai Zhang 0001, Lanqing Guo, Siyu Huang, Bihan Wen |
ACM Multimedia | 3 |
| 2021 | Analyzing travel time belief reliability in road network under uncertain random environment
Yi Yang 0006, Siyu Huang, Meilin Wen |
Soft Comput. | 2 |
| 2020 | An Investigation of Containment Measures Against the COVID-19 Pandemic in Mainland ChinaabstractAs the recent COVID-19 outbreak rapidly expands all over the world, various containment measures have been carried out to fight against the COVID-19 pandemic. In Mainland China, the containment measures consist of three types, i.e., Wuhan travel ban, intra-city quarantine and isolation, and intercity travel restriction. In order to carry out the measures, local economy and information acquisition play an important role. In this paper, we investigate the correlation of local economy and the information acquisition on the execution of containment measures to fight against the COVID-19 pandemic in Mainland China. First, we use a parsimonious model, i.e., SIR-X model to estimate the parameters, which represent the execution of intra-city quarantine and isolation in major cities of Mainland China. In order to understand the execution of intra-city quarantine and isolation, we analyze the correlation between the representative parameters including local economy, mobility, and information acquisition. To this end, we collect the data of Gross Domestic Product (GDP), the inflows from Wuhan and outflows, and the COVID-19 related search frequency from a widely-used Web mapping service, i.e., Baidu Maps, and Web search engine, i.e., Baidu Search Engine, in Mainland China. Based on the analysis, we confirm the strong correlation between the local economy and the execution of information acquisition in major cities of Mainland China. We further evidence that, although the cities with high GDP per capita attract more inflows from Wuhan, people are more likely to conduct the quarantine measure and to reduce travelling to other cities. Finally, the correlation analysis using search data shows that well-informed individuals are likely to carry out containment measures. Ji Liu 0003, Xiakai Wang, Haoyi Xiong, Jizhou Huang, Siyu Huang, Haozhe An, Dejing Dou, Haifeng Wang 0001 |
IEEE BigData | 5 |
| 2020 | Neighbours Matter: Image Captioning with Similar Images
Qingzhong Wang, Jiuniu Wang, Antoni B. Chan, Siyu Huang, Haoyi Xiong, Xingjian Li 0002, Dejing Dou |
BMVC | 4 |
| 2020 | TP-LSD: Tri-Points Based Line Segment Detector
Siyu Huang, Fangbo Qin, Pengfei Xiong, Yijia He, Xiao Liu 0042 |
ECCV (27) | 1 |
| 2020 | Stacked Pooling for Boosting Scale Invariance of Crowd CountingabstractIn this work, we take insight into the dense crowd counting problem by exploring the phenomenon of cross-scale visual similarity caused by perspective distortions. It is a quite common phenomenon in crowd scenarios, suggesting the crowd counting model to enable a good performance of scale invariance. Existing deep crowd counting approaches mainly focus on the multi-scale techniques over convolutional layers to capture scale-adaptive features, resulting in high computing costs. In this paper, we propose simple but effective pooling variants, i.e., multi-kernel pooling and stacked pooling, to take place of the vanilla pooling layers in convolutional neural networks (CNNs) for boosting the scale invariance. Our proposed pooling modules do not introduce extra parameters and can be easily implemented in practice. Empirical studies on two benchmark crowd counting datasets show that the proposed pooling modules beat the vanilla pooling layer in most experimental cases. Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001 |
ICASSP | 1 |
| 2020 | Generating Person Images with Appearance-aware Pose StylizerabstractGeneration of high-quality person images is challenging, due to the sophisticated entanglements among image factors, e.g., appearance, pose, foreground, background, local details, global structures, etc. In this paper, we present a novel end-to-end framework to generate realistic person images based on given person poses and appearances. The core of our framework is a novel generator called Appearance-aware Pose Stylizer (APS) which generates human images by coupling the target pose with the conditioned person appearance progressively. The framework is highly flexible and controllable by effectively decoupling various complex person image factors in the encoding phase, followed by re-coupling them in the decoding phase. In addition, we present a new normalization method named adaptive patch normalization, which enables region-specific normalization and shows a good performance when adopted in person image generation model. Experiments on two benchmark datasets show that our method is capable of generating visually appealing and realistic-looking results using arbitrary image and pose inputs. Siyu Huang, Haoyi Xiong, Zhi-Qi Cheng, Qingzhong Wang, Xingran Zhou, Bihan Wen, Jun Huan, Dejing Dou |
IJCAI | 1 |
| 2020 | SBAT: Video Captioning with Sparse Boundary-Aware TransformerabstractIn this paper, we focus on the problem of applying the transformer structure to video captioning effectively. The vanilla transformer is proposed for uni-modal language generation task such as machine translation. However, video captioning is a multimodal learning problem, and the video features have much redundancy between different time steps. Based on these concerns, we propose a novel method called sparse boundary-aware transformer (SBAT) to reduce the redundancy in video representation. SBAT employs boundary-aware pooling operation for scores from multihead attention and selects diverse features from different scenarios. Also, SBAT includes a local correlation scheme to compensate for the local information loss brought by sparse operation. Based on SBAT, we further propose an aligned cross-modal encoding scheme to boost the multimodal interaction. Experimental results on two benchmark datasets show that SBAT outperforms the state-of-the-art methods under most of the metrics. Tao Jin 0004, Siyu Huang, Yingming Li, Zhongfei Zhang |
IJCAI | 2 |
| 2019 | Text Guided Person Image SynthesisabstractThis paper presents a novel method to manipulate the visual appearance (pose and attribute) of a person image according to natural language descriptions. Our method can be boiled down to two stages: 1) text guided pose generation and 2) visual appearance transferred image synthesis. In the first stage, our method infers a reasonable target human pose based on the text. In the second stage, our method synthesizes a realistic and appearance transferred person image according to the text in conjunction with the target pose. Our method extracts sufficient information from the text and establishes a mapping between the image space and the language space, making generating and editing images corresponding to the description possible. We conduct extensive experiments to reveal the effectiveness of our method, as well as using the VQA Perceptual Score as a metric for evaluating the method. It shows for the first time that we can automatically edit the person image from the natural language descriptions. Xingran Zhou, Siyu Huang, Bin Li 0038, Yingming Li, Zhongfei Zhang |
CVPR | 2 |
| 2019 | Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video CaptioningabstractTao Jin, Siyu Huang, Yingming Li, Zhongfei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Jin 0004, Siyu Huang, Yingming Li, Zhongfei Zhang |
EMNLP/IJCNLP (1) | 2 |
| 2019 | User-Ranking Video Summarization With Multi-Stage Spatio-Temporal RepresentationabstractVideo summarization is a challenging task, mainly due to the difficulties in learning complicated semantic structural relations between videos and summaries. In this paper, we present a novel supervised video summarization scheme based on threestage deep neural networks. The scheme takes a divide-andconquer strategy to resolve the complicated task of 3D video summarization into a set of easy and flexible computational subtasks, and then to sequentially perform 2D CNNs, 1D CNNs, and LSTM to address the subtasks in an hierarchical fashion. The hierarchical modeling of spatio-temporal structure leads to high performance and efficiency. In addition, we propose a simple but effective user-ranking method to cope with the labeling subjectivity problem of user-created video summarization, leading to the labeling quality refinement for robust supervised learning. Experimental results show that our approach outperforms the state-of-the-art video summarization methods on two benchmark datasets. Siyu Huang, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Junwei Han 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | TVT: Two-View Transformer Network for Video CaptioningabstractVideo captioning is a task of automatically generating the natural text description of a given video. There are two main challenges in video captioning under the context of an encoder-decoder framework: 1) How to model the sequential information; 2) How to combine the modalities including video and text. For challenge 1), the recurrent neural networks (RNNs) based methods are currently the most common approaches for learning temporal representations of videos, while they suffer from a high computational cost. For challenge 2), the features of different modalities are often roughly concatenated together without insightful discussion. In this paper, we introduce a novel video captioning framework, i.e., Two-View Transformer (TVT). TVT comprises of a backbone of Transformer network for sequential representation and two types of fusion blocks in decoder layers for combining different modalities effectively. Empirical study shows that our TVT model outperforms the state-of-the-art methods on the MSVD dataset and achieves a competitive performance on the MSR-VTT dataset under four common metrics. Yingming Li, Zhongfei Zhang, Siyu Huang |
ACML | 4 |
| 2018 | Attribute Reduction of Set-Valued Decision Information System
Siyu Huang |
IPMU (2) | 2 |
| 2018 | Learning to Transfer: Generalizable Attribute Learning with Multitask Neural Model SearchabstractAs attribute leaning brings mid-level semantic properties for objects, it can benefit many traditional learning problems in multimedia and computer vision communities. When facing the huge number of attributes, it is extremely challenging to automatically design a generalizable neural network for other attribute learning tasks. Even for a specific attribute domain, the exploration of the neural network architecture is always optimized by a combination of heuristics and grid search, from which there is a large space of possible choices to be searched. In this paper, Generalizable Attribute Learning Model (GALM) is proposed to automatically design the neural networks for generalizable attribute learning. The main novelty of GALM is that it fully exploits the Multi-Task Learning and Reinforcement Learning to speed up the search procedure. With the help of parameter sharing, GALM is able to transfer the pre-searched architecture to different attribute domains. In experiments, we comprehensively evaluate GALM on 251 attributes from three domains: animals, objects, and scenes. Extensive experimental results demonstrate that GALM significantly outperforms the state-of-the-art attribute learning approaches and previous neural architecture search methods on two generalizable attribute learning scenarios. Zhi-Qi Cheng, Xiao Wu 0001, Siyu Huang, Jun-Xiu Li, Alex Hauptmann 0001, Qiang Peng |
ACM Multimedia | 3 |
| 2018 | GNAS: A Greedy Neural Architecture Search Method for Multi-Attribute LearningabstractA key problem in deep multi-attribute learning is to effectively discover the inter-attribute correlation structures. Typically, the conventional deep multi-attribute learning approaches follow the pipeline of manually designing the network architectures based on task-specific expertise prior knowledge and careful network tunings, leading to the inflexibility for various complicated scenarios in practice. Motivated by addressing this problem, we propose an efficient greedy neural architecture search approach (GNAS) to automatically discover the optimal tree-like deep architecture for multi-attribute learning. In a greedy manner, GNAS divides the optimization of global architecture into the optimizations of individual connections step by step. By iteratively updating the local architectures, the global tree-like architecture gets converged where the bottom layers are shared across relevant attributes and the branches in top layers more encode attribute-specific features. Experiments on three benchmark multi-attribute datasets show the effectiveness and compactness of neural architectures derived by GNAS, and also demonstrate the efficiency of GNAS in searching neural architectures. Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2018 | Body Structure Aware Deep Crowd CountingabstractCrowd counting is a challenging task, mainly due to the severe occlusions among dense crowds. This paper aims to take a broader view to address crowd counting from the perspective of semantic modeling. In essence, crowd counting is a task of pedestrian semantic analysis involving three key factors: pedestrians, heads, and their context structure. The information of different body parts is an important cue to help us judge whether there exists a person at a certain position. Existing methods usually perform crowd counting from the perspective of directly modeling the visual properties of either the whole body or the heads only, without explicitly capturing the composite body-part semantic structure information that is crucial for crowd counting. In our approach, we first formulate the key factors of crowd counting as semantic scene models. Then, we convert the crowd counting problem into a multi-task learning problem, such that the semantic scene models are turned into different sub-tasks. Finally, the deep convolutional neural networks are used to learn the sub-tasks in a unified scheme. Our approach encodes the semantic nature of crowd counting and provides a novel solution in terms of pedestrian semantic analysis. In experiments, our approach outperforms the state-of-the-art methods on four benchmark crowd counting data sets. The semantic structure information is demonstrated to be an effective cue in scene of crowd counting. Siyu Huang, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Shenghua Gao, Rongrong Ji, Junwei Han 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Cyber-physical system enabled nearby traffic flow modelling for autonomous vehiclesabstractWe propose a nearby traffic flow modelling solution based on built-in Cyber-Physical System (CPS) sensors of autonomous vehicles. Our goal is to enhance the offline route planning and driving decision adjustment based on the first-hand traffic information, especially during poor Internet connection moments. Specifically, our model helps to select the optimal speed on a road, the optimal distance for timing to brake, and the safe distance from other vehicles to keep. Moreover, our model can also assist neighboring autonomous vehicles by communicating required information through Ad-Hoc network communications or through a centralized cloud. In detail, we first focus on the unique characteristic of traffic flow (such as traffic rule, avoid collision behaviours), and then build a comprehensive model to handle multiple scenarios. Technically, our model uses density functions of velocities, the differential equation of traffic flows, and the traffic viscosity with information collected from the traffic flow, the distances between vehicles, the amount and density of vehicle, the instant velocity, the speed limit, and the momentum to analysis the the driving scene. We evaluate our model with real traffic data collected by in-vehicle CPS sensors to the proposed nearby traffic flow model. Results show that our work can accurately conduct offline estimation on nearby traffic signal influence, and reveal the correlations among velocity, density and (spatial and temporal) location to adjust route during runtime. Zhengyu Yang 0001, Siyu Huang, Xianzhi Du, Janki Bhimani, Ningfang Mi |
IPCCC | 3 |
| 2016 | Deep Learning Driven Visual Path Prediction From a Single ImageabstractCapabilities of inference and prediction are the significant components of visual systems. Visual path prediction is an important and challenging task among them, with the goal to infer the future path of a visual object in a static scene. This task is complicated as it needs high-level semantic understandings of both the scenes and underlying motion patterns in video sequences. In practice, cluttered situations have also raised higher demands on the effectiveness and robustness of models. Motivated by these observations, we propose a deep learning framework, which simultaneously performs deep feature learning for visual representation in conjunction with spatiotemporal context modeling. After that, a unified path-planning scheme is proposed to make accurate path prediction based on the analytic results returned by the deep context models. The highly effective visual representation and deep context models ensure that our framework makes a deep semantic understanding of the scenes and motion patterns, consequently improving the performance on visual path prediction task. In experiments, we extensively evaluate the model's performance by constructing two large benchmark datasets from the adaptation of video tracking datasets. The qualitative and quantitative experimental results show that our approach outperforms the state-of-the-art approaches and owns a better generalization capability. Siyu Huang, Xi Li 0001, Zhongfei Zhang, Zhouzhou He, Fei Wu 0001, Wei Liu 0005, Jinhui Tang 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 1 |