EDBT 2026 Demo / reviewers in the wild / expert
Weizhan Zhang
dblp:17/6858
· DBLP profile ↗
68ranked-venue papers
10as first author
37since 2021 · last 2026
0000-0003-0330-5435ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 13 since 2021Computer networks · 14 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Recovering Coherent Affective Patterns: Addressing Modality Missing in Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) seeks to decode human emotions by integrating heterogeneous modalities. However, real-world scenarios often involve missing or misaligned data due to sensor failures or transmission errors, leading to disrupted temporal dynamics and degraded cross-modal correlations. To address these challenges, we propose RECAP (REcovery of Coherent Affective Patterns), a robust two-stage framework to restore temporal and structural emotional integrity under modality incompleteness. The first stage employs a causality-aware adversarial generator for multi-granularity temporal reconstruction, complemented by a contrastive mutual information factorization module that disentangles shared and modality-specific semantics. The second stage introduces a mutual information-guided attention fusion mechanism with a ranking-based objective, enabling adaptive integration of complementary signals for refined prediction. Extensive experiments on MOSI, MOSEI, and SIMS under various missing-modality conditions demonstrate that RECAP consistently outperforms state-of-the-art methods. Notably, it improves ACC-7 on MOSI by 2.71 percentage points and F1 on SIMS by 6.38 percentage points. These results verify the performance of RECAP in terms of capturing fine-grained emotional cues and robustness. Huiting Huang, Tieliang Gong, Kai He 0001, Wen Wen 0013, Weizhan Zhang, Mengling Feng |
AAAI | 5 |
| 2026 | InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment TransferabstractRecently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that fine-tune CLIP for segmentation on limited seen categories often lead to overfitting and degrade the pretrained vision-language alignment. To stabilize modality alignment during fine-tuning, we propose InfoCLIP, which leverages an information-theoretic perspective to transfer alignment knowledge from pretrained CLIP to the segmentation task. Specifically, this transfer is guided by two novel objectives grounded in mutual information. First, we compress the pixel-text modality alignment from pretrained CLIP to reduce noise arising from its coarse-grained local semantic representations learned under image-text supervision. Second, we maximize the mutual information between the alignment knowledge of pretrained CLIP and the fine-tuned model to transfer compact local semantic relations suited for the segmentation task. Extensive evaluations across various benchmarks validate the effectiveness of InfoCLIP in enhancing CLIP fine-tuning for open-vocabulary semantic segmentation, demonstrating its adaptability and superiority in asymmetric transfer. Muyao Yuan, Yuanhong Zhang, Weizhan Zhang, Jiangyong Ying, Yudeng Xin |
AAAI | 3 |
| 2026 | Unifying Granularity and Reliability: A Robust and Efficient Framework for Text-based Person RetrievalabstractText-based person retrieval (TPR) has become a crucial task in cross-modal retrieval due to its broad application in fields such as public safety and criminal investigation. Existing TPR methods typically rely on fully fine-tuning large-scale pretrained vision-language models like CLIP, which incurs high computational costs and tends to exhibit poor generalization in unseen domains due to overfitting. Fortunately, Parameter-Efficient Transfer Learning (PETL) has emerged as a lightweight alternative. However, applying PETL to TPR remains challenging, as its limited adaptation capacity struggles to capture intricate identity cues and becomes highly susceptible to gradient interference from unreliable image-text pairs. To address these challenges, we present a PETL-based framework named UniGR that unifies granularity and reliability for robust and efficient TPR. Specifically, we design a multi-granularity relational adapter (MRA) to capture both coarse-grained global and fine-grained local relational features among tokens, equipping the generic backbone with the task-specific, precise understanding needed for TPR. To combat the noise sensitivity of PETL, a reliability-aware reweighting strategy (RRS) is introduced to adaptively down-weight unreliable samples during training. Furthermore, we propose a parameter-free cross-modal cyclic verification (CMCV) module to mitigate ambiguities in cross-modal matching computations and refine retrieval ranking further. Experiments on benchmarks corroborate the superiority of UniGR among parameter-efficient methods. Remarkably, with only 4.5% of trainable parameters, UniGR outperforms most fully fine-tuned methods while maintaining strong generalization. Jingchen Hao, Zhen Peng 0005, Yuting Zhang 0007, Zhongjiang He, Weizhan Zhang, Hao Sun 0038 |
SIGIR | 6 |
| 2026 | Nyström-aware approximations for matrix-based Rényi's entropy
Tieliang Gong, Wen Wen 0013, Yuxin Dong 0003, Zeyu Gao 0001, Weizhan Zhang |
Neural Networks | 5 |
| 2026 | VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection
Caixia Yan, Muyan Jiao, Nuohan Xue, Weizhan Zhang, Jiahao Wang 0004, Xiaojun Chang, Feng Tian 0002 |
Neural Networks | 4 |
| 2026 | A Concern-Decoupled Architecture for Scenario-Optimized Congestion Control in Large-Scale Live Video CDNsabstractOptimizing quality of experience (QoE) for live video streaming (LVS) remains a long-standing challenge for content delivery network (CDN) providers. Today’s CDNs predominantly employ static congestion control (CC) configurations, yet the heterogeneity of LVS scenarios undermines the efficacy of a uniform CC solution applicable to all users, as evidenced by our production measurements. While learning-based CC approaches show promise, they often suffer from limited generalizability, non-transparent design, and high computational overhead, struggling in large-scale CDNs. In this paper, we propose BIFROST, a new CC architecture grounded in a concerndecoupling paradigm. BIFROST separates fixed control logic from scenario-specific parameter optimization, enabling CDN systems to adapt dynamically to diverse LVS scenarios. To materialize it at CDN scale, BIFROST introduces a bilateral collaboration mechanism that identifies the LVS scenarios of each session by extracting client-side characteristics from viewing requests. It further employs offline model-based optimization to derive more effective control parameters for each scenario. We have deployed BIFROST on Alibaba Cloud’s production CDN for nearly a year, serving a commercial LVS application. BIFROST has markedly improved QoE metrics by 2.7% to 32.1%, reinforcing the competitive edge in the CDN market. We also share our experiences and lessons learned from its large-scale deployment. Danfu Yuan, Yubing Qiu, Weizhan Zhang, Haipeng Du |
IEEE Trans. Netw. | 3 |
| 2026 | CollabVisAdapt: Spatio-Temporal Context-Aware Adaptation of Shared Object Visualization for MR TelecollaborationabstractMixed Reality (MR) telecollaboration aims to enable users to share local objects as real-time synchronized virtual replicas to remote partners and collaborate on physical tasks as if they were co-located. However, in everyday scenarios with mobile and easy-to-setup MR devices, visualizing shared objects in a single modality, ranging from 2D images to 3D reconstruction, struggles to simultaneously optimize all the aspects of Spatiality, Fidelity, and Real-time performance. To overcome this issue, existing methods explore integrating multiple visualization modalities to leverage their respective advantages in subsets of the three aspects. However, they focus on fixed modality combinations without considering user-centered task contexts and workflow, where users may prioritize different aspects of the visualization across task phases. Moreover, they lack support for switching or require manual switching across modalities, which could become disruptive and tiring. In this paper, we propose adapting object visualization based on spatiotemporal contexts in telecollaboration. Specifically, we first couple task type with the user's relative viewing distance as the spatial context, and examine its impact on users' prioritized visualization aspects, and the corresponding switching thresholds. With differing generation speeds of modalities, we then explore temporal switching schemes when the preferred modality is not immediately available. With the obtained design choices, we implement CollabVisAdapt, a proof-of-concept prototype that supports automatic adaptation of object visualization based on spatiotemporal contexts in MR telecollaboration. A user study in remote maintenance verifies the effectiveness of the proposed workflow with adaptive visualization and the usability of the system. Xuanyu Wang 0001, Weizhan Zhang, Shuaichen Guo, Caixia Yan, Shuming Yang, Haipeng Du, Wangdu Chen, Qi Wang 0180 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | SpotActor: Training-Free Layout-Controlled Consistent Image GenerationabstractText-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of each subject across images. For these issues, we pioneer a novel task, Layout-to-Consistent-Image (L2CI) generation, which produces consistent and compositional images in accordance with the given layout conditions and text prompts. To accomplish this challenging task, we present a new formalization of dual energy guidance with optimization in a dual semantic-latent space and thus propose a training-free pipeline, SpotActor, which features a layout-conditioned optimizing stage and a consistent sampling stage. In the optimizing stage, we innovate a nuanced layout energy function to mimic the attention activations with a sigmoid-like objective. While in the sampling stage, we design Regional Interconnection Self-Attention (RISA) and Semantic Fusion Cross-Attention (SFCA) mechanisms that allow mutual interactions across images. To evaluate the performance, we present ActorBench, a specified benchmark with hundreds of reasonable prompt-box pairs stemming from object detection datasets. Comprehensive experiments are conducted to demonstrate the effectiveness of our method. The results prove that SpotActor fulfills the expectations of this task and showcases the potential for practical applications with superior layout alignment, subject consistency, prompt conformity and background diversity. Jiahao Wang 0004, Caixia Yan, Weizhan Zhang, Haonan Lin, Mengmeng Wang 0005, Guang Dai, Tieliang Gong, Hao Sun 0015, Jingdong Wang 0001 |
AAAI | 3 |
| 2025 | Efficient Real-Time On-Mobile Video Super-Resolution with Automatic Evolutionary Neural Architecture Search
Xuncheng Liu, Weizhan Zhang, Caixia Yan, Haipeng Du |
ICANN (2) | 2 |
| 2025 | DynamicID: Zero-Shot Multi-ID Image Personalization With Flexible Facial Editability
Xirui Hu, Jiahao Wang 0004, Weizhan Zhang, Benqi Wang, Haishun Nan |
ICCV | 4 |
| 2025 | InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic PerspectiveabstractThe Segment Anything Model (SAM), a vision foundation model, exhibits impressive zero-shot capabilities in general tasks but struggles in specialized domains. Parameter-efficient fine-tuning (PEFT) is a promising approach to unleash the potential of SAM in novel scenarios. However, existing PEFT methods for SAM neglect the domain-invariant relations encoded in the pre-trained model. To bridge this gap, we propose InfoSAM, an information-theoretic approach that enhances SAM fine-tuning by distilling and preserving its pre-trained segmentation knowledge. Specifically, we formulate the knowledge transfer process as two novel mutual information-based objectives: (i) to compress the domain-invariant relation extracted from pre-trained SAM, excluding pseudo-invariant information as possible, and (ii) to maximize mutual information between the relational knowledge learned by the teacher (pre-trained SAM) and the student (fine-tuned model). The proposed InfoSAM establishes a robust distillation framework for PEFT of SAM. Extensive experiments across diverse benchmarks validate InfoSAM’s effectiveness in improving SAM family’s performance on real-world tasks, demonstrating its adaptability and superiority in handling specialized scenarios. The code and models are available at https://muyaoyuan.github.io/InfoSAM_Page. Yuanhong Zhang, Muyao Yuan, Weizhan Zhang, Tieliang Gong, Wen Wen 0013, Jiangyong Ying |
ICML | 3 |
| 2025 | EchoShot: Multi-Shot Portrait Video GenerationabstractVideo diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge multiple shots with identity consistency and flexible content controllability. In this work, we propose EchoShot, a native and scalable multi-shot framework for portrait customization built upon a foundation video diffusion model. To start with, we propose shot-aware position embedding mechanisms within the video diffusion transformer architecture to model inter-shot variations and establish intricate correspondence between multi-shot visual content and their textual descriptions. This simple yet effective design enables direct training on multi-shot video data without introducing additional computational overhead. To facilitate model training within multi-shot scenarios, we construct PortraitGala, a large-scale and high-fidelity human-centric video dataset featuring cross-shot identity consistency and fine-grained captions such as facial attributes, outfits, and dynamic motions. To further enhance applicability, we extend EchoShot to perform reference image-based personalized multi-shot generation and long video synthesis with infinite shot counts. Extensive evaluations demonstrate that EchoShot achieves superior identity consistency as well as attribute-level controllability in multi-shot portrait video generation. Notably, the proposed framework demonstrates potential as a foundational paradigm for general multi-shot video modeling. Project page: https://johnneywang.github.io/EchoShot-webpage. Jiahao Wang 0004, Hualian Sheng, Sijia Cai, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, Jieping Ye |
NeurIPS | 4 |
| 2025 | Leveraging differentiable NAS and abstract genetic algorithms for optimizing on-mobile VSR performance
Xuncheng Liu, Weizhan Zhang, Tieliang Gong, Caixia Yan |
Mach. Learn. | 2 |
| 2025 | Understanding Operational CDN Live Streaming: A Measurement Study on Performance, Costs, and EnhancementsabstractThe escalating need for live video streaming has emerged as a significant catalyst for the business expansion of today’s content delivery networks (CDN). Selecting the right CDN live streaming architecture is fundamentally important in achieving the objective of enhancing users’ quality of experience (QoE) while reducing bandwidth costs. Regrettably, a limited number of studies have been conducted to systematically measure and compare the current typical solutions at production scale. Consequently, the performance and costs of different streaming architectures remain myths. This paper aims to address the existing research gap by undertaking a large-scale measurement study of three representative CDN live streaming architectures, defined by streaming protocol and overlay topology choices, currently running on Alibaba Cloud’s production video delivery network. By analyzing the results of over 500 million video plays over two months on a large live streaming platform hosted on Alibaba Cloud’s CDN, we reveal the impact of architectural compositions and operational factors on live streaming performance and bandwidth costs. In particular, our study reveals the trade-offs between QoE metrics and bandwidth costs for operational streaming architectures. Drawing upon the insights of this study, we further develop and deploy pragmatic strategies that yield remarkable real-world impact—our design saves over 17% bandwidth costs while maintaining the QoE. Danfu Yuan, Weizhan Zhang, Haiyu Huang 0005, Xuan Zeng 0002, Hongfei Yan, Yubing Qiu, Jinghui Zhong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | How Does Distribution Matching Help Domain Generalization: An Information-Theoretic AnalysisabstractDomain generalization aims to learn invariance across multiple source domains, thereby enhancing generalization against out-of-distribution data. While gradient or representation matching algorithms have achieved remarkable success in domain generalization, these methods generally lack generalization guarantees or depend on strong assumptions, leaving a gap in understanding the underlying mechanism of distribution matching. In this work, we formulate domain generalization from a novel probabilistic perspective, ensuring robustness while avoiding overly conservative solutions. Through comprehensive information-theoretic analysis, we provide key insights into the roles of gradient and representation matching in promoting generalization. Our results reveal the complementary relationship between these two components, indicating that existing works focusing solely on either gradient or representation alignment are insufficient to solve the domain generalization problem. In light of these theoretical findings, we introduce IDM to simultaneously align the inter-domain gradients and representations. Integrated with the proposed PDM method for complex distribution matching, IDM achieves superior performance over various baseline methods. Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Shuangyong Song, Weizhan Zhang, Chen Li 0011 |
IEEE Trans. Inf. Theory | 5 |
| 2025 | Lightweight Configuration Adaptation With Multi-Teacher Reinforcement Learning for Live Video AnalyticsabstractThe proliferation of video data and advancements in Deep Neural Networks (DNNs) have greatly boosted live video analytics, driven by the growing video capture capabilities of mobile devices. However, resource limitations necessitate the transmission of endpoint-collected videos to servers for inference. To meet real-time requirements and ensure accurate inference, it is essential to adjust video configurations at the endpoint. Traditional methods rely on deterministic strategies, posing difficulties in adapting to dynamic networks and video content. Meanwhile, emerging learning-based schemes suffer from trial-and-error exploration mechanisms, resulting in a concerning long-tail effect on upload latency. In this paper, we propose a novel lightweight and robust configuration adaptation policy (LCA), which fuses heuristic and RL-based agents using multi-teacher knowledge distillation (MKD) theory. Firstly, we propose a content-sensitive and bandwidth-adaptive RL agent and introduce a Lyapunov-based optimization agent for ensuring latency robustness. To leverage both agents' strengths, we design a feature-guided multi-teacher distillation network to transfer their advantages to the student. The experimental results across two vision tasks (pose estimation and semantic segmentation) demonstrate that LCA significantly reduces transmission latency compared to prior work (average reduction of 47.11%-89.55%, 95-percentile reduction of 27.63%-88.78%) and computational overhead while maintaining comparable inference accuracy. Yuanhong Zhang, Weizhan Zhang, Muyao Yuan, Caixia Yan, Tieliang Gong, Haipeng Du |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Data Quality-Aware Mixed-Precision Quantization via Hybrid Reinforcement LearningabstractMixed-precision quantization mostly predetermines the model bit-width settings before actual training due to the non-differential bit-width sampling process, obtaining suboptimal performance. Worse still, the conventional static quality-consistent training setting, i.e., all data is assumed to be of the same quality across training and inference, overlooks data quality changes in real-world applications which may lead to poor robustness of the quantized models. In this article, we propose a novel data quality-aware mixed-precision quantization framework, dubbed DQMQ, to dynamically adapt quantization bit-widths to different data qualities. The adaption is based on a bit-width decision policy that can be learned jointly with the quantization training. Concretely, DQMQ is modeled as a hybrid reinforcement learning (RL) task that combines model-based policy optimization with supervised quantization training. By relaxing the discrete bit-width sampling to a continuous probability distribution that is encoded with few learnable parameters, DQMQ is differentiable and can be directly optimized end-to-end with a hybrid optimization target considering both task performance and quantization benefits. Trained on mixed-quality image datasets, DQMQ can implicitly select the most proper bit-width for each layer when facing uneven input qualities. Extensive experiments on various benchmark datasets and networks demonstrate the superiority of DQMQ against existing fixed/mixed-precision quantization methods. Yingchun Wang 0002, Song Guo 0001, Jingcai Guo, Yuanhong Zhang, Weizhan Zhang, Jie Zhang 0076 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Real-and-Present: Investigating the Use of Life-Size 2D Video Avatars in HMD-Based AR TeleconferencingabstractAugmented Reality (AR) teleconferencing allows spatially distributed users to interact with each other in 3D through agents in their own physical environments. Existing methods leveraging volumetric capturing and reconstruction can provide a high-fidelity experience but are often too complex and expensive for everyday use. Other solutions target mobile and effortless-to-setup teleconferencing on AR Head Mounted Displays (HMD). They directly transplant the conventional video conferencing onto an AR-HMD platform or use avatars to represent remote participants. However, they can only support either a high fidelity or a high level of co-presence. Moreover, the limited Field of View (FoV) of HMDs could further degrade users' immersive experience. To achieve a balance between fidelity and co-presence, we explore using life-size 2D video-based avatars (video avatars for short) in AR teleconferencing. Specifically, with the potential effect of FoV on users' perception of proximity, we first conducted a pilot study to explore the local-user-centered optimal placement of video avatars in small-group AR conversations. With the placement results, we then implement a proof-of-concept prototype of video-avatar-based teleconferencing. We conduct user evaluations with our prototype to verify its effectiveness in balancing fidelity and co-presence. Following the indication in the pilot study, we further quantitatively explore the effect of FoV size on the video avatar's optimal placement through a user study involving more FoV conditions in a VR-simulated environment. We regress placement models to serve as references for computationally determining video avatar placements in such teleconferencing applications on various existing AR HMDs and future ones with bigger FoVs. Xuanyu Wang 0001, Weizhan Zhang, Christian Sandor, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | SAUI: Scale-Aware Unseen Imagineer for Zero-Shot Object DetectionabstractZero-shot object detection (ZSD) aims to localize and classify unseen objects without access to their training annotations. As a prevailing solution to ZSD, generation-based methods synthesize unseen visual features by taking seen features as reference and class semantic embeddings as guideline. Although previous works continuously improve the synthesis quality, they fail to consider the scale-varying nature of unseen objects. The generation process is preformed over a single scale of object features and thus lacks scale-diversity among synthesized features. In this paper, we reveal the scale-varying challenge in ZSD and propose a Scale-Aware Unseen Imagineer (SAUI) to lead the way of a novel scale-aware ZSD paradigm. To obtain multi-scale features of seen-class objects, we design a specialized coarse-to-fine extractor to capture features through multiple scale-views. To generate unseen features scale by scale, we innovate a Series-GAN synthesizer along with three scale-aware contrastive components to imagine separable, diverse and robust scale-wise unseen features. Extensive experiments on PASCAL VOC, COCO and DIOR datasets demonstrate SAUI's better performance in different scenarios, especially for scale-varying and small objects. Notably, SAUI achieves the new state-of-the art performance on COCO and DIOR. Jiahao Wang 0004, Caixia Yan, Weizhan Zhang, Hao Sun 0015 |
AAAI | 3 |
| 2024 | SFP: Spurious Feature-Targeted Pruning for Out-of-Distribution GeneralizationabstractRecent studies reveal that even highly biased dense networks can contain an invariant substructure with superior out-of-distribution (OOD) generalization. While existing works commonly seek these substructures using global sparsity constraints, the uniform imposition of sparse penalties across samples with diverse levels of spurious contents renders such methods suboptimal. The precise adaptation of model sparsity, specifically tailored for spurious features, remains a significant challenge. Motivated by the insight that in-distribution (ID) data containing spurious features may exhibit lower experiential risk, we propose a novel Spurious Feature-targeted Pruning framework, dubbed SFP, to induce the authentic invariant substructures without referring to the above concerns. Specifically, SFP distinguishes spurious features within ID instances during training by a theoretically validated threshold. It then penalizes the corresponding feature projections onto the model space, steering the optimization towards subspaces spanned by those invariant factors. Moreover, we also conduct detailed theoretical analysis to provide a rationality guarantee and a proof framework for OOD structures based on model sparsity. Experiments on various OOD datasets show that SFP can significantly outperform both structure-based and non-structure-based OOD generalization state-of-the-art (SOTA) methods by large margins. Yingchun Wang 0001, Jingcai Guo, Song Guo 0001, Yi Liu 0057, Jie Zhang 0076, Weizhan Zhang |
ACM Multimedia | 6 |
| 2024 | Accelerating Non-Maximum Suppression: A Graph Theory PerspectiveabstractNon-maximum suppression (NMS) is an indispensable post-processing step in object detection. With the continuous optimization of network models, NMS has become the ``last mile'' to enhance the efficiency of object detection. This paper systematically analyzes NMS from a graph theory perspective for the first time, revealing its intrinsic structure. Consequently, we propose two optimization methods, namely QSI-NMS and BOE-NMS. The former is a fast recursive divide-and-conquer algorithm with negligible mAP loss, and its extended version (eQSI-NMS) achieves optimal complexity of $\mathcal{O}(n\log n)$. The latter, concentrating on the locality of NMS, achieves an optimization at a constant level without an mAP loss penalty. Moreover, to facilitate rapid evaluation of NMS methods for researchers, we introduce NMS-Bench, the first benchmark designed to comprehensively assess various NMS methods. Taking the YOLOv8-N model on MS COCO 2017 as the benchmark setup, our method QSI-NMS provides $6.2\times$ speed of original NMS on the benchmark, with a $0.1\%$ decrease in mAP. The optimal eQSI-NMS, with only a $0.3\%$ mAP decrease, achieves $10.7\times$ speed. Meanwhile, BOE-NMS exhibits $5.1\times$ speed with no compromise in mAP. King-Siong Si, Weizhan Zhang, Tieliang Gong, Jiahao Wang 0004, Hao Sun 0015 |
NeurIPS | 3 |
| 2024 | OneActor: Consistent Subject Generation via Cluster-Conditioned GuidanceabstractText-to-image diffusion models benefit artists with high-quality image generation. Yet their stochastic nature hinders artists from creating consistent images of the same subject. Existing methods try to tackle this challenge and generate consistent content in various ways. However, they either depend on external restricted data or require expensive tuning of the diffusion model. For this issue, we propose a novel one-shot tuning paradigm, termed OneActor. It efficiently performs consistent subject generation solely driven by prompts via a learned semantic guidance to bypass the laborious backbone tuning. We lead the way to formalize the objective of consistent subject generation from a clustering perspective, and thus design a cluster-conditioned model. To mitigate the overfitting challenge shared by one-shot tuning pipelines, we augment the tuning with auxiliary samples and devise two inference strategies: semantic interpolation and cluster guidance. These techniques are later verified to significantly improve the generation quality. Comprehensive experiments show that our method outperforms a variety of baselines with satisfactory subject consistency, superior prompt conformity as well as high image quality. Our method is capable of multi-subject generation and compatible with popular diffusion extensions. Besides, we achieve a $4\times$ faster tuning speed than tuning-based baselines and, if desired, avoid increasing the inference time. Furthermore, our method can be naturally utilized to pre-train a consistent subject generation network from scratch, which will implement this research task into more practical applications. (Project page: https://johnneywang.github.io/OneActor-webpage/) Jiahao Wang 0004, Caixia Yan, Haonan Lin, Weizhan Zhang, Mengmeng Wang 0005, Tieliang Gong, Guang Dai, Hao Sun 0015 |
NeurIPS | 4 |
| 2024 | An Experimental Study on Half-Closed TCP Connections in Public Cloud GatewaysabstractMany TCP-based services, such as social media and e-commerce platforms, depend on Load Balancer(LB) or Network Address Translation(NAT) gateways to manage end-user connections. In this paper, we conduct an in-depth study on the prevalence and impact of half-closed TCP connections in public cloud LB gateways. Our trace data analysis reveals that, on average, these connections account for 9.26% of new connections and consume 16.14% more memory resources compared to fully active connections. Through a series of experiments, we investigate the role of popular HTTP libraries in causing half-closed connections and identify specific implementation patterns that contribute to this behavior. To address this, we propose feasible approaches and demo code to reduce the number of such half-closed connections on both the gateway and client sides. Zhuang Yuan, Kejing Xu, Weizhan Zhang |
TrustCom | 6 |
| 2024 | A3RT: Attention-Aware AR Teleconferencing with Life-Size 2.5D Video AvatarsabstractAugmented Reality (AR) teleconferencing aims to enable remotely separated users to meet with each other in their own physical spaces as if they are face-to-face. Among all solutions, the video-avatar-based approach has the advantage of balancing fidelity and the sense of co-presence using easy-to-setup devices, including only a camera and an AR Head-Mounted Display (HMD). However, non-verbal cues indicating “who is looking at whom” are always lost or misdelivered in multiparty teleconferencing experiences. To make users aware of such non-verbal cues, existing solutions explore screen-based visualizations, incorporate additional hardware, or alter to use a virtual avatar representation. However, they lack immersion, are less feasible for everyday usage due to complex installations, or lose the fidelity of remote users’ authentic appearances. In this paper, we decompose such attention awareness into the awareness of being looked at and the awareness of attention between other users and address them in a decoupled process. Specifically, through a user study, we first find an unobtrusive and reasonable layout “Attention Circle” to retarget a looker’s head gaze to the one being looked at. We then conduct the second user study to find an effective and intuitive “rotatable 2.5D video avatar with attention thumbnail” visualization to aid users in being aware of other users’ attention. With the design choice distilled from the studies, we implement A3RT, a proof-of-concept prototype system that empowers attention-aware 2.5D-video-avatar-based multiparty AR teleconferencing in an easy, everyday setup. Ablation and usability studies on the prototype verify the effectiveness of our proposed components and the full system. Xuanyu Wang 0001, Weizhan Zhang, Hongbo Fu 0001 |
VR | 2 |
| 2024 | Adaptive token selection for efficient detection transformer with dual teacher supervision
Muyao Yuan, Weizhan Zhang, Caixia Yan, Tieliang Gong, Yuanhong Zhang, Jiangyong Ying |
Knowl. Based Syst. | 2 |
| 2024 | Towards performance-maximizing neural network pruning via global channel attention
Yingchun Wang 0001, Song Guo 0001, Jingcai Guo, Jie Zhang 0076, Weizhan Zhang, Caixia Yan, Yuanhong Zhang |
Neural Networks | 5 |
| 2024 | Context-Aware Cross-Layer Congestion Control for Large-Scale Live StreamingabstractLive video streaming has come to dominate today’s Internet traffic. Content Delivery Network (CDN) providers, responsible for hosting outsourced live streaming services, are now striving to ensure an enhanced quality of experience (QoE) to meet the ever-increasing user expectations. Existing congestion control (CC) schemes in the kernel, however, suffer from unsatisfactory performance for live video delivery due to disparities in traffic characteristics and differentiated optimization goals between generic traffic and live video traffic. In this paper, we propose XCC, a streaming context-aware CC approach that helps achieve better QoE for the live streaming services from CDN provider. The core of XCC is to adaptively coordinate the transmission strategy and frame rate through a cross-layer feedback framework, responding to the fluctuating traffic dynamics and network conditions in the short term. Further, XCC matches the long-term traffic characteristics (i.e., two-stage delivery mode) by employing a task-specific state transition mechanism as the underlying TCP. XCC has been implemented in the Linux kernel’s TCP stack and media engine and has been fully deployed in Alibaba Cloud’s production service. Evaluation in experimental environments and A/B testing serving tens of millions of sessions demonstrate that XCC is competitive in streaming delay against the most prevalent TCP in today’s Operating Systems, while reducing startup delay by 9.9%, stall time by 36.4%, and stall frequency by 42.5% on average in deployment. Danfu Yuan, Weizhan Zhang, Yubing Qiu, Haiyu Huang 0005, Hongfei Yan, Yaming He |
IEEE/ACM Trans. Netw. | 2 |
| 2024 | FHVAC: Feature-Level Hybrid Video Adaptive Configuration for Machine-Centric Live StreamingabstractWith the widespread deployment of edge computing, the focus has shifted to machine-centric live video streaming, where endpoint-collected videos are transmitted over networks to edge servers for analysis. Unlike maximizing user's Quality of Experience (QoE), machine-centric video streaming optimizes the machine's Quality of Inference (QoI) by balancing the inference accuracy, inference delay, and transmission latency with video adaptive configuration. Traditional heuristic configuration adaption methods are reliable but unable to respond to erratic network fluctuations. Reinforcement learning (RL) based algorithms exhibit superior flexibility but suffer from exploration mechanisms, resulting in long-tail effects on upload latency. In this paper, we propose FHVAC, which dynamically selects video encoding parameters for live streaming by coherently fusing rule-based and RL-based agent at the feature level. We initially develop a robust rule-based approach for ensuring the low latency in transmission, and employ imitation learning to convert it into a neural network equivalently. Subsequently, we design a novel module to combine the two approaches and assess various fusion mechanisms. Our evaluation of FHVAC across two vision tasks (pose estimation and semantic segmentation) in two scenarios (trace-driven simulation and testbed-based experiment) shows that FHVAC enhances the average QoI, and reduces 10.61%-65.27% latency tail performance compared to prior work. Yuanhong Zhang, Weizhan Zhang, Haipeng Du, Caixia Yan, Li Liu 0036 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Towards Fairer and More Efficient Federated Learning via Multidimensional Personalized Edge ModelsabstractFederated learning (FL) is an emerging technique that trains massive and geographically distributed edge data while maintaining privacy. However, FL has inherent challenges in terms of fairness and computational efficiency due to the rising heterogeneity of edges, and thus usually results in sub-optimal performance in recent state-of-the-art (SOTA) solutions. In this paper, we propose a Customized Federated Learning (CFL) system to eliminate FL heterogeneity from multiple dimensions. Specifically, CFL tailors personalized models from the specially designed global model for each client jointly guided by an online trained model-search helper and a novel aggregation algorithm. Extensive experiments demonstrate that CFL has full-stack advantages for both FL training and edge reasoning and significantly improves the SOTA performance w.r.t. model accuracy (up to 7.2% in the non-heterogeneous environment and up to 21.8% in the heterogeneous environment), efficiency, and FL fairness. Jingcai Guo, Jie Zhang 0076, Song Guo 0001, Weizhan Zhang |
IJCNN | 5 |
| 2023 | Tile Classification Based Viewport Prediction with Multi-modal Fusion TransformerabstractViewport prediction is a crucial aspect of tile-based 360° video streaming system. However, existing trajectory based methods lack of robustness, also oversimplify the process of information construction and fusion between different modality inputs, leading to the error accumulation problem. In this paper, we propose a tile classification based viewport prediction method with Multi-modal Fusion Transformer, namely MFTR. Specifically, MFTR utilizes transformer-based networks to extract the long-range dependencies within each modality, then mine intra- and inter-modality relations to capture the combined impact of user historical inputs and video contents on future viewport selection. In addition, MFTR categorizes future tiles into two categories: user interested or not, and selects future viewport as the region that contains most user interested tiles. Comparing with predicting head trajectories, choosing future viewport based on tile's binary classification results exhibits better robustness and interpretability. To evaluate our proposed MFTR, we conduct extensive experiments on two widely used PVS-HM and Xu-Gaze dataset. MFTR shows superior performance over state-of-the-art methods in terms of average prediction accuracy and overlap ratio, also presents competitive computation efficiency. Weizhan Zhang, Caixia Yan, Qi Wang 0180, Wangdu Chen |
ACM Multimedia | 3 |
| 2022 | Deep Reinforcement Learning Based Adaptive 360-degree Video Streaming with Field of View Joint PredictionabstractWith the development of 360-degree video and HTTP adaptive streaming (HAS), tile-based adaptive 360-degree video streaming has become a promising paradigm for reducing the bandwidth consumption of delivering the panoramic video content. However, there are two main challenges for the adaptive 360-degree video streaming, accurate long-term prediction of the future field of view (Fo V) and optimal adaptive bitrate (ABR) transmission strategy. In this paper, we propose an attention-based multi-user Fo V joint prediction approach to improve the accuracy, establishing a probability model of watching video tiles for users and applying Long Short-Term Memory (LSTM) network and DBSCAN clustering method. Furthermore, we present an adaptive 360-degree video streaming approach based on deep reinforcement learning (DRL), using A3C algorithm to optimize the QoE. The real-world trace-driven experiments demonstrate that our approach achieves about 8 % gains on user Fo V prediction precision and an increase at least 20 % on user QoE compared with the benchmarks. Yuanhong Zhang, Junquan Liu, Haipeng Du, Weizhan Zhang |
ISCC | 6 |
| 2022 | PRIOR: deep reinforced adaptive video streaming with attention-based throughput predictionabstractVideo service providers have deployed dynamic video bitrate adaptation services to fulfill user demands for higher video quality. However, fluctuations and instability of network conditions inhibit the performance promotion of adaptive bitrate (ABR) algorithms. Existing rule-based approaches fail to guarantee accurate throughput estimates, and learning-based algorithms are considerably sensitive to the variability of network. Therefore, how to gain effective and stable throughput estimates has become one of the critical challenges to further enhancing ABR methods. To eliminate this concern, we propose PRIOR, an ABR algorithm that fuses an effective throughput prediction module and a state-of-the-art multi-agent reinforcement learning method to provide a high quality of experience (QoE). PRIOR aims to maximize the QoE metric by straightforwardly utilizing accurate throughput estimates rather than past throughput measurements. Specifically, PRIOR employs a light-weighted prediction module with attention mechanism to obtain effective future throughput. Considering the excellent features introduced by the HTTP/3 protocol, we apply PRIOR to trace-driven simulations and real-world scenarios over HTTP/1.1 and HTTP/3. Trace-driven emulation illustrates that PRIOR outperforms existing ABR schemes over HTTP/1.1 and HTTP/3, and our prediction module can also reinforce the performance of other ABR algorithms. Extensive results on real-world evaluation demonstrate the superiority of PRIOR over existing state-of-the-art ABR schemes. Danfu Yuan, Yuanhong Zhang, Weizhan Zhang, Xuncheng Liu, Haipeng Du |
NOSSDAV | 3 |
| 2022 | Improving unsupervised image clustering with spatial consistency
Rui Zhao 0028, Jianfei Ruan, Bo Dong 0001, Weizhan Zhang |
Knowl. Based Syst. | 5 |
| 2022 | Predict-and-Drive: Avatar Motion Adaption in Room-Scale Augmented Reality Telepresence with Heterogeneous SpacesabstractAvatar-mediated symmetric Augmented Reality (AR) telepresence has emerged with the ability to empower users located in different remote spaces to interact with each other in 3D through avatars. However, different spaces have heterogeneous structures and features, which bring difficulties in synchronizing avatar motions with real user motions and adapting avatar motions to local scenes. To overcome these issues, existing methods generate mutual movable spaces or retarget the placement of avatars. However, these methods limit the telepresence experience in a small sub-area space, fix the positions of users and avatars, or adjust the beginning/ending positions of avatars without presenting smooth transitions. Moreover, the delay between the avatar retargeting and users' real transitions can break the semantic synchronization between users' verbal conversation and perceived avatar motion. In this paper, we first examine the impact of the aforementioned transition delay and explore the preferred transition style with the existence of such delay through user studies. With the results showing a significant negative effect of avatar transition delay and providing the design choice of the transition style, we propose a Predict-and-Drive controller to diminish the delay and present the smooth transition of the telepresence avatar. We also introduce a grouping component as an upgrade to immediately calculate a coarse virtual target once the user initiates a transition, which could further eliminate the avatar transition delay. Once having the coarse virtual target or an exactly predicted target, we find the corresponding target for the avatar according to the pre-constructed mapping of objects of interest between two spaces. The avatar control component maintains an artificial potential field of the space and drives the avatar towards the target while respecting the obstacles in the physical environment. We further conduct ablation studies to evaluate the effectiveness of our proposed components. Xuanyu Wang 0001, Christian Sandor, Weizhan Zhang, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Adaptive Video Streaming Using Dynamic Server Push over HTTP/2abstractWith the increasing popularity of video services, HTTP adaptive streaming (HAS) has become the mainstream technology for media streaming distribution. In traditional HAS over HTTP/1.1, the HAS server responds to each request from the client individually. This process adds additional round-trip time, resulting in underestimation of available bandwidth. As a result, the HAS client chooses a lower bitrate, which reduces network utilization and the user's quality of experience. In recent years, the HTTP/2 protocol has emerged, which allows server to actively push multiple data segments to the client. Pushing multiple segments can reduce the negative impact of network latency on estimating available bandwidth, thereby increasing the user's request bitrate and video quality. However, when the network is unstable, the more video segments that are pushed by the server, the more challenges the client encounters in responding to network fluctuations i n time, causing playback stalling and poor user experience. Therefore, this paper proposes a dynamic server push algorithm over HTTP/2, which chooses a different number of segments for server push according to network fluctuations. For the evaluation results, relative to its benchmarks, the proposed approach improves the average video request bitrate while minimizing the probability of playback stalling. Shouqin Huang, Weizhan Zhang, Haipeng Du |
CSCWD | 3 |
| 2021 | Dynamic Push for HTTP Adaptive Streaming with Deep Reinforcement Learning
Haipeng Du, Danfu Yuan, Weizhan Zhang |
ICPADS | 3 |
| 2021 | QoE-driven HAS Live Video Channel Placement in the Media CloudabstractHTTP adaptive streaming (HAS) technology has been increasingly employed by video service providers (VSPs) due to its prominent benefits such as reducing interruptions of video playback and achieving higher bandwidth utilization and outstanding quality of experience (QoE). And many VSPs have deployed HAS applications in the media cloud to provide large-scale video streaming services. At present, research into the media cloud typically focuses on the management and optimization of cloud resources, such as the placement and migration of virtual machines in media cloud data centers. However, considering the HASlive video streamingservice, existing related works have not adequately discussed the specific impact of the consumption of computing and bandwidth resources of media cloud servers on the user experience (QoE), particularly under the resource constraints in the media cloud. In this paper, we first investigate and formulate the computing and bandwidth resource consumption characteristics of HAS live video streaming with different frame rates and resolutions, and we further establish a resources-aware QoE model to quantify the user experience oflive video channels(i.e., programs). Then, based on the model, we present a QoE-driven HAS live video channel placement approach (including a placement algorithmHCPand a rescheduling algorithmHCR) to optimize the channel allocation in media cloud servers, aiming to maximize the average user QoE. We abstract the maximization problem into an MMKP problem, and employ a heuristic solution to address this problem. The experimental results demonstrate the effectiveness of our proposed approach compared with benchmark solutions. Junquan Liu, Weizhan Zhang, Shouqin Huang, Haipeng Du |
IEEE Trans. Multim. | 2 |
| 2020 | Deploying Fused Sharable Video Interaction Channels in Mobile CloudabstractA plethora of mobile video applications involving all aspects of social life are becoming increasingly prevalent. However, for resource-hungry and delay-sensitive multi-view video applications such as multi-channel video conferencing and 3D videos, the hardware resources of mobile terminals becomes a bottleneck in concurrently decoding multiple videos. Fusing multiple views into a single-view video stream in the cloud before transmission can unload the computation of mobile terminals. But the delay increment caused by video fusing in this cloudbased multi-view video streaming makes the strategy of deploying such applications a new area to examine. In this paper, in the interest of deployment with higher resource utilization and less latency, first, together with a load model of the fused sharable video interaction process, a channel admission control algorithm that targets maximizing the supportive capacity of the cloud is introduced. Subsequently, a channel deployment algorithm based on load coordination between the cloud and mobile clients is proposed to ensure that the delay caused by cloud processing is acceptable while minimizing the terminal computing load. Xuanyu Wang 0001, Haipeng Du, Weizhan Zhang |
GLOBECOM | 3 |
| 2020 | EmotionTracker: A Mobile Real-time Facial Expression Tracking System with the Assistant of Public AI-as-a-ServiceabstractPublic AI-as-a-Service (AIaaS) is a promising next-generation computing paradigm that attracts resource-limited mobile users to outsource their machine learning tasks. However, the time delay between cloud/edge servers and end users makes it hard for real-time mobile artificial intelligence applications. In this demonstration, we present EmotionTracker, a real-time mobile facial expression tracking system combining AIaaS and mobile local auxiliary computing, including facial expression tracking and the corresponding task offloading. Mobile facial expression tracking iteratively estimates the facial expression with the help of sparse optical flow and neural network. Task offloading dynamically estimate the moment of task offloading with machine learning method. According to the results in a real-world environment, EmotionTracker successfully fulfills the mobile real-time facial expression tracking requirements. Xuncheng Liu, Weizhan Zhang, Xuanya Li |
ACM Multimedia | 3 |
| 2020 | AvatarMeeting: An Augmented Reality Remote Interaction System With Personalized AvatarsabstractTo further enhance the immersion perception of remote interaction, avatars can be involved harnessing Head Mounted Display (HMD) based Augmented Reality (AR). In our demonstration, we present an avatar based remote interaction system AvatarMeeting, enabling users to meet with remote peers through interactive personalized avatars just like face to face. Specifically, we propose a novel framework including a consumer-grade set-up, a complete transmission scheme and a processing pipeline, which consists of prescan modeling, pose detection and action reconstruction. And an angle based reconstruction approach is introduced to empower the AR avatars to perform the same actions as each remote real person do in real time smoothly while keeping a good avatar shape. Xuanyu Wang 0001, Weizhan Zhang |
ACM Multimedia | 4 |
| 2019 | MUCH: Priority-based Collaborative Multi-Channel HTTP Adaptive StreamingabstractHTTP adaptive streaming (HAS) provides an effective means to deploy media streaming applications over today's Internet. A major challenge in developing such an application is how to deal with bandwidth competition among multiple video channels. The client-driven nature of HAS makes it difficult to provide server-side resource management, and the dynamic adaptability of HAS complicates the situation. Therefore, in this paper, we propose a priority-based multichannel collaborative HAS scheme to provide differentiated HAS services among channels. Different from the start-of-the-art client-oriented approaches, the solution refocuses on the server and employs a dynamic HAS request response algorithm on the server side. The algorithm guarantees the pre-allocated bandwidth for each channel by the request management and allows flexible bandwidth occupation among different channels to collaboratively utilize the server bandwidth resources. At the same time, it still obeys the client-driven nature of HAS. The rate adaptation decisions of each client are still made locally to ensure an automatic reactive system. The scheme has been carefully evaluated in both simulation and real-world environments with different client adaptive algorithms. For the evaluation results, compared with its benchmarks, the scheme both achieves the channel-specific priority-based quality assurance and improves the server resource utilization. Bingfang Qi, Weizhan Zhang, Chunmeng Yang |
CSCWD | 2 |
| 2019 | Work-in-Progress: Version-Aware Video Caching Strategy for Multi-version VoD SystemsabstractRecently, many video-on-demand (VoD) providers store multiple versions of the same videos to offer multiple-quality video services with different bitrates to users, called as multi-version VoD. To decrease the start-up delay for users, it is a good idea to cache videos at caching server that is in close proximity. However, how to decide which versions of which videos should be cached and replaced in caching server is still one major challenge for multi-version VoD systems because of limited caching storage. In this paper, we propose a version-aware video caching strategy for multi-version VoD systems, which aims to reduce start-up delay and improve cache hit ratio. First, we take into account the transcoding delay among versions and transmit delay from content server to caching server to calculate version-aware caching profit when caching a certain version or multiple versions of a video. It is the basis for the following caching replacement algorithm. Second, we propose version-aware video caching (VaVC) algorithm to decide which versions of which videos will be replaced based on the version-aware caching profit dynamically. In this way, VaVC can reduce start-up delay and improve the cache hit ratio. Our simulation results have shown that VaVC outperforms the others in both the start-up delay and the cache hit ratio. Hui Zhao 0003, Zili Wu, Quan Wang 0006, Jing Wang 0028, Weizhan Zhang |
RTSS | 5 |
| 2018 | Resource Allocation for Virtual Streaming Media Server Cluster in Cloud-based Multi-version VoDabstractWith the rapid development of mobile Internet and smart devices, VoD (video on demand) providers build media cloud to offer multi-bitrate video streaming services to users at a reduced cost, called as cloud-based multi-version VoD. In cloud-based multi-version VoD, we need to solve the problem of allocating appropriate resources for virtual streaming media server cluster with the aim of optimizing the user experience and reducing the service cost. To address this problem, a resource allocation for virtual streaming media server cluster in cloud-based multi-version VoD is proposed in this paper. We firstly analyze the user historical learning logs to mine the user behavior characteristics, including the average user request arrival rate, the video playing time distribution, and the video popularity distribution, etc. Then, based on the user behavior characteristics and the queueing theory, a resource allocation model for the virtual streaming media server cluster is introduced. It predicts the user arrival rate at first and then allocates appropriate resources dynamically to solve the resources allocation irrationality problem. Simulation results have proved the proposed method can allocate appropriate resources for virtual streaming media server cluster, which can ensure the user experience satisfaction and improve the resources utilization. Hui Zhao 0003, Jing Wang 0028, Quan Wang 0006, Nan Luo, Weizhan Zhang |
CSCWD | 6 |
| 2018 | Integrated Bandwidth Variation Pattern Differentiation for HTTP Adaptive Streaming over 4G Cellular NetworksabstractHTTP adaptive streaming (HAS) is the state-of-the-art technology for improving the quality of user experience under conditions of time-varying available bandwidth. Developing the bitrate adaptation algorithm becomes more challenging with the transition to 4G cellular networks. The features of bandwidth variation caused by changes in radio channel quality are significantly different from the pattern caused by changes in radio channel resources. In this paper, we propose a bitrate adaptation algorithm for 4G cellular networks with bandwidth variation pattern differentiation. By investigating the bandwidth data profiles collected from an actual 4G cellular network with the field test approach, we distinguish the pattern of bandwidth capacity variations as sustained fluctuations and instantaneous hopping. With bandwidth variation pattern differentiation, the proposed algorithm performs a smoothed bitrate adaptation to sustained fluctuations and an instant bitrate adaptation to instantaneous hopping. Performance evaluations obtained on a 4G cellular network testbed demonstrate that the algorithm achieves reductions in bitrate switching frequency and playback stalls while increasing the average bitrate. Haipeng Du, Weizhan Zhang, Xuanyu Wang 0001 |
IPCCC | 2 |
| 2018 | Mining temporal characteristics of behaviors from interval events in e-learning
Tao Xie 0007, Weizhan Zhang |
Inf. Sci. | 3 |
| 2018 | Prediction-Based and Locality-Aware Task Scheduling for Parallelizing Video Transcoding Over Heterogeneous MapReduce ClusterabstractMapReduce is a popular programming model in cloud computing to deal with the high computational task, such as video transcoding. It splits the video (task) into multiple segments (subtasks) and transcodes them in parallel in cluster. Due to the complexity of video transcoding and the poor performance of heterogeneous MapReduce cluster, scheduling these subtasks to minimize the total transcoding time is still a challenge. In this paper, we propose a prediction-based and locality-aware task scheduling (PLTS) method for parallelizing video transcoding over heterogeneous MapReduce cluster. First, we analyze video decoding and encoding technologies and predict the segment transcoding complexity, which can provide a foundational base for the following scheduling. Second, we attempt to schedule subtasks on machines that contain the related input data, which are referred to as data locality, so as to reduce large-scale data movement and data transfer during the mapping phase. Third, we formulate the scheduling as a job shop scheduling problem and propose a heuristic PLTS algorithm. It combines the benefits of two traditional heuristic scheduling algorithms, Max-Min and Min-Min, to make load balancing in cluster and short the total transcoding time. The experimental results also show the efficiency of our algorithm. Hui Zhao 0003, Weizhan Zhang, Jing Wang 0028 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Power-Aware and Performance-Guaranteed Virtual Machine Placement in the CloudabstractCloud service providers offer virtual machines (VMs) as services to users over Internet. As VMs are running on physical machines (PMs), PM power consumption needs to be considered. Meanwhile, VMs running on the same PM share physical resources, and there exists great resource contention, which results in VM performance degradation. Therefore, how to place VMs to reduce PM power consumption and guarantee VM performance is still one major challenge. However, existing VMPs did not study VM performance degradation, so they could not guarantee VM performance. To solve the high power consumption and VMs performance degradation problems, this paper explores the balance between saving PM power and guaranteeing VM performance, and proposes a power-aware and performance-guaranteed VMP (PPVMP). First, we investigate the relationship between power consumption and CPU utilization to build a non-linear power model, which is helpful for the following VMP. Second, we construct VM performance models to present the VM performance degradation trend. Third, based on these models, we formulate VMP as a bi-objective optimization problem, which tries to minimize PM power consumption and guarantee VM performance. We then propose an algorithm based on ant colony optimization to solve it. Finally, the results show the efficiency of our algorithm. Hui Zhao 0003, Jing Wang 0028, Quan Wang 0006, Weizhan Zhang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | Recognizing physical contexts of mobile video learners via smartphone sensors
Tao Xie 0007, Weizhan Zhang |
Knowl. Based Syst. | 3 |
| 2017 | LTE-EMU: A High Fidelity LTE Cellar Network Testbed for Mobile Video Streaming
Haipeng Du, Weizhan Zhang, Yunhui Huang |
Mob. Networks Appl. | 3 |
| 2017 | A Segment-Based Storage and Transcoding Trade-off Strategy for Multi-version VoD Systems in the CloudabstractMulti-version video-on-demand (VoD) providers either store multiple versions of the same video or transcode video to multiple versions in real time to offer multiple-bitrate streaming services to heterogeneous clients. However, this could incur tremendous storage cost or transcoding computation cost. There have been some works regarding trading off between transcoding and storing whole videos, but they did not take into account video segmentation and internal popularity. As a result, they were not cost-efficient. This paper introduces video segmentation and proposes a segment-based storage and transcoding trade-off strategy for multi-version VoD systems in the cloud. First, we split each video into multiple segments depending on the video internal popularity. Second, we describe the transcoding relationships among versions using a transcoding weighted graph, which can be used to calculate the version-aware transcoding cost from one version to another. Third, we take the video segmentation, version-aware transcoding weighted graph, and video internal popularity into account to propose a storage and transcoding trade-off strategy, which stores multiple versions of popular segments and transcodes unpopular segments. We then formulate it as an optimization problem and present a heuristic divide-and-conquer algorithm to get an approximate optimal solution. Finally, we conduct extensive simulations to evaluate the solution; the results show that it can significantly lower the storage and transcoding cost of multi-version VoD systems. Hui Zhao 0003, Weizhan Zhang, Biao Du, Haifei Li 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Cluster-Aware Virtual Machine Collaborative Migration in Media CloudabstractMedia cloud has become a promising paradigm for deploying large-scale streaming media applications at a reduced cost. Due to dynamic and diverse demands of users, media cloud presents two crucial characteristics: high resource consumption and dynamic traffic among media servers. Consequently, Virtual Machine (VM) migration in media cloud is highly required to suit varying resource requirements and the dynamic traffic patterns. Moreover, migration of such bandwidth-intensive media applications in media cloud needs cautious handling, especially for the internal traffic of Data Center Networks (DCN). However, existing media cloud resource management schemes or traffic-aware VM deployment approaches are insufficient for media cloud, ignoring the characteristics of either cloud infrastructure or media streaming requirements. In this paper, we propose a cluster-aware VM collaborative migration scheme for media cloud, tightly integrating clustering, placement, and dynamic migration process. The scheme employs a clustering algorithm and a placement algorithm to obtain ideal migration strategies for newly perceived media server clusters, and a migration algorithm to effectively accomplish the migration process of media servers. Evaluation results demonstrate that our scheme can effectively migrate virtual media servers in media cloud, while reducing the total internal traffic in DCN under the resource consumption constraints of media streaming applications. Weizhan Zhang, Zhichao Mo, Zongqing Lu 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | A priority-based adaptive scheme for multi-view live streaming over HTTP
Weizhan Zhang, Shuyan Ye, Hui Zhao 0003 |
Comput. Commun. | 1 |
| 2016 | A behavioral sequence analyzing framework for grouping students in an e-learning system
Tao Xie 0007, Weizhan Zhang |
Knowl. Based Syst. | 3 |
| 2016 | MSC: a multi-version shared caching for multi-bitrate VoD services
Hui Zhao 0003, Weizhan Zhang, Haifei Li 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Towards Information Diffusion in Mobile Social NetworksabstractThe emerging of mobile social networks opens opportunities for viral marketing. However, before fully utilizing mobile social networks as a platform for viral marketing, many challenges have to be addressed. In this paper, we address the problem of identifying a small number of individuals through whom the information can be diffused to the network as soon as possible, referred to as thediffusion minimizationproblem. Diffusion minimization under the probabilistic diffusion model can be formulated as an asymmetric$k$-center problem which is NP-hard, and the best known approximation algorithm for the asymmetric$k$-center problem has approximation ratio of$\log ^*n$and time complexity$O(n^5)$. Clearly, the performance and the time complexity of the approximation algorithm are not satisfiable in large-scale mobile social networks. To deal with this problem, we propose a community based algorithm and a distributed set-cover algorithm. The performance of the proposed algorithms is evaluated by extensive experiments on both synthetic networks and a real trace. The results show that the community based algorithm has the best performance in both synthetic networks and the real trace compared to existing algorithms, and the distributed set-cover algorithm outperforms the approximation algorithm in the real trace in terms of diffusion time. Zongqing Lu 0002, Yonggang Wen 0001, Weizhan Zhang, Guohong Cao |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Virtual machine placement based on the VM performance models in cloudabstractCloud service providers can offer users virtual machines (VMs) on-demand as a service over the Internet. VMs running on top of a physical machine (PM) share physical resources (CPU, memory, and/or bandwidth), and there may be a great resource contention among them, which results in VMs performance degradation. To prevent this, cloud providers need to study how to place VMs on PMs efficiently. However, the existing virtual machine placement (VMP) methods mainly tried to optimize the cloud resources instead of the VM performance. In this paper, we propose a VMP method based on the VM performance models in cloud. Firstly, with a real OpenStack cloud platform, we study the virtualization resource scheduling principle, analyze the interaction among VMs with shared hardware, consider the relationship between VMs and the host PM, and then we introduce the VM performance models to present the VM performance degradation problem. Secondly, to choose an appropriate PM for placing VM, we take into consideration the application-aware resource consumption characteristic, the VM resource requirement and the VM performance models, so as to minimize the PM performance degradation and guarantee the VM performance. Finally, we take the streaming media services for examples, and conduct some experiments to evaluate our method. The results show it works better than others and guarantees the VM performance significantly. Hui Zhao 0003, Weizhan Zhang, Yunhui Huang |
IPCCC | 3 |
| 2015 | A version-aware computation and storage trade-off strategy for multi-version VoD systems in the cloudabstractNowdays, many Video-on-Demand (VoD) providers offer multiple-quality video streaming services to heterogeneous clients, called as multi-version VoD. Some researches focus on video transcoding in real-time or video layered encoding/decoding, but they are not widely used in VoD industry. Storing multiple versions of the same video is an easy solution, but it consumes lots of storage space. Although there are also a few works about trading-off between transcoding and storage, they did not utilize the transcoding relationships among different versions and took the video popularity into account, which bring that they may have little cost-efficiency for multi-version VoD systems. To minimize the cost, in this paper, we propose a version-aware transcoding computation and storage trade-off strategy for multi-version VoD systems in the cloud. Firstly, it utilizes the transcoding weight graph to describe the transcoding relationships among different versions of a video. According to the graph, the transcoding computation cost from one version to another version can be calculated. Secondly, it takes the video popularity of different versions, the prices of storage and computation resources in the cloud into account to decide which versions of which videos should be stored or transcoded. We then formulate it as an optimization problem and present a heuristic approximate optimal solution. Finally, we conduct extensive simulations to evaluate our strategy and solution, and the results show that they can significantly lower the cost of multi-version VoD systems. Hui Zhao 0003, Weizhan Zhang, Biao Du |
ISCC | 3 |
| 2015 | Workload modeling for virtual machine-hosted application
Weizhan Zhang, Jun Liu 0002, Wei Zhang 0053 |
Expert Syst. Appl. | 1 |
| 2014 | Recognizing and regulating e-learners' emotions based on interactive Chinese texts in e-learning systems
Feng Tian 0002, Pengda Gao, Longzhuang Li, Weizhan Zhang, Huijun Liang, Ya-nan Qian, Ruomeng Zhao |
Knowl. Based Syst. | 4 |
| 2014 | CBC: Caching for cloud-based VOD systems
Weizhan Zhang, Zhichao Mo |
Multim. Tools Appl. | 1 |
| 2012 | An overlay multicast protocol for live streaming and delay-guaranteed interactive media
Weizhan Zhang, Haifei Li 0001, Feng Tian 0002 |
J. Netw. Comput. Appl. | 1 |
| 2011 | Multi-channel live streaming in service overlay network
Weizhan Zhang |
Multim. Tools Appl. | 1 |
| 2009 | Prototype Demonstration: Trojan Detection and Defense SystemabstractThis paper presents a novel Trojan detection and defense system. The prototype searches the important files which contain users' confidential information on the disk. And then, these files will be monitored to find which processes will access them by capturing and analyzing the IRPs (I/O request packets). The processes of Trojans will be distinguished from regular ones by evaluating their API-calls with several machine-learning models, rather than traditional signature-based mechanism. Testing results show that this prototype could detect and defend the unknown Trojans quickly and accurately. Ting Liu 0002, Xiaohong Guan, Yuanfeng Song, Weizhan Zhang |
CCNC | 6 |
| 2009 | A Multi-Tree Construction Algorithm for Multi-Channel Live Media Delivery on Overlay Service NetworkabstractIn this paper, a multi-tree construction algorithm based on overlay service network is proposed, which is tailored toward support of multi-channel live media delivery. The multiple source-specific trees are built to deliver the multiple sources respectively on shared overlay network resources. Unlike the legacy schemes, the construction of trees is not from the view of a single tree, but considering the shared overlay nodes by multiple trees holistically. The relativity of the different tree sessions is well considered. From the simulation results, a better- balanced performance is achieved among different sessions. Moreover, the proposed method makes it possible for the overlay service provider to gain a better control of the multiple sessions, and provides service differentiation on overlay service network. Weizhan Zhang, Xinyan Jia |
CCNC | 1 |
| 2009 | RealClass: An Interactive-Enable Multi-Channel Live Teaching System on Overlay Service NetworkabstractIn this demonstration, a prototype of live teaching system on overlay service network is presented. The system evolves the following novel characteristics: (1) Firstly, an infrastructure-based overlay multicast structure is introduced, providing multi-channel live media delivery with service differentiation on shared overlay service network. (2) Secondly, the system can support two-way interactive media as well as oneway live media at the same time. The enabling technology is a dynamic user interaction adjustment based on pre-reserved resources in the top layers of the tree. In the demonstration, diversified real-time services will be presented, including audio, video, instant message delivery for multi-channel live teaching scenario, and flexible interactive communication for group discussion. Weizhan Zhang, Yanze Lian, Xingzhuo Liu, Jixin Hou |
CCNC | 1 |
| 2008 | LDCOM: a Layered degree-constrained overlay multicast for interactive mediaabstractIn this paper, a layered degree-constrained overlay multicast (LDCOM) for interactive media is presented, which is tailored toward support of two-way interactive media as well as one-way live media at the same time. LDCOM distinguishes itself from previous research with the following characteristics: (1) the multicast tree is organized in two hierarchies, the layered degree-constrained core tree for interactive media and the extended clustering tree for large-scale live media delivery. (2) A dynamic user interaction adjustment algorithm based on the layered degree-constrained tree is introduced, which is especially beneficial for the interactive scenario. The simulation results show that LDCOM builds a scalable overlay multicast tree, and is able to cope with the interactive media. Weizhan Zhang, Yi Che |
ISCC | 1 |
| 2007 | IALM: an Interaction-enable Application Layer Multicast Protocol for Live TeachingabstractIn this paper, an interaction-enable application layer multicast protocol is introduced for a scalable live teaching system. With a simple and effective interaction-enable algorithm, the proposed protocol abates the additional application layer relay time to enable the video and audio interactive behavior, which takes an important part in a collaborative learning scenario. Besides, the peculiar join request distribution of the learner in live teaching is also considered, and a new client admission algorithm is adopted in the protocol to enhance the judgment efficiency. Based on the above approaches, a real live teaching system is realized. Weizhan Zhang, Wenjiang Liu, Yi Che |
CSCWD | 1 |
| 2006 | A Speech Enhancement Approach for Live Teaching System Using a Novel Hybrid Double Talk DetectorabstractAcoustic echo is a very trouble question in the live teaching system where teacher's speech will be echoed in the classroom after a moment. This paper presents a speech enhancement framework shaping for live teaching system based on a novel hybrid double-talk detector (DTD). The proposed double-talk detection algorithm consists of two decision stages. First, double-talk is detected based on the classic energy-based approach. While a further judgment for silence situation is made in the second stage. If silence is announced, the microphone channel is shut down to reduce the bandwidth requirement. After simulation and experiment, it is found that not only the proposed approach is effective in enhancing the speech quality but also can it save the network bandwidth for e-learning system Weizhan Zhang, Sha Gong, Yiqin Yu, Geli Lv |
CSCWD | 2 |