Yuanxing Zhang

dblp:194/7059 · DBLP profile ↗
← Back
52ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-1460-8124ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 13 since 2021Computer networks · 17 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
abstract
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
Lei Li 0039, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Lingpeng Kong, Qi Liu 0049, Xu Sun 0001
AAAI6
2026 ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
abstract
Yuhang Li, Chenchen Zhang, Ruilin Lv, Ao Liu, Ken Deng, Yuanxing Zhang, Jiaheng Liu, Bo Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruilin Lv, Ken Deng, Yuanxing Zhang
ACL (1)6
2025 HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
abstract
Recent Multi-modal Large Language Models (MLLMs) have made great progress in video understanding. However, their performance on videos involving human actions is still limited by the lack of high-quality data. To address this, we introduce a two-stage data annotation pipeline. First, we design strategies to accumulate videos featuring clear human actions from the Internet. Second, videos are annotated in a standardized caption format that uses human attributes to distinguish individuals and chronologically details their actions and interactions. Through this pipeline, we curate two datasets, namely HAICTrain and HAICBench. HAICTrain comprises 126K video-caption pairs generated by Gemini-Pro and verified for training purposes. Meanwhile, HAICBench includes 412 manually annotated video-caption pairs and 2,000 QA pairs, for a comprehensive evaluation of human action understanding. Experimental results demonstrate that training with HAICTrain not only significantly enhances human understanding abilities across 4 benchmarks, but can also improve text-to-video generation results. Both the HAICTrain and HAICBench will be made open-source to facilitate further research.
Xiao Wang 0056, Jingyun Hua, Weihong Lin, Yuanxing Zhang, Jianlong Wu, Di Zhang 0026, Liqiang Nie
ACL (1)4
2025 RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
abstract
Yuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan, Xu Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yuchi Wang, Yishuo Cai, Shuhuai Ren, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan 0001, Xu Sun 0001
EMNLP7
2025 MIO: A Foundation Model on Multimodal Tokens
abstract
Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jessie Jiashuo Wang, Ning Shi, Haoran Que, Zhaoxiang Zhang 0001, Yuanxing Zhang, Ge Zhang 0009, Ke Xu 0001, Jie Fu 0001, Wenhao Huang 0001
EMNLP13
2025 SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
abstract
Yuanyang Yin, Yaqi Zhao, Yajie Zhang, Yuanxing Zhang, Ke Lin, Jiahao Wang, Xin Tao, Pengfei Wan, Wentao Zhang, Feng Zhao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yuanyang Yin, Yuanxing Zhang, Xin Tao 0001, Pengfei Wan 0001, Wentao Zhang 0001
EMNLP4
2025 Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
abstract
Long-context video understanding in Multimodal Large Language Models (MLLMs) faces a critical challenge: balancing computational efficiency with the retention of fine-grained spatio-temporal patterns. Existing approaches (e.g., sparse sampling, dense sampling with low resolution, and token compression) suffer from significant information loss in temporal dynamics, spatial details, or subtle interactions, particularly in videos with complex motion or varying resolutions. To address this, we propose Mavors, a novel framework that introduces Multi-granularity video representation for holistic long-video modeling. Specifically, Mavors directly encodes raw video content into latent representations through two core components: 1) an Intra-chunk Vision Encoder (IVE) that preserves high-resolution spatial features via 3D convolutions and Vision Transformers, and 2) an Inter-chunk Feature Aggregator (IFA) that establishes temporal coherence across chunks using transformer-based dependency modeling with chunk-level rotary position encodings. Moreover, the framework unifies image and video understanding by treating images as single-frame videos via sub-image decomposition. Experiments across diverse benchmarks demonstrate Mavors' superiority in maintaining both spatial fidelity and temporal continuity, significantly outperforming existing methods in tasks requiring fine-grained spatio-temporal reasoning.
Yang Shi 0009, Yushuo Guan, Yuanxing Zhang, Weihong Lin, Jingyun Hua, Xinlong Chen, Bohan Zeng, Wentao Zhang 0001, Wenjing Yang 0002, Di Zhang 0026
ACM Multimedia5
2025 TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Linli Yao, Yuancheng Wei, Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Lingpeng Kong, Qi Liu 0049, Yuanxing Zhang, Xu Sun 0001
ACM Multimedia13
2025 EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
abstract
Diffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and mask-and-inpainting. Among these, instruction-based editing stands out for its convenience and effectiveness in following human instructions across diverse scenarios. However, it still focuses on simple editing operations like adding, replacing, or deleting, and falls short of understanding aspects of world dynamics that convey the realistic dynamic nature in the physical world. Therefore, this work EditWorld introduces a new editing task, namely world-instructed image editing, which defines and categorizes the instructions grounded by various world scenarios. We curate a new image editing dataset with world instructions using a set of large pretrained models (e.g., GPT, Video-LLava and SDXL). To enable sufficient simulation of world dynamics for image editing, our EditWorld trains model in the curated dataset, and improves instruction-following ability with designed post-edit strategy. Extensive experiments demonstrate our method significantly outperforms existing editing methods in this new task. https://github.com/YangLing0818/EditWorld
Bohan Zeng, Ling Yang 0006, Yuanxing Zhang, Pengfei Wan 0001, Wentao Zhang 0001, Shuicheng Yan
ACM Multimedia5
2025 MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
abstract
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos.The benchmark will be made publicly available to foster future research.
Yuanxing Zhang, Noah Wang, Ge Zhang 0009, Jian Yang 0037, Yanghai Wang, Xintao Wang 0002, Houyi Li, Wei Ji 0011, Pengfei Wan 0001, Wenhao Huang 0001, Zhaoxiang Zhang 0001
NeurIPS3
2025 MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
abstract
Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves only an accuracy of 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios.
Yang Shi 0009, Huanqian Wang, Wulin Xie, Huanyao Zhang, Lijie Zhao, Yifan Zhang 0004, Xinfeng Li, Chaoyou Fu, Zhuoer Wen, Zhuoran Zhang 0003, Xinlong Chen, Bohan Zeng, Yushuo Guan, Zhang Zhang 0001, Liang Wang 0001, Haoxuan Li 0001, Zhouchen Lin, Yuanxing Zhang, Pengfei Wan 0001, Haotian Wang 0001, Wenjing Yang 0002
NeurIPS20
2025 Null Space Projection-Based Hybrid Beamforming for Multi-User Massive MIMO
abstract
This study employs an array-of-subarrays (AoSA) hybrid beamforming (HBF) architecture in ultra-massive multiple-input multiple-output (UM-MIMO) systems to enhance the total achievable rate. Our primary objective is to mitigate the strong multi-user interference (MUI) through the design of null-space projection (NSP)-based HBF scheme, which involves two stages: (i) RF beamforming based on introducing beam perturbations to steer beams in the null space of interfering users, and (ii) baseband MU precoding stage based on the instantaneous effective channel to mitigate the residual MUI by a regularized zero-forcing (RZF) technique. To solve this challenging non-convex optimization problem, we propose a swarm intelligence-based sequential optimization solution that finds the optimal beam perturbations while adhering to the directivity degradation constraints for the beams in each user direction. The illustrative results depict the high achievable rate by using the proposed NSP scheme over maximum-directivity beamforming (MBF) irrespective of the users' angular locations, which can be a promising beamforming solution in future sub-Terahertz (sub- THz) UM-MIMO systems.
Mobeen Mahmood, Yuanxing Zhang, Tho Le-Ngoc
VTC2025-Spring2
2024 DDK: Distilling Domain Knowledge for Efficient Large Language Models
abstract
Despite the advanced intelligence abilities of large language models (LLMs) in various applications, they still face significant computational and storage demands. Knowledge Distillation (KD) has emerged as an effective strategy to improve the performance of a smaller LLM (i.e., the student model) by transferring knowledge from a high-performing LLM (i.e., the teacher model). Prevailing techniques in LLM distillation typically use a black-box model API to generate high-quality pretrained and aligned datasets, or utilize white-box distillation by altering the loss function to better transfer knowledge from the teacher LLM. However, these methods ignore the knowledge differences between the student and teacher LLMs across domains. This results in excessive focus on domains with minimal performance gaps and insufficient attention to domains with large gaps, reducing overall performance. In this paper, we introduce a new LLM distillation framework called DDK, which dynamically adjusts the composition of the distillation dataset in a smooth manner according to the domain performance differences between the teacher and student models, making the distillation process more stable and effective. Extensive evaluations show that DDK significantly improves the performance of student models, outperforming both continuously pretrained baselines and existing knowledge distillation methods by a large margin.
Yuanxing Zhang, Haoran Que, Ken Deng, Zhiqi Bai, Jie Liu 0047, Ge Zhang 0009, Jiakai Wang, Congnan Liu, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007
NeurIPS4
2024 D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
abstract
Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model’s fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important question is how to choose the optimal mixture ratio between the general-corpus (e.g., Dolma, Slim-pajama) and the downstream domain-corpus. Existing methods usually adopt laborious human efforts by grid-searching on a set of mixture ratios, which require high GPU training consumption costs. Besides, we cannot guarantee the selected ratio is optimal for the specific domain. To address the limitations of existing methods, inspired by the Scaling Law for performance prediction, we propose to investigate the Scaling Law of the Domain-specific Continual Pre-Training (D-CPT Law) to decide the optimal mixture ratio with acceptable training costs for LLMs of different sizes. Specifically, by fitting the D-CPT Law, we can easily predict the general and downstream performance of arbitrary mixture ratios, model sizes, and dataset sizes using small-scale training costs on limited experiments. Moreover, we also extend our standard D-CPT Law on cross-domain settings and propose the Cross-Domain D-CPT Law to predict the D-CPT law of target domains, where very small training costs (about 1\% of the normal training costs) are needed for the target domains. Comprehensive experimental results on six downstream domains demonstrate the effectiveness and generalizability of our proposed D-CPT Law and Cross-Domain D-CPT Law.
Haoran Que, Ge Zhang 0009, Xingwei Qu, Yinghao Ma, Feiyu Duan, Zhiqi Bai, Jiakai Wang, Yuanxing Zhang, Xu Tan 0003, Jie Fu 0001, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007
NeurIPS10
2024 Adaptive Modulus RF Beamforming for Enhanced Self-Interference Suppression in Full-Duplex Massive MIMO Systems
abstract
This study employs a uniform rectangular array (URA) sub-connected hybrid beamforming (SC-HBF) architecture to provide a novel self-interference (SI) suppression scheme in a full-duplex (FD) massive multiple-input multiple-output (mMIMO) system. Our primary objective is to mitigate the strong SI through the design of RF beamforming stages for uplink and downlink transmissions that utilize the spatial degrees of freedom provided due to the use of large array structures. We propose a non-constant modulus RF beamforming (NCM-BF-SIS) scheme that incorporates the gain controllers for both transmit (Tx) and receive (Rx) RF beamforming stages and optimizes the uplink and downlink beam directions jointly with gain controller coefficients. To solve this challenging non-convex optimization problem, we propose a swarm intelligence-based algorithmic solution that finds the optimal beam perturbations while also adjusting the Tx/Rx gain controllers to alleviate SI subject to the directivity degradation constraints for the beams. The data-driven analysis based on the measured SI channel in an anechoic chamber shows that the proposed NCM-BF-SIS scheme can suppress SI by around 80 dB in FD mMIMO systems.
Mobeen Mahmood, Yuanxing Zhang, Robert Morawski, Tho Le-Ngoc
WCNC2
2023 TAG: Joint Triple-Hierarchical Attention and GCN for Review-Based Social Recommender System
abstract
Recommender systems across many Internet services have become a critical part of online businesses, as consumers would refer to them before making decisions. However, the lack of explicit ratings for items on many services makes it challenging to capture user preferences and item characteristics. Both academia and the industry have drawn attention to rating predications as a fundamental problem in recommendation systems. With the emergence of social networks, social recommender systems have been proposed to utilize the relationship between users and items to alleviate the data sparsity problem for rating predictions. However, they either concentrate on the opinion mining for each user and item, or consider the connections between users only. In this paper, we present an effective framework, Triple-hierarchical Attention Graph-based social rating prediction (TAG), to exploit the social relationships between users, the user-item interest relationships, the correlation relationships between items, and reviews for rating predictions. In order to consider opinions from reviews and these complex relationships, we first employ two triple-hierarchical attention to extract user and item features from reviews. We then design an inductive GNN, which generates effective embedding for users and items. Experiments over Yelp show that TAG outperforms state-of-the-art methods across RMSE, MAE, and NDCG metrics.
Pengpeng Qiao, Zhiwei Zhang 0002, Zhetao Li, Yuanxing Zhang, Kaigui Bian, Yanzhou Li, Guoren Wang
IEEE Trans. Knowl. Data Eng.4
2022 PICASSO: Unleashing the Potential of GPU-centric Training for Wide-and-deep Recommender Systems
abstract
The development of personalized recommendation has significantly improved the accuracy of information matching and the revenue of e-commerce platforms. Recently, it has two trends: 1) recommender systems must be trained timely to cope with ever-growing new products and ever-changing user interests from online marketing and social network; 2) state-of-the-art recommendation models introduce deep neural network (DNN) modules to improve prediction accuracy. Traditional CPU-based recommender systems cannot meet these two trends, and GPU-centric training has become a trending approach. However, we observe that GPU devices in training recommender systems are underutilized, and they cannot attain an expected throughput improvement as what it has achieved in Computer Vision (CV) and Neural Language Processing (NLP) areas. This issue can be explained by two characteristics of these recommendation models: First, they contain up to a thousand of input feature fields, introducing fragmentary and memory-intensive operations; Second, the multiple constituent feature interaction submodules introduce substantial small-sized compute kernels. To remove this roadblock to the development of recommender systems, we propose a novel framework named PICASSO to accelerate the training of recommendation models on commodity hardware. Specifically, we conduct a systematic analysis to reveal the bottlenecks encountered in training recommendation models. We leverage the model structure and data distribution to unleash the potential of hardware through our packing, interleaving, and caching optimization. Experiments show that PICASSO increases the hardware utilization by an order of magnitude on the basis of state-of-the-art baselines and brings up to 6× throughput improvement for a variety of industrial recommendation models. Using the same hardware budget in production, PICASSO on average shortens the walltime of daily training tasks by 7 hours, significantly reducing the delay of continuous delivery.
Yuanxing Zhang, Langshi Chen, Siran Yang, Man Yuan, Huimin Yi, Jie Zhang 0135, Jiamang Wang, Jianbo Dong, Yong Li 0045, Di Zhang 0026, Wei Lin 0016, Lin Qu, Bo Zheng 0007
ICDE1
2022 GBA: A Tuning-free Approach to Switch between Synchronous and Asynchronous Training for Recommendation Models
abstract
High-concurrency asynchronous training upon parameter server (PS) architecture and high-performance synchronous training upon all-reduce (AR) architecture are the most commonly deployed distributed training modes for recommendation models. Although synchronous AR training is designed to have higher training efficiency, asynchronous PS training would be a better choice for training speed when there are stragglers (slow workers) in the shared cluster, especially under limited computing resources. An ideal way to take full advantage of these two training modes is to switch between them upon the cluster status. However, switching training modes often requires tuning hyper-parameters, which is extremely time- and resource-consuming. We find two obstacles to a tuning-free approach: the different distribution of the gradient values and the stale gradients from the stragglers. This paper proposes Global Batch gradients Aggregation (GBA) over PS, which aggregates and applies gradients with the same global batch size as the synchronous training. A token-control process is implemented to assemble the gradients and decay the gradients with severe staleness. We provide the convergence analysis to reveal that GBA has comparable convergence properties with the synchronous training, and demonstrate the robustness of GBA the recommendation models against the gradient staleness. Experiments on three industrial-scale recommendation tasks show that GBA is an effective tuning-free approach for switching. Compared to the state-of-the-art derived asynchronous training, GBA achieves up to 0.2% improvement on the AUC metric, which is significant for the recommendation models. Meanwhile, under the strained hardware resource, GBA speeds up at least 2.4x compared to synchronous training.
Wenbo Su, Yuanxing Zhang, Yufeng Cai, Kaixu Ren, Pengjie Wang 0002, Huimin Yi, Hongbo Deng, Jian Xu 0015, Lin Qu, Bo Zheng 0007
NeurIPS2
2021 AMEIR: Automatic Behavior Modeling, Interaction Exploration and MLP Investigation in the Recommender System
abstract
Recently, deep learning models have been widely explored in recommender systems. Though having achieved remarkable success, the design of task-aware recommendation models usually requires manual feature engineering and architecture engineering from domain experts. To relieve those efforts, we explore the potential of neural architecture search (NAS) and introduce AMEIR for Automatic behavior Modeling, interaction Exploration and multi-layer perceptron (MLP) Investigation in the Recommender system. Specifically, AMEIR divides the complete recommendation models into three stages of behavior modeling, interaction exploration, MLP aggregation, and introduces a novel search space containing three tailored subspaces that cover most of the existing methods and thus allow for searching better models. To find the ideal architecture efficiently and effectively, AMEIR realizes the one-shot random search in recommendation progressively on the three stages and assembles the search results as the final outcome. The experiment over various scenarios reveals that AMEIR outperforms competitive baselines of elaborate manual design and leading algorithmic complex NAS methods with lower model complexity and comparable time cost, indicating efficacy, efficiency, and robustness of the proposed method.
Kecheng Xiao, Yuanxing Zhang, Kaigui Bian, Wei Yan 0007
IJCAI3
2021 Enhanced review-based rating prediction by exploiting aside information and user influence
Shiwen Wu, Yuanxing Zhang, Wentao Zhang 0001, Kaigui Bian, Bin Cui 0001
Knowl. Based Syst.2
2021 EPASS360: QoE-Aware 360-Degree Video Streaming Over Mobile Devices
abstract
The 360-degree video streaming system delivers a monocular panoramic video surrounding the user, and the user can change the viewing direction of mobile devices to see different parts of the video through the “viewport”. Due to the limited network bandwidth, playbacks of high-resolution 360-degree videos often suffer from rebuffering, while too much bandwidth is wasted in delivering those out-of-viewport parts that the user never watches. In this article, we present an Ensemble Prediction and Allocation based Streaming System, named as EPASS360, for delivering high Quality of Experience (QoE) 360-degree videos. The prediction model takes advantages of ensemble learning, providing high accuracy on the prediction of viewports. The allocation model divides a video into tiles, and allocates high resolution to tiles where a user's viewpoint may appear in the future by solving the QoE-aware optimization problem. Trace-driven emulation on real-world datasets shows that EPASS360 enhances the QoE in various scenarios compared to state-of-the-art streaming approaches. Experiments on the head-mounted device and the hand-held device over real-world Internet confirm the high user experience of EPASS360.
Yuanxing Zhang, Yushuo Guan, Kaigui Bian, Yunxin Liu 0001, Hu Tuo, Lingyang Song, Xiaoming Li 0001
IEEE Trans. Mob. Comput.1
2020 CFP: A Cross-layer Recommender System with Fine-grained Preloading for Short Video Streaming at Network Edge
abstract
Nowadays, short video feed has attracted billions of mobile users all around the world to interact with content effortlessly, yielding an explosive growth of short video commerce. Typically, users watch full-screen short videos of a few seconds one-by-one in a watch-list generated by recommender systems, skipping those they are not interested in. However, the recommender system at the cloud makes a user-interest-specific decision mostly based on the users' behavior data collected within the application itself (e.g., users' view history), without examining the lower-layer network and communication statistics. When the playback choked due to the limited network bandwidth, the user will probably skip the video, leading to a waste of bandwidth and degradation of the user's quality of experience (QoE). Meanwhile, the excessive number of user requests to video contents raises a heavy computational load and communication cost for the recommender system at the cloud to determine which videos to be recommended and delivered to each user in a real-time manner. The advance of edge computing provides a promising avenue of deploying edge nodes with caches (e.g., household devices) beyond cloud and edge servers, such that the recommender system in the cloud can place popular video contents closer to client users, and meanwhile the contents are delivered to client users with good network condition. In this paper, we propose CFP, a cross-layer recommender system for short video streaming with fine-grained preloading technique at the network edge. CFP jointly optimizes the recommendation effect of the video application and the content preloading efficiency under various network conditions at the network edge. CFP takes a two-stage approach: the cloud server first seeks to perform edge-wise instead of user-interest-specific recommendation with neural collaborative filtering recommender, preloading a list of candidate videos to edge nodes, and each edge node, deploying the GRU with attention, then delivers the proper video contents to the client user device according to the user's preference. Trace-driven emulations demonstrate the efficiency of the proposed CFP scheme.
Dezhi Ran, Yuanxing Zhang, Kaigui Bian
CLOUD2
2020 Spherical Criteria for Fast and Accurate 360° Object Detection
abstract
With the advance of omnidirectional panoramic technology, 360◦ imagery has become increasingly popular in the past few years. To better understand the 360◦ content, many works resort to the 360◦ object detection and various criteria have been proposed to bound the objects and compute the intersection-over-union (IoU) between bounding boxes based on the common equirectangular projection (ERP) or perspective projection (PSP). However, the existing 360◦ criteria are either inaccurate or inefficient for real-world scenarios. In this paper, we introduce a novel spherical criteria for fast and accurate 360◦ object detection, including both spherical bounding boxes and spherical IoU (SphIoU). Based on the spherical criteria, we propose a novel two-stage 360◦ detector, i.e., Reprojection R-CNN, by combining the advantages of both ERP and PSP, yielding efficient and accurate 360◦ object detection. To validate the design of spherical criteria and Reprojection R-CNN, we construct two unbiased synthetic datasets for training and evaluation. Experimental results reveal that compared with the existing criteria, the two-stage detector with spherical criteria achieves the best mAP results under the same inference speed, demonstrating that the spherical criteria can be more suitable for 360◦ object detection. Moreover, Reprojection R-CNN outperforms the previous state-of-the-art methods by over 30% on mAP with competitive speed, which confirms the efficiency and accuracy of the design.
Ansheng You, Yuanxing Zhang, Jiaying Liu 0001, Kaigui Bian, Yunhai Tong
AAAI3
2020 Differentiable Feature Aggregation Search for Knowledge Distillation
Yushuo Guan, Bingxuan Wang, Yuanxing Zhang, Cong Yao, Kaigui Bian, Jian Tang 0008
ECCV (17)4
2020 Preference-Aware Mask for Session-Based Recommendation with Bidirectional Transformer
abstract
User profiles are not always visible in E-commerce scenarios, in which case the recommender systems can only summarize users' preferences through sessions of historical records. However, the items in a session might be irrelevant to users' preferences or become the disturbances for modelling the users' portraits, and thus degrade the performance of the recommender systems. In this paper, we propose the preference-aware mask to capture user preferences over the items within the sessions, which adapts to the preference-irrelevant items within the sessions and provides explainable evidence for the recommendation. Evaluation over three real-world datasets verifies that MBTREC performs well on the new-item recommendation task, and outperforms several state-of-the-art recommender systems on the general metrics.
Yuanxing Zhang, Yushuo Guan, Lin Chen 0003, Kaigui Bian, Lingyang Song, Bin Cui 0001, Xiaoming Li 0001
ICASSP1
2020 TSSRGCN: Temporal Spectral Spatial Retrieval Graph Convolutional Network for Traffic Flow Forecasting
abstract
Traffic flow forecasting is of great significance for improving the efficiency of transportation systems and preventing emergencies. Due to the highly non-linearity and intricate evolutionary patterns of short-term and long-term traffic flow, existing methods often fail to take full advantage of spatial-temporal information, especially the various temporal patterns with different period shifting and the characteristics of road segments. Besides, the globality representing the absolute value of traffic status indicators and the locality representing the relative value have not been considered simultaneously. This paper proposes a neural network model that focuses on the globality and locality of traffic networks as well as the temporal patterns of traffic data. The cycle-based dilated deformable convolution block is designed to capture different time-varying trends on each node accurately. Our model can extract both global and local spatial information since we combine two graph convolutional network methods to learn the representations of nodes and edges. Experiments on two real-world datasets show that the model can scrutinize the spatial-temporal correlation of traffic data, and its performance is better than the compared state-of-the-art methods. Further analysis indicates that the locality and globality of the traffic networks are critical to traffic flow prediction and the proposed TSSRGCN model can adapt to the various temporal traffic patterns.
Xu Chen 0022, Yuanxing Zhang, Lun Du, Zheng Fang 0007, Kaigui Bian, Kunqing Xie
ICDM2
2020 MA360: Multi-Agent Deep Reinforcement Learning Based Live 360-Degree Video Streaming on Edge
abstract
The mobile edge caching has made video service providers deliver live 360-degree videos worldwide. However, these services still suffer from the huge network traffic on the core network due to the spherical nature and the diverse requests generated from large user populations. It is challenging to optimize the Quality of Experience (QoE) and the bandwidth consumption simultaneously under the significant number of users as well as dynamic network and playback status. In this paper, we propose a Multi-Agent deep reinforcement learning based 360-degree video streaming system, named MA360, to tackle this multi-user live 360-degree video streaming problem in the context of the edge cache network. Specifically, MA360 employs the Mean Field Actor-Critic (MFAC) algorithm to make clients collaboratively and distributively request tiles aiming at maximizing the overall QoE while minimizing the total bandwidth consumption. Experiments over real-world datasets show that MA360 can improve the QoE while significantly reducing the bandwidth consumption compared with several state-of-the-art edge-assisted 360-degree video streaming strategies.
Yixuan Ban, Yuanxing Zhang, Haodan Zhang, Xinggong Zhang, Zongming Guo
ICME2
2020 Adversarial Oracular Seq2seq Learning for Sequential Recommendation
abstract
Recently, sequential recommendation has become a significant demand for many real-world applications, where the recommended items would be displayed to users one after another and the order of the displays influences the satisfaction of users. An extensive number of models have been developed for sequential recommendation by recommending the next items with the highest scores based on the user histories while few efforts have been made on identifying the transition dependency and behavior continuity in the recommended sequences. In this paper, we introduce the Adversarial Oracular Seq2seq learning for sequential Recommendation (AOS4Rec), which formulates the sequential recommendation as a seq2seq learning problem to portray time-varying interactions in the recommendation, and exploits the oracular learning and adversarial learning to enhance the recommendation quality. We examine the performance of AOS4Rec over RNN-based and Transformer-based recommender systems on two large datasets from real-world applications and make comparisons with state-of-the-art methods. Results indicate the accuracy and efficiency of AOS4Rec, and further analysis verifies that AOS4Rec has both robustness and practicability for real-world scenarios.
Tianxiao Shui, Yuanxing Zhang, Kecheng Xiao, Kaigui Bian
IJCAI3
2020 PERM: Neural Adaptive Video Streaming with Multi-path Transmission
abstract
The multi-path transmission techniques enable multiple paths to maximize resource usage and increase throughput in transmission, which have been installed over mobile devices in recent years. For video streaming applications, compared to the single-path transmission, the multi-path techniques can establish multiple subflows simultaneously to extend the available bandwidth for streaming high-quality videos in mobile devices. Existing adaptive video streaming systems have difficulty in harnessing multi-path scheduling and balancing the tradeoff between the quality of experience (QoE) and quality of service (QoS) concerns. In this paper, we propose an actor-critic network based on Periodical Experience Replay for Multi-path video streaming (PERM). Specifically, PERM employs two actor modules and a critic module: the two actor modules respectively assign the path usage of each subflow and select bitrates for the next chunk of the video, while the critic module predicts the overall objectives. We conduct trace-driven emulation and real-world testbed experiment to examine the performance of PERM, and results show that PERM outperforms state-of-the-art multi-path and single path streaming systems, with an improvement of 10%- 15% on the QoE and QoS metrics.
Yushuo Guan, Yuanxing Zhang, Bingxuan Wang, Kaigui Bian, Xiaoliang Xiong, Lingyang Song
INFOCOM2
2020 Improving Quality of Experience by Adaptive Video Streaming with Super-Resolution
abstract
Given high-speed mobile Internet access today, audiences are expecting much higher video quality than before. Video service providers have deployed dynamic video bitrate adaptation services to fulfill such user demands. However, legacy video bitrate adaptation techniques are highly dependent on the estimation of dynamic bandwidth, and fail to integrate the video quality enhancement techniques, or consider the heterogeneous computing capabilities of client devices, leading to low quality of experience (QoE) for users. In this paper, we present a super-resolution based adaptive video streaming (SRAVS) framework, which applies a Reinforcement Learning (RL) model for integrating the video super-resolution (VSR) technique with the video streaming strategy. The VSR technique allows clients to download low bitrate video segments, reconstruct and enhance them to high-quality video segments while making the system less dependent on estimating dynamic bandwidth. The RL model investigates both the playback statistics and the distinguishing features related to the client-side computing capabilities. Trace-driven emulations over real-world videos and bandwidth traces verify that SRAVS can significantly improve the QoE for users compared to the state-of-the-art video streaming strategies with or without involving VSR techniques.
Yinjie Zhang, Yuanxing Zhang, Bill Tao, Kaigui Bian, Pan Zhou 0001, Lingyang Song, Hu Tuo
INFOCOM2
2020 SSR: Joint Optimization of Recommendation and Adaptive Bitrate Streaming for Short-form Video Feed
abstract
Short-form video feed has become one of the most popular ways for billions of users to interact with content, where users watch short-form videos of a few seconds one-by-one in a session. The common solution to improve the quality of experience (QoE) for short-form video feed is to treat it as a common sequential item recommendation problem and maximize its click-through rate prediction. However, the QoE of short-form video streaming under dynamic network conditions is jointly determined by both recommendation accuracy and streaming efficiency, and thus merely considering recommendation will lead to the degradation of the QoE of the streaming system for the audience. In this paper, we propose SSR, namely the short-form video streaming and recommendation system, which consists of a Transformer-based recommendation module and a reinforcement learning (RL) based bitrate adaptation streaming module. Specifically, we use Transformer to encode the session into a representation vector and recommend proper short-form videos based on the user's recent interest and the timeliness characteristics of short-form video contents. Then, the RL module combines the representation of session and other observations within the playback, and yields the appropriate bitrate allocation for the next short-form video to optimize a given QoE objective. Trace-driven emulations verify the efficiency of SSR compared to several state-of-the-art recommender systems and streaming strategies with at least 10%-15% QoE improvement under various QoE objectives.
Dezhi Ran, Yuanxing Zhang, Wenhan Zhang 0004, Kaigui Bian
MSN2
2020 GARG: Anonymous Recommendation of Point-of-Interest in Mobile Networks by Graph Convolution Network
abstract
Abstract The advances of mobile equipment and localization techniques put forward the accuracy of the location-based service (LBS) in mobile networks. One core issue for the industry to exploit the economic interest of the LBSs is to make appropriate point-of-interest (POI) recommendation based on users’ interests. Today, the LBS applications expect the recommender systems to recommend the accurate next POI in an anonymous manner, without inquiring users’ attributes or knowing the detailed features of the vast number of POIs. To cope with the challenge, we propose a novel attentive model to recommend appropriate new POIs for users, namely Geographical Attentive Recommendation via Graph (GARG), which takes full advantage of the collaborative, sequential and content-aware information. Unlike previous strategies that equally treat POIs in the sequence or manually define the relationships between POIs, GARG adaptively differentiates the relevance of POIs in the sequence to the prediction, and automatically identifies the POI-wise correlation. Extensive experiments on three real-world datasets demonstrate the effectiveness of GARG and reveal a significant improvement by GARG on the precision, recall and mAP metrics, compared to several state-of-the-art baseline methods.
Shiwen Wu, Yuanxing Zhang, Chengliang Gao, Kaigui Bian, Bin Cui 0001
Data Sci. Eng.2
2019 Multi-view Moments Embedding Network for 3D Shape Recognition
abstract
Benefited from rapid developments of deep learning, 3D shape recognition has become a remarkable subject in computer vision systems.The existing methods of multi-perspective views have shown competitive performance in 3D shape recognition.However, they have not yet fully exploited the information among all views of projection.In this paper, we propose a novel Multi-view Moments Embedding Network(MMEN) for capturing multiple moments information.MMEN obtains the similarity between different views and retains the description of the original view by generating moments matrix for representing the general features of the 3D shape.Additionally, we apply the matrix square-root layer to perform a non-linear scaling to the eigenvalues of the moment embedding matrix.We compare the performance of our proposed network with several state-of-the-art models on the ModelNet datasets, and the results of the average instance/class accuracy demonstrate the promising performance of MMEN on 3D shape recognition.
Yuanxing Zhang, Kecheng Xiao, Kaigui Bian, Chunli Zhang, Wei Yan 0007
CIKM2
2019 Adversarial Learning of Transitive Semantic Features for Cross-Domain Recommendation
abstract
In the era of big data, recommender systems have become the key part of many Internet applications. One successful recommendation strategy is to jointly recommend items from different domains where the system can model an accurate portrait of user behaviors. However, it is still challenging to identify the correlation among various domains and make efficient utilization of features from each domain. In this paper, we propose a novel framework, called Domain Adversarial Cross-Domain Recommendation (DACDR), to learn the implicit transitive semantic features among various information relevant domains. The framework automatically retrieves semantic features from both the source and the target domains, and adaptively learns the transitive latent factors to connect the two domains. The user behaviors are then modelled by the learnt latent factors, based on which DACDR can provide an accurate recommendation. Evaluation over real-world dataset verifies that the proposed framework outperforms the state-of-the-art algorithms in terms of F1, NDCG and MRR metrics.
Zhetao Li, Pengpeng Qiao, Yuanxing Zhang, Kaigui Bian
GLOBECOM3
2019 Learning Multiple Temporal Relational Network Embeddings via Graph Convolutional Network
abstract
In the era of big data, information on relationships changes along with time, and the graphs of relationships captured at consecutive timestamps form the multiple temporal relational (MTR) network. To identify the relations in the network while preserving the network structure, a common solution is to learn the network representations through network embedding methods, and then build the relations upon the similarity among these representations. However, the existing network embedding methods either focus on a single relation or ignore the correlation between the heterogeneous and homogeneous relations, and thus it is difficult to investigate the multiple temporal features in the network. In this paper, we propose a novel network embedding method, named Homo- Hetero Network Embedding (HHNE), for the MTR networks. The HHNE utilizes the Graph Convolutional Network (GCN) to extract homogeneous features from each temporal relational network and then generates the homo-hetero network embeddings by fusing the single temporal relational features through a Multi- Layer Perceptron (MLP). Therefore, HHNE could capture the multi- dimensional characteristics in the network, including both intra-relation information and inter-relation information. To show the efficiency of HHNE, we conduct experiments in a real-world dataset on predicting new relations in the MTR networks. The result reveals that our method could outperform several legacy network embedding methods and state- of-the-art multi-relational network embedding methods in the task of relation prediction, demonstrating that the proposed embedding method is more suitable for the MTR network.
Kecheng Xiao, Yuanxing Zhang, Yanzhou Li, Kaigui Bian, Wei Yan 0007
GLOBECOM3
2019 DenXFPN: Pulmonary Pathologies Detection Based on Dense Feature Pyramid Networks
abstract
Computer-aided detection and diagnosis (CAD) have been applied to many departments of medical institutions, and early detection of diseases can prevent serious health loss. Pulmonary diseases generate negative effects on human health, even leading to death. The chest X-ray is a common examination for diagnosis of pulmonary diseases. The experienced radiologist can quickly infer patients' symptoms by screening the chest X-ray image. While in some developing countries or remote rural areas, due to the lack of experienced radiologists or doctors, patients may be misdiagnosed. Many efforts have been spent on developing an effective auxiliary detection system to provide medical workers with evidence on diseases. In particular, detecting pulmonary complication via chest X-ray images is one of the most challenging tasks. In this paper, we transform the pulmonary complications detection task into a multi-binary classification task for each pulmonary pathology, and propose a new classification model, DenXFPN (for X-ray). DenXFPN combines multiple feature maps at different scales extracted through a densely convolutional neural network. Our model achieves 0.827 on the area under the receiver operating characteristic curve (AUC) metric on average, which outperforms the state-of-the-art results on most of all pathologies in the Chest X-ray14 dataset.
Yuanxing Zhang, Kaigui Bian, Guopeng Zhou, Wei Yan 0007
ICASSP2
2019 LadderNet: Knowledge Transfer Based Viewpoint Prediction in 360◦ Video
abstract
In the past few years, virtual reality (VR) has become an enabling technique, not only for enriching our visual experience but also for providing new channels for businesses. Untethered mobile devices are the main players for watching 360-degree content, thereby the precision of predicting the future viewpoints is one key challenge to improve the quality of the playbacks. In this paper, we investigate the image features of the 360-degree videos and the contextual information of the viewpoint trajectories. Specifically, we design ladder convolution to adapt for the distorted image, and propose LadderNet to transfer the knowledge from the pre-trained model and retrieve the features from the distorted image. We then combine the image features and the contextual viewpoints as the inputs for long short-term memory (LSTM) to predict the future viewpoints. Our approach is compared with several state-of-the-art viewpoint prediction algorithms over two 360-degree video datasets. Results show that our approach can improve the Intersection over Union (IoU) by at least 5% and meeting the requirements of the playback of 360-degree video on mobile devices.
Yuanxing Zhang, Kaigui Bian, Hu Tuo, Lingyang Song
ICASSP2
2019 DRL360: 360-degree Video Streaming with Deep Reinforcement Learning
abstract
360-degree videos have gained more popularity in recent years, owing to the great advance of panoramic cameras and head-mounted devices. However, as 360-degree videos are usually in high resolution, transmitting the content requires extremely high bandwidth. To protect the Quality of Experience (QoE) of users, researchers have proposed tile-based 360-degree video streaming systems that allocate high/low bit rates to selected tiles of video frames for streaming over the limited bandwidth. It is challenging to determine which tiles should be allocated with a high/low rate, because (1) the video playbacks include too many features that dynamically change over time when making the rate allocation; (2) most of the state-of-the-art systems focus on a fixed set of heuristics to optimize a specific QoE objective, while users may have various QoE objectives that need to be optimized in different ways. This paper presents a Deep Reinforcement Learning (DRL) based framework for 360-degree video streaming, named DRL360. The DRL360 framework helps improve the system performance by jointly optimizing multiple QoE objectives across a broad set of dynamic features. The DRL-based model adaptively allocates rates for the tiles of the future video frames based on the observations collected by client video players. We compare the proposed DRL360 to the existing systems by trace-driven evaluations as well as conducting a realworld experiment over a wide variety of network conditions. Evaluation results reveal that DRL360 can adapt to all considered scenarios, and outperform the state-of-the-art approaches by 20%-30% on average given different QoE objectives.
Yuanxing Zhang, Kaigui Bian, Yunxin Liu 0001, Lingyang Song, Xiaoming Li 0001
INFOCOM1
2019 Addressing the Conflict of Negative Feedback and Sampling for Online Ad Recommendation in Mobile Social Networks
abstract
Online advertisement (ad) recommendation in the mobile social network (MSN) is an uprising interest of research. Compared to traditional recommendation systems, one of its major difference is the presence of explicit negative feedback from users (e.g., a user does not click an ad, or she/he does not like it). On the other hand, most methods utilize negative sampling (e.g., randomly sampling an item that a user never interacts with to avoid overfitting, that is, she/he is assumed to dislike it) while training conventional recommendation systems. This may lead to a conflict between negative feedback and sampling, as they should be treated differently, but they are considered as the same if traditional methods are directly applied for online ad recommendation. In this paper, we present AdRec, a novel framework of online ad recommendation in MSN to address this conflict. We introduce an auxiliary output and modify the loss function to assign different weights to negative samples and feedbacks. A theoretical analysis is applied to show the efficiency of our design, and experiments on real world datasets demonstrate that our proposed method outperforms several state-of-the-art approaches.
Bill Tao, Yuanxing Zhang, Jianing Lin, Kaigui Bian
MSN2
2018 Competitive Influence Blocking in Online Social Networks: A Case Study on WeChat
abstract
Rapid development of online social networks further facilitates the expansion of information over the Internet. In many scenarios, several pieces of different or even opposite information could diffuse competitively at the same time. Recently, some messenger APPs such as WeChat arose, where users can send links in their "WeChat Moments (WM)", which is called the messenger-based social network (Msg-SN). The unique characteristics in Msg-SN may lead to great difference on the network topology compared to conventional social networks. It is impossible to find the key opinion leaders with millions of followers to help block the rumors or to locate the source of rumors. Thus, the business often tries to select and inject a set of users in the network with the truth to block negative influence diffusion. We call this problem as the competitive influence blocking (CIB) problem. In this paper, we study the CIB problem in Msg-SN under the competitive linear threshold (CLT) model. We propose a fast heuristic algorithm based on eigenvector centrality to optimize the negative influence reduction by selecting a positive seed set. The experimental results using real-world WeChat Moments data show that our algorithm performs better than the effective CLDAG algorithm and also runs faster.
Yuanxing Zhang, Kaigui Bian, Lingyang Song
APCC2
2018 SIGN: War-Driving Free Indoor Navigation Using Coded Visual Tags
abstract
Recent advance in Internet-of-Things (IoT) brings consumer- level smart mobile robot to our life. Indoor navigation is one of the most critical challenges for mobile robots. Existing approaches using wireless signal fingerprinting (e.g., WiFi fingerprint), computer vision techniques, require extensive war-driving of the indoor environment to collect sufficient environmental data. In this paper, we present SIGN, a lightweight, visual-tag based, indoor navigation approach that is free of indoor war-driving. The approach deploys a set of coded visual tags in the environment, and allows the robot to autonomously decide the moving direction by recognizing nearby tags and leveraging the geometry information. The proposed approach is robust to the change of the environment such as unexpected obstacles. Experiments in two indoor spaces under various scenarios show that SIGN helps the mobile robot using the off-the- shelf camera to self-navigate in indoor environment with the deployment of coded visual tags.
Yuanxing Zhang, Zhuojin Li, Chengxu Yang, Kaigui Bian, Lingyang Song, Xiaoming Li 0001
GLOBECOM1
2018 Optimal Trajectory Planning of Drones for 3D Mobile Sensing
abstract
Mobile sensing is challenging in 3D space, as there are many inaccessible places where people rarely venture. Unmanned aerial vehicle (UAV), commonly known as drone, has greatly extended the scope of mobile sensing in 3D space, and pushed forward a variety of 3D mobile sensing applications, such as aerial photo- or video-graphy, 3D wireless signal survey, and air quality monitoring. However, the short battery life of drones has largely restricted the wide adoption of these applications. In this paper, we study the trajectory planning problem for optimizing the flight route in a given sensing space. We first divide the 3D space into an infinite three-dimensional network of observation locations (OLs), and model the sensing scope as a finite subgraph of 3D OL network. We formulate the problem as finding the optimal trajectory in the sensing scope. We propose an algorithm that finds trajectory in each divided 3D grid of the sensing scope by generating a nearly optimal dominating path, and finding the minimum dominating set in the dominating path. Then, we concatenate obtained trajectories in 3D grids to a nearly optimal trajectory in the sensing scope. Experimental results show that the proposed algorithm takes 24% less time to complete sensing the given space, and during the battery life it can cover 19% more sensing scope, than existing solutions.
Yuzhe Yang 0003, Yuanxing Zhang, Kaigui Bian, Lingyang Song, Pengpeng Qiao, Zhetao Li
GLOBECOM3
2018 On Lifecycle of Interactive Web Apps in WeChat
abstract
WeChat is the largest mobile instant messaging service in China, where users can send messages to friends or post them over their walls (a.k.a. friend circle, or WeChat Moments). Interactive web apps are quite attractive for businesses, institutes, or individuals to promote products or events. In this paper, we analyze the diffusion statistics of interactive web apps in WeChat and conduct an empirical measurement study over a dataset with 54 million users and 20 thousand web apps crawled. We discover the lifecycle of interactive web apps varies drastically, which is largely dependent on the content, date, time upon the first release, and the social influence of viewers and senders. Meanwhile, we develop a model based on the matrix factorization method to extract latent features of interactive web apps and an app lifecycle model that characterizes how the features affect apps' lifecycle, achieving the mean absolute error (MAE) of 2.32 days in predicting app's lifecycle. Our results hold the promise of helping businesses to promote their marketing information dissemination through long-lived interactive web apps at the right timing and with the appropriate content.
Chengliang Gao, Yuanxing Zhang, Kaigui Bian, Shaoling Dong, Lingyang Song
ICC2
2018 ATDPS: An Adaptive Time-Dependent Push Strategy in Hybrid CDN-P2P VoD System
abstract
Online video service has become an emerging application of Internet, playing an important and indispensable role in our daily life. Aiming at relieving the heavy burden on CDN servers, many video service providers deploy a hybrid CDN-P2P system. To minimize the cost of the "last mile delivery" of CDN servers, the system should be able to dispatch video files that receive a large amount of view requests during peak hours to P2P network in advance. However, to accurately predict whether a video will gain popularity is difficult due to the uncertainty of people's viewing tendency. In this paper, we reveal the correlation between the popularity of videos in the daytime and the popularity in peak hours at night based on statistics collected in a real-world large commercial VoD system. Based on our findings, we propose Adaptive Time-Dependent Push Strategy (ATDPS), a scalable push strategy that implements adaptive time-dependent techniques in the hybrid CDN-P2P system. A deep neural network is leveraged to predict the demand of hot videos during peak hours at night. Besides, the system adaptively adjusts the time-dependent strategy of pushing the content of hot videos to the P2P network according to the seed scarcity of the videos in different time periods. Simulations over historical data and pilot deployment on smart routers show that our ATDPS can significantly decrease the addressed bandwidth consumption at peak hours.
Yangze Guo, Yuanxing Zhang, Zhi Yang 0001, Kaigui Bian, Hu Tuo, Yafei Dai
ICC2
2018 An Optimal Design and Analysis of a Hybrid Power Charging Station for Electric Vehicles Considering Uncertainties
abstract
The charging problem becomes prominent with the increasing number of electric vehicles. It is necessary to built charging station (CS), like the gasoline station, to satisfy the recharging and be convenient for the drivers. In this paper, a new type of charging station integrated with renewable energy source was studied. A hierarchical energy management strategy oriented to real-time application was proposed to handle the uncertainties. To determine the optimal size of the CS by considering multiobjective including economic, environment and battery energy storage system degradation, Monte Carlo simulation was adopted to solve the problem with many uncertainties. We treated battery degradation as a specific objective function. And we obtained the optimal Pareto set. The result demonstrated the optimal decision variable for CS sizing can compromise the objectives as well as realize the reasonable resource dispatch.
Taoyong Li, Yuanxing Zhang, Linru Jiang, Dongxiang Yan, Chengbin Ma
IECON3
2018 A Hierarchical Distributed Energy Management for Multiple PV-Based EV Charging Stations
abstract
A hierarchical distributed energy management for multiple photovoltaic (PV) based electric vehicle (EV) charging stations (PV-CSs) is proposed and analyzed in this paper. In the station level, PV-CSs are modelled as independent players with objectives to stabilize their average available capacity (AAC) of the storage battery tank. Meanwhile, in EV level, EV owners are modeled as players with objectives to maximize their charging power. Then a two level power distribution game is utilized to model the power distribution problem in both station and EV level. Through utilizing a consensus network based learning algorithm, a cooperative and a generalized Stackelberg equilibrium are achieved in station and EV level through a distributed fashion. One case studies, i.e., two station case, is implemented in simulation to verify the performance and effectiveness of the proposed strategy. The simulation results show that the proposed energy management has an excellent performance in both cases and comparing against stations without management.
Yuanxing Zhang, Taoyong Li, Linru Jiang, He Yin, Chengbin Ma
IECON2
2018 Towards Reading Comprehension for Long Documents
abstract
Machine reading comprehension has gained attention from both industry and academia. It is a very challenging task that involves various domains such as language comprehension, knowledge inference, summarization, etc. Previous studies mainly focus on reading comprehension on short paragraphs, and these approaches fail to perform well on the documents. In this paper, we propose a hierarchical match attention model to instruct the machine to extract answers from a specific short span of passages for the long document reading comprehension (LDRC) task. The model takes advantages from hierarchical-LSTM to learn the paragraph-level representation, and implements the match mechanism (i.e., quantifying the relationship between two contexts) to find the most appropriate paragraph that includes the hint of answers. Then the task can be decoupled into reading comprehension task for short paragraph, such that the answer can be produced. Experiments on the modified SQuAD dataset show that our proposed model outperforms existing reading comprehension models by at least 20% regarding exact match (EM), F1 and the proportion of identified paragraphs which are exactly the short paragraphs where the original answers locate.
Yuanxing Zhang, Yangbin Zhang, Kaigui Bian, Xiaoming Li 0001
IJCAI1
2018 Proactive Video Push for Optimizing Bandwidth Consumption in Hybrid CDN-P2P VoD Systems
abstract
Decentralizing content delivery to edge devices has become a popular solution for saving the bandwidth consumption of CDN when the CDN bandwidth is expensive. One successful realization is the hybrid CDN-P2P VoD system, where a client is allowed to request video content from a number of seeds (seed clients) in the P2P network. However, the seed scarcity problem may arise for a video resource when there are an insufficient number of seeds to satisfy requests to the video. To alleviate this problem, many commercial VoD systems have employed a video push mechanism that directly sends the recent scarce video resources to randomly-chosen seeds to serve more requests. However, the current video push mechanism fails to consider which videos will become scarce in the future, or differentiate the uploading capability of different seeds. In this paper, we propose Proactive-Push, a video push mechanism that lowers the bandwidth consumption of CDN by predicting future scarce videos and proactively sending them to competent seeds with strong uploading capabilities. Proactive-Push trains neural network models to correctly predict 80% of future scarce video resources, and identify over 90% of competent seeds. We evaluate Proactive-Push using a trace-driven emulation and a real-world pilot deployment over a commercial VoD system. Results show that Proactive-Push can further reduce the proportion of direct download from CDN by 21%, and save the CDN bandwidth cost at peak time by 18%.
Yuanxing Zhang, Chengliang Gao, Yangze Guo, Kaigui Bian, Xin Jin 0008, Zhi Yang 0001, Lingyang Song, Jiangang Cheng, Hu Tuo, Xiaoming Li 0001
INFOCOM1
2017 Trajectory-Matching Prediction for Friend Recommendation in Anonymous Social Networks
abstract
People connect to each other over conventional online social networks (OSNs) based on many parameters (common interests, experiences, locations), while the anonymous social networks (ASNs) recommend candidate friends to a user mainly by the location proximity at a coarse-granularity. In this paper, we formulate a fine-grained trajectory-matching prediction problem for friend recommendation in ASNs-what is the likelihood for two users to encounter with each other in the future based on their historical trajectory data? We define the serendipity of two trajectories in both spatial and temporal domains to quantify the similarity between two users' trajectories, and propose an algorithm that recommends candidate friends to a user by determining the similarity of their trajectories. Experiments show that our proposed algorithm can predict the encounter of users quite accurately and it outperforms the conventional algorithms in terms of the precision and consumed time.
Yichun Duan, Yuanxing Zhang, Chengliang Gao, Meng Tong, Kaigui Bian, Wei Yan 0007
GLOBECOM2
2017 Holiday syndrome: A measurement study of mobile social network use during holidays
abstract
Businesses are interested in marketing over the mobile social network (MSN) during holiday seasons to expand their holiday sales. Understanding the “holiday syndrome” - how people behave during holidays - over the MSN is important to create a successful marketing campaign that delights customers. In this paper, we conduct an empirical measurement study of WeChat Moments (the most popular social network of the mobile messaging app WeChat in China) during the holiday season of Chinese Spring Festival (with 137 million users involved and 329 thousand applications crawled), and present a comprehensive view of the MSN's impact on the social ties among users, user interests, as well as users' migration patterns before, during, and after the holiday. Our research findings suggest that the MSN is predominantly used during holiday seasons for holiday-atmosphere building and experience sharing. It is revealed that there exist strong correlations between the timing and popular topics, and between holiday migration and regional distribution. Our results hold the promise of helping businesses to promote their marketing information dissemination to targeted groups of customers, at the right region, with appealing words, during the holiday season.
Chengliang Gao, Yuanxing Zhang, Kaigui Bian, Zhuojin Li, Yichong Bai, Xuanzhe Liu
ICC2
2017 Reliability research of DC charging module
abstract
LLC resonant converter is most frequently effective DC/DC converter to improve the power density and efficiency of DC charging module in electric vehicle power supply. Soft-switching technique can be applied in the topology to reduce the switch losses. However, DC charging module brings about 27% of failures in Electric vehicles (EV) charge device. The aim of this paper is to evaluate the reliability of a 3.8 kW LLC resonant converter as DC charging module in both component level and system level. The power losses of mainly power components are deduced and the thermal model together with the lifetime model are built to evaluate the reliability of the LLC resonant converter in component level. The converter level reliability is obtained by the component level reliability, Weibull distribution and Reliability Block Diagram (RBD).
Dingjun Zeng, Yuanxing Zhang, Taoyong Li, Mingxuan Qi, Erjie Qi, Guorong Zhu
IECON2
2016 Influence Maximization in Messenger-Based Social Networks
abstract
Many online social networks have provided a messenger app (e.g., facebook messenger, direct message on Twitter) to facilitate communication between strong- tied friends. Meanwhile, some messenger apps (WeChat) also start to offer social-networking services ("WeChat Moments" (WM), a.k.a. friend circle) that allow users to post pictures, texts, links of webpages, on their walls, which is called the messenger-based social network (Msg-SN). In online social networks, Key Opinion Leaders (KOLs) with millions of followers are easy to identify for helping viral marketing/advertising. However, most users of a messenger app have a small number of friends (e.g., hundreds of friends), which makes it challenging to detect a KOL in Msg-SN by only counting the number of his/her friends. In this paper, we study the influence maximization problem in the Msg-SN of finding the set of most influential KOL nodes that maximize the spread of information. We develop a novel efficient approximation algorithm that calculates the influence by looking at the user's local contribution to the information diffusion process, which scales to large datasets with provable near-optimal performance. Experiment results using the real-world WeChat Moments data (on January 14th, 2016, 100 thousand users) show that our algorithms can identify the set of KOL nodes with a low time complexity.
Yuanxing Zhang, Yichong Bai, Lin Chen 0003, Kaigui Bian, Xiaoming Li 0001
GLOBECOM1