VLDB 2026 Research / reviewers in the wild / expert
Di Fu
dblp:135/7221
· DBLP profile ↗
29ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neutral by Default? Replicating User Vocal Responses to Negative Affective Cues in Conversational AgentsabstractConversational agents (CAs) increasingly detect users’ emotions, yet deciding how to respond, especially to negative affect, remains a central design challenge. We conducted a role-switching study in which participants reply as the CAs to simulated users expressing anger, sadness, or fear. Results reveal systematic, gender-linked patterns: most male participants favored a neutral, affect-balanced stance and prioritized clarification or task progress, whereas most female participants produced a wider range of non-neutral responses, more often using explicit empathy, reassurance, and reflective listening. We also observe differences in de-escalation phrasing, validation timing, and follow-up questioning across scenarios. These findings indicate that strategies for handling negative emotions vary with user characteristics and context. Based on these findings, we argue for adaptive CA response policies that calibrate first-turn acknowledgment and information-gathering, tailoring prosody and wording to emotional context in order to support de-escalation, perceived understanding, and user trust. Yong Ma 0003, Yuchong Zhang 0001, Di Fu, Stephanie Zubicueta Portales, Morten Fjeld |
HRI | 3 |
| 2025 | Vote & Mix: Plug-and-Play Token Reduction for Efficient Vision TransformerabstractDespite the remarkable success of Vision Transformers (ViTs) in various visual tasks, they are often hindered by substantial computational cost. In this work, we introduce Vote&Mix (VoMix), a plug-and-play and parameter-free token reduction method, which can be readily applied to off-the-shelf ViT models without any training. VoMix tackles the computational redundancy of ViTs by identifying tokens with high homogeneity through a layer-wise token similarity voting mechanism. Subsequently, the selected tokens are mixed into the retained set, thereby preserving visual information. Experiments demonstrate VoMix significantly improves the speed-accuracy tradeoff of ViTs on both images and videos. Without any training, VoMix achieves a 2× increase in throughput of existing ViT-H on ImageNet-1K and a 2.4× increase in throughput of existing ViT-L on Kinetics-400 video dataset, with a mere 0.3% drop in top-1 accuracy. Shuai Peng, Di Fu, Baole Wei, Liangcai Gao, Zhi Tang 0001 |
ICME | 2 |
| 2025 | SeedLoRA: A Fusion Approach to Efficient LLM Fine-TuningabstractDespite Low-Rank Adaptation (LoRA)’s popularity for fine-tuning large models, it often exhibits a noticeable performance gap compared to full fine-tuning, particularly in complex tasks such as mathematical reasoning and code generation. Motivated by this discrepancy, we propose a novel fusion approach for LoRA fine-tuned models. Our key insight is that LoRA models trained with different random seeds on the same task often exhibit complementary strengths. In contrast to existing research that typically focuses on fusing models trained on diverse tasks, we explore the potential of combining multiple LoRA models fine-tuned on the same task with different random seeds. This intra-task fusion method aims to leverage the strengths of various fine-tuned models to create a more robust and effective adaptation. To validate our approach, we conducted comprehensive experiments across three key areas: mathematical reasoning, code generation, and general instruction-tuning tasks. The results demonstrate that our fusion method significantly enhances LoRA’s performance, outperforming both standalone LoRA models and current fusion methods. Notably, this advancement substantially narrows the gap between LoRA and full fine-tuning, thus offering a more effective approach to model adaptation without the GPU memory burden of full parameter fine-tuning. Yong Liu 0020, Di Fu, Shenggan Cheng, Minhao Cheng, Cho-Jui Hsieh, Yang You 0001 |
ICML | 2 |
| 2025 | Hri-Free: Cognitive Robotic Simulation for Evaluating Embodied Social Attention ModelsabstractScaling social robot studies is constrained due to the need for human interaction, making large participant recruitment impractical. Robotics simulators help mitigate this limitation but generally lack the realism to accurately simulate social cues. We introduce a cognitive robotic simulation scheme to evaluate social attention models in physical environments. By projecting ground-truth priority maps to a simulated environment, we can directly compare predicted maps using common saliency metrics. Using the iCub robot, we assess a dynamic scanpath model that predicts attention targets, simulating human scanpaths. Evaluations with the FindWho and MVVA datasets show strong correlations between robotcaptured metrics and direct-streamed video metrics. Our results indicate robustness of the social attention model to noise and real-world conditions, suggesting its practical usability for predicting personalized scanpaths in real settings. This approach reduces the need for extensive human-robot interaction studies in the early stages of study design, enabling the scalability and reproducibility of social robot evaluations. Fares Abawi, Di Fu |
ICRA | 2 |
| 2025 | Direct Preference Optimization of Video Large Multimodal Models from Language Model RewardabstractRuohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander G Hauptmann, Yonatan Bisk, Yiming Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alex Hauptmann 0001, Yonatan Bisk, Yiming Yang 0002 |
NAACL (Long Papers) | 7 |
| 2025 | Influence of Robots' Voice Naturalness on Trust and ComplianceabstractWith the increasing performance of text-to-speech systems and their generated voices indistinguishable from natural human speech, the use of these systems for robots raises ethical and safety concerns. A robot with a natural voice could increase trust, which might result in over-reliance despite evidence for robot unreliability. To estimate the influence of a robot’s voice on trust and compliance, we design a study that consists of two experiments. In a pre-study ( \(N_{1}=60\) ) the most suitable natural and mechanical voice for the main study are estimated and selected for the main study. Afterward, in the main study ( \(N_{2}=68\) ), the influence of a robot’s voice on trust and compliance is evaluated in a cooperative game of Battleship with a robot as an assistant. During the experiment, the acceptance of the robot’s advice and response time are measured, which indicate trust and compliance, respectively. The results show that participants expect robots to sound human-like and that a robot with a natural voice is perceived as safer. Additionally, a natural voice can affect compliance. Despite repeated incorrect advice, the participants are more likely to rely on the robot with the natural voice. The results do not show a direct effect on trust. Natural voices provide increased intelligibility, and while they can increase compliance with the robot, the results indicate that natural voices might not lead to over-reliance. The results highlight the importance of incorporating voices into the design of social robots to improve communication, avoid adverse effects, and increase acceptance and adoption in society. Dennis Becker, Lukas Braach, Lennart Clasmeier, Teresa Kaufmann, Oskar Ong, Kyra Ahrens, Connor Gaede, Erik Strahl, Di Fu, Stefan Wermter |
ACM Trans. Hum. Robot Interact. | 9 |
| 2024 | Wrapyfi: A Python Wrapper for Integrating Robots, Sensors, and Applications across Multiple MiddlewareabstractMessage oriented and robotics middleware play an important role in facilitating robot control, abstracting complex functionality, and unifying communication patterns between sensors and devices. However, using multiple middleware frameworks presents a challenge in integrating different robots within a single system. To address this challenge, we present Wrapyfi, a Python wrapper supporting multiple message oriented and robotics middleware, including ZeroMQ, YARP, ROS, and ROS 2. Wrapyfi also provides plugins for exchanging deep learning framework data, without additional encoding or preprocessing steps. Using Wrapyfi eases the development of scripts that run on multiple machines, thereby enabling cross-platform communication and workload distribution. We finally present the three communication schemes that form the cornerstone of Wrapyfi's communication model, along with examples that demonstrate their applicability. Fares Abawi, Philipp Allgeuer, Di Fu, Stefan Wermter |
HRI | 3 |
| 2024 | Reward Shaping for Reinforcement Learning with An Assistant Reward AgentabstractReward shaping is a promising approach to tackle the sparse-reward challenge of reinforcement learning by reconstructing more informative and dense rewards. This paper introduces a novel dual-agent reward shaping framework, composed of two synergistic agents: a policy agent to learn the optimal behavior and a reward agent to generate auxiliary reward signals. The proposed method operates as a self-learning approach, without reliance on expert knowledge or hand-crafted functions. By restructuring the rewards to capture future-oriented information, our framework effectively enhances the sample efficiency and convergence stability. Furthermore, the auxiliary reward signals facilitate the exploration of the environment in the early stage and the exploitation of the policy agent in the late stage, achieving a self-adaptive balance. We evaluate our framework on continuous control tasks with sparse and delayed rewards, demonstrating its robustness and superiority over existing methods. Haozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu, Tze-Yun Leong |
ICML | 4 |
| 2023 | The Emotional Dilemma: Influence of a Human-like Robot on Trust and CooperationabstractIncreasing anthropomorphic robot behavioral design could affect trust and cooperation positively. However, studies have shown contradicting results and suggest a task-dependent relationship between robots that display emotions and trust. Therefore, this study analyzes the effect of robots that display human-like emotions on trust, cooperation, and participants’ emotions. In the between-group study, participants play the coin entrustment game with an emotional and a non-emotional robot. The results showthat the robot that displays emotions induces more anxiety than the neutral robot. Accordingly, the participants trust the emotional robot less and are less likely to cooperate. Furthermore, the perceived intelligence of a robot increases trust, while a desire to outcompete the robot can reduce trust and cooperation. Thus, the design of robots expressing emotions should be task dependent to avoid adverse effects that reduce trust and cooperation. Dennis Becker, Diana Rueda, Felix Beese, Brenda Scarleth Gutierrez Torres, Myriem Lafdili, Kyra Ahrens, Di Fu, Erik Strahl, Tom Weber, Stefan Wermter |
RO-MAN | 7 |
| 2023 | The Robot in the Room: Influence of Robot Facial Expressions and Gaze on Human-Human-Robot CollaborationabstractRobot facial expressions and gaze are important factors for enhancing human-robot interaction (HRI), but their effects on human collaboration and perception are not well understood, for instance, in collaborative game scenarios. In this study, we designed a collaborative triadic HRI game scenario where two participants worked together to insert objects into a shape sorter. One participant assumed the role of a guide. The guide instructed the other participant, who played the role of an actor, to place occluded objects into the sorter. A humanoid robot issued instructions, observed the interaction, and displayed social cues to elicit changes in the two participants’ behavior. We measured human collaboration as a function of task completion time and the participants’ perceptions of the robot by rating its behavior as intelligent or random. Participants also evaluated the robot by filling out the Godspeed questionnaire. We found that human collaboration was higher when the robot displayed a happy facial expression at the beginning of the game compared to a neutral facial expression. We also found that participants perceived the robot as more intelligent when it displayed a positive facial expression at the end of the game. The robot’s behavior was also perceived as intelligent when directing its gaze toward the guide at the beginning of the interaction, not the actor. These findings provide insights into how robot facial expressions and gaze influence human behavior and perception in collaboration. Di Fu, Fares Abawi, Stefan Wermter |
RO-MAN | 1 |
| 2023 | Smart-Contract-Based Agricultural Service Platform for Drone Plant Protection Operation OptimizationabstractThe platform-based agricultural service is receiving popularity in small-scale farming and shows significant advantages in gathering dispersed service requests and matching supply and demand. However, it also generates new challenges, including service traceability, denial and fraud, information security, and privacy issues. Blockchain is an emerging technology that provides a secure and trusted environment to track and manage the service process. In this study, we propose a blockchain-based service platform for efficient agricultural service operations with the support of Internet of Things technology. Following the smart platform, we use the drone plant protection service as an example and develop a new execution procedure for smart contract-based agricultural services. In the proposed procedure, we focus on integrating optimization methods to deal with multiple service requests as well as potential disruption events and establishing detailed interactions among service plans, smart contracts, and physical services. Moreover, we formulate the drone plant protection issue using a mixed-integer linear programming model to obtain the optimal service plan and develop a recovery model to deal with potential disruptions of new order arrival. Finally, we design detailed smart contract terms for drone plant protection services. Results of numerical experiments demonstrate the effectiveness of the developed optimization model in obtaining the optimal service plan before and after disruptions. Also, we verify the applicability and security of the smart contract on the Ethereum platform based on a three-phase functional test and a comprehensive security test. Qianqian Zheng 0001, Na Lin 0003, Di Fu, Tianjun Liu, Yuchun Zhu, Xiaochun Feng, Junhu Ruan |
IEEE Internet Things J. | 3 |
| 2022 | Self-reconstructive evidential clustering for high-dimensional dataabstractAlthough many algorithms have been presented to tackle the curse of dimensionality in high-dimensional clustering, most of these algorithms require prior knowledge of the number of clusters. Besides, these existing algorithms create only a hard or fuzzy partition for high-dimensional objects, which are often located in highly overlapping areas. The adoption of hard/fuzzy partition ignores the ambiguity in the assignment of objects and may lead to performance degradation. To address these issues, we propose a novel self-reconstructive evidential clustering (SREC) algorithm. After learning the correlations between objects from a self-reconstruction process, SREC provides a human-readable chart. Through this chart, users can select several objects existing in the dataset as the cluster centers, instead of just detecting the number of clusters. Under the framework of evidence theory, SREC derives a more flexible credal partition that improves the fault tolerance of clustering. Ablation study demonstrates the benefits of the self-reconstruction and evidence theory. Comparison experiments on real-world datasets show that SREC consumes competitive running time and performs better than other state-of-the-art algorithms. We also apply SREC in a real-world application scenario to illustrate the rationality of selecting cluster centers by human intervention. Chaoyu Gong, Di Fu, Yong Liu 0020, Pei-hong Wang, Yang You 0001 |
ICDE | 3 |
| 2022 | Compute Like Humans: Interpretable Step-by-step Symbolic Computation with Deep Neural NetworkabstractNeural network capability in symbolic computation has emerged in much recent work. However, symbolic computation is always treated as an end-to-end blackbox prediction task, where human-like symbolic deductive logic is missing. In this paper, we argue that any complex symbolic computation can be broken down to a sequence of finite Fundamental Computation Transformations (FCT), which are grounded as certain mathematical expression computation transformations. The entire computation sequence represents a full human understandable symbolic deduction process. Instead of studying on different end-to-end neural network applications, this paper focuses on approximating FCT which further build up symbolic deductive logic. To better mimic symbolic computations with math expression transformations, we propose a novel tree representation learning architecture GATE (Graph Aggregation Transformer Encoder) for math expressions. We generate a large-scale math expression transformation dataset for training purpose and collect a real-world dataset for validation. Experiments demonstrate the feasibility of producing step-by-step human-like symbolic deduction sequences with the proposed approach, which outperforms other neural network approaches and heuristic approaches. Shuai Peng, Di Fu, Yijun Liang, Gu Xu, Liangcai Gao, Zhi Tang 0001 |
KDD | 2 |
| 2021 | MusicBERT: A Self-supervised Learning of Music RepresentationabstractMusic recommendation has been one of the most used information retrieval services on internet. Finding suitable music for users' demands from tens of millions of music relies on the understanding of music content. Traditional studies usually focus on music representation based on massive user behavioral data and music meta-data, which ignore the audio characteristic of music. However, it is found that the melodic characteristics of music themselves can be further used to understand music. Moreover, how to utilize large-scale audio data to learn music representation is not well explored. To this end, we propose a self-supervised learning model for music representation. We firstly utilize a beat-level music pre-training model to learn the structure of music. Then, we use a multi-task learning framework to model music self-representation and co-relations between music, concurrently. Besides, we propose several downstream tasks to evaluate music representation, including music genre classification, music highlight, and music similarity retrieval. Extensive experiments on multiple music datasets demonstrate our model's superiority over baselines on learning music representation. Hongyuan Zhu 0001, Ye Niu, Di Fu, Hao Wang 0005 |
ACM Multimedia | 3 |
| 2021 | Reversible data hiding in encrypted medical DICOM image
Ping Kong, Di Fu, Chuan Qin 0001 |
Multim. Syst. | 2 |
| 2018 | Expectation Learning and Crossmodal Modulation with a Deep Adversarial NetworkabstractThe human brain is able to learn, generalize, and predict crossmodal stimuli which help us to understand the world around us. Some characteristics of crossmodal learning inspired some computational models but most of the solutions only go as far as to implement strategies for early or late crossmodal fusion. In this paper, we propose the use of two mechanisms from behavioral psychology to enhance the capabilities of a deep adversarial network to learn crossmodal stimuli: the unity assumption modulation and expectation learning. We use real-world data to train and evaluate our model in a set of experiments and demonstrate how these mechanisms affect the learning behavior of the model and how they contribute to making it learn crossmodal coincident stimuli. Our experiments show that the addition of these two mechanisms modulates the crossmodal binding capabilities of the model and improves the learning of unisensory descriptors. Pablo V. A. Barros, German Ignacio Parisi, Di Fu, Xun Liu 0001, Stefan Wermter |
IJCNN | 3 |
| 2018 | A Neurorobotic Experiment for Crossmodal Conflict Resolution in Complex EnvironmentsabstractCrossmodal conflict resolution is crucial for robot sensorimotor coupling through the interaction with the environment, yielding swift and robust behaviour also in noisy conditions. In this paper, we propose a neurorobotic experiment in which an iCub robot exhibits human-like responses in a complex crossmodal environment. To better understand how humans deal with multisensory conflicts, we conducted a behavioural study exposing 33 subjects to congruent and incongruent dynamic audio-visual cues. In contrast to previous studies using simplified stimuli, we designed a scenario with four animated avatars and observed that the magnitude and extension of the visual bias are related to the semantics embedded in the scene, i.e., visual cues that are congruent with environmental statistics (moving lips and vocalization) induce the strongest bias. We implement a deep learning model that processes stereophonic sound, facial features, and body motion to trigger a discrete behavioural response. After training the model, we exposed the iCub to the same experimental conditions as the human subjects, showing that the robot can replicate similar responses in real time. Our interdisciplinary work provides important insights into how crossmodal conflict resolution can be modelled in robots and introduces future research directions for the efficient combination of sensory observations with internally generated knowledge and expectations. German Ignacio Parisi, Pablo V. A. Barros, Di Fu, Sven Magg, Haiyan Wu, Xun Liu 0001, Stefan Wermter |
IROS | 3 |
| 2018 | Deep Predictive Coding Network with Local Recurrent Processing for Object RecognitionabstractInspired by "predictive coding" - a theory in neuroscience, we develop a bi-directional and dynamic neural network with local recurrent processing, namely predictive coding network (PCN). Unlike feedforward-only convolutional neural networks, PCN includes both feedback connections, which carry top-down predictions, and feedforward connections, which carry bottom-up errors of prediction. Feedback and feedforward connections enable adjacent layers to interact locally and recurrently to refine representations towards minimization of layer-wise prediction errors. When unfolded over time, the recurrent processing gives rise to an increasingly deeper hierarchy of non-linear transformation, allowing a shallow network to dynamically extend itself into an arbitrarily deep network. We train and test PCN for image classification with SVHN, CIFAR and ImageNet datasets. Despite notably fewer layers and parameters, PCN achieves competitive performance compared to classical and state-of-the-art models. Further analysis shows that the internal representations in PCN converge over time and yield increasingly better accuracy in object recognition. Errors of top-down prediction also reveal visual saliency or bottom-up attention. Kuan Han, Haiguang Wen, Yizhen Zhang 0004, Di Fu, Eugenio Culurciello, Zhongming Liu |
NeurIPS | 4 |
| 2017 | On Energy-Efficient Offloading in Mobile Cloud for Real-Time Video ApplicationsabstractBatteries of modern mobile devices remain severely limited in capacity, which makes energy consumption a key concern for mobile applications, particularly for the computation-intensive video applications. Mobile devices can save energy by offloading computation tasks to the cloud, yet the energy gain must exceed the additional communication cost for cloud migration to be beneficial. The situation is further complicated by real-time video applications that have stringent delay and bandwidth constraints. In this paper, we closely examine the performance and energy efficiency of representative mobile cloud applications under dynamic wireless network channels and state-of-the-art mobile platforms. We identify the unique challenges of and opportunities for offloading real-time video applications and develop a generic model for energy-efficient computation offloading accordingly in this context. We propose a scheduling algorithm that makes adaptive offloading decisions in fine granularity in dynamic wireless network conditions and verify its effectiveness through trace-driven simulations. We further present case studies with advanced mobile platforms and practical applications to demonstrate the superiority of our solution and the substantial gain of our approach over baseline approaches. Lei Zhang 0066, Di Fu, Jiangchuan Liu, Edith C. H. Ngai, Wenwu Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Live Broadcast With Community Interactions: Bottlenecks and OptimizationsabstractRecent years have witnessed the rapid growth of new live broadcast services, represented by Twitch.tv and YouTube live events, where videos are crowdsourced from amateur users (e.g., game players), rather than from commercial and professional TV broadcaster or content providers. The viewers also actively contribute to the content through embedded open-chat channels. Such community interactions among viewers, or even between broadcasters and viewers, make content generation highly diversified and engaging, particularly for the young generation. In this context, cross-viewer synchronization is highly desirable; otherwise the viewers with shorter broadcast latency may act as spoilers, significantly affecting the user experience of other viewers. In this paper, we show that the end-to-end delay has a dramatically amplified impact on the broadcast latency for individual viewers. We suggest smart rate adaptation to achieve cross-viewer synchronization, and develop distributed algorithms based on dual decomposition. We further extend our solution to the cloud environment, and present the concept of ShadowCast, which moves broadcasters to the cloud to provide high-quality streams beyond broadcasters' network bandwidth constraint. Its practicability and effectiveness is demonstrated by our implementation and test bed experiments. Xiaoqiang Ma, Cong Zhang 0002, Jiangchuan Liu, Ryan Shea, Di Fu |
IEEE Trans. Multim. | 5 |
| 2016 | Load-aware hybrid scheduling in large compute clustersabstractWith the increasing of workloads in large scale heterogeneous compute clusters, distributed scheduling has won support from the academia and industry because of its inherent scalability and flexibility. However, existing schedulers cannot guarantee that all the jobs are acceptable and the average latency is extremely large. In particular, when the schedulers apply gang scheduling and non-preemptive policy, the serious job starvation problem will be triggered especially in heavily loaded clusters. In this paper, we introduce a novel hierarchical hybrid design of schedulers to address this problem, called En-Omega. In En-Omega, we enhance the fully distributed schedulers with a central scheduler, which can provide global fairness to the jobs from different schedulers and simultaneously reduce the average latency of all the jobs sharply. To reduce the overhead, in our En-Omega design, we activate the central scheduler only when the cluster is heavily loaded. Furthermore, the cache used for central queuing and the scoring policy used in central scheduling are all load-aware. We evaluate En-Omega based on Google trace and experimental results show that, compared to the baseline design, our method can reduce the average latency of starving jobs up to 90% with reasonable overhead. Di Fu, Jiahai Yang 0001, Hui Zhang 0052 |
ISCC | 1 |
| 2015 | Improving SVM based multi-label classification by using label relationshipabstractThis paper proposes an improved SVM based multi-label classification method by using relationship among labels. Following a traditional multi-label solution, binary relevance (BR) method is first used to decompose the multi-label classification problem into multiple binary classification sub-problems, each of which is solved by an SVM classifier. By using Platt's sigmoid technique, each SVM classifier gives probability output for the following correction. A probability model is introduced to estimate the relationship among labels. The extracted label relationship is then applied to correct the outputs of SVM classifiers, in which a dynamic weight strategy is further introduced. Numerical experiments on widely used benchmark datasets show that the proposed method can improve the accuracy of multi-label classification when compared with traditional BR method and some other conventional multi-label classification methods. Di Fu, Bo Zhou 0016, Jinglu Hu |
IJCNN | 1 |
| 2015 | A Transductive SVM with quasi-linear kernel based on cluster assumption for semi-supervised classificationabstractThis paper presents a Transductive Support Vector Machine (TSVM) with quasi-linear kernel based on a clustering assumption for semi-supervised classification. Since the potential separating boundary is located in low density area between classes, a modified density clustering method by considering label information is firstly introduced to extract the information of potential separating boundary in low density region between different classes. Then the information is used to compose a quasi-linear kernel for the TSVM. The optimization of TSVM is further speeded up by developing a pairwise label switching method on minimal sets. Experiment results on benchmark datasets show that the proposed method is effective and improves classification performances. Bo Zhou 0016, Di Fu, Jinglu Hu |
IJCNN | 2 |
| 2015 | Rhizome: utilizing the public cloud to provide 3D gaming infrastructureabstractMotivated by our systematic study on the diverse aspects of migrating gaming services to a virtualized cloud environment, we designed and implemented a fully virtualized cloud gaming platform, Rhizome, utilizing the latest hardware support for both remote servers and local clients. Our platform takes the first step towards bridging online gaming systems and public clouds. To accomplish ultra-low latency and a low power consumption gaming experience, we further optimized Rhizome's thin-client configuration and its interaction modules. In this proposed demo, we demonstrate that gaming over a virtualized cloud can be made possible with careful optimization and integration of different modules. It also helps us reveal the critical challenges towards full-fledged deployment of gaming services over public virtualized cloud. Ryan Shea, Di Fu, Jiangchuan Liu |
MMSys | 2 |
| 2015 | Towards bridging online game playing and live broadcasting: design and optimizationabstractRecent years have witnessed the emergence and growth of Cloud Gaming, where players interact with the remote game instance and receive rendered game scenes in video stream. Meanwhile, broadcasting and viewing games through live streaming platforms, e.g., Twitch.tv, have become increasingly popular. The interaction and performance of the many modules involved in this new generation of gaming and streaming platforms have yet to be closely investigated. In this paper, we present an initial experiment-based performance study, in which we profile the architecture of realworld gaming and streaming platforms, namely the Open Broadcast Software (OBS) module and its connection to the Twitch server. Our investigation shows that the recording operation can greatly increase the CPU utilization and the power consumption can increase over 60% on the game streaming computer. The use of advanced hardware encoding found on modern GPUs can greatly alleviate these performance issues. Yet, through profiling, we show that hardware encoding can introduce remarkable delays to the whole pipeline. We track this to a complicated interplay between the CPUs power saving methods and the implementation of hardware encoders. Ryan Shea, Di Fu, Jiangchuan Liu |
NOSSDAV | 2 |
| 2015 | Cloud Gaming: Understanding the Support From Advanced Virtualization and HardwareabstractExisting cloud gaming platforms have mainly focused on private nonvirtualized environments with proprietary hardware. Modern public cloud platforms heavily rely on virtualization for efficient resource sharing, the potentials of which have yet to be explored. Migrating gaming to a public cloud is nontrivial, however, particularly considering the overhead for virtualization and that the graphics processing units (GPUs) for game rendering has long been an obstacle in virtualization. This paper takes a first step toward bridging the online gaming system and the public cloud platforms. We present the design and implementation of a fully virtualized cloud gaming platform with the latest hardware support for both remote servers and local clients. We explore many critical design issues inherent in cloud gaming, including the choice of hardware or software video encoding, and the configuration and the detailed power consumption of thin client. We demonstrate that with the latest hardware and virtualization support, gaming over virtualized cloud can be made possible with careful optimization and integration of the different modules. We also highlight critical challenges toward full-fledged deployment of gaming services over the public virtualized cloud. Ryan Shea, Di Fu, Jiangchuan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Towards Optimal Collaboration of Policies in the Two-Phase Scheduling of Cloud Tasks
Jiahai Yang 0001, Di Fu, Hui Zhang 0052 |
NPC | 3 |
| 2014 | Multi-Label Classification Based on Multi-Objective OptimizationabstractMulti-label classification refers to the task of predicting potentially multiple labels for a given instance. Conventional multi-label classification approaches focus on single objective setting, where the learning algorithm optimizes over a single performance criterion (e.g., Ranking Loss ) or a heuristic function. The basic assumption is that the optimization over one single objective can improve the overall performance of multi-label classification and meet the requirements of various applications. However, in many real applications, an optimal multi-label classifier may need to consider the trade-offs among multiple inconsistent objectives, such as minimizing Hamming Loss while maximizing Micro F1 . In this article, we study the problem of multi-objective multi-label classification and propose a novel solution (called M oml ) to optimize over multiple objectives simultaneously. Note that optimization objectives may be inconsistent, even conflicting, thus one cannot identify a single solution that is optimal on all objectives. Our M oml algorithm finds a set of non-dominated solutions which are optimal according to different trade-offs among multiple objectives. So users can flexibly construct various predictive models from the solution set, which provides more meaningful classification results in different application scenarios. Empirical studies on real-world tasks demonstrate that the M oml can effectively boost the overall performance of multi-label classification by optimizing over multiple objectives simultaneously. Chuan Shi 0001, Xiangnan Kong, Di Fu, Philip S. Yu, Bin Wu 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2013 | A link clustering based overlapping community detection algorithm
Chuan Shi 0001, Yanan Cai, Di Fu, Yuxiao Dong, Bin Wu 0001 |
Data Knowl. Eng. | 3 |