EDBT 2026 Demo / reviewers in the wild / expert
Cam-Tu Nguyen
dblp:14/5079
· DBLP profile ↗
43ranked-venue papers
6as first author
21since 2021 · last 2026
0009-0006-9484-6876ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 8 since 2021Computer networks · 14 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text RepresentationabstractText representation plays a critical role in tasks like clustering, retrieval, and other downstream applications. With the emergence of large language models (LLMs), there is increasing interest in harnessing their capabilities for this purpose. However, most of the LLMs are inherently causal and optimized for next-token prediction, making them suboptimal for producing holistic representations. To address this, recent studies introduced pretext tasks to adapt LLMs for text representation. Most of these tasks, however, rely on token-level prediction objectives, such as the masked next-token prediction (MNTP) used in LLM2Vec. In this work, we explore the untapped potential of context compression as a pretext task for unsupervised adaptation of LLMs. During compression pre-training, the model learns to generate compact memory tokens, which substitute the whole context for downstream sequence prediction. Experiments demonstrate that a well-designed compression objective can significantly enhance LLM-based text representations, outperforming models trained with token-level pretext tasks. Further improvements through contrastive learning produce a strong representation model (LLM2Comp) that outperforms contemporary LLM-based text encoders on a wide range of tasks while being more sample-efficient, requiring significantly less training data. Yeqin Zhang, Yizheng Zhao, Binxing Jiao, Daxin Jiang, Ruihang Miao, Cam-Tu Nguyen |
AAAI | 7 |
| 2026 | Thermometer of Thoughts: Enhancing LLM's Exploration via Attention Temperature ModulationabstractZhiyuan Yu, Shijian Xiao, Cam-Tu Nguyen, Zhangyue Yin, Lekai Xing, Wenzhong Li, Sanglu Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shijian Xiao, Cam-Tu Nguyen, Zhangyue Yin, Lekai Xing, Sanglu Lu |
ACL (1) | 3 |
| 2026 | GPU-Centric Stateless LLM Serving With GIGANETSabstractGiganetes is a GPU-centric architecture for stateless LLM serving that externalizes KV caches to a disaggregated remote memory pool via GPUDirect RDMA. By treating remote memory as a GPU-addressable tier via GPUDirect RDMA, Giganetes eliminates session affinity constraints: any GPU can serve any request, enabling near-linear horizontal scaling in a Kubernetes-native deployment. A Scatter/Gather I/O interface bypasses the CPU and host memory entirely, achieving 52.4 GB/s application-level read throughput on our 4×200 Gbps RDMA testbed. A session-level metadata abstraction and proactive readahead mechanism reduce GPU bubbles by overlapping remote KV fetches with prefill computation and scheduling slack. On a 4-node H800 cluster, Giganetes delivers 33% higher throughput (QPS 2.4 vs. 1.8) and 1.75× lower P95 TPOT than PD-disaggregation with sticky sessions, with the gain driven by scheduling flexibility rather than faster transport alone. Xiaoliang Wang 0001, Zhenwei Pi, Cam-Tu Nguyen |
SIGCOMM | 4 |
| 2026 | Fine-Grained Scheduling of In-Network Aggregation Resources for Efficient Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001, Guihai Chen |
IEEE Trans. Netw. | 10 |
| 2025 | Momentum Posterior Regularization for Multi-hop Dense RetrievalabstractMulti-hop question answering (QA) often requires sequential retrieval (multi-hop retrieval), where each hop retrieves missing knowledge based on information from previous hops. To facilitate more effective retrieval, we aim to distill knowledge from a posterior retrieval, which has access to posterior information like an answer, into a prior retrieval used during inference when such information is unavailable. Unfortunately, current methods for knowledge distillation in one-time retrieval are ineffective for multi-hop QA due to two issues: 1) posterior information is often defined as the response (i.e. answers), which may not clearly connect to the query without intermediate retrieval; and 2) the large knowledge gap between prior and posterior retrievals makes distillation using existing methods unstable, even resulting in performance loss. As such, we propose MoPo (Momentum Posterior Regularization) with two key innovations: 1) Posterior information of one hop is defined as a query-focus summary from the golden knowledge of the previous and current hops; 2) We develop an effective training strategy where the posterior retrieval is updated along with the prior retrieval via momentum moving average method, allowing smoother and effective distillation. Experiments on HotpotQA and StrategyQA demonstrate that MoPo outperforms existing baselines in both retrieval and downstream QA tasks. Zehua Xia, Yuyang Wu, Yiyun Xia, Cam-Tu Nguyen |
COLING | 4 |
| 2025 | Mina: Fine-Grained In-network Aggregation Resource Scheduling for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001 |
INFOCOM | 10 |
| 2025 | SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM InferenceabstractKV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns during decoding (the saliency shift problem), and (2) they treat both marginally important tokens and truly unimportant tokens uniformly, despite the collective significance of marginal tokens to model performance (the marginal information over-compression problem). To address these issues, we design two compensation mechanisms based on the high similarity of attention matrices between LLMs with different scales. We propose SmallKV, a small model assisted compensation method for KV cache compression. SmallKV can maintain attention matching between different-scale LLMs to: 1) assist the larger model in perceiving globally important information of attention; and 2) use the smaller model’s attention scores to approximate those of marginal tokens in the larger model. Extensive experiments on benchmarks including GSM8K, BBH, MT-Bench, and LongBench demonstrate the effectiveness of SmallKV. Moreover, efficiency evaluations show that SmallKV achieves 1.75 - 2.56 times higher throughput than baseline methods, highlighting its potential for efficient and performant LLM inference in resource constrained environments. Yajuan Peng, Cam-Tu Nguyen, Zuchao Li, Xiaoliang Wang 0001, Hai Zhao 0001, Xiaoming Fu 0001 |
NeurIPS | 3 |
| 2025 | Accelerating Network Features Deployment With Heterogeneous PlatformsabstractEnhancing the networking system with appropriate functions is a longstanding goal. Unfortunately, in today’s large-scale high-speed data centers, the feature velocity of network functions is slow because it is hard to verify the function in realistic scenarios. Recent advances in programmable switching ASICs have enabled the network data plane to move beyond its traditional role of packet forwarding. However, the current compromise between performance and flexibility results in limitations such as restricted memory/computation resources and programmable models. These limitations make it challenging for programmable switches to offer more features and to be deployed in large-scale production environments. In response, we present CLIP, a framework that works in collaboration with programmable devices and commodity servers to enhance the validation and deployment velocity of features. CLIP defines a cross-platform function definition framework and provides a set of tools to reduce the complexity of manually writing cross-platform programs. We propose an automatic traffic placement and scaling mechanism to coordinate packet processing performance across heterogeneous devices. Compared with software-based Network Functions (NFs), CLIP achieves a throughput ranging from$1.36\times $to$16.06\times $under different realistic traffic loads. Through the development and deployment of three self-defined functions within a realistic testbed, we demonstrate the feasibility and efficiency of CLIP. Xiaoliang Wang 0001, Chen Tian 0001, Yun Xiong, Sanglu Lu, Cam-Tu Nguyen |
IEEE Trans. Netw. | 7 |
| 2024 | Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence RegularizationabstractIn open-domain Question Answering (QA), dense text retrieval is crucial for finding relevant passages to generate answers. Typically, contrastive learning is used to train a retrieval model, which maps passages and queries to the same semantic space, making similar ones closer and dissimilar ones further apart. However, training such a system is challenging due to the false negative problem, where relevant passages may be missed during data annotation. Hard negative sampling, commonly used to improve contrastive learning, can introduce more noise in training. This is because hard negatives are those close to a given query, and thus more likely to be false negatives. To address this, we propose a novel contrastive confidence regularizer for Noise Contrastive Estimation (NCE) loss, a commonly used contrastive loss. Our analysis shows that the regularizer helps make the dense retrieval model more robust against false negatives with a theoretical guarantee. Additionally, we propose a model-agnostic method to filter out noisy negative passages in the dataset, improving any downstream dense retrieval models. Through experiments on three datasets, we demonstrate that our method achieves better retrieval performance in comparison to existing state-of-the-art dense retrieval systems. Shiqi Wang 0003, Yeqin Zhang, Cam-Tu Nguyen |
AAAI | 3 |
| 2024 | Retrospex: Language Agent Meets Offline Reinforcement Learning CriticabstractLarge Language Models (LLMs) possess extensive knowledge and commonsense reasoning capabilities, making them valuable for creating powerful agents.However, existing LLM agent frameworks have not fully utilized past experiences for improvement.This work introduces a new LLM-based agent framework called Retrospex , which addresses this challenge by analyzing past experiences in depth.Unlike previous approaches, Retrospex does not directly integrate experiences into the LLM's context.Instead, it combines the LLM's action likelihood with action values estimated by a Reinforcement Learning (RL) Critic, which is trained on past experiences through an offline "retrospection" process.Additionally, Retrospex employs a dynamic action rescoring mechanism that increases the importance of experience-based values for tasks that require more interaction with the environment.We evaluate Retrospex in ScienceWorld, ALFWorld and Webshop environments, demonstrating its advantages over strong, contemporary baselines 1 . Yufei Xiang, Yiqun Shen, Yeqin Zhang, Cam-Tu Nguyen |
EMNLP | 4 |
| 2024 | Sample Efficiency Matters: Training Multimodal Conversational Recommendation Systems in a Small Data SettingabstractWith the increasing prevalence of virtual assistants, multimodal conversational recommendation systems (multimodal CRS) becomes essential for boosting customer engagement, improving conversion rates, and enhancing user satisfaction. Yet conversational samples, as training data for such a system, are difficult to obtain in large quantities, particularly in new platforms. To effectively train multimodal CRS in a small data setting, we enhance data quality to make up for the small data quantity by augmenting conversations with dialogue states. We then devise an effective dialogue state encoder to bridge the semantic gap between conversation and product representations for recommendation. To further reduce the cost of dialogue state annotation, a semi-supervised learning method is developed to effectively train the dialogue state encoder with a small set of labeled conversations. In addition, we design a correlation regularisation that leverages knowledge in the multimodal product database to help align textual and visual modalities. Experiments on the dataset MMD demonstrate the effectiveness of our method. Particularly, with only 5% of the MMD training set, our method (namely SeMANTIC) obtains better NDCG scores than those of baseline models trained on the full MMD training set. Wenzhe Du, Xiaoliang Wang 0001, Cam-Tu Nguyen |
ACM Multimedia | 4 |
| 2024 | CyberStar: Simple, Elastic and Cost-Effective Network Functions Management in Cloud Network at Scale
Bengbeng Xue, Yang Song 0031, Xiaoxin Peng, Yilong Lyu, Xiaoliang Wang 0001, Chen Tian 0001, Cam-Tu Nguyen, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu |
USENIX ATC | 10 |
| 2023 | MINA: Auto-scale In-network Aggregation for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Wei Wang 0002, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Yongqiang Xiong, Chen Tian 0001, Cam-Tu Nguyen, Xiaoliang Wang 0001 |
APNet | 13 |
| 2023 | Diversify Question Generation with Retrieval-Augmented Style TransferabstractGiven a textual passage and an answer, humans are able to ask questions with various expressions, but this ability is still challenging for most question generation (QG) systems.Existing solutions mainly focus on the internal knowledge within the given passage or the semantic word space for diverse content planning.These methods, however, have not considered the potential of external knowledge for expression diversity.To bridge this gap, we propose RAST, a framework for Retrieval-Augmented Style Transfer, where the objective is to utilize the style of diverse templates for question generation.For training RAST, we develop a novel Reinforcement Learning (RL) based approach that maximizes a weighted combination of diversity reward and consistency reward.Here, the consistency reward is computed by a Question-Answering (QA) model, whereas the diversity reward measures how much the final output mimics the retrieved template.Experimental results show that our method outperforms previous diversity-driven baselines on diversity while being comparable in terms of consistency scores.Our code is available at https://github.com/gouqi666/RAST. Qi Gou, Zehua Xia, Bowen Yu 0002, Haiyang Yu 0003, Fei Huang 0002, Yongbin Li 0001, Cam-Tu Nguyen |
EMNLP | 7 |
| 2023 | Coarse-To-Fine Knowledge Selection for Document Grounded DialogsabstractMulti-document grounded dialogue systems (DGDS) belong to a class of conversational agents that answer users’ requests by finding supporting knowledge from a collection of documents. Most previous studies aim to improve the knowledge retrieval model or propose more effective ways to incorporate external knowledge into a parametric generation model. These methods, however, focus on retrieving knowledge from mono-granularity language units (e.g. passages, sentences, or spans in documents), which is not enough to effectively and efficiently capture precise knowledge in long documents. This paper proposes Re3G, which aims to optimize both coarse-grained knowledge retrieval and fine-grained knowledge extraction in a unified framework. Specifically, the former efficiently finds relevant passages in a retrieval-and-reranking process, whereas the latter effectively extracts finer-grain spans within those passages to incorporate into a parametric answer generation model (BART, T5). Experiments on DialDoc Shared Task demonstrate the effectiveness of our method. Yeqin Zhang, Haomin Fu, Cheng Fu 0003, Haiyang Yu 0003, Yongbin Li 0001, Cam-Tu Nguyen |
ICASSP | 6 |
| 2023 | Long Short-Term Planning for Conversational Recommendation Systems
Hongguang Shi, Yeqin Zhang, Xubin Li, Cam-Tu Nguyen |
ICONIP (6) | 6 |
| 2023 | Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational RecommendationabstractMultimodal Conversational Recommendation aims to find appropriate products based on a multi-turn dialogue, where user requests and products can be presented in both visual and textual modalities. While previous studies have focused on understanding user preferences from conversational contexts, the task of product modeling has been relatively unexplored. This study targets to fill this gap and demonstrates that information from multiple product views and cross-view interactions are essential for recommendation, along with dialog information. To this end, a product image is first encoded using a gated multi-view image encoder, and representations for the global and local views are obtained. On the textual side, two views are considered: the structure view (product attributes) and the sequence view (product description/reviews). Two forms of inter-modal interactions for product representation are then modeled: interactions between the global image view and the textual structure view, and interactions between the local image view and the textual sequence view. Furthermore, the representation is enhanced to attend to the latest user request in the dialog context, resulting in query-aware product representation. The experimental results indicate that our method, named Enteract, achieves state-of-the-art performance on two well-known datasets (MMD and SIMMC). Wenzhe Du, Cam-Tu Nguyen |
ACM Multimedia | 3 |
| 2022 | Tuning Target Delay for RTT-based Congestion ControlabstractThe congestion control strategy plays an essential role in the high-speed datacenter network. It aims to deliver low latency, high throughput network service. RTT-based congestion control leverages advanced NIC hardware to identify accumulated queuing delay of the end-to-end path. Sender adjusts the sending rate or congestion window if the delay exceeds a predetermined value, i.e., target delay. Therefore, setting the target delay is the key for RTT-based congestion control strategies. We provide a comprehensive study of the impact of target delay on recent RTT-based congestion control strategies, and demonstrate that a fixed inappropriate target value can lead to low bandwidth utilization or high latency. We then propose a practical queuing target updating approach to solve this problem. The proposed method maintains a shared near-optimal queuing target at the receiving host. We leverage the widely supported ECN flag to estimate the empty state of switch queue instead of indicating congestion, which requires no complicated threshold configuration. We have integrated the dynamic queuing target updating approach into the state-of-the-art RTT-based congestion strategy, SWIFT, and named the design RET. Test-bed experiments and simulations in the large-scale network with synthesized traffic of real workloads show that RET can achieve up to 1.5x and 3.6x lower tail latency than SWIFT and DCQCN, respectively. This paper provides a deep understanding on tuning target delay for RTT-based congestion control algorithms in datacenter networks. Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu |
ICNP | 3 |
| 2021 | Maximizing the Benefit of RDMA at End HostsabstractRDMA is increasingly deployed in data center to meet the demands of ultra-low latency, high throughput and low CPU overhead. However, it is not easy to migrate existing applications from the TCP/IP stack to the RDMA. The developers usually need to carefully select communication primitives and manually tune the parameters for each single-purpose system. After operating the high-speed RDMA network, we identify multiple hidden costs which may cause degraded and/or unpredictable performance of RDMA-based applications. We demonstrate these hidden costs including the combination of complicated parameter settings, scalability of Reliable Connections, two-sided memory management and page alignment, resource contention among diverse traffics, etc. Furthermore, to address these problems, we introduce Nem, a suite that allows developers to maximize the benefit of RDMA by i) eliminating the resource contention at NIC cache through asynchronous resource sharing; ii) introducing hybrid page management based on messages sizes; iii) isolating flows of different traffic classes based hardware features. We implement the prototype of Nem and verify its effectiveness by rebuilding the RPC message service, which demonstrates the high throughput for large messages, low latency for small messages without compromising the low CPU utilization and good scalability performance for a large number of active connections. Xiaoliang Wang 0001, Hexiang Song, Cam-Tu Nguyen, Dongxu Cheng, Tiancheng Jin |
INFOCOM | 3 |
| 2021 | Gray Failures Detection for Shared Bicycles
Hangfan Zhang, Mingchao Zhang, Cam-Tu Nguyen, Sheng Zhang 0001, Xiaoliang Wang 0001 |
WASA (1) | 3 |
| 2021 | SAKE: Estimating Katz Centrality Based on Sampling for Large-Scale Social NetworksabstractKatz centrality is a fundamental concept to measure the influence of a vertex in a social network. However, existing approaches to calculating Katz centrality in a large-scale network are unpractical and computationally expensive. In this article, we propose a novel method to estimate Katz centrality based on graph sampling techniques, which object to achieve comparable estimation accuracy of the state-of-the-arts with much lower computational complexity. Specifically, we develop a Horvitz–Thompson estimate for Katz centrality by using a multi-round sampling approach and deriving an unbiased mean value estimator. We further propose SAKE , a S ampling-based A lgorithm for fast K atz centrality E stimation. We prove that the estimator calculated by SAKE is probabilistically guaranteed to be within an additive error from the exact value. Extensive evaluation experiments based on four real-world networks show that the proposed algorithm can estimate Katz centralities for partial vertices with low sampling rate, low computation time, and it works well in identifying high influence vertices in social networks. Mingkai Lin, Lynda Jiwen Song, Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu |
ACM Trans. Knowl. Discov. Data | 4 |
| 2020 | App trajectory recognition over encrypted internet traffic based on deep neural network
Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu |
Comput. Networks | 4 |
| 2020 | Construction of Subexponential-Size Optical Priority Queues With Switches and Fiber Delay LinesabstractAll-optical switching has been considered as a natural choice to keep pace with growing fiber link capacity. One key research issue of all-optical switching is the design of optical buffers for packet contention resolution. One of the most general buffering schemes is optical priority queue, where every packet is associated with a unique priority upon its arrival and departs the queue in order of priority, and the packet with the lowest priority is always dropped when a new packet arrives but the buffer is full. In this paper, we focus on the feedback construction of an optical priority queue with a single (M + 2) × (M + 2) optical crossbar Switch and M fiber Delay Lines (SDL) connecting M inputs and M outputs of the switch. We propose a novel construction of an optical priority queue with buffer 2Θ(√M), which improves substantially over all previous constructions that only have buffers of O(Mc) size for constant integer c. The key ideas behind our construction include (i) the use of first in first out multiplexers, which admit efficient SDL constructions, for feeding back packets to the switch instead of fiber delay lines, and (ii) the use of a routing policy that is similar to self-routing, where each packet entering the switch is routed to some multiplexer mainly determined by the current ranking of its priority. Bin Tang 0002, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu |
IEEE/ACM Trans. Netw. | 3 |
| 2019 | Sampling Based Katz Centrality Estimation for Large-Scale Social Networks
Mingkai Lin, Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu |
ICA3PP (2) | 3 |
| 2019 | ActiveTracker: Uncovering the Trajectory of App Activities over Encrypted Internet Traffic StreamsabstractDespite the increasing popularity of mobile applications and the widespread adoption of encryption techniques, mobile devices are still susceptible to security and privacy risks. In this paper, we propose ActiveTracker, a new type of sniffing attack that can reveal the fine-grained trajectory of user’s mobile app usage from a sniffed encrypted Internet traffic stream. It firstly adopts a sliding window based approach to divide the encrypted traffic stream into a sequence of segments corresponding to different app activities. Then each traffic segment is represented by a normalized temporal-spacial traffic matrix and a traffic spectrum vector. Based on the normalized representation, a deep neural network (DNN) classification algorithm is developed to recognize the crucial activities conducted with different apps by the user. We show by extensive experiments on real-world app usage traffic collected from volunteers that the proposed approach achieves up to 78.5% accuracy in recognizing app trajectory over encrypted traffic streams. Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu |
SECON | 4 |
| 2018 | Optimizing User Experience through Implicit Content-aware Network Service in the Home EnvironmentabstractThere has always been a gap between Internet Service Providers (ISPs) and end users when considering the performance of network-based application. On one hand, ISPs keep raising the investment on infrastructures to speed up the data transportation. On the other hand, users are not satisfied with the perceived quality of experience (QoE). This happens mainly due to the inflexible network flow management, where only the function of rate limiting is provided for home users in the shared network environment. In this paper, we focus on the optimization of users experience by customizing bandwidth allocation for user specified preferences while maintaining high bandwidth utilization. We introduce implicit content-aware bandwidth allocation to minimize the involvement of users on complicated network setting. By leveraging the technique of software-defined networking (SDN), a prototype of content-aware traffic scheduling, Conan, is developed to verify the effectiveness of our design. Experiments show that Conan can reduce the average task completion time of interactive applications by 30-40%. During heavy traffic load, Conan can ensure stable bandwidth for each video streaming flow and greatly reduce the average stall duration. Haixiang Yang, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu |
GROUP | 3 |
| 2018 | Nem: Toward Fine-grained Load Balancing through RNIC EC OffloadingabstractModern datacenter networks employ Load-balancing (LB) in the large-scale multi-tier topology to ensure high network utilization as well as low flow completion time. This paper presents the design and evaluation of Nem, a robust Erasure Coding (EC) based load balancing scheme at end-host to spread data across multiple paths. Our design is based on two key insights. First, both theory and implementation have shown that redundancy is a powerful technique to reduce latency in networked system. Second, the commercial RDMA network interface card supports EC offload which can dramatically reduce the CPU consumption. Nem is an optimal user-level LB design, which leveraging redundant fine-grained data blocks and high speed lossless RDMA network to realize effective load balancing transmission. Evaluation over many workloads shows that Nem is adaptive to the asymmetric networks, and achieves better performance compared to the state-of-art host-based load balancing mechanism. Xiaoliang Wang 0001, Cam-Tu Nguyen, Zhuzhong Qian, Bin Tang 0002, Sanglu Lu |
HPSR | 2 |
| 2017 | A Virtual Middleboxes Network Placement Algorithm in Multi-tenant Datacenter NetworksabstractHardware middleboxes are widely used in current cloud datacenter to provide network functions such as firewalls, intrusion detection system, load balancers, etc. Unfortunately, they are expensive and unable to offer customized functions for individual tenant. To overcome this issue, there is an increasing interest in deploying software middleboxes to enable flexible security, network access functionality. This paper addresses the software middleboxes placement problem with minimum bandwidth guarantee. We first specify the model of tenants' requirement that specifies the need for virtual machines of application and middleboxes, as well as communication traffic. A virtual middlebox placement algorithm called MISSILE is then proposed to offer predictable network performance for each accepted tenant, and minimize datacenter bandwidth utilization. Extensive simulation results based on current large-scale datacenter networks verify that MISSILE is effective and provides network performance guarantee for tenants. Xiaoliang Wang 0001, Cam-Tu Nguyen, Jian Wang 0038, Zhuzhong Qian, Sanglu Lu |
ICPADS | 3 |
| 2017 | AGRA: An Analysis-Generation-Ranking Framework for Automatic Abbreviation from Paper TitlesabstractPeople sometimes choose word-like abbreviations to refer to items with a long description. These abbreviations usually come from the descriptive text of the item and are easy to remember and pronounce, while preserving the key idea of the item. Coming up with a nice abbreviation is not an easy job, even for human. Previous assistant naming systems compose names by applying hand-written rules, which may not perform well. In this paper, we propose to view the naming task as an artificial intelligence problem and create a data set in the domain of academic naming. To generate more delicate names, we propose a three-step framework, including description analysis, candidate generation and abbreviation ranking, each of which is parameterized and optimizable. We conduct experiments to compare different settings of our framework with several analysis approaches from different perspectives. Compared to online or baseline systems, our framework could achieve the best results. Shujian Huang, Cam-Tu Nguyen, Xiaoliang Wang 0001, Xinyu Dai, Jiajun Chen 0001, Yang Yu 0001 |
IJCAI | 4 |
| 2017 | SilentTalk: Lip reading through ultrasonic sensing on mobile phonesabstractThe recently enhanced computing capability and rich sensing functionality on mobile devices lead to the ubiquitous application of speech recognition. Traditional speech recognition records acoustic signals or visual images to interpret speech. However, the acoustic based scheme has many drawbacks. It is easily affected by the environmental noise when users are in the factory or market, and can not be used in a place where people need to be quite such as library. Specifically the current design is not suitable for people with speaking or hearing difficulties. Unfortunately, the visual-based approach is sensitive to fight conditions which shows poor performance in the dark area. As a result, it is necessary to provide an new human-computer interaction channel to assist speech recognition. This paper presents SilentTalk, a non-invasive lip reading system based on ultrasonic Doppler effect The main idea is to generate ultrasonic signals from a mobile phone, then capture the reflections and analyze the fine-grained frequency shift caused by mouth movements. A Frequency Shift Detection Model (FSDM) is proposed to quantify the correlation between frequency variations and mouth movements that form different syllables. SilentTalk then applies a Continuous Lip Reading Model (CLRM) on top of FSDM to realize continuous lip reading. Based on Markov assumption, CLRM effectively combines pronunciation rules and context knowledge to connect isolated syllables to words and sentences. Experiments show that SilentTalk can identify 12 basic mouth motions up to 95.4% accuracy in English. The system can also recognize short sentences up to six words with an average accuracy of 74.8%. Jiayao Tan, Cam-Tu Nguyen, Xiaoliang Wang 0001 |
INFOCOM | 2 |
| 2017 | Predicting Happiness State Based on Emotion Representative Mining in Online Social Networks
Xiao Zhang 0015, Hong Huang 0001, Cam-Tu Nguyen, Xu Chen 0004, Xiaoliang Wang 0001, Sanglu Lu |
PAKDD (1) | 4 |
| 2016 | Constructing sub-exponentially large optical priority queues with switches and fiber delay linesabstractOptical switching has been considered as a natural choice to keep pace with growing fiber link capacity. One key research issue of all-optical switching is the design of optical queues by using optical crossbar switches and fiber delay lines (SDLs). In this paper, we focus on the construction of an optical priority queue with a single (M+2)×(M+2) crossbar switch and M fiber delay lines, and evaluate it in terms of the buffer size of the priority queue. Currently, the best known upper bound of the buffer size is O(2M), while existing methods can only construct a priority queue with buffer O(M3). In this paper, we make a great step towards closing the above huge gap. We propose a very efficient construction of priority queues with buffer 2Θ(√M). We use 4-to-1 multiplexers with different buffer sizes, which can be constructed efficiently with SDL, as intermediate building blocks to simplify the design. The key idea in our construction is to route each packet entering the switch to some group of four 4-to-1 multiplexers according to its current priority, which is shown to be collision-free. Bin Tang 0002, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu |
ISIT | 3 |
| 2016 | An Efficient Walking Safety Service for Distracted Mobile UsersabstractThere is a growing number of incidents related to using cell phones while walking on the street. To address this issue, this paper proposes a new system based on tactile paving detection on the sidewalk to alert distracted mobile users to avoid traffic hazard. Our system (namely Inspector) is deployed as an application on an off-the-shelf normal mobile phone equipped with back camera. Inspector plays as a third eye to alert users when they step out of the safe zones where no tactile paving is detected. In order to obtain reliable yet effective decision results, we exploit a lightweight image processing method, simple classifiers along with a smart sampling strategy. The main idea is that the application capture more images of the surrounding environments when we doubt that the result from one image is not sufficient. The sampling interval is adjusted dynamically so that we can save the energy and thus extend the working time of the system. Real-scenario tests show that Inspector can detect whether a mobile user is walking along a blind sidewalk with an accuracy of 92.72%, 98.78%, and 99.44% when the detection algorithm sample 2 times, 3 times, and 4 times continuous detection respectively. The reaction time, which is measured by the difference between the alert time and the time when a user steps out of a safety zone, is 0.52 seconds early with a sampling interval of 2 seconds. Maozhi Tang, Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu |
MASS | 2 |
| 2016 | Conan: Content-aware Access Network Flow Scheduling to Improve QoE of Home UsersabstractThere has always been a gap of perception between Internet Service Providers (ISPs) and their customers when considering the performance of network service. On one hand, ISPs invest to increase downstream speed of access network infrastructure. On the other hand, users cannot achieve perceived quality of experience (QoE). This paper addresses this problem by introducing a system, Conan, which enables content-aware flow scheduling to improve the QoE of users. Conan exploits to satisfy users' requirements in the access network (LAN), which is the performance bottleneck actually. By leveraging the technique of software defined networking (SDN), Conan are able to specify the expected network capacity for different applications. Automatic application identification is deployed at home gateway to improve the scalability, and flexible bandwidth allocation is realized at LAN for specified applications. Using video streaming service optimization as an example, we demonstrate that our system can automatically allocate bandwidth for video flows. Haixiang Yang, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu |
SIGCOMM | 3 |
| 2014 | Labeling Complicated Objects: Multi-View Multi-Instance Multi-Label LearningabstractMulti-Instance Multi-Label (MIML) is a learning framework where an example is associated with multiple labels and represented by a set of feature vectors (multiple instances). In the formalization of MIML learning, instances come from a single source (single view). To leverage multiple information sources (multi-view), we develop a multi-view MIML framework based on hierarchical Bayesian Network, and derive an effective learning algorithm based on variational inference. The model can naturally deal with examples in which some views could be absent (partial examples). On multi-view datasets, it is shown that our method is better than other multi-view and single-view approaches particularly in the presence of partial examples. On single-view benchmarks, extensive evaluation shows that our method is highly competitive or better than other MIML approaches on labeling examples and instances. Moreover, our method can effectively handle datasets with a large number of labels. Cam-Tu Nguyen, Xiaoliang Wang 0001, Jing Liu 0001, Zhi-Hua Zhou |
AAAI | 1 |
| 2013 | Multi-Modal Image Annotation with Multi-Instance Multi-Label LDA
Cam-Tu Nguyen, De-Chuan Zhan, Zhi-Hua Zhou |
IJCAI | 1 |
| 2013 | A feature-word-topic model for image annotation and retrievalabstractImage annotation is a process of finding appropriate semantic labels for images in order to obtain a more convenient way for indexing and searching images on the Web. This article proposes a novel method for image annotation based on combining feature-word distributions, which map from visual space to word space, and word-topic distributions, which form a structure to capture label relationships for annotation. We refer to this type of model as Feature-Word-Topic models. The introduction of topics allows us to efficiently take word associations, such as {ocean, fish, coral} or {desert, sand, cactus}, into account for image annotation. Unlike previous topic-based methods, we do not consider topics as joint distributions of words and visual features, but as distributions of words only. Feature-word distributions are utilized to define weights in computation of topic distributions for annotation. By doing so, topic models in text mining can be applied directly in our method. Our Feature-word-topic model, which exploits Gaussian Mixtures for feature-word distributions, and probabilistic Latent Semantic Analysis (pLSA) for word-topic distributions, shows that our method is able to obtain promising results in image annotation and retrieval. Cam-Tu Nguyen, Natsuda Kaothanthong, Takeshi Tokuyama, Xuan-Hieu Phan |
ACM Trans. Web | 1 |
| 2011 | A Hidden Topic-Based Framework toward Building Applications with Short Web DocumentsabstractThis paper introduces a hidden topic-based framework for processing short and sparse documents (e.g., search result snippets, product descriptions, book/movie summaries, and advertising messages) on the Web. The framework focuses on solving two main challenges posed by these kinds of documents: 1) data sparseness and 2) synonyms/homonyms. The former leads to the lack of shared words and contexts among documents while the latter are big linguistic obstacles in natural language processing (NLP) and information retrieval (IR). The underlying idea of the framework is that common hidden topics discovered from large external data sets (universal data sets), when included, can make short documents less sparse and more topic-oriented. Furthermore, hidden topics from universal data sets help handle unseen data better. The proposed framework can also be applied for different natural languages and data domains. We carefully evaluated the framework by carrying out two experiments for two important online applications (Web search result classification and matching/ranking for contextual advertising) with large-scale universal data sets and we achieved significant results. Xuan-Hieu Phan, Cam-Tu Nguyen, Dieu-Thu Le, Minh Le Nguyen 0001, Susumu Horiguchi, Quang-Thuy Ha |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | A feature-word-topic model for image annotationabstractImage annotation is to automatically associate semantic labels with images in order to obtain a more convenient way for indexing and searching images on the Web. This paper proposes a novel method for image annotation based on feature-word and word-topic distributions. The introduction of topics enables us to efficiently take word associations, such as {ocean, fish, coral}, into image annotation. Feature-word distributions are utilized to define weights in computation of topic distributions for annotation. By doing so, topic models in text mining can be applied directly in our method. Experiments show that our method is able to obtain promising improvements over the state-of-the-art method - Supervised Multiclass Labeling (SML) Cam-Tu Nguyen, Natsuda Kaothanthong, Xuan-Hieu Phan, Takeshi Tokuyama |
CIKM | 1 |
| 2009 | Web Search Clustering and Labeling with Hidden TopicsabstractWeb search clustering is a solution to reorganize search results (also called “snippets”) in a more convenient way for browsing. There are three key requirements for such post-retrieval clustering systems: (1) the clustering algorithm should group similar documents together; (2) clusters should be labeled with descriptive phrases; and (3) the clustering system should provide high-quality clustering without downloading the whole Web page. This article introduces a novel framework for clustering Web search results in Vietnamese which targets the three above issues. The main motivation is that by enriching short snippets with hidden topics from huge resources of documents on the Internet, it is able to cluster and label such snippets effectively in a topic-oriented manner without concerning whole Web pages. Our approach is based on recent successful topic analysis models, such as Probabilistic-Latent Semantic Analysis, or Latent Dirichlet Allocation. The underlying idea of the framework is that we collect a very large external data collection called “universal dataset,” and then build a clustering system on both the original snippets and a rich set of hidden topics discovered from the universal data collection. This can be seen as a richer representation of snippets to be clustered. We carry out careful evaluation of our method and show that our method can yield impressive clustering quality. Cam-Tu Nguyen, Xuan-Hieu Phan, Susumu Horiguchi, Thu-Trang Nguyen, Quang-Thuy Ha |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2008 | Word Segmentation of Vietnamese Texts: a Comparison of Approaches
Quang Thang Dinh, Hong Phuong Le, Nguyên Thi Minh Huyên, Cam-Tu Nguyen, Mathias Rossignol, Xuân Luong Vu |
LREC | 4 |
| 2008 | Matching and Ranking with Hidden Topics towards Online Contextual AdvertisingabstractIn online contextual advertising, ad messages are displayed related to the content of the target Web page. It leads to the problem in information retrieval community: how to select the most relevant ad messages given the content of a page. To deal with this problem, we propose a framework that takes advantage of large scale external datasets. This framework provides a mechanism to discover the semantic relations between Web pages and ad messages by analyzing topics for them. This helps overcome the problem of mismatch due to unimportant words and the difference in vocabularies between Web pages and ad messages. The framework has been evaluated through a number of experiments. It shows a significant improvement in accuracy over word/lexicon-based matching and ranking methods. Dieu-Thu Le, Cam-Tu Nguyen, Quang-Thuy Ha, Xuan-Hieu Phan, Susumu Horiguchi |
Web Intelligence | 2 |
| 2006 | Vietnamese Word Segmentation with CRFs and SVMs: An Investigation
Cam-Tu Nguyen, Trung-Kien Nguyen, Xuan-Hieu Phan, Minh Le Nguyen 0001, Quang-Thuy Ha |
PACLIC | 1 |