跳到主要导航 跳到搜索 跳到主要内容

EC-LLM: Low-cost, Low-latency, High-quality Large Language Model Inference Based on Edge-cloud Collaboration

  • National Engineering Laboratory for Big Data Analytics (NEL-BDA)
  • School of Mathematics and Statistics
  • Xi'an Jiaotong University
  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

摘要

Large language models (LLMs) have driven sub-stantial performance gains in natural language processing (NLP), opening up the possibility of nascent intelligent applications like chatbots. However, current LLM-based applications confront with the challenges of considerable inference cost, unstable response latency and restricted application effectiveness. In this paper, we propose EC-LLM, an edge-cloud collaborative LLM cascading scheme to achieve low-cost and low-latency LLM inference while maintaining high quality. Specifically, the light-weight LLM deployed at the edge (termed edge-LLM) is first fine-tuned on frequently asked user queries for achieving comparable performance to the complicated LLM at the cloud (termed cloud-LLM) within certain specialized domains. Then, we introduce two early-stop triggers (i.e., foresight and hindsight) to avoid unnecessary edge-LLM inference when the current query is deemed beyond the edge-LLM’s capabilities. Particularly, the foresight trigger acts on user queries to identify the out-of-domain queries, the hindsight trigger works during the output generation with a novel double-value head (i.e., confidence and importance head) model architecture. Extensive experiments based on our real-world edge-cloud collaborative prototype and typical NLP benchmarks demonstrate that, compared to state-of-the-art solutions, EC-LLM achieves at least 21% cost saving and up to 54% latency reduction, with no more than 4% accuracy loss. Moreover, benefiting from our early-stop triggers, EC-LLM exhibits 2%-10% latency less with the same answer accuracy compared to existing intuitive LLM cascading approaches.

源语言英语
期刊IEEE Internet of Things Journal
DOI
出版状态已接受/待刊 - 2026
已对外发布

学术指纹

探究 'EC-LLM: Low-cost, Low-latency, High-quality Large Language Model Inference Based on Edge-cloud Collaboration' 的科研主题。它们共同构成独一无二的学术指纹。

引用此