Abstract
Large language models (LLMs) have driven sub-stantial performance gains in natural language processing (NLP), opening up the possibility of nascent intelligent applications like chatbots. However, current LLM-based applications confront with the challenges of considerable inference cost, unstable response latency and restricted application effectiveness. In this paper, we propose EC-LLM, an edge-cloud collaborative LLM cascading scheme to achieve low-cost and low-latency LLM inference while maintaining high quality. Specifically, the light-weight LLM deployed at the edge (termed edge-LLM) is first fine-tuned on frequently asked user queries for achieving comparable performance to the complicated LLM at the cloud (termed cloud-LLM) within certain specialized domains. Then, we introduce two early-stop triggers (i.e., foresight and hindsight) to avoid unnecessary edge-LLM inference when the current query is deemed beyond the edge-LLM’s capabilities. Particularly, the foresight trigger acts on user queries to identify the out-of-domain queries, the hindsight trigger works during the output generation with a novel double-value head (i.e., confidence and importance head) model architecture. Extensive experiments based on our real-world edge-cloud collaborative prototype and typical NLP benchmarks demonstrate that, compared to state-of-the-art solutions, EC-LLM achieves at least 21% cost saving and up to 54% latency reduction, with no more than 4% accuracy loss. Moreover, benefiting from our early-stop triggers, EC-LLM exhibits 2%-10% latency less with the same answer accuracy compared to existing intuitive LLM cascading approaches.
| Original language | English |
|---|---|
| Journal | IEEE Internet of Things Journal |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- cost-quality trade-off
- edge-cloud collaboration
- Large language model
- low-latency inference
- model cascading
Fingerprint
Dive into the research topics of 'EC-LLM: Low-cost, Low-latency, High-quality Large Language Model Inference Based on Edge-cloud Collaboration'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver