Multi-Tenant Edge-Cloud Hybrid LLM Serving using Speculative Decoding and Quantization
This paper proposes a multi-tenant edge-cloud hybrid LLM serving architecture that combines W4A16-quantized speculative decoding with asynchronous communication and a multi-tenant verification orchestrator to simultaneously address privacy, latency, and resource utilization challenges.