Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
Orthrus is a memory-efficient dual-architecture framework that integrates a lightweight diffusion module with a frozen autoregressive LLM to enable parallel token generation while guaranteeing lossless inference fidelity through a shared KV cache and exact consensus mechanism.