返回顶部
返回首页 会员充值 我的足迹 返回上一页

杨珂_Mooncake:解耦式架构和以存换算优化大模型推理.pdf

2026-06-13
文档编号:1268205
文档页数:50
文档大小:12.42MB
下载积分:VIP专享
文档格式:PDF

1、杨珂 趋境科技技术专家|Mooncake 核心贡献者MooncakeMooncake:Ke Yang,Approaching.AI Tech Expert|Mooncake Core Contributor解耦式架构和以存换算,优化大模型推理解耦式架构和以存换算,优化大模型推理目 录CONTENTS Background:LLM Inference in Long-contex xt EraMooncake:A KVCache-centric Disaggregated ArchitectureMooncake LL M Ecosystem CollaborationCurrent Parad

2、igm:Data+Algorithm+Hardware=IntelligenceAlgorithm-Transformer is all we need?Data Big Data is EverywhereHardware Huangs Law Take OverIntelligence AI Become Everywhere TooThe Old Scaling Law is Slowing downBUT,who use it?LargerModelMoreDataGrowingComputing PowerThe Old Scaling LawThe Old Scaling LawP

3、erformance gains from adding more parameters are increasingly limited.It is becoming difficult to gather enough high-quality data to feed ultra-large models.Everyone is Talking about Scaling Law But the Real Question is What to Scale?https:/ Data+Larger Model+Longer Context=Higher IntelligenceIn Jan

4、uary 2025,DeepSeek R1 quickly rose to become one of the most renowned large model services for its strong reasoning(long-output)capability.Long input-KimiLong output DeepSeek R1In March 2024,Kimi became one of the leading large model services thanks to its strong long-context(long-input)processing c

5、apability.More Data+Larger Model+Longer Context=Higher IntelligenceChain-of-ThoughtMore Data+Larger Model+Longer Context=Higher IntelligenceAI applications are evolving from simple chat to complex agent-based systems.Single-turn,short inputs/outputsMulti-turn,complex execution topologies,long inputs

6、/outputs.More Data+Larger Model+Longer Context=heavier workloadHiger Inference CostLonger Response TimeLack of Computing and Memory ResourcesOne of the key bottlenecks in the long-context era:Inference costs are skyrocketingAmazon reports that over 90%of costs come from inference rather than trainin

杨珂_Mooncake:解耦式架构和以存换算优化大模型推理.pdf_第1页
杨珂_Mooncake:解耦式架构和以存换算优化大模型推理.pdf_第2页
杨珂_Mooncake:解耦式架构和以存换算优化大模型推理.pdf_第3页
杨珂_Mooncake:解耦式架构和以存换算优化大模型推理.pdf_第4页
杨珂_Mooncake:解耦式架构和以存换算优化大模型推理.pdf_第5页

点击查看更多