sskj/docs/SLO_STANDARDS.md

34 lines
1.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 推理服务 SLO 标准
> 记录各模型层级Tier的延迟目标。TTFT 以 **P95** 为主要关注指标。
## 模型层级定义
| 层级 | 名称 | 参数规模 | 典型模型 | 应用场景 | 核心目标 |
|---|---|---|---|---|---|
| S0 | 轻量层 | < 30B | Llama 3.1 8B Instruct | 高并发轻推理 | 高并发轻推理 |
| S1 | 实时交互层 | [30B, 100B) | Llama 3.3 70B Instruct | Chat / Copilot / Agent | 极低延迟 |
| S2 | 均衡服务层 | [100B, 500B) | DeepSeek V4 FlashMiniMax M2.7 | 企业 API | 性价比平衡 |
| S3 | 深度推理层 | [500B, 1T) | DeepSeek V3.2 | 高复杂推理 | 稳定吞吐 |
| S4 | 超级推理层 | 1T | DS V4 Pro | 深度思考 | 智能优先 |
## TTFT SLOP95 为主要关注指标)
| 指标 | S0 | S1 | S2 | S3 | S4 |
|---|---|---|---|---|---|
| P50 | < 500ms | < 800ms | < 1.5s | < 2s | < 1.5s |
| P95 | < 0.4s | < 1s | < 3s | < 5s | < 3s |
## TPOT SLO
| 指标 | S0 | S1 | S2 | S3 | S4 |
|---|---|---|---|---|---|
| TPOT | < 20ms | < 40ms | < 50ms | < 50ms | < 50ms |
## 说明
- **TTFT**Time To First Token从请求到达服务端到首个 token 返回的延迟
- **TPOT**Time Per Output Token除首 token 平均每输出一个 token 的耗时
- 当前重点关注 **TTFT P95** 95% 请求的 TTFT 应满足对应层级的阈值
- _benchmark 输出应同时给出 P50 / P90 / P95 / P99 TTFT但评估以 P95 为准_