evalstone/EVALUATION_GUIDE.md
2026-07-21 09:32:49 +00:00

267 lines
8.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# DeepSeek-V4-Flash-INT8 评测经验总结
## 1. 环境配置
### 1.1 硬件
- GPU: 8× P800 OAM, 96GB each (XPU)
- 模型: DeepSeek-V4-Flash-INT8
- 服务: sglang at `http://localhost:30000/v1`
- 上下文长度: 1M tokens (`--context-length 1048576`)
### 1.2 sglang 启动参数
```bash
sglang.launch_server \
--host 0.0.0.0 --port 30000 \
--model-path /data1/models/DeepSeek-V4-Flash-INT8 \
--attention-backend nsa --nsa-prefill klxdsa --nsa-decode klxdsa \
--trust-remote-code --disable-custom-all-reduce \
--tensor-parallel-size 8 --ep-size 8 \
--disable-shared-experts-fusion --page-size 64 \
--mem-fraction-static 0.75 --quantization w8a8_int8 \
--kv-cache-dtype float16 --disable-piecewise-cuda-graph \
--cuda-graph-max-bs 16 --context-length 1048576 \
--watchdog-timeout 3000000 \
--tool-call-parser deepseekv4 --reasoning-parser deepseek-v4 \
--constrained-json-disable-any-whitespace \
--enable-metrics --enable-request-time-stats-logging
```
### 1.3 Python 环境
- conda env: `evalscope` (Python 3.12)
- 正确路径: `/data1/miniconda3/envs/evalscope/bin/python`
- 注意: `conda run -n evalscope` 会覆盖 PYTHONPATH导致 ms_enclave 找不到
## 2. 评测脚本
### 2.1 主运行脚本
- `/data1/sora/evalscope/bash/run.py` - 主评测驱动
- `/data1/sora/evalscope/bash/run_limit30.py` - limit=100 快速 benchmark
- `/data1/sora/evalscope/bash/run_swe_bench.py` - swe_bench 系列 (limit=3)
- `/data1/sora/evalscope/bash/auto_run.py` - 自动顺序运行
- `/data1/sora/evalscope/bash/run_longbench_oom_test.py` - 长文本 OOM 测试
### 2.2 配置文件
- `/data1/sora/evalscope/bash/dpv4-int8_nothinking.yaml` - 各数据集生成配置
### 2.3 关键参数
```python
# run.py 顶部可调参数
DATASETS = [...] # 要评测的数据集列表
ENABLE_THINKING = False # 是否启用 thinking 模式
BATCH_SIZE_LIST = [4] # 评测 batch size
LIMIT = 10 # 每个子集最大样本数 (None=全部)
SEED = 42 # 随机种子
SHUFFLE = True # 是否打乱样本
```
## 3. 各数据集配置
### 3.1 温度设置
- 所有数据集: `temperature: 0.0`(确定性输出)
- temperature=1.0 需要多次运行取平均
### 3.2 max_tokens 设置
| 类别 | 数据集 | max_tokens |
|------|--------|-----------|
| 代码/工程 | terminal_bench_v2 | 262144 |
| 代码/工程 | live_code_bench, bigcodebench, humaneval | 32768-16384 |
| 推理/数学 | aime24/25/26, hmmt26, competition_math, gsm8k | 32768 |
| 推理/数学 | imo_answerbench | 98304 |
| 推理/数学 | gpqa_diamond, hle, super_gpqa | 8192 |
| 知识 | mmlu_pro, simple_qa, arc, bbh, cmmlu, drop, hellaswag, mmlu, trivia_qa, winogrande | 8192-32768 |
| 长上下文 | longbench_v2, openai_mrcr | 8192 |
| 智能体/工具 | browsecomp, mcp_atlas, tau2_bench | 16384-32768 |
| 智能体/工具 | swe_bench 系列 | 32768 |
### 3.3 Agent 配置
- swe_bench 系列: `max_steps: 300`
- tau2_bench: `max_steps: 50`
- terminal_bench_v2: `max_turns: 300`, `timeout_multiplier: 6.0` (1小时)
### 3.4 沙箱配置
```python
# 需要沙箱的数据集
SANDBOX_DATASETS = {'humaneval', 'bigcodebench'}
# 沙箱配置
SandboxTaskConfig(
enabled=True,
engine='docker',
default_config={
'image': 'python:3.11-slim',
'tools_config': {
'shell_executor': {},
'python_executor': {}
}
}
)
```
## 4. 长文本截断
### 4.1 当前配置
```python
TRUNCATION_CONFIG = {
'longbench_v2': 32768*4, # 131072 tokens (128K)
'openai_mrcr': 32768*4, # 131072 tokens (128K)
}
```
### 4.2 截断方式
- Middle-truncation: 保留头部和尾部,截断中间
- 适用于长上下文 benchmark保留问题和上下文的关键部分
### 4.3 OOM 测试
- 测试脚本: `run_longbench_oom_test.py`
- 测试 limit: 128K, 256K, 512K, 1M
- 输出目录: `output_longbench_128k/`, `output_longbench_256k/`
## 5. Docker 镜像
### 5.1 需要预拉的镜像
```bash
# 代码评测
docker pull docker.1ms.run/bigcodebench/bigcodebench-evaluate:latest
# Python 沙箱
docker pull python:3.11-slim
```
### 5.2 按样本拉取的镜像(无法预准备)
- **swe_bench 系列**: 每个样本需要不同的 `jefzda/sweap-images:{tag}`
- 每个镜像 2-5GB
- limit=3 需要 ~10GB
- 完整 500 样本需要 ~1.75TB(不现实)
- 建议: limit=1-3 或跳过
- **terminal_bench_v2**: 使用统一基础镜像,但任务环境不同
- 镜像在运行时自动构建/拉取
- 超时问题常见,需要 `timeout_multiplier: 6.0`
### 5.3 镜像源
- Docker Hub 被墙,使用镜像站:
- `docker.1ms.run`
- `docker.m.daocloud.io`
- `docker.1panel.dev`
### 5.4 磁盘空间管理
- sglang 镜像: ~58GB不能删
- bigcodebench: ~25GB跑完可删
- sweap-images: ~21GB按需删除
- terminal_bench: ~2GB可删
## 6. 常见问题与解决方案
### 6.1 tau2_bench 空响应问题
**问题**: `UserMessage must have either content or tool_calls`
**解决**: 修改 `tau2-bench/src/tau2/user/user_simulator.py`,增加重试逻辑:
```python
max_retries = 5
for attempt in range(max_retries):
assistant_message = generate(...)
user_response = assistant_message.content
if user_response and user_response.strip():
break
if attempt < max_retries - 1:
logger.warning(f"Empty response, retrying...")
```
### 6.2 terminal_bench_v2 超时问题
**问题**: `EnvironmentStartTimeoutError: Environment start timed out after 600.0 seconds`
**解决**: 增加 `timeout_multiplier: 6.0` (1小时)
### 6.3 swe_bench 镜像拉取慢
**问题**: 网络慢,镜像大,磁盘满
**解决**:
- 使用国内镜像站
- 限制 limit=1-3
- 定期清理 Docker 镜像
### 6.4 模型服务重启
**问题**: sglang 服务可能崩溃
**解决**: `/data1/restart_model_service.sh`
## 7. 结果提取与报告
### 7.1 提取脚本
- `/data1/sora/evalscope/bash/extract_results.py`
- 从 JSON 报告提取 score、perf metrics、percentiles
### 7.2 Excel 报告
- `/data1/sora/P800模型能力评测结果_updated.xlsx`
- 包含: score, 时间, tokens, TTFT/TPOT percentiles
### 7.3 结果合并策略
- 优先使用 `output_limit100/` 的结果(样本更多)
- 然后使用 `output_swe_bench/` 的结果
- 最后使用 `output/` 的结果
## 8. 评测策略建议
### 8.1 快速评测 (1天)
- 使用 limit=100 跑快速 benchmark
- 跳过 swe_bench 系列
- 长文本使用 128K 截断
### 8.2 完整评测 (1周+)
- 所有 benchmark 全量运行
- swe_bench limit=3-5
- 长文本测试 256K/512K/1M 边界
### 8.3 关键 benchmark
- **推理/数学**: aime24/25/26, competition_math, gpqa_diamond
- **代码**: humaneval, live_code_bench, bigcodebench
- **知识**: mmlu, mmlu_pro, cmmlu
- **长文本**: longbench_v2, openai_mrcr
- **智能体**: browsecomp, tau2_bench
## 9. 文件清单
```
/data1/sora/evalscope/bash/
├── run.py # 主运行脚本
├── run_limit30.py # limit=100 快速运行
├── run_swe_bench.py # swe_bench 系列
├── auto_run.py # 自动顺序运行
├── run_longbench_oom_test.py # 长文本 OOM 测试
├── extract_results.py # 结果提取
├── dpv4-int8_nothinking.yaml # 数据集配置
├── generate_excel.py # Excel 生成
├── generate_dashboard.py # Dashboard 生成
/data1/sora/evalscope/output/ # 主输出目录
/data1/sora/evalscope/output_limit100/ # limit=100 输出
/data1/sora/evalscope/output_swe_bench/ # swe_bench 输出
/data1/sora/evalscope/output_longbench_*/ # 长文本测试输出
/data1/sora/P800模型能力评测结果_updated.xlsx # Excel 报告
```
## 10. 运行命令汇总
```bash
# 激活环境
source /data1/miniconda3/bin/activate evalscope
# 运行主评测
python /data1/sora/evalscope/bash/run.py
# 运行 limit=100 快速评测
python /data1/sora/evalscope/bash/run_limit30.py
# 运行 swe_bench
python /data1/sora/evalscope/bash/run_swe_bench.py
# 自动顺序运行
python /data1/sora/evalscope/bash/auto_run.py
# 长文本 OOM 测试
python /data1/sora/evalscope/bash/run_longbench_oom_test.py
# 提取结果
python /data1/sora/evalscope/bash/extract_results.py
# 重启模型服务
bash /data1/restart_model_service.sh
```