This website requires JavaScript.
97c208cfdb
b300eq: TP2PP4 MRR64 retest (pass-1 + ordered v2 + MRR48 control) — c64 gains +27%/+19%, TTFT collapse, -14% 1K c32 config cost attributed
main
yy-fighting
2026-09-10 21:29:05 +08:00
b9e8c0c4a2
CURRENT.md 60.1: two campaign windows restored; D-upgrade recipe is D-config-only; A-baseline verdicts (image/autotune negative, IPC@TP8 real but patch-coupled)
yy-fighting
2026-09-10 19:13:44 +08:00
2370d33c73
glm53 dsv4-migration A-baseline (TP8+EAGLE) campaign: image negative, autotune split-negative, IPC@TP8 real win +4.6~9.1% but net-neutral vs production (patch/image coupling); eager-trap profile + day-drift finding (60.1, 09-10)
yy-fighting
2026-09-10 19:04:28 +08:00
c5d91ceafe
b300-equivalent matrix: E7b high-concurrency retest (MRR64 + decode-graph buckets 1-64) - graph-drop cliff fixed, report numbers overwritten in place
yy-fighting
2026-09-10 18:52:54 +08:00
d063bd63c9
b300-equivalent matrix: add provenance.md (gpu inventory, vram peaks, qg verdicts, in-service preserve/restore record)
yy-fighting
2026-09-10 17:14:38 +08:00
21fcca5d16
b300-equivalent matrix: dual-plan (TP2PP4-D vs E7b) full B300 scenario replication on 60.8 - 41 valid points, divide at C=8, E7b usable window <=C8 (MRR16+graph-drop), boundary 256K/512K/896K TP2PP4-only, 4-5x absolute gap vs B300 narrowing to ~2x at boundary prefill; hicache host-layer cold-cache pitfall documented; in-service container preserved-renamed-restored and verified (09-10)
yy-fighting
2026-09-10 17:13:50 +08:00
35512db505
[Docs] add standalone B300 DeepSeek-V4-Flash report
hzy
Zhiyi Hong
2026-09-10 16:11:01 +08:00
81d17407bc
[Artifacts] finalize B300 matrix at reclaim cutoff
Zhiyi Hong
2026-09-10 15:56:57 +08:00
9652bfdb9d
[Artifacts] update B300 matrix through 15:23
Zhiyi Hong
2026-09-10 15:29:00 +08:00
7984c25586
[Artifacts] archive B300 DSV4 and GLM-5.3 matrix snapshot
Zhiyi Hong
2026-09-10 10:52:03 +08:00
e3476aff86
glm53 dsv4-migration bench: autotune+latest-image winner (+1.2~3.2% all 5 pts, A/B/A confirmed), PCIe-IPC pack negative (-0.4~-3.5%), page-mark kernel N/A for GLM (60.1, 09-09)
yy-fighting
2026-09-10 03:13:02 +08:00
f15003d2f5
hit90 bench: add aggregated summaries (41 scored points) + md5 manifest (70 raw logs on 60.8)
yy-fighting
2026-09-09 17:11:37 +08:00
ad1853f49b
hit90 scenario bench: TP4PP2-nomtp@0.90 winner, DP-attention/DCP verdicts (60.8 serving, 60.5 v3 delivered, 09-09)
yy-fighting
2026-09-09 17:11:04 +08:00
0d929948a5
60.4 redeploy: TP8+EAGLE+custom-AR 1stage (E7b recipe) replaces TP2PP4 (09-09)
yy-fighting
2026-09-09 15:54:52 +08:00
92517e87f0
128k low-cc capacity topology: TP2PP4-nomtp winner (60.8 serving, 60.5 v2 delivered, 09-09)
yy-fighting
2026-09-09 12:43:11 +08:00
498e333730
CURRENT.md: 60.4 now serving TP2PP4 GLM-5.3-NVFP4 (D-recipe replica, 09-09)
yy-fighting
2026-09-09 10:55:59 +08:00
0c89fd4fbe
r37 addendum: nomtp sweep raw logs (primary evidence, force-added past *.log ignore)
yy-fighting
2026-09-09 09:24:43 +08:00
4d6dac010c
r37 addendum: B'(nomtp) vs PP+MTP same-profile input/output throughput comparison (i16k/o512 random-ids cc8-64)
yy-fighting
2026-09-09 09:24:25 +08:00
911a9a2fb0
r37 addendum: i16k/o512 cc 8-64 sweep (nreq=2cc) on serving container; throughput saturates 130-135 tok/s at cc16+, accept stable; rc=1 root-caused (missing bs_results dir, never existed)
yy-fighting
2026-09-09 08:29:51 +08:00
a3b8f1b2ab
r37: fix PP+MTP verify CUDA graph (pre-planned path missing pp_proxy fill); beats A16 by 40% on killer, now serving on 60.8
yy-fighting
2026-09-09 01:24:29 +08:00
ffda226f5a
PP+MTP r36 de-GLOO: fix + verdict (PP+MTP loses i8k 1.81x to B', graph mode correctness-broken, A16 restored)
yy-fighting
2026-09-08 22:21:20 +08:00
123023b6ae
feat(pro6000/GLM-5.3): i8k/o1k/c16 三方案对比压测入库(A/A16/B 轮换实测 + 60.8 转 A16 在役 + CURRENT.md 台账修正 60.2/60.3/60.6/60.8)
yy-fighting
2026-09-08 19:04:41 +08:00
5c749cda03
feat(pro6000/GLM-5.3): 方案 D/E/F 部署资产入库(TP2PP4 生产配方 / TP8+DFlash2 / PD 分离四角色链)
yy-fighting
2026-09-08 16:36:15 +08:00
b3165a1d3c
feat(pro6000/GLM-5.3): 部署方案入库(TP8+EAGLE 生产标准 / 场景二高并发变体 / TP4PP2+IndexCache / E7b CAR 实验补丁 + deploy profiles)
yy-fighting
2026-09-08 11:38:06 +08:00
f0ab17c561
feat(pro6000/GLM-5.3): 双场景压测标准入库(bench_corpus 真实语料工具链 + run-id 窗口纪律 + 质量门禁 + 语料构建链)
yy-fighting
2026-09-08 11:37:54 +08:00
9e56401384
PP+MTP deepdive r3: race bisect (mask tooling, r34 candidate), decode-round quantification (gloo rendezvous stalls + AR spin dominate; 3-source hypothesis refuted), bench-profile crash forensics
yy-fighting
2026-09-07 02:43:31 +08:00
5ed30006a5
[Experiment] GLM-5.3-NVFP4 TP4PP2 round-2 optimization: all config-level quick wins refuted (2026-09-07)
yy-fighting
2026-09-07 01:04:37 +08:00
1e8c36b7d1
[Experiment] GLM-5.3-NVFP4 TP4PP2 torch-profiler profile on 174.1.60.5 (2026-09-06)
yy-fighting
2026-09-07 00:07:04 +08:00
75ab8865de
test: benchmark DP2 TP32 with DFlash
hzy-kimi3-pd-pp8-standard
Zhiyi Hong
2026-09-02 15:18:11 +08:00
411802e4b1
benchmark Kimi-K3 DP2 TP32 and PP8 without speculation
Zhiyi Hong
2026-09-02 13:40:58 +08:00
c9305fb830
profile Kimi-K3 PD DFlash 16K512 C8
Zhiyi Hong
2026-09-02 11:16:53 +08:00
3fbfb84e12
fix: validate Kimi PP8 DFlash PD on GSM8K C1 and C8
Zhiyi Hong
2026-08-31 18:29:06 +08:00
0e33209ae5
fix: derive hybrid PD transfer bounds from full-attention layers
Zhiyi Hong
2026-08-31 17:53:30 +08:00
91ab3fe32d
feat: add Kimi-K3 PP8 DFlash PD integration and warmup regression
Zhiyi Hong
2026-08-31 17:10:00 +08:00
8189942353
[Feature] Add Kimi-K3 standard PD deployment
Zhiyi Hong
2026-08-27 14:02:33 +08:00
c06ef5fd61
[Feature] Add Kimi-K3 standard PD deployment
Zhiyi Hong
2026-08-27 14:02:33 +08:00
a58b931cc8
[Profiling] Explain Kimi-K3 Deep PP scaling
hzy-kimi-k3-sm120-flashinfer-mxfp4
Zhiyi Hong
2026-08-21 18:05:42 +08:00
0433fcc3ee
[Benchmark] Add Kimi-K3 PP16 prefill results
Zhiyi Hong
2026-08-21 17:08:26 +08:00
d790df39b2
[Docs] Explain Kimi-K3 EP32 and EP4 execution
Zhiyi Hong
2026-08-21 15:19:12 +08:00
96ffea6d37
[Docs] Explain Kimi-K3 Deep PP Prefill optimization
Zhiyi Hong
2026-08-21 14:53:44 +08:00
d7381abe84
[Test] Add Kimi-K3 Prefill PP baseline search
Zhiyi Hong
2026-08-21 14:19:45 +08:00
60f77cd4ef
[Profile] Attribute Kimi-K3 Prefill communication
Zhiyi Hong
2026-08-20 15:56:58 +08:00
9cdbc1fd22
Revert "feat(p800): parameterize GLM5.2 single-node tuning"
Spike
2026-08-20 07:33:38 +00:00
ba20973bea
feat(p800): parameterize GLM5.2 single-node tuning
Spike
2026-08-20 07:27:02 +00:00
6fac5ad567
[Docs] Attribute Kimi-K3 Prefill collectives
Zhiyi Hong
2026-08-20 13:54:18 +08:00
2a3b12fa78
[Docs] Reject Kimi-K3 Prefill TP Reduce Scatter path
Zhiyi Hong
2026-08-20 10:25:37 +08:00
08a35066d7
[Test] Complete Kimi-K3 Prefill MoE backend report
Zhiyi Hong
2026-08-19 16:54:45 +08:00
74ec19dd48
[Docs] Close Kimi SM120 delivery audit
Zhiyi Hong
2026-08-19 14:25:03 +08:00
0fdcab9927
[Docs] Rebase Kimi SM120 Draft onto synced main
Zhiyi Hong
2026-08-19 14:02:35 +08:00
a5248ed80e
[Docs] Record exact Draft patch verification
Zhiyi Hong
2026-08-19 13:31:56 +08:00
fc336a3c7b
[Test] Finalize Kimi SM120 PR representative benchmark
Zhiyi Hong
2026-08-19 13:11:21 +08:00
63f2327a90
[Fix] Validate per-request benchmark errors correctly
Zhiyi Hong
2026-08-19 11:54:34 +08:00
a90c898683
[Fix] Use official kernel version-check override for validation
Zhiyi Hong
2026-08-19 11:29:55 +08:00
ab9a5422f6
[Fix] Persist FlashInfer JIT cache across services
Zhiyi Hong
2026-08-19 11:16:07 +08:00
7f67dfe6b3
[Fix] Keep ABI-matched SGLang kernel in PR image
Zhiyi Hong
2026-08-19 11:00:50 +08:00
e8ff3ce1e8
[Test] Add exact Kimi SM120 PR validation point
Zhiyi Hong
2026-08-19 10:47:26 +08:00
eeec56c2cf
update(P800/GLM5.2): 单节点部署测试脚本
Spike
2026-08-19 02:24:17 +00:00
d28db48e4b
[Docs] Scope SGLang draft to compatibility
Zhiyi Hong
2026-08-19 00:12:38 +08:00
39f692caae
[Docs] Keep draft checklist evidence-based
Zhiyi Hong
2026-08-19 00:10:20 +08:00
a9206ff105
[Docs] Add Kimi SM120 completion audit
Zhiyi Hong
2026-08-18 23:58:37 +08:00
a11c80b703
[Docs] Finalize Kimi SM120 SGLang draft PR
Zhiyi Hong
2026-08-18 23:47:34 +08:00
e01df16667
[Docs] Prepare Kimi SM120 SGLang draft PR
Zhiyi Hong
2026-08-18 23:16:53 +08:00
ec7b604a50
[Docs] Record Kimi EP4 MoE backend acceptance
Zhiyi Hong
2026-08-18 18:38:59 +08:00
27b8be09cb
[Fix] Keep Kimi benchmark tokenizer offline
Zhiyi Hong
2026-08-18 15:12:24 +08:00
be9d6bfe3a
[Fix] Materialize FlashInfer MXFP8 input layout
Zhiyi Hong
2026-08-18 14:50:13 +08:00
d0863501ca
[Fix] Support legacy Kimi MoE runner config
Zhiyi Hong
2026-08-18 14:33:42 +08:00
13944079fa
[Test] Add EP4 maximum-pressure capacity probe
Zhiyi Hong
2026-08-18 13:57:31 +08:00
e3974e2352
[Fix] Patch Kimi image for SM120 FlashInfer MXFP4
Zhiyi Hong
2026-08-18 13:05:07 +08:00
b50de8fe99
[Fix] Keep Kimi image dependency baseline for Phase 5
Zhiyi Hong
2026-08-18 12:45:44 +08:00
c14f8aa43a
[Fix] Preserve FlashInfer wheel filename in image build
Zhiyi Hong
2026-08-18 12:39:30 +08:00
daeffd147b
[Fix] Use built-in random IDs for Kimi Prefill matrix
Zhiyi Hong
2026-08-18 12:36:47 +08:00
5454fb984e
[Test] Add Kimi SM120 real-serving MoE backend matrix
Zhiyi Hong
2026-08-18 12:30:46 +08:00
c8f30ab7dc
[Perf] Profile Kimi SM120 FlashInfer MXFP4 MoE
Zhiyi Hong
2026-08-18 11:04:06 +08:00
6493798ad5
[Feature] Complete Kimi SM120 FlashInfer MXFP4 integration
Zhiyi Hong
2026-08-17 14:58:52 +08:00
a1c18d736b
[Test] Add Kimi SM120 MXFP4 correctness matrix
Zhiyi Hong
2026-08-17 12:08:06 +08:00
dac1bb652d
[Test] Reproduce Kimi SM120 SiTU contract gap
Zhiyi Hong
2026-08-14 17:09:13 +08:00
0684d269df
[Docs] Audit Kimi-K3 SM120 FlashInfer MXFP4 gap
Zhiyi Hong
2026-08-14 15:52:14 +08:00
987f1db4b0
feat(pro6000): Kimi-K3 DP=2 部署(方案 B:两个独立 TP32×EP32 实例 + router 负载均衡)
shishi
2026-08-11 17:55:00 +08:00
c69831f258
feat(pd): PD 长上下文 adaptive concurrency bench(SLO 方案 A,并发 +16)
shishi
2026-08-11 14:45:36 +08:00
fe375e1307
docs(pd): 补充运维一键部署完整步骤(干净环境从零到跑通)
shishi
2026-08-11 11:55:39 +08:00
ddf807d458
feat(bench/pd): 支持 --flush-cache 透传 + PD profile 加 --disable-radix-cache
shishi
2026-08-11 11:42:43 +08:00
da0e1b4372
fix(pd): config.env 加 DOCKER_CLIENT_IMAGE(bench 用 docker client 复用 kimi-k3 镜像)
shishi
2026-08-11 11:19:42 +08:00
04dfa31583
fix(pd): PATCH_MOUNTS 补充 flashkda wheel 挂载(BOOTSTRAP 需要 /flash_kda-*.whl)
shishi
2026-08-11 10:50:43 +08:00
d72dbff689
feat(pro6000): Kimi-K3 PD 分离部署(MoonCake RDMA)- 8 节点 P/D 双 profile + deploy_pd.sh 编排 + 文档
shishi
2026-08-11 10:36:09 +08:00
3bd04698bb
feat(pro6000): Kimi-K3 TP32×EP32 部署 profile、sm_120 补丁与运维手册 - 4 节点 RoCE 部署 + bench 实验
shishi
2026-08-10 10:37:13 +08:00
d9a2e3b5e5
fix(910c/glm52): 镜像与 client 修正 - 部署在 910c.2 的 GLM5.2-tuned 镜像
shishi
2026-08-03 17:22:04 +08:00
f33c5f1d3d
fix(deploy): DP_FLAG 未达 BOOTSTRAP 启动命令 + bench 的 model/tokenizer 分离
shishi
2026-08-03 16:41:29 +08:00
ea8302561e
fix(910c): bench client 可用化 - 本地 vllm-ascend-sglang 镜像 + torch_npu 自动加载禁用
shishi
2026-08-03 15:43:45 +08:00
6ba04325d3
feat(910c): 部署解耦 - vLLM-Ascend profile 与 deploy 层接管服务启停
shishi
2026-08-03 15:36:11 +08:00
9acf9fdfdb
feat(pro6000): 部署/测试解耦 - deploy 层支持多节点与 vLLM,新增 6 个 profile
Zhiyi Hong
2026-08-03 15:17:41 +08:00
1aa6c0f911
[Docs] narrow Phase 3 artifact scope
Zhiyi Hong
2026-08-03 10:12:44 +08:00
783c9325ae
[Artifacts] archive Phase 3 Nsight reports
Zhiyi Hong
2026-08-03 10:06:57 +08:00
56d286bb4e
[Docs] define Phase 3 raw artifact archive
Zhiyi Hong
2026-08-03 10:03:35 +08:00
3761d75b00
feat(ops): add unified bench/deploy layers and P800 profile
yy-fighting
2026-08-02 15:36:28 +08:00
41e2b000a4
[Docs] finalize Phase 3 timeline profiling results
Zhiyi Hong
2026-08-02 01:16:49 +08:00
82b7d91ac0
[BugFix] capture mixed trace after prefill admission
Zhiyi Hong
2026-08-02 00:30:46 +08:00
bc491eeeed
[BugFix] align mixed Phase 3 capture with prefill injection
Zhiyi Hong
2026-08-02 00:05:54 +08:00
4628d49755
[Feat] finalize Phase 3 SGLang timeline capture
Zhiyi Hong
2026-08-01 19:38:36 +08:00
e1719bd575
[Docs] finalize Phase 2.5 RDMA demand model
Zhiyi Hong
2026-08-01 15:29:09 +08:00
c5fa700c50
[Feat] add Phase 2.5 RDMA demand modeling
Zhiyi Hong
2026-08-01 02:40:06 +08:00