Commit Graph

  • 97c208cfdb b300eq: TP2PP4 MRR64 retest (pass-1 + ordered v2 + MRR48 control) — c64 gains +27%/+19%, TTFT collapse, -14% 1K c32 config cost attributed main yy-fighting 2026-09-10 21:29:05 +08:00
  • b9e8c0c4a2 CURRENT.md 60.1: two campaign windows restored; D-upgrade recipe is D-config-only; A-baseline verdicts (image/autotune negative, IPC@TP8 real but patch-coupled) yy-fighting 2026-09-10 19:13:44 +08:00
  • 2370d33c73 glm53 dsv4-migration A-baseline (TP8+EAGLE) campaign: image negative, autotune split-negative, IPC@TP8 real win +4.6~9.1% but net-neutral vs production (patch/image coupling); eager-trap profile + day-drift finding (60.1, 09-10) yy-fighting 2026-09-10 19:04:28 +08:00
  • c5d91ceafe b300-equivalent matrix: E7b high-concurrency retest (MRR64 + decode-graph buckets 1-64) - graph-drop cliff fixed, report numbers overwritten in place yy-fighting 2026-09-10 18:52:54 +08:00
  • d063bd63c9 b300-equivalent matrix: add provenance.md (gpu inventory, vram peaks, qg verdicts, in-service preserve/restore record) yy-fighting 2026-09-10 17:14:38 +08:00
  • 21fcca5d16 b300-equivalent matrix: dual-plan (TP2PP4-D vs E7b) full B300 scenario replication on 60.8 - 41 valid points, divide at C=8, E7b usable window <=C8 (MRR16+graph-drop), boundary 256K/512K/896K TP2PP4-only, 4-5x absolute gap vs B300 narrowing to ~2x at boundary prefill; hicache host-layer cold-cache pitfall documented; in-service container preserved-renamed-restored and verified (09-10) yy-fighting 2026-09-10 17:13:50 +08:00
  • 35512db505 [Docs] add standalone B300 DeepSeek-V4-Flash report hzy Zhiyi Hong 2026-09-10 16:11:01 +08:00
  • 81d17407bc [Artifacts] finalize B300 matrix at reclaim cutoff Zhiyi Hong 2026-09-10 15:56:57 +08:00
  • 9652bfdb9d [Artifacts] update B300 matrix through 15:23 Zhiyi Hong 2026-09-10 15:29:00 +08:00
  • 7984c25586 [Artifacts] archive B300 DSV4 and GLM-5.3 matrix snapshot Zhiyi Hong 2026-09-10 10:52:03 +08:00
  • e3476aff86 glm53 dsv4-migration bench: autotune+latest-image winner (+1.2~3.2% all 5 pts, A/B/A confirmed), PCIe-IPC pack negative (-0.4~-3.5%), page-mark kernel N/A for GLM (60.1, 09-09) yy-fighting 2026-09-10 03:13:02 +08:00
  • f15003d2f5 hit90 bench: add aggregated summaries (41 scored points) + md5 manifest (70 raw logs on 60.8) yy-fighting 2026-09-09 17:11:37 +08:00
  • ad1853f49b hit90 scenario bench: TP4PP2-nomtp@0.90 winner, DP-attention/DCP verdicts (60.8 serving, 60.5 v3 delivered, 09-09) yy-fighting 2026-09-09 17:11:04 +08:00
  • 0d929948a5 60.4 redeploy: TP8+EAGLE+custom-AR 1stage (E7b recipe) replaces TP2PP4 (09-09) yy-fighting 2026-09-09 15:54:52 +08:00
  • 92517e87f0 128k low-cc capacity topology: TP2PP4-nomtp winner (60.8 serving, 60.5 v2 delivered, 09-09) yy-fighting 2026-09-09 12:43:11 +08:00
  • 498e333730 CURRENT.md: 60.4 now serving TP2PP4 GLM-5.3-NVFP4 (D-recipe replica, 09-09) yy-fighting 2026-09-09 10:55:59 +08:00
  • 0c89fd4fbe r37 addendum: nomtp sweep raw logs (primary evidence, force-added past *.log ignore) yy-fighting 2026-09-09 09:24:43 +08:00
  • 4d6dac010c r37 addendum: B'(nomtp) vs PP+MTP same-profile input/output throughput comparison (i16k/o512 random-ids cc8-64) yy-fighting 2026-09-09 09:24:25 +08:00
  • 911a9a2fb0 r37 addendum: i16k/o512 cc 8-64 sweep (nreq=2cc) on serving container; throughput saturates 130-135 tok/s at cc16+, accept stable; rc=1 root-caused (missing bs_results dir, never existed) yy-fighting 2026-09-09 08:29:51 +08:00
  • a3b8f1b2ab r37: fix PP+MTP verify CUDA graph (pre-planned path missing pp_proxy fill); beats A16 by 40% on killer, now serving on 60.8 yy-fighting 2026-09-09 01:24:29 +08:00
  • ffda226f5a PP+MTP r36 de-GLOO: fix + verdict (PP+MTP loses i8k 1.81x to B', graph mode correctness-broken, A16 restored) yy-fighting 2026-09-08 22:21:20 +08:00
  • 123023b6ae feat(pro6000/GLM-5.3): i8k/o1k/c16 三方案对比压测入库(A/A16/B 轮换实测 + 60.8 转 A16 在役 + CURRENT.md 台账修正 60.2/60.3/60.6/60.8) yy-fighting 2026-09-08 19:04:41 +08:00
  • 5c749cda03 feat(pro6000/GLM-5.3): 方案 D/E/F 部署资产入库(TP2PP4 生产配方 / TP8+DFlash2 / PD 分离四角色链) yy-fighting 2026-09-08 16:36:15 +08:00
  • b3165a1d3c feat(pro6000/GLM-5.3): 部署方案入库(TP8+EAGLE 生产标准 / 场景二高并发变体 / TP4PP2+IndexCache / E7b CAR 实验补丁 + deploy profiles) yy-fighting 2026-09-08 11:38:06 +08:00
  • f0ab17c561 feat(pro6000/GLM-5.3): 双场景压测标准入库(bench_corpus 真实语料工具链 + run-id 窗口纪律 + 质量门禁 + 语料构建链) yy-fighting 2026-09-08 11:37:54 +08:00
  • 9e56401384 PP+MTP deepdive r3: race bisect (mask tooling, r34 candidate), decode-round quantification (gloo rendezvous stalls + AR spin dominate; 3-source hypothesis refuted), bench-profile crash forensics yy-fighting 2026-09-07 02:43:31 +08:00
  • 5ed30006a5 [Experiment] GLM-5.3-NVFP4 TP4PP2 round-2 optimization: all config-level quick wins refuted (2026-09-07) yy-fighting 2026-09-07 01:04:37 +08:00
  • 1e8c36b7d1 [Experiment] GLM-5.3-NVFP4 TP4PP2 torch-profiler profile on 174.1.60.5 (2026-09-06) yy-fighting 2026-09-07 00:07:04 +08:00
  • 75ab8865de test: benchmark DP2 TP32 with DFlash hzy-kimi3-pd-pp8-standard Zhiyi Hong 2026-09-02 15:18:11 +08:00
  • 411802e4b1 benchmark Kimi-K3 DP2 TP32 and PP8 without speculation Zhiyi Hong 2026-09-02 13:40:58 +08:00
  • c9305fb830 profile Kimi-K3 PD DFlash 16K512 C8 Zhiyi Hong 2026-09-02 11:16:53 +08:00
  • 3fbfb84e12 fix: validate Kimi PP8 DFlash PD on GSM8K C1 and C8 Zhiyi Hong 2026-08-31 18:29:06 +08:00
  • 0e33209ae5 fix: derive hybrid PD transfer bounds from full-attention layers Zhiyi Hong 2026-08-31 17:53:30 +08:00
  • 91ab3fe32d feat: add Kimi-K3 PP8 DFlash PD integration and warmup regression Zhiyi Hong 2026-08-31 17:10:00 +08:00
  • 8189942353 [Feature] Add Kimi-K3 standard PD deployment Zhiyi Hong 2026-08-27 14:02:33 +08:00
  • c06ef5fd61 [Feature] Add Kimi-K3 standard PD deployment Zhiyi Hong 2026-08-27 14:02:33 +08:00
  • a58b931cc8 [Profiling] Explain Kimi-K3 Deep PP scaling hzy-kimi-k3-sm120-flashinfer-mxfp4 Zhiyi Hong 2026-08-21 18:05:42 +08:00
  • 0433fcc3ee [Benchmark] Add Kimi-K3 PP16 prefill results Zhiyi Hong 2026-08-21 17:08:26 +08:00
  • d790df39b2 [Docs] Explain Kimi-K3 EP32 and EP4 execution Zhiyi Hong 2026-08-21 15:19:12 +08:00
  • 96ffea6d37 [Docs] Explain Kimi-K3 Deep PP Prefill optimization Zhiyi Hong 2026-08-21 14:53:44 +08:00
  • d7381abe84 [Test] Add Kimi-K3 Prefill PP baseline search Zhiyi Hong 2026-08-21 14:19:45 +08:00
  • 60f77cd4ef [Profile] Attribute Kimi-K3 Prefill communication Zhiyi Hong 2026-08-20 15:56:58 +08:00
  • 9cdbc1fd22 Revert "feat(p800): parameterize GLM5.2 single-node tuning" Spike 2026-08-20 07:33:38 +00:00
  • ba20973bea feat(p800): parameterize GLM5.2 single-node tuning Spike 2026-08-20 07:27:02 +00:00
  • 6fac5ad567 [Docs] Attribute Kimi-K3 Prefill collectives Zhiyi Hong 2026-08-20 13:54:18 +08:00
  • 2a3b12fa78 [Docs] Reject Kimi-K3 Prefill TP Reduce Scatter path Zhiyi Hong 2026-08-20 10:25:37 +08:00
  • 08a35066d7 [Test] Complete Kimi-K3 Prefill MoE backend report Zhiyi Hong 2026-08-19 16:54:45 +08:00
  • 74ec19dd48 [Docs] Close Kimi SM120 delivery audit Zhiyi Hong 2026-08-19 14:25:03 +08:00
  • 0fdcab9927 [Docs] Rebase Kimi SM120 Draft onto synced main Zhiyi Hong 2026-08-19 14:02:35 +08:00
  • a5248ed80e [Docs] Record exact Draft patch verification Zhiyi Hong 2026-08-19 13:31:56 +08:00
  • fc336a3c7b [Test] Finalize Kimi SM120 PR representative benchmark Zhiyi Hong 2026-08-19 13:11:21 +08:00
  • 63f2327a90 [Fix] Validate per-request benchmark errors correctly Zhiyi Hong 2026-08-19 11:54:34 +08:00
  • a90c898683 [Fix] Use official kernel version-check override for validation Zhiyi Hong 2026-08-19 11:29:55 +08:00
  • ab9a5422f6 [Fix] Persist FlashInfer JIT cache across services Zhiyi Hong 2026-08-19 11:16:07 +08:00
  • 7f67dfe6b3 [Fix] Keep ABI-matched SGLang kernel in PR image Zhiyi Hong 2026-08-19 11:00:50 +08:00
  • e8ff3ce1e8 [Test] Add exact Kimi SM120 PR validation point Zhiyi Hong 2026-08-19 10:47:26 +08:00
  • eeec56c2cf update(P800/GLM5.2): 单节点部署测试脚本 Spike 2026-08-19 02:24:17 +00:00
  • d28db48e4b [Docs] Scope SGLang draft to compatibility Zhiyi Hong 2026-08-19 00:12:38 +08:00
  • 39f692caae [Docs] Keep draft checklist evidence-based Zhiyi Hong 2026-08-19 00:10:20 +08:00
  • a9206ff105 [Docs] Add Kimi SM120 completion audit Zhiyi Hong 2026-08-18 23:58:37 +08:00
  • a11c80b703 [Docs] Finalize Kimi SM120 SGLang draft PR Zhiyi Hong 2026-08-18 23:47:34 +08:00
  • e01df16667 [Docs] Prepare Kimi SM120 SGLang draft PR Zhiyi Hong 2026-08-18 23:16:53 +08:00
  • ec7b604a50 [Docs] Record Kimi EP4 MoE backend acceptance Zhiyi Hong 2026-08-18 18:38:59 +08:00
  • 27b8be09cb [Fix] Keep Kimi benchmark tokenizer offline Zhiyi Hong 2026-08-18 15:12:24 +08:00
  • be9d6bfe3a [Fix] Materialize FlashInfer MXFP8 input layout Zhiyi Hong 2026-08-18 14:50:13 +08:00
  • d0863501ca [Fix] Support legacy Kimi MoE runner config Zhiyi Hong 2026-08-18 14:33:42 +08:00
  • 13944079fa [Test] Add EP4 maximum-pressure capacity probe Zhiyi Hong 2026-08-18 13:57:31 +08:00
  • e3974e2352 [Fix] Patch Kimi image for SM120 FlashInfer MXFP4 Zhiyi Hong 2026-08-18 13:05:07 +08:00
  • b50de8fe99 [Fix] Keep Kimi image dependency baseline for Phase 5 Zhiyi Hong 2026-08-18 12:45:44 +08:00
  • c14f8aa43a [Fix] Preserve FlashInfer wheel filename in image build Zhiyi Hong 2026-08-18 12:39:30 +08:00
  • daeffd147b [Fix] Use built-in random IDs for Kimi Prefill matrix Zhiyi Hong 2026-08-18 12:36:47 +08:00
  • 5454fb984e [Test] Add Kimi SM120 real-serving MoE backend matrix Zhiyi Hong 2026-08-18 12:30:46 +08:00
  • c8f30ab7dc [Perf] Profile Kimi SM120 FlashInfer MXFP4 MoE Zhiyi Hong 2026-08-18 11:04:06 +08:00
  • 6493798ad5 [Feature] Complete Kimi SM120 FlashInfer MXFP4 integration Zhiyi Hong 2026-08-17 14:58:52 +08:00
  • a1c18d736b [Test] Add Kimi SM120 MXFP4 correctness matrix Zhiyi Hong 2026-08-17 12:08:06 +08:00
  • dac1bb652d [Test] Reproduce Kimi SM120 SiTU contract gap Zhiyi Hong 2026-08-14 17:09:13 +08:00
  • 0684d269df [Docs] Audit Kimi-K3 SM120 FlashInfer MXFP4 gap Zhiyi Hong 2026-08-14 15:52:14 +08:00
  • 987f1db4b0 feat(pro6000): Kimi-K3 DP=2 部署(方案 B:两个独立 TP32×EP32 实例 + router 负载均衡) shishi 2026-08-11 17:55:00 +08:00
  • c69831f258 feat(pd): PD 长上下文 adaptive concurrency bench(SLO 方案 A,并发 +16) shishi 2026-08-11 14:45:36 +08:00
  • fe375e1307 docs(pd): 补充运维一键部署完整步骤(干净环境从零到跑通) shishi 2026-08-11 11:55:39 +08:00
  • ddf807d458 feat(bench/pd): 支持 --flush-cache 透传 + PD profile 加 --disable-radix-cache shishi 2026-08-11 11:42:43 +08:00
  • da0e1b4372 fix(pd): config.env 加 DOCKER_CLIENT_IMAGE(bench 用 docker client 复用 kimi-k3 镜像) shishi 2026-08-11 11:19:42 +08:00
  • 04dfa31583 fix(pd): PATCH_MOUNTS 补充 flashkda wheel 挂载(BOOTSTRAP 需要 /flash_kda-*.whl) shishi 2026-08-11 10:50:43 +08:00
  • d72dbff689 feat(pro6000): Kimi-K3 PD 分离部署(MoonCake RDMA)- 8 节点 P/D 双 profile + deploy_pd.sh 编排 + 文档 shishi 2026-08-11 10:36:09 +08:00
  • 3bd04698bb feat(pro6000): Kimi-K3 TP32×EP32 部署 profile、sm_120 补丁与运维手册 - 4 节点 RoCE 部署 + bench 实验 shishi 2026-08-10 10:37:13 +08:00
  • d9a2e3b5e5 fix(910c/glm52): 镜像与 client 修正 - 部署在 910c.2 的 GLM5.2-tuned 镜像 shishi 2026-08-03 17:22:04 +08:00
  • f33c5f1d3d fix(deploy): DP_FLAG 未达 BOOTSTRAP 启动命令 + bench 的 model/tokenizer 分离 shishi 2026-08-03 16:41:29 +08:00
  • ea8302561e fix(910c): bench client 可用化 - 本地 vllm-ascend-sglang 镜像 + torch_npu 自动加载禁用 shishi 2026-08-03 15:43:45 +08:00
  • 6ba04325d3 feat(910c): 部署解耦 - vLLM-Ascend profile 与 deploy 层接管服务启停 shishi 2026-08-03 15:36:11 +08:00
  • 9acf9fdfdb feat(pro6000): 部署/测试解耦 - deploy 层支持多节点与 vLLM,新增 6 个 profile Zhiyi Hong 2026-08-03 15:17:41 +08:00
  • 1aa6c0f911 [Docs] narrow Phase 3 artifact scope Zhiyi Hong 2026-08-03 10:12:44 +08:00
  • 783c9325ae [Artifacts] archive Phase 3 Nsight reports Zhiyi Hong 2026-08-03 10:06:57 +08:00
  • 56d286bb4e [Docs] define Phase 3 raw artifact archive Zhiyi Hong 2026-08-03 10:03:35 +08:00
  • 3761d75b00 feat(ops): add unified bench/deploy layers and P800 profile yy-fighting 2026-08-02 15:36:28 +08:00
  • 41e2b000a4 [Docs] finalize Phase 3 timeline profiling results Zhiyi Hong 2026-08-02 01:16:49 +08:00
  • 82b7d91ac0 [BugFix] capture mixed trace after prefill admission Zhiyi Hong 2026-08-02 00:30:46 +08:00
  • bc491eeeed [BugFix] align mixed Phase 3 capture with prefill injection Zhiyi Hong 2026-08-02 00:05:54 +08:00
  • 4628d49755 [Feat] finalize Phase 3 SGLang timeline capture Zhiyi Hong 2026-08-01 19:38:36 +08:00
  • e1719bd575 [Docs] finalize Phase 2.5 RDMA demand model Zhiyi Hong 2026-08-01 15:29:09 +08:00
  • c5fa700c50 [Feat] add Phase 2.5 RDMA demand modeling Zhiyi Hong 2026-08-01 02:40:06 +08:00