Commit Graph

  • 119701a417 [BugFix] align Phase 3 captures with active decode Zhiyi Hong 2026-08-01 00:50:45 +08:00
  • 405608ad23 [BugFix] make Phase 3 artifact collection fail closed Zhiyi Hong 2026-07-31 18:53:19 +08:00
  • 771b868757 [Docs] link Phase 2 results to server evidence Zhiyi Hong 2026-07-31 18:43:30 +08:00
  • 3b7291e8a0 [Feat] add Phase 3 SGLang timeline profiling Zhiyi Hong 2026-07-31 18:37:12 +08:00
  • 1dc5612e3a [Docs] finalize Phase 2 hardware attribution Zhiyi Hong 2026-07-31 17:38:36 +08:00
  • 5096661ce3 [Docs] gate phase archives on completed results Zhiyi Hong 2026-07-31 17:05:19 +08:00
  • 7be3062c51 [Docs] unify phase experiment archive naming Zhiyi Hong 2026-07-31 17:00:22 +08:00
  • 5f24b7d22f [Docs] record Phase 2 worker staging smoke test Zhiyi Hong 2026-07-31 16:29:03 +08:00
  • e9c5f85500 [Docs] document Phase 2 worker staging fix Zhiyi Hong 2026-07-31 16:20:53 +08:00
  • 39fc2ba565 [BugFix] stage Phase 2 communication tool on workers Zhiyi Hong 2026-07-31 16:19:23 +08:00
  • 66d1db8581 [Docs] link Phase 2 code walkthrough Zhiyi Hong 2026-07-31 16:01:57 +08:00
  • 4892c0b14d [Docs] document final Phase 2 attribution workflow Zhiyi Hong 2026-07-31 15:47:10 +08:00
  • 30664faa41 [Feat] finalize Phase 2 hardware attribution pipeline Zhiyi Hong 2026-07-31 15:14:44 +08:00
  • a583c337ba [Docs] restore Phase 2 metric command guide Zhiyi Hong 2026-07-31 14:22:35 +08:00
  • 72bae06576 [Docs] clarify Phase 2 worker preparation Zhiyi Hong 2026-07-31 14:00:58 +08:00
  • 337195254a [Docs] summarize Phase 2 hardware attribution Zhiyi Hong 2026-07-31 13:50:27 +08:00
  • 3964b3d210 [Docs] add Phase 1 and Phase 2 code walkthroughs Zhiyi Hong 2026-07-31 13:13:57 +08:00
  • ca1f2f6337 [Fix] ignore Phase 2 runtime artifacts Zhiyi Hong 2026-07-31 12:24:00 +08:00
  • daa4221128 [Feat] add DSV4-Pro two-node SGLang hardware attribution Zhiyi Hong 2026-07-31 12:18:34 +08:00
  • 451782457d [Docs] record two-node sysstat monitoring setup Zhiyi Hong 2026-07-31 10:48:48 +08:00
  • ae85217225 [Docs] record DSV4-Pro long-decode results Zhiyi Hong 2026-07-31 00:15:43 +08:00
  • 06b017483c [Feat] add DSV4-Pro long-decode cases Zhiyi Hong 2026-07-30 23:41:08 +08:00
  • 25418ec174 [Docs] record completed DSV4-Pro Phase 1 quick map Zhiyi Hong 2026-07-30 23:10:22 +08:00
  • 75182c6ede [Feat] add Phase 1 sanity gate Zhiyi Hong 2026-07-30 18:59:14 +08:00
  • 0d3dd86519 [BugFix] enforce RDMA transport for DSV4-Pro TP16 quick map Zhiyi Hong 2026-07-30 18:17:33 +08:00
  • 595bdde5d7 [Docs] audit DSV4-Pro TP16 TTFT benchmark semantics Zhiyi Hong 2026-07-30 16:41:11 +08:00
  • d5d96bd7e4 [Feat] add DSV4-Pro two-node SGLang quick map Zhiyi Hong 2026-07-30 14:40:16 +08:00
  • e885fd0dc2 feat(adaptive): support tiered per-ISL TTFT SLO via get_ttft_slo_ms() shishi 2026-07-30 10:47:56 +08:00
  • ad4fd2b878 fix: INVALID_WORKLOAD 只跳过当前shape而不中止所有配置 shishi 2026-07-30 10:05:15 +08:00
  • e53b2c7e4c fix: JSONL parsing in run_batch, add results/ to gitignore Zhiyi Hong 2026-07-29 17:51:47 +08:00
  • c839230c0b feat: add EAGLE speculative decoding experiment (dsv4_pro6000_sglang_tp16_eagle) Zhiyi Hong 2026-07-29 17:51:33 +08:00
  • 4914ff4041 fix(dsv4): disable MTP speculative decoding for fair H20 comparison shishi 2026-07-29 16:50:13 +08:00
  • 455a78161b fix(910c/glm52): 对齐官方A3教程参数(修DP die分配问题,同dsv4 99a22f0) shishi 2026-07-29 15:28:54 +08:00
  • 99a22f05b8 fix(dsv4): align launch params with official A3 tutorial (fixes DP die allocation) shishi 2026-07-29 15:09:27 +08:00
  • c222ed98b2 feat(910c/glm52): 并行配置改为 4 4 / 8 2 / 16 1 对标H20的 2 4 / 4 2 / 8 1 shishi 2026-07-29 13:44:40 +08:00
  • a65849b77d fix(dsv4): use all 16 dies with TP4/DP4 + TP8/DP2 + TP16/DP1 shishi 2026-07-29 13:37:26 +08:00
  • 63ab41b65a docs: consolidate project docs (dedup, relocate, expand 910C client guide) shishi 2026-07-29 11:51:24 +08:00
  • 70c5c57f8f fix(910c/glm52): start_vllm_docker.sh 加per-TP参数覆盖 + 修printf转义破坏JSON shishi 2026-07-29 10:59:53 +08:00
  • 6d3338244b fix: cuda-graph-max-bs=64 (256 OOMs during capture) Zhiyi Hong 2026-07-28 17:31:52 +08:00
  • cd362c5ca0 fix: max_running_requests=256, cuda_graph_max_bs=256 to match max concurrency 128 Zhiyi Hong 2026-07-28 17:20:25 +08:00
  • da1d4d3d78 add tp = 16 deepseek v4 pro sglang bench Zhiyi Hong 2026-07-28 17:06:21 +08:00
  • 0d091769d9 chore: remove unused capacity_validation_20260721 directory shishi 2026-07-28 17:03:23 +08:00
  • cd7b2d4a62 docs(datasets): add README with ShareGPT download instructions shishi 2026-07-28 17:02:19 +08:00
  • d6e00d61dc fix(dsv4): sync glm52 add16 fixes (sglang 0.5.2 client + health timeout) shishi 2026-07-28 16:53:34 +08:00
  • 6c81183fd7 feat(910c/glm52): 并行配置改为只测 TP=8 和 TP=16,并按TP区分服务参数 shishi 2026-07-28 16:51:14 +08:00
  • 98cdb67b66 fix(910c/glm52): 修复sglang0.5.2解析兼容性+TP=16设备挂载+health超时 shishi 2026-07-28 16:43:05 +08:00
  • 4197e2738d feat(dsv4): make DSV4-Flash 910C experiment runnable (verified TP4/DP2) shishi 2026-07-28 16:36:53 +08:00
  • 46e79d63e7 feat(platform): add Ascend 910C NPU platform support shishi 2026-07-27 22:00:05 +08:00
  • d13f61f7b8 Add vLLM TP×DP matrix benchmarking scripts and Docker support qqtang 2026-07-23 02:06:47 +00:00
  • c234468cd6 docs: add new-experiment quick start and forbid raw_requests in results.json yy-fighting 2026-07-22 04:08:51 +00:00
  • 3b0297516e fix: replace 5 buggy parse_results.py (raw_requests bloat) with shared parse_backend.py wrapper yy-fighting 2026-07-22 03:47:22 +00:00
  • 489c5a5e61 fix(lib.sh): avoid cd side-effect in git_commit/git_dirty and safe JSON in write_metadata_json yy-fighting 2026-07-22 03:44:18 +00:00
  • b6dab8e139 build: add pyproject.toml with ruff lint/format config and requirements-dev yy-fighting 2026-07-22 03:40:39 +00:00
  • 3ba4968971 chore(gitignore): ignore envs venvs and skills-lock.json, keep envs docs yy-fighting 2026-07-22 03:38:52 +00:00
  • 6286733f3c chore: gitignore dummy_sharegpt.json dataset sskj-agent 2026-07-22 00:22:06 +08:00
  • aee25d4088 feat(pro6000): add qwen3_235b_pro6000_sglang_tp8 experiment (code + README + report.md) Quantong Qiu 2026-07-21 23:54:19 +08:00
  • 059eb2521c !19 fix(parse_results): 不再把 raw_requests 嵌入 results.json,避免 8.5MB 膨胀 Quantong Qiu 2026-07-21 10:07:16 +00:00
  • 31a21631b8 fix(parse_results): 不再把 raw_requests 嵌入 results.json,避免 8.5MB 膨胀 yy-fighting 2026-07-21 10:04:26 +00:00
  • 3bd8f60474 !18 chore(repo): 瘦身规范 - 忽略编译产物/JIT缓存/raw_outputs,取消跟踪618个垃圾文件 Quantong Qiu 2026-07-21 09:49:31 +00:00
  • 080e095ece chore(repo): 瘦身规范 - 忽略编译产物/JIT缓存/raw_outputs,取消跟踪618个垃圾文件 yy-fighting 2026-07-21 09:44:41 +00:00
  • 7231536db2 !17 [Feat] use nightly SGLang FlashInfer MoE Quantong Qiu 2026-07-21 09:09:44 +00:00
  • 8c737b840b [Feat] use nightly SGLang FlashInfer MoE Quantong Qiu 2026-07-21 15:22:47 +08:00
  • d0319504e4 !16 feat(p800): add Qwen3-235B-A22B SGLang TP=8 benchmark experiment Quantong Qiu 2026-07-21 06:58:53 +00:00
  • 6087798262 !15 [Fix] skip longer contexts after C1 OOM Quantong Qiu 2026-07-21 06:58:33 +00:00
  • 3f28edd1a6 feat(p800): add Qwen3-235B-A22B SGLang TP=8 benchmark experiment yy-fighting 2026-07-21 06:34:57 +00:00
  • 83f5ea2197 [Fix] skip longer contexts after C1 OOM Quantong Qiu 2026-07-21 14:10:17 +08:00
  • 8e21d8a34f !14 [Feat] add 6000D DSV4 tiny adaptive benchmarks Quantong Qiu 2026-07-21 05:30:58 +00:00
  • f263695e1b Merge branch 'main' of gitee.com:yy-fighting/sskj into auto/main/11665927/9c484989-1 Quantong Qiu 2026-07-21 05:30:44 +00:00
  • 5e864d2393 !13 [Feat] apply validated deployment capacity caps Quantong Qiu 2026-07-21 05:26:41 +00:00
  • 7d79a1a7ec !12 6000D Results Quantong Qiu 2026-07-21 05:25:41 +00:00
  • 680f2c31d6 !11 feat(p800): enable operator-level timing in profiling experiment Quantong Qiu 2026-07-21 05:24:54 +00:00
  • 0c2256ca1c [Feat] add 6000D DSV4 tiny adaptive benchmarks Quantong Qiu 2026-07-21 11:54:04 +08:00
  • db16057962 [Feat] apply validated deployment capacity caps Quantong Qiu 2026-07-21 11:50:23 +08:00
  • db25bc7e5f 6000D Results Quantong Qiu 2026-07-21 11:11:16 +08:00
  • ac5a01dbd0 fix(sglang): cap SM120 CUDA graph prefill Quantong Qiu 2026-07-20 16:05:33 +08:00
  • 41a3ff0a6b feat(p800): enable operator-level timing in profiling experiment yy 2026-07-21 02:02:50 +00:00
  • bf03201168 feat(p800): auto-load timing module in every TP worker process yy 2026-07-21 02:02:06 +00:00
  • f81c647d16 feat(p800): add manual timing module to bypass broken PyTorch Profiler yy 2026-07-21 01:59:38 +00:00
  • 039a2fa7b9 !10 update Quantong Qiu 2026-07-20 15:23:45 +00:00
  • 2f15d09877 update iiGray 2026-07-20 15:21:35 +00:00
  • eaf005845a !9 update VLLM + GLM + H20 Quantong Qiu 2026-07-20 15:18:16 +00:00
  • 3f50608aa7 update VLLM + GLM + H20 iiGray 2026-07-20 15:13:09 +00:00
  • 866af70760 add VLLM + H20 + GLM iiGray 2026-07-20 14:16:44 +00:00
  • a8ed016f1e !6 fix(h20): capture decode graphs through batch 128 Quantong Qiu 2026-07-20 09:59:34 +00:00
  • 105267b1b8 fix(h20): capture decode graphs through batch 128 Quantong Qiu 2026-07-20 17:56:58 +08:00
  • ca965b6d26 !5 refactor(paths): derive repository files from root Quantong Qiu 2026-07-20 09:38:55 +00:00
  • 359df6a7b6 refactor(paths): derive repository files from root Quantong Qiu 2026-07-20 17:31:16 +08:00
  • 580b6c1c56 !2 feat(p800): add sglang profiling experiment with PROFILE_REPORT Quantong Qiu 2026-07-20 08:11:59 +00:00
  • f0c93a7d72 !3 feat: add P800 vs H20 diagnosis plan for performance gap analysis Quantong Qiu 2026-07-20 08:11:18 +00:00
  • ac821cf0f5 !4 fix(sglang): cap SM120 CUDA graph prefill Quantong Qiu 2026-07-20 08:09:16 +00:00
  • 8562c26524 fix(sglang): cap SM120 CUDA graph prefill Quantong Qiu 2026-07-20 16:05:33 +08:00
  • 0afd552ad3 feat: add P800 vs H20 diagnosis plan for performance gap analysis yy-fighting 2026-07-20 05:23:49 +00:00
  • a7e2037471 feat(p800): add sglang profiling experiment with PROFILE_REPORT yy-fighting 2026-07-20 04:27:07 +00:00
  • fad6336d2c !1 docs: test protected-branch PR automation Quantong Qiu 2026-07-20 04:28:59 +00:00
  • 19f3894de6 feat(p800): add sglang profiling experiment with PROFILE_REPORT yy-fighting 2026-07-20 04:27:07 +00:00
  • fa39145e11 docs: test protected-branch PR automation Quantong Qiu 2026-07-20 12:27:04 +08:00
  • 5ffbd0a21a fix(bench): use default limits and back off after OOM Quantong Qiu 2026-07-20 11:37:18 +08:00
  • ceb1170969 Add adaptive results and configuration files for sglang experiment SSKJ Dev 2026-07-19 04:03:47 +00:00
  • 958c5778f0 Add new kernels and autotune configurations for DeepseekV4 model Quantong Qiu 2026-07-18 10:31:24 +08:00
  • 8beed2411f update 8k and 32k Quantong Qiu 2026-07-17 16:47:25 +08:00