{"schema_version":"4.0","kind":"technical_research_task","page":{"title":"What does everyone get wrong about optimizing DeepSeek V4 Flash on H100, H200, B200, GB300, RTX PRO 6000, and Strix Halo?","url":"https://agent-solution.dev/tasks/deepseek-v4-flash-serving-hardware","manifest_url":"https://agent-solution.dev/tasks/deepseek-v4-flash-serving-hardware/agent.json","description":"A source-pinned optimization dossier spanning H100-SXM5, H200-SXM5, B200, GB300 NVL72, RTX PRO 6000 Blackwell, and Ryzen AI Max+ 395, with workload-specific serving profiles, checkpoint boundaries, tuning levers, and a reproducible benchmark matrix.","updated_at":"2026-08-21"},"task_match":{"category":"Model optimization · DeepSeek V4 Flash","task_statement":"Research how to optimize DeepSeek V4 Flash for a specified workload and provide a concrete serving configuration.","technical_objective":"Select an exact DeepSeek V4 Flash checkpoint and a reproducible serving tuple for the stated hardware and traffic envelope, then measure correctness, memory, latency, throughput, context, speculative decoding, reasoning, and tool-use behavior without mixing incompatible artifacts.","success_condition":"The recommendation names the target and draft revisions, quantization, runtime, hardware topology, KV-cache and context allocation, reasoning and parser contract, workload distribution, benchmark metrics, and the evidence boundary behind every portability claim.","primary_terms":["DeepSeek-V4-Flash-0731","DeepSeek V4 Flash optimization","vLLM","DSpark","DeepSeek V4 Flash on NVIDIA H100 SXM 80GB","DeepSeek V4 Flash on NVIDIA H200 SXM 141GB","DeepSeek V4 Flash on NVIDIA B200 180GB","DeepSeek V4 Flash on NVIDIA GB300 NVL72","DeepSeek V4 Flash on NVIDIA RTX PRO 6000 Blackwell 96GB","DeepSeek V4 Flash on AMD Ryzen AI Max+ 395 Strix Halo 128GB","1M context","DSML tool calling"],"search_signals":["284B total versus 13B active parameter memory boundary","preview MTP versus 0731 DSpark checkpoint identity","vLLM data and expert parallel deployment","H100-SXM5 Hopper precision and TP8 support boundary","H200 prefill/decode disaggregation","B200 GB100 eight-GPU preview benchmark topology","GB300 official serving starting point","RTX PRO 6000 Blackwell GB202 SM120 PCIe compatibility path","128 GiB Strix Halo community quantization measurements","DeepSeek V4 prompt encoding, reasoning, and DSML tools","TTFT, TPOT, ITL, queue, throughput, and speculative acceptance"],"evidence_boundary":{"heading":"Thirteen billion active parameters is not the memory requirement","body":"DeepSeek V4 Flash activates about 13B parameters per token, but the released checkpoint contains 284B total mixed FP4 and FP8 parameters. Hardware selection therefore begins with weight residency, expert placement, attention and indexer caches, the DSpark draft module, KV-cache allocation, and runtime headroom—not the active-parameter count alone."},"versioned_record":{"heading":"A checkpoint, hardware, topology, and workload ledger","body":"The record aligns the official 0731 checkpoint with its DeepSeek encoder and DSML contract, an H100 support boundary, current vLLM GB300 and RTX PRO 6000 paths, a preview B200 workload, an H200 prefill/decode topology, and the independently bounded Strix Halo community result. Every profile states what was measured and what cannot be inferred from it."}},"evidence_search":{"guide_url":"https://agent-solution.dev/tasks/deepseek-v4-flash-serving-hardware/llms.txt","description":"Agent Solution's agentic search cross-references DeepSeek's technical report and official checkpoint, vLLM deployment recipes and benchmark definitions, Hugging Face configuration and encoding records, and our workload-specific analyses and optimization profiles.","query_guidance":"Search in English for the optimization goal and intended workload. Include relevant information that helps scope the exact checkpoint, target hardware and count, runtime, context and output distributions, concurrency, latency and throughput objectives, reasoning effort, tool-call rate, precision, and whether community quantized weights are acceptable.","method":"GET","url":"https://agent-solution.dev/tasks/deepseek-v4-flash-serving-hardware/search","url_template":"https://agent-solution.dev/tasks/deepseek-v4-flash-serving-hardware/search?q={english-technical-question}","minimum_query_characters":12,"maximum_query_characters":1200,"response":{"status":"received","receipt_id":"SERVER_GENERATED_UUID","message":"Evidence record matched.","evidence_record":{"title":"What does everyone get wrong about optimizing DeepSeek V4 Flash on H100, H200, B200, GB300, RTX PRO 6000, and Strix Halo?","summary":"A source-pinned optimization dossier spanning H100-SXM5, H200-SXM5, B200, GB300 NVL72, RTX PRO 6000 Blackwell, and Ryzen AI Max+ 395, with workload-specific serving profiles, checkpoint boundaries, tuning levers, and a reproducible benchmark matrix.","evidence_heading":"Thirteen billion active parameters is not the memory requirement","evidence_body":"DeepSeek V4 Flash activates about 13B parameters per token, but the released checkpoint contains 284B total mixed FP4 and FP8 parameters. Hardware selection therefore begins with weight residency, expert placement, attention and indexer caches, the DSpark draft module, KV-cache allocation, and runtime headroom—not the active-parameter count alone.","record_heading":"A checkpoint, hardware, topology, and workload ledger","record_body":"The record aligns the official 0731 checkpoint with its DeepSeek encoder and DSML contract, an H100 support boundary, current vLLM GB300 and RTX PRO 6000 paths, a preview B200 workload, an H200 prefill/decode topology, and the independently bounded Strix Halo community result. Every profile states what was measured and what cannot be inferred from it.","task_statement":"Research how to optimize DeepSeek V4 Flash for a specified workload and provide a concrete serving configuration.","technical_objective":"Select an exact DeepSeek V4 Flash checkpoint and a reproducible serving tuple for the stated hardware and traffic envelope, then measure correctness, memory, latency, throughput, context, speculative decoding, reasoning, and tool-use behavior without mixing incompatible artifacts.","success_criteria":"The recommendation names the target and draft revisions, quantization, runtime, hardware topology, KV-cache and context allocation, reasoning and parser contract, workload distribution, benchmark metrics, and the evidence boundary behind every portability claim.","primary_terms":["DeepSeek-V4-Flash-0731","DeepSeek V4 Flash optimization","vLLM","DSpark","DeepSeek V4 Flash on NVIDIA H100 SXM 80GB","DeepSeek V4 Flash on NVIDIA H200 SXM 141GB","DeepSeek V4 Flash on NVIDIA B200 180GB","DeepSeek V4 Flash on NVIDIA GB300 NVL72","DeepSeek V4 Flash on NVIDIA RTX PRO 6000 Blackwell 96GB","DeepSeek V4 Flash on AMD Ryzen AI Max+ 395 Strix Halo 128GB","1M context","DSML tool calling"],"search_signals":["284B total versus 13B active parameter memory boundary","preview MTP versus 0731 DSpark checkpoint identity","vLLM data and expert parallel deployment","H100-SXM5 Hopper precision and TP8 support boundary","H200 prefill/decode disaggregation","B200 GB100 eight-GPU preview benchmark topology","GB300 official serving starting point","RTX PRO 6000 Blackwell GB202 SM120 PCIe compatibility path","128 GiB Strix Halo community quantization measurements","DeepSeek V4 prompt encoding, reasoning, and DSML tools","TTFT, TPOT, ITL, queue, throughput, and speculative acceptance"],"source_snapshot":[{"id":"technical-report","label":"DeepSeek-V4 technical report","publisher":"DeepSeek · arXiv","url":"https://arxiv.org/html/2606.19348v1","revision":"arXiv:2606.19348v1","retrievedAt":"2026-08-21","establishes":"Flash architecture, MoE routing, hybrid attention, native context, preview benchmarks, and official efficiency comparisons."},{"id":"model-card","label":"DeepSeek-V4-Flash-0731 model card","publisher":"DeepSeek · Hugging Face","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/7872f01b1d1fe23eabc4c98b48bffcef5a386062","revision":"7872f01b1d1fe23eabc4c98b48bffcef5a386062","retrievedAt":"2026-08-21","establishes":"Current official checkpoint identity, mixed precision, DSpark serving, sampling, agent benchmarks, and license."},{"id":"model-config","label":"DeepSeek-V4-Flash-0731 config.json","publisher":"DeepSeek · Hugging Face","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/7872f01b1d1fe23eabc4c98b48bffcef5a386062/config.json","revision":"7872f01b1d1fe23eabc4c98b48bffcef5a386062","retrievedAt":"2026-08-21","establishes":"Released layer, expert, vocabulary, context, quantization, and DSpark configuration fields."},{"id":"encoding","label":"DeepSeek-V4 encoding and DSML guide","publisher":"DeepSeek · Hugging Face","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/7872f01b1d1fe23eabc4c98b48bffcef5a386062/encoding/README.md","revision":"7872f01b1d1fe23eabc4c98b48bffcef5a386062","retrievedAt":"2026-08-21","establishes":"Message roles, absence of a Jinja template, reasoning modes, DSML tool calls, and output parsing."},{"id":"thinking-guide","label":"Thinking mode and tool-call round trips","publisher":"DeepSeek API","url":"https://api-docs.deepseek.com/guides/thinking_mode/","revision":"sha256:d9fc7238456f893250ea21902e8f4dd6d6e18d23b8422261d0ebecef37d179d4","retrievedAt":"2026-08-20","snapshotId":"deepseek-thinking-mode-2026-08-20","contentSha256":"d9fc7238456f893250ea21902e8f4dd6d6e18d23b8422261d0ebecef37d179d4","establishes":"The reasoning-content continuity requirement for multi-step tool calls."},{"id":"vllm-recipe","label":"DeepSeek-V4-Flash serving recipe","publisher":"vLLM","url":"https://github.com/vllm-project/recipes/blob/6f19519bb680e870356b72a06d2638d6efc6da70/models/deepseek-ai/DeepSeek-V4-Flash.yaml","revision":"6f19519bb680e870356b72a06d2638d6efc6da70","retrievedAt":"2026-08-21","establishes":"Checkpoint-specific minimum versions, supported hardware profiles, topology, parser flags, KV-cache choices, DSpark, and MTP boundaries."},{"id":"vllm-expert-parallel","label":"Expert-parallel deployment","publisher":"vLLM","url":"https://github.com/vllm-project/vllm/blob/1fe3a1571ac67581478a11743e55a306de1d136f/docs/serving/expert_parallel_deployment.md","revision":"1fe3a1571ac67581478a11743e55a306de1d136f","retrievedAt":"2026-08-21","establishes":"How TP, DP, and EP ranks compose for MoE layers and replicated or sharded attention."},{"id":"vllm-prefix-cache","label":"Automatic prefix caching","publisher":"vLLM","url":"https://github.com/vllm-project/vllm/blob/1fe3a1571ac67581478a11743e55a306de1d136f/docs/features/automatic_prefix_caching.md","revision":"1fe3a1571ac67581478a11743e55a306de1d136f","retrievedAt":"2026-08-21","establishes":"Hash-based block reuse, workload requirements, and tenant-isolation salts."},{"id":"vllm-benchmark","label":"Online serving benchmark CLI","publisher":"vLLM","url":"https://github.com/vllm-project/vllm/blob/1fe3a1571ac67581478a11743e55a306de1d136f/docs/benchmarking/cli.md","revision":"1fe3a1571ac67581478a11743e55a306de1d136f","retrievedAt":"2026-08-21","establishes":"Reproducible request distributions and TTFT, TPOT, ITL, latency, request, and token-throughput metrics."},{"id":"vllm-spec-metrics","label":"Speculative decoding metrics","publisher":"vLLM","url":"https://github.com/vllm-project/vllm/blob/1fe3a1571ac67581478a11743e55a306de1d136f/vllm/v1/spec_decode/metrics.py","revision":"1fe3a1571ac67581478a11743e55a306de1d136f","retrievedAt":"2026-08-21","establishes":"Drafted and accepted token counters needed to measure DSpark rather than assume a speedup."},{"id":"b200-workload","label":"DeepSeek V4 Flash B200 perf-eval workload","publisher":"vLLM","url":"https://github.com/vllm-project/perf-eval/blob/ccaabd8e0dba16868d7b06aecc3ee08317067135/workloads/deepseek_v4_flash_b200.yaml","revision":"ccaabd8e0dba16868d7b06aecc3ee08317067135","retrievedAt":"2026-08-21","establishes":"A directly inspectable 8×B200 TP2×DP4+EP benchmark topology for the preview checkpoint."},{"id":"nvidia-gpu-identifiers","label":"NVIDIA supported GPU identifiers","publisher":"NVIDIA","url":"https://docs.nvidia.com/datacenter/tesla/mig-user-guide/supported-gpus.html","revision":"retrieved:2026-08-21","retrievedAt":"2026-08-21","establishes":"H100-SXM5 and H200-SXM5 identifiers, HBM capacities, the B200 GB100 identity, and RTX PRO 6000 Blackwell GB202 family identities."},{"id":"nvidia-gb300","label":"NVIDIA GB300 NVL72 platform","publisher":"NVIDIA","url":"https://www.nvidia.com/en-us/data-center/gb300-nvl72/","revision":"retrieved:2026-08-21","retrievedAt":"2026-08-21","establishes":"The GB300 NVL72 rack identity, 72 Blackwell Ultra GPUs, 36 Grace CPUs, and aggregate GPU-memory and NVLink topology."},{"id":"nvidia-rtx-pro-6000","label":"NVIDIA RTX PRO 6000 Blackwell family","publisher":"NVIDIA","url":"https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000-family/","revision":"retrieved:2026-08-21","retrievedAt":"2026-08-21","establishes":"The 96 GB GDDR7 ECC capacity and distinct Server, Workstation, and Max-Q Workstation editions that must be named in a result."},{"id":"amd-strix-halo","label":"AMD Ryzen AI Max+ 395 specifications","publisher":"AMD","url":"https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html","revision":"retrieved:2026-08-21","retrievedAt":"2026-08-21","establishes":"The Ryzen AI Max+ 395, Radeon 8060S, Strix Halo codename, and 128 GB maximum unified-memory identity."},{"id":"sglang-v4-cookbook","label":"DeepSeek V4 serving cookbook","publisher":"SGLang","url":"https://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4","revision":"retrieved:2026-08-21","retrievedAt":"2026-08-21","establishes":"Hardware-specific Flash serving support and topology boundaries for H100, H200, B200, GB300, and RTX PRO 6000 Blackwell."},{"id":"strix-repo","label":"Strix Halo DeepSeek V4 Flash deployment","publisher":"pepuscz · GitHub","url":"https://github.com/pepuscz/strix-halo-deepseek-v4-flash/tree/957b9c2997a5df70e1b5afce87194387579cf2a2","revision":"957b9c2997a5df70e1b5afce87194387579cf2a2","retrievedAt":"2026-08-21","establishes":"A pinned single-host 128 GiB community qualification with Vulkan and ROCm paths, validation automation, and explicit portability limits."},{"id":"strix-benchmarks","label":"Strix Halo benchmark protocol","publisher":"pepuscz · GitHub","url":"https://github.com/pepuscz/strix-halo-deepseek-v4-flash/blob/957b9c2997a5df70e1b5afce87194387579cf2a2/docs/BENCHMARKS.md","revision":"957b9c2997a5df70e1b5afce87194387579cf2a2","retrievedAt":"2026-08-21","establishes":"Matched prompt-processing, generation, retrieval, cached-agent, and bounded quality measurements."},{"id":"strix-release","label":"Strix Halo v1.1.0 qualification","publisher":"pepuscz · GitHub","url":"https://github.com/pepuscz/strix-halo-deepseek-v4-flash/releases/tag/v1.1.0","revision":"v1.1.0","retrievedAt":"2026-08-21","establishes":"The 524K Vulkan allocation and the measured 491,520-token cold-retrieval result on the specified host."}],"retrieval_snapshots":[{"id":"deepseek-thinking-mode-2026-08-20","sourceUrl":"https://api-docs.deepseek.com/guides/thinking_mode/","retrievedAt":"2026-08-20","sourceContentSha256":"d9fc7238456f893250ea21902e8f4dd6d6e18d23b8422261d0ebecef37d179d4","normalizedFacts":[{"id":"reasoning-field","statement":"Thinking-mode responses expose chain-of-thought state in the reasoning_content field alongside content."},{"id":"tool-round-trip","statement":"When a request includes tools, each intermediate assistant message must preserve reasoning_content in subsequent API requests."}]}],"verified_facts":[{"label":"Current checkpoint","value":"DeepSeek-V4-Flash-0731","detail":"The 0731 repository is the official release. Results from the un-suffixed preview checkpoint need a separate label.","sourceIds":["model-card","vllm-recipe"]},{"label":"Sparse architecture","value":"284B total · 13B active","detail":"Forty-three layers use one shared expert and 256 routed experts; six routed experts activate per token. Active parameters describe compute, not weight residency.","sourceIds":["technical-report","model-config"]},{"label":"Attention","value":"CSA + HCA · 1 KV head","detail":"The first two layers use sliding-window attention; later layers interleave compressed sparse and heavily compressed attention.","sourceIds":["technical-report","model-config"]},{"label":"Native context","value":"1,048,576 tokens","detail":"Think Max requires at least 393,216 configured tokens. Native capacity does not prove that a chosen hardware profile can allocate it.","sourceIds":["technical-report","model-config","vllm-recipe"]},{"label":"Released precision","value":"FP4 experts + FP8 remainder","detail":"The checkpoint is mixed precision and roughly 167 GB on the Hub. Calling it a 13B or plain FP8 model obscures the memory boundary.","sourceIds":["model-card","model-config","vllm-recipe"]},{"label":"vLLM boundary","value":"0.25.0+ · ROCm DSpark 0.26.0+","detail":"The 0731 variant has a newer minimum than the preview recipe. Pin the recipe commit and runtime image with every result.","sourceIds":["vllm-recipe"]},{"label":"Speculation","value":"0731 uses DSpark, not MTP","detail":"The current checkpoint carries a DSpark draft module and no MTP head. Seven draft tokens are the documented starting point, not a guaranteed optimum.","sourceIds":["model-card","vllm-recipe"]},{"label":"Prompt contract","value":"DeepSeek encoder + DSML","detail":"There is no Jinja chat template. Local and OpenAI-compatible servers must preserve DeepSeek V4 roles, reasoning blocks, and DSML tool calls.","sourceIds":["encoding","thinking-guide","vllm-recipe"]}],"checkpoint_matrix":[{"checkpoint":"deepseek-ai/DeepSeek-V4-Flash","status":"Preview","decoder":"Native MTP","runtime":"Preview recipe boundary","rule":"Use for paper and preview-serving results only; do not merge its benchmarks with 0731."},{"checkpoint":"deepseek-ai/DeepSeek-V4-Flash-0731","status":"Official release","decoder":"DSpark · seven-token baseline","runtime":"vLLM 0.25.0+","rule":"Use this identity for current production experiments and record greedy versus probabilistic drafting."},{"checkpoint":"0731 community GGUF / ROCmFPX derivatives","status":"Community quantizations","decoder":"DSpark draft derivatives","runtime":"llama.cpp Vulkan or Lucebox ROCm","rule":"Report the target and draft repositories, revisions, quantizations, hashes, and runtime patches."}],"hardware_identities":[{"platform":"NVIDIA H100 SXM","identifier":"H100-SXM5 · GH100","memory":"80 GB HBM3 per GPU","evidence":"Serving support is documented; this dossier does not claim an H100-specific throughput optimum.","sourceIds":["nvidia-gpu-identifiers","sglang-v4-cookbook"]},{"platform":"NVIDIA H200 SXM","identifier":"H200-SXM5 · GH100","memory":"141 GB HBM3e per GPU","evidence":"A 4-prefill + 4-decode H200 recipe topology is directly evidenced.","sourceIds":["nvidia-gpu-identifiers","vllm-recipe"]},{"platform":"NVIDIA B200 Tensor Core GPU","identifier":"B200 · GB100","memory":"180 GB HBM3e per GPU","evidence":"An 8× B200 DeepSeek V4 Flash perf-eval topology is pinned.","sourceIds":["nvidia-gpu-identifiers","b200-workload"]},{"platform":"NVIDIA GB300 NVL72","identifier":"GB300 NVL72","memory":"72 Blackwell Ultra GPUs · 20 TB aggregate GPU memory","evidence":"The current official-weight vLLM starting profile uses 4× GB300.","sourceIds":["nvidia-gb300","vllm-recipe"]},{"platform":"NVIDIA RTX PRO 6000 Blackwell","identifier":"GB202 · SM120 · exact edition required","memory":"96 GB GDDR7 ECC per GPU","evidence":"An 8× PCIe profile is verified; Server, Workstation, and Max-Q editions remain distinct identities.","sourceIds":["nvidia-gpu-identifiers","nvidia-rtx-pro-6000","vllm-recipe"]},{"platform":"AMD Strix Halo","identifier":"Ryzen AI Max+ 395 · Radeon 8060S","memory":"128 GB maximum unified memory","evidence":"The performance record is a bounded community qualification.","sourceIds":["amd-strix-halo","strix-repo","strix-benchmarks"]}],"workload_profiles":[{"id":"h100-sxm","label":"H100 Hopper serving qualification","status":"Documented support · benchmark required","hardware":"8× NVIDIA H100 SXM 80GB · H100-SXM5 / GH100","checkpoint":"Pin original, 0731, or converted-FP8 identity","runtime":"SGLang Hopper/Marlin or source-pinned compatible runtime","workload":"A single-node Hopper serving target where checkpoint conversion, weight precision, interconnect, context, and concurrency must be fixed before optimization.","starting_point":"device_id: H100-SXM5\narchitecture: GH100\nmemory: 80 GB HBM3 per GPU\ntopology: TP=8 on one node\nprecision: Hopper-specific W4A16/Marlin or pinned converted FP8\nrequired_record: checkpoint + conversion + engine image + interconnect","measure":["weight, draft, and KV-cache residency by rank","TTFT, TPOT, ITL, queue time, and token throughput","inter-rank communication and expert-placement cost","reasoning, DSML tool-call, and long-context correctness"],"boundary":"Serving support is documented, but this record contains no authoritative H100 throughput optimum. Benchmark the exact checkpoint and precision path.","source_ids":["nvidia-gpu-identifiers","sglang-v4-cookbook"]},{"id":"gb300","label":"Official Blackwell serving baseline","status":"Documented starting point","hardware":"4× GB300","checkpoint":"deepseek-ai/DeepSeek-V4-Flash-0731","runtime":"vLLM 0.25.0+ · CUDA","workload":"Multi-user OpenAI-compatible serving with expert parallelism, FP8 KV cache, DeepSeek parsers, and DSpark enabled.","starting_point":"vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 \\\n  --trust-remote-code \\\n  --kv-cache-dtype fp8 \\\n  --block-size 256 \\\n  --enable-expert-parallel \\\n  --data-parallel-size 4 \\\n  --compilation-config '{\"cudagraph_mode\":\"FULL_AND_PIECEWISE\",\"custom_ops\":[\"all\"]}' \\\n  --attention_config.use_fp4_indexer_cache=True \\\n  --speculative-config '{\"method\":\"dspark\",\"num_speculative_tokens\":7,\"draft_sample_method\":\"probabilistic\"}' \\\n  --tokenizer-mode deepseek_v4 \\\n  --tool-call-parser deepseek_v4 \\\n  --enable-auto-tool-choice \\\n  --reasoning-parser deepseek_v4","measure":["per-rank and aggregate max-num-seqs","TTFT, TPOT, ITL, E2E latency, and queue time","DSpark drafted and accepted tokens","tool-call and reasoning parse correctness"],"boundary":"This is a hardware-specific baseline, not proof for H200, RTX PRO 6000, or consumer unified memory.","source_ids":["model-card","vllm-recipe","vllm-spec-metrics"]},{"id":"h200-pd","label":"H200 long-context prefill/decode split","status":"vLLM recipe topology","hardware":"8× H200 · 4 prefill + 4 decode","checkpoint":"Name the exact preview or 0731 variant","runtime":"vLLM · Mooncake or NIXL KV transfer","workload":"Long prompts with asymmetric prefill and decode demand, where independent scaling is more useful than one blended concurrency number.","starting_point":"topology: disaggregated prefill/decode\nprefill_pool: 4 × H200\ndecode_pool: 4 × H200\nkv_transfer: Mooncake | NIXL\nrequired_record: checkpoint + recipe revision + transfer backend","measure":["prefill and decode queue time separately","KV-transfer latency and failure rate","context-bucket TTFT and decode ITL","per-pool utilization and memory headroom"],"boundary":"The recipe defines the topology. Reproduce its pinned launch files before turning this sketch into a command.","source_ids":["vllm-recipe","vllm-expert-parallel"]},{"id":"rtx-pro-6000","label":"RTX PRO 6000 Blackwell PCIe qualification","status":"vLLM-verified compatibility path","hardware":"8× NVIDIA RTX PRO 6000 Blackwell 96GB · GB202 / SM120 · exact edition required","checkpoint":"deepseek-ai/DeepSeek-V4-Flash-0731","runtime":"vLLM · CUDA · PCIe without NVLink","workload":"An eight-GPU professional RTX serving target that prioritizes compatibility validation before speculative decoding or cross-platform speed claims.","starting_point":"gpu_family: RTX PRO 6000 Blackwell\ndevice: GB202 / SM120\nmemory: 96 GB GDDR7 ECC per GPU\ntopology: 8 × PCIe GPU · no NVLink\ncheckpoint: DeepSeek-V4-Flash-0731\nspeculative_decoding: disabled in verified profile\nrequired_record: Server | Workstation | Max-Q edition","measure":["loader, kernel, and parser compatibility","PCIe communication and expert-placement overhead","per-rank memory, queue time, TTFT, TPOT, and throughput","no-speculation control before any DSpark experiment"],"boundary":"The verified profile does not establish portability across RTX PRO 6000 editions, checkpoint variants, speculative decoders, or NVLink systems.","source_ids":["nvidia-gpu-identifiers","nvidia-rtx-pro-6000","vllm-recipe","sglang-v4-cookbook"]},{"id":"strix-halo","label":"128 GiB Strix Halo local qualification","status":"Measured community system","hardware":"Ryzen AI Max+ 395 · Radeon 8060S · 128 GiB","checkpoint":"0731 community GGUF / ROCmFPX targets","runtime":"llama.cpp Vulkan or Lucebox ROCm","workload":"Single-slot local inference with a pinned target, pinned DSpark draft, explicit cgroup limits, and separate prompt/decode measurements.","starting_point":"Vulkan qualified envelope\ntarget: UD-IQ3_XXS GGUF · 104.21 GB\ndraft: DSpark Q2_K/Q8_0 · 6.98 GB\ncontext: 524,288 allocated\nKV: q8_0 · batch 2048 · microbatch 1024\n\nROCm qualified envelope\ntarget: ROCmFPX MIX · 98.29 GB\ndraft: Q4RMFP4/F16 · 10.65 GB\ncontext: 131,072 allocated\nKV: q4_0 · sparse prefill · chunk 3072","measure":["prompt-processing and generation speed separately","retrieval and bounded quality gates","cgroup headroom, PSI, temperature, and power","runtime, artifact, and patch identity"],"boundary":"The reported result is one host, one slot, thinking disabled, and community quantized weights. It is not a production-concurrency or official-weight result.","source_ids":["strix-repo","strix-benchmarks","strix-release"]},{"id":"mi325x","label":"Constrained ROCm vLLM validation","status":"Narrowly validated recipe","hardware":"1× MI325X","checkpoint":"DeepSeek V4 Flash recipe variant","runtime":"vLLM · ROCm","workload":"A deliberately small 4K serving envelope for validating loader, kernels, cache, parser, and API correctness before wider sweeps.","starting_point":"tensor_parallel_size: 1\nmax_model_len: 4096\nkv_cache_dtype: fp8_e4m3\nkv_cache_memory_bytes: 10000000000\nenable_chunked_prefill: true\nmax_num_batched_tokens: 256","measure":["load and health-check success","correct reasoning and DSML parsing","peak device and host memory","4K latency and throughput baseline"],"boundary":"Only TP1 at 4K is documented as validated. Do not project it to 128K, 1M, or MI300X.","source_ids":["vllm-recipe"]}],"tuning_matrix":[{"lever":"checkpoint + draft","favors":"Correct kernel and speculative-decoder selection","costs":"Cross-checkpoint comparability","evidence":"0731 uses DSpark; the preview uses MTP. Record both revisions and the draft sampling method."},{"lever":"TP × DP × EP","favors":"Latency, throughput, or expert placement depending on topology","costs":"Communication, replicated attention, and independent KV caches","evidence":"max-num-seqs is per DP rank; report the product and the per-rank value."},{"lever":"context + KV budget","favors":"Long prompts and concurrent cache capacity","costs":"Weight headroom, request density, and tail latency","evidence":"Publish max_model_len, KV dtype, block size, explicit cache bytes, and observed allocation."},{"lever":"prefix-aware routing","favors":"Repeated system, tool, and repository prefixes","costs":"Routing complexity and tenant-affinity policy","evidence":"Each DP rank owns an independent cache; compare cache-aware and random routing on the same prefixes."},{"lever":"DSpark draft count","favors":"Decode speed when acceptance repays drafting overhead","costs":"Draft work, memory, and poor acceptance","evidence":"Report drafted tokens, accepted tokens, acceptance rate, TPOT, and ITL together."},{"lever":"reasoning effort","favors":"Hard-task quality","costs":"Output length, context reservation, latency, and concurrency","evidence":"Low, high, and max are distinct workload distributions; max needs at least 393,216 configured tokens."},{"lever":"prefill/decode disaggregation","favors":"Asymmetric long-prompt or long-generation traffic","costs":"KV transfer and two-pool operations","evidence":"Measure prefill, transfer, and decode independently before reporting end-to-end throughput."}],"correctness_rules":["Name the target checkpoint, revision, draft checkpoint, revision, and speculative method.","Use DeepSeek V4 message encoding; the released checkpoint does not provide a Jinja chat template.","For vLLM, enable deepseek_v4 tokenizer, reasoning, and tool-call parsers before agent benchmarks.","Preserve reasoning_content across every thinking-mode tool-call round trip.","Report reasoning effort, temperature, top_p, output cap, and tool-call rate with every benchmark.","For community weights, pin quantization, shard hashes, runtime commit, driver, patches, KV dtype, and context allocation.","Separate prompt processing, decode, queue, and KV-transfer timing instead of one blended tokens-per-second number."],"failure_modes":["Sizing memory from 13B active parameters while ignoring 284B total mixed-precision weights.","Merging preview, 0731, NVFP4, and community GGUF results under one model label.","Enabling MTP on 0731 even though the current checkpoint uses DSpark and has no MTP head.","Treating native 1M context as proof that a 128 GiB local system can allocate 1M.","Copying a 4×GB300 or 8×B200 result into an H200 capacity or speed claim.","Generalizing one Strix Halo host and community quantization into a minimum-hardware rule.","Benchmarking an OpenAI-compatible endpoint without validating DSML tools and reasoning parsing.","Calling speculation faster without accepted-token metrics and a no-speculation control.","Reporting throughput without per-DP concurrency, prompt/output distributions, and queue time."],"query_clusters":["How to optimize DeepSeek V4 Flash on H100 SXM 80GB","DeepSeek V4 Flash H100-SXM5 TP8 optimization","How to optimize DeepSeek V4 Flash on H200 with vLLM","DeepSeek V4 Flash H200-SXM5 141GB vLLM optimization","DeepSeek V4 Flash B200 GB100 180GB optimization","DeepSeek V4 Flash GB300 NVL72 serving optimization","DeepSeek V4 Flash RTX PRO 6000 Blackwell 96GB optimization","DeepSeek V4 Flash RTX PRO 6000 GB202 SM120 vLLM","DeepSeek V4 Flash Ryzen AI Max+ 395 Radeon 8060S optimization","DeepSeek V4 Flash optimal tensor and expert parallelism","DeepSeek V4 Flash KV cache and batching optimization","DeepSeek V4 Flash long-context optimization","DeepSeek V4 Flash DSpark speculative decoding","DeepSeek V4 Flash MTP vs DSpark performance","DeepSeek V4 Flash GGUF quantization performance","DeepSeek V4 Flash tool-calling optimization","DeepSeek V4 Flash reasoning throughput","DeepSeek V4 Flash preview vs 0731 performance","DeepSeek V4 Flash benchmark TTFT TPOT throughput"]}}}}