b200
Here are 14 public repositories matching this topic...
Serving and benchmarking 9 frontier MoE checkpoints (490 GiB+) with vLLM on a single 8x B200 node, plus a pre-download HBM fit checker
-
Updated
Aug 18, 2026 - Python
A Flexible and High-Performance Inference Serving Engine for Diffusion Language Models
-
Updated
Aug 31, 2026 - Python
Deploy DeepSeek-V4-Flash-0731 on dual NVIDIA RTX PRO 6000 Blackwell GPUs with vLLM PR #41834 (jasl fork) and DSpark speculative decoding, achieving ~200-227 tok/s in no-overseas-network environments.
-
Updated
Aug 31, 2026
Burst-serving Qwen3.8-27B FP8 on an on-demand RunPod B200, joined to a tailnet as b200 — no public endpoint. Measured: what MTP speculative decoding is worth, what it costs, and what is still unmeasured.
-
Updated
Aug 21, 2026 - HTML
Reproducible vLLM-Omni serving and benchmark reference for NVIDIA Cosmos3-Super on H200 and B200 GPUs
-
Updated
Aug 31, 2026 - Python
One-B200 DeepSeek V4 Flash deployment, correctness, and reproducible benchmark harness
-
Updated
Aug 5, 2026 - Python
Prebuilt spconv wheels for NVIDIA Blackwell / RTX 50-series (sm_120), CUDA 12.8 — fixes "no kernel image is available for execution on the device"
-
Updated
Sep 2, 2026 - Shell
Spheron — independent third-party profile of a public API surface, by API Evangelist. Spheron Network is a decentralized GPU and cloud compute marketplace that aggregates enterprise-grade NVIDIA GPU capacity from certified Tier 3/4 data centers worldwide and exposes it through a single on-demand, per-minute billed interface.
-
Updated
Sep 1, 2026
Add this topic to your repo
To associate your repository with the b200 topic, visit your repo's landing page and select "manage topics."