GLM-5.3-Flash is a mixture-of-experts (MoE) model with 320B total parameters, 18B of them active per token. Officially it requires NVIDIA Hopper or newer, and its reference sparse-attention and indexer kernels only run on Hopper. This report covers an effort to serve it on a single node of eight NVIDIA L40s, which are Ada-generation GPUs, inside a VM where the GPUs have no peer-to-peer path to each other.
The first build that worked decoded at 7.2 tokens per second and had enough KV cache for one 256k-token session. After a series of profiled changes, the same node holds eight coding sessions of roughly 230k tokens each and decodes 704 tokens per second across them. Swapping in a draft model that is licensed for non-commercial use only raises that to 798. Several of the changes are lossy, and one of them makes prefill slower, so the costs are covered alongside the gains.
Results at a glance
Metric | First working build | Final build | Change |
|---|---|---|---|
Combined decode at ~240k context | 7.2 tok/s | 704 tok/s (798 with DFlash2) | 98× |
Single-session decode | 7.2 tok/s | 191 tok/s (203 with DFlash2) | 26× |
256k-token sessions held in KV cache | 1 | 8 of 8 slots | 8× |
KV cache capacity | 295k tokens | 2.85M tokens | 9.7× |
Time to first token, 256k prompt | 298 s | 50 s | 6× |
Perplexity vs. exact FP8 model | n/a | +0.26% | n/a |
Some qualifications apply. The final KV pool could hold 10.9 sessions of 256k tokens, but the server was configured with eight slots. The 50-second prefill comes from the FP8-weight build. The 4-bit build, which is the one that produces the eight-session numbers, takes 56 to 72 seconds. Final decode figures are steady-state measurements at about 240k context. A 250k-token needle-in-a-haystack test also passed.
Test environment
- 8× NVIDIA L40, 46 GB each (Ada Lovelace,
sm_89) - KVM guest, with all GPUs behind the host bridge and GPU peer-to-peer (P2P) disabled
- Measured NCCL bandwidth of about 1 GB/s
- vLLM 0.30 with a custom plugin, no fork of vLLM
The three constraints
Missing kernels. The model's sparse attention and its indexer ship with kernels written for Hopper. Running on Ada at all meant writing replacements. These were written in Triton and cover both the sparse-attention and indexer kernels, with FP8 KV cache support and split-K decode.
No fast path between GPUs. A model this size has to be split across all eight GPUs with tensor parallelism, which puts collective operations (all-reduce, all-gather) on the critical path of every layer. With P2P disabled, NCCL ran at about 1 GB/s. Long prefills, which move large activations between GPUs, were hit hardest.
Memory. At FP8, 320B parameters take up roughly 320 GB, against 368 GB of total GPU memory on the node. After weights and runtime overhead, the first build had room for 294,677 tokens of KV cache, or a single 256k session.
Step-by-step changes
Each change was profiled, applied and re-measured. In the table below, prefill is the time to first token for a 256k-token prompt, and decode is a single session at about 64k context. Steps 0 through 4 were measured on random-token prompts and later steps on real coding-agent prompts, so the two halves are not directly comparable.
Step | Change | 256k prefill | Decode (tok/s) | KV tokens | 256k sessions that fit |
|---|---|---|---|---|---|
0 | Baseline: Ada kernels for the Hopper-only sparse attention, eager mode, NCCL | 298 s | 7.2 | 294,677 | 1 |
1 | CUDA graphs | 298 s | 70.7 | 294,677 | 1 |
2 | Host-staged all-reduce, bypassing the 1 GB/s NCCL path | 77.7 s | 70 | 294,677 | 1 |
3 | INT8-compressed prefill all-reduce | 56.1 s | 70 | 294,677 | 1 |
4 | FP8 KV cache | 57 s | 70 | 561,895 | 2 |
5 | MTP speculative decoding (3 tokens) | 57 s | 163 | 310,490 | 1 |
6 | Tensor-parallel MoE, plus a fix for a vLLM out-of-bounds read | ~50 s | 175 | 310,490 | 1 |
7 | FP8 linear-attention projections | ~50 s | 182 | 382,472 | 1 |
8 | Host-staged logits all-gather | ~50 s | 194 | 382,472 | 1 |
9 | 4-bit EXL3 experts running inside vLLM (new MoE kernel path) | 57–72 s | 186 | 2,847,055 | 8 |
10 | DFlash2 drafter in place of MTP, CUDA graphs up to 48 tokens | 56–72 s | 195 | 2,825,772 | 8 |
"Sessions that fit" is KV tokens divided by 262,144, rounded down and capped at the eight configured slots.
Launch overhead (step 1)
Capturing decode in CUDA graphs took single-session decode from 7.2 to 70.7 tokens per second, close to a 10× gain, with no change to prefill. A jump that large indicates the eager-mode baseline was spending most of its decode time on CPU-side kernel launches rather than on GPU work.
Collectives through host memory (steps 2, 3 and 8)
All-reduce and all-gather were moved off NCCL and staged through pinned host memory instead, using the copy directions that are fast inside this VM. The implementation is safe to use under CUDA graphs and can optionally send INT8 payloads.
Routing all-reduce this way cut 256k prefill from 298 s to 77.7 s, a 3.8× improvement. Compressing the prefill all-reduce to INT8 brought it down to 56.1 s, about 5.3× faster than the baseline. Decode stayed at roughly 70 tok/s through both steps. Later, sending the logits all-gather through host memory as well raised decode from 182 to 194 tok/s.
KV cache and weight precision (steps 4 and 7)
Storing the KV cache in FP8 nearly doubled capacity, from 294,677 to 561,895 tokens, which is enough for two 256k sessions. Converting the linear-attention projections to FP8 shrank those weights. The KV pool grew from 310,490 to 382,472 tokens, and decode went from 175 to 182 tok/s.
Speculative decoding (steps 5 and 10)
Multi-token prediction (MTP) with three speculative tokens took single-session decode from 70 to 163 tok/s, a 2.3× gain. It also cut the KV pool from 561,895 to 310,490 tokens, back below what two 256k sessions need. That is why the best FP8-weight result at long context, covered in the next section, runs without speculation.
Both MTP and DFlash2 had to be wired into this model's KV cache layout. In the final build, the drafter's cache is aliased onto existing KV blocks, so it only consumes the blocks it actually uses. Step 10 replaces MTP with the DFlash2 drafter and extends CUDA graph capture to 48 tokens, which raised single-session decode at 64k from 186 to 195 tok/s.
Tensor-parallel MoE (step 6)
Running the expert layers tensor-parallel, together with a fix for an out-of-bounds read in vLLM, raised decode from 163 to 175 tok/s and brought 256k prefill down to about 50 s.
4-bit experts (step 9)
The largest capacity gain came from quantizing the expert weights to 4 bits. A new EXL3 MoE path runs ExLlamaV3's fused kernel on each tensor-parallel rank, with no host synchronization. Expert weights dropped from 39 GiB to 23 GiB per GPU, and the freed memory went to the KV cache, which grew 7.4×, from 382,472 to 2,847,055 tokens. This is the step that makes eight concurrent 256k sessions possible.
It came at a cost. Single-session decode at 64k slipped from 194 to 186 tok/s, and 256k prefill rose from about 50 s to between 57 and 72 s.
Throughput at long context
The table below shows combined steady-state decode at about 240k context for each major configuration. All runs used real coding-agent prompts made up of a system prompt, tool schemas and repository source files.
Configuration | Sessions resident | Combined decode (tok/s) |
|---|---|---|
Baseline, FP8 weights | 1 | 7.2 |
Tuned, FP8 weights, no speculation | 2 | 148 |
Tuned, 4-bit experts + MTP | 8 | 704 |
Tuned, 4-bit experts + DFlash2 | 8 | 798 |
The jump from 148 to 704 tok/s comes from two things the FP8-weight build could not have at the same time at this context length: speculative decoding and eight resident sessions. Both depend on the memory freed by 4-bit experts.
Scaling from one to eight sessions
These runs use 4-bit experts and keep all eight sessions resident in KV cache. The FP8-weight build could not run eight sessions at 64k context or above. Values are combined steady-state decode in tokens per second. MTP-3 is multi-token prediction with three speculative tokens, and DFlash2-5 is the DFlash2 drafter.
Context | Drafter | 1 session | 2 sessions | 4 sessions | 8 sessions |
|---|---|---|---|---|---|
~22k | MTP-3 | 172 | 245 | 395 | 646 |
~22k | DFlash2-5 | 202 | 263 | 439 | 644 |
~64k | MTP-3 | 186 | 244 | 430 | 642 |
~64k | DFlash2-5 | 195 | 301 | 534 | 713 |
~240k | MTP-3 | 191 | 276 | 439 | 704 |
~240k | DFlash2-5 | 203 | 242 | 608 | 798 |
A few patterns stand out:
- Combined throughput rises with concurrency but less than linearly. At about 240k with MTP-3, eight sessions deliver 3.7× the throughput of one, or about 88 tok/s per session.
- Throughput does not fall as context grows. With MTP-3 at eight sessions, the results were 646, 642 and 704 tok/s at about 22k, 64k and 240k. The report does not isolate why. Sparse attention, which limits how much of the cache each decode step reads, and differences in speculative acceptance between prompt sets are both plausible factors, but neither was tested directly.
- DFlash2 leads at eight sessions for 64k and 240k context, by 11% and 13%, and is level with MTP at 22k. At lower concurrency the results are mixed. At about 240k with two sessions, DFlash2 trailed MTP, 242 to 276 tok/s.
Quality
Four of the changes trade precision for speed or memory: 4-bit experts, the FP8 KV cache, INT8 all-reduce and FP8 linear-attention projections. Two checks compared the final build against the exact FP8 model:
- Perplexity on 30k tokens of source code was 1.5420, against 1.5380 for the exact model, an increase of 0.26%.
- A 250k-token needle-in-a-haystack test passed at all three depths tested.
These results suggest the lossy changes did not seriously degrade the model, but the evaluation is narrow. Perplexity over 30k tokens and a single retrieval test do not measure how well the model completes coding-agent tasks, which is the workload these sessions represent. A task-level evaluation would be needed before treating the quantized build as equivalent to the reference model.
Licensing
The two drafters come with different terms. The DFlash2 drafter is licensed CC BY-NC-ND 4.0, which rules out commercial serving. For commercial use, the configuration is 4-bit experts with MTP, which delivered 704 tok/s across eight sessions at about 240k context.
The 4-bit expert checkpoint is published under the ShapleyMcg License 1.0. It allows commercial use as long as ShapleyMcg (Brandon M. Music) is credited and the license notice is kept. As with any third-party weights, read the license text directly before deploying.
Measurement notes
- "Steady state" decode is read from the server's generated-token counter while every session is decoding.
- Prompts in the ~240k rows ranged from 216k to 239k tokens.
- Long prompts queue behind each other at high concurrency. With eight sessions of about 227k tokens, time to first token for the last session reached about 220 s while the other prefills finished.
- Steps 0–4 of the step table used random-token prompts. Everything after used coding-agent prompts. Speculative decoding gains depend on how predictable the output is, so other workloads may see different numbers.
What still limits performance
Prefill is now bound by the VM's host memory bandwidth of about 100 GB/s, because every cross-GPU transfer passes through host memory. GPUs with P2P or NVLink would remove that ceiling. On this node it shows up in two places: 256k prefill takes 50 s or more, and time to first token grows quickly when several long prompts arrive together.
The 4-bit build also favors capacity over latency. It is the only configuration that fits eight long sessions, but it prefills more slowly than the FP8-weight build. For workloads with a few latency-sensitive long prompts rather than many concurrent ones, the FP8-weight build may be the better fit.
Summary
The first working build of GLM-5.3-Flash on eight L40s ran, but at 7 tokens per second with room for one long session, it was not practical for long-context serving. Most of the gap was closed in software. CUDA graphs gave close to 10× on decode, host-staged collectives made 256k prefill about 5× faster, speculative decoding added 2.3× on decode, and 4-bit experts increased KV capacity 7.4×. The result serves eight coding sessions of about 230k tokens at 704 tok/s combined, with a stack that permits commercial use, on hardware the model does not officially support and without forking vLLM. What remains is the interconnect, which software can work around on this VM but not remove.