---
title: "Introducing Kiln: LLM Inference on AWS Trainium"
description: "Kiln serves GLM-5.3-Flash on one trn1.32xlarge at $4.86 per 1M output tokens, 11-16% below vLLM on 8 x H200 at spot. How it was measured, and what it loses."
author: "Foxl Team"
date_published: "2026-10-05"
date_modified: "2026-10-05"
canonical_url: "https://foxl.ai/blog/introducing-kiln"
markdown_url: "https://foxl.ai/blog/introducing-kiln/index.md"
image: "https://foxl.ai/blog/introducing-kiln-cover.png"
social_image: "https://foxl.ai/blog/social/introducing-kiln.png"
content_type: "Deep dive"
topics: ["AWS Trainium","LLM inference","Mixture of experts","NKI kernels","Cost per token","Open source","Measurement"]
products: ["Kiln"]
---

<!-- Generated by scripts/blog-publishing.mjs; do not edit. -->

# Introducing Kiln: LLM Inference on AWS Trainium

> Kiln, Foxl's open-source LLM inference engine for AWS Trainium, serves GLM-5.3-Flash on one trn1.32xlarge spot instance at $4.86 per million output tokens at 64 concurrent requests: 11-16% below vLLM on eight H200 GPUs priced at their own spot band, and 73-75% below at 16. How it was measured, the changes that took the model from 19.4 to 122.9 output tokens a second, the ideas that did not pay, the bugs only the device showed, and what Kiln does not win: latency, concurrency 128, trn2 and on-demand prices.

- Author: Foxl Team
- Published: 2026-10-05
- Last reviewed: 2026-10-05
- Reading time: 25 minutes
- Canonical HTML: [https://foxl.ai/blog/introducing-kiln](<https://foxl.ai/blog/introducing-kiln>)

![An isometric illustration of Kiln's compile path: a stack of CPU servers, the stack of compiled graphs they produced, and the trn1.32xlarge the graphs load onto, drawn as a board of 16 Trainium chips with two NeuronCores on each.](<https://foxl.ai/blog/introducing-kiln-cover.png>)

Kiln is Foxl's LLM inference engine for AWS Trainium, and as of today it is open source under the Apache-2.0 license at [github.com/foxl-ai/kiln](<https://github.com/foxl-ai/kiln>). Serving zai-org/GLM-5.3-Flash on one trn1.32xlarge at its spot price of $2.15 an hour, Kiln produces output tokens for **$4.86 per million** at 64 concurrent requests of 8,192 tokens in and 256 out. vLLM on eight H200 GPUs, an ml.p5en.48xlarge, serves requests of the same shape at $5.48-5.77 per million output tokens when that instance is priced at its own spot band. That makes Kiln 11-16% cheaper per output token at concurrency 64 and 73-75% cheaper at concurrency 16. It is also far slower per request: at concurrency 64 the first token arrives after 6.9 seconds at the median, against 634 milliseconds on the GPUs.

This post explains what Kiln is, how the comparison was measured and priced, the changes that took this model from 19.4 to 122.9 output tokens a second on one box, the ideas that were measured and did not pay, the bugs only a device could show, and what Kiln still does not win. Every number comes from the measurement log published with the code, with its conditions stated beside it.

Code: [github.com/foxl-ai/kiln](<https://github.com/foxl-ai/kiln>), Apache-2.0 ([LICENSE](<https://github.com/foxl-ai/kiln/blob/main/LICENSE>)). Every run, with its exact command: [docs/price-performance.md](<https://github.com/foxl-ai/kiln/blob/main/docs/price-performance.md>). Engineering notes, including the negative results: [docs/neuron-notes.md](<https://github.com/foxl-ai/kiln/blob/main/docs/neuron-notes.md>). Design: [DESIGN.md](<https://github.com/foxl-ai/kiln/blob/main/DESIGN.md>).

## Highlights

| zai-org/GLM-5.3-Flash, 8,192 in / 256 out, measured 2026-10-05 | Kiln, trn1.32xlarge | vLLM, 8 x H200 |
| --- | --- | --- |
| Instance price, spot | $2.15/h | $28.77-30.29/h |
| Output tokens a second at concurrency 16 / 32 / 64 | 87.3 / 112.3 / 122.9 | 311 / 842 / 1,458 |
| Cost per 1M output tokens, concurrency 16 | **$6.84** (73-75% lower) | $25.7-27.1 |
| Cost per 1M output tokens, concurrency 32 | **$5.32** (44-47% lower) | $9.49-9.99 |
| Cost per 1M output tokens, concurrency 64 | **$4.86** (11-16% lower) | $5.48-5.77 |
| The same, with Kiln's opt-in decode kernels | **$4.36** (20-24% lower) |  |
| First token at concurrency 64, median | 6.9 s | 634 ms |
| Time between tokens at concurrency 64, median | 452 ms | 32.3 ms |
| Concurrency 128 | does not fit (16 GiB per NeuronCore) | $3.39-3.57 |
| Cost per request against a $0.15 / $0.50 list price ($0.001357) | $0.001244 (8% below) |  |

> **Cost per output token.** The instance's hourly price divided by the output tokens it produces in that hour at the measured throughput. At concurrency 64, $2.15 an hour over 122.9 × 3,600 = 442,440 output tokens an hour is $4.86 per million. The 8,192 prompt tokens of every request are paid for inside the same hour, so on a workload with 32 prompt tokens per output token this number is mostly the price of prefill.

## What is Kiln?

Kiln serves large open models on Trainium without AWS's managed Neuron inference layer, NxD Inference and vllm-neuron. From the Neuron stack it keeps only what cannot be replaced: the driver, the runtime, the collectives library and the neuronx-cc compiler, plus NKI, the kernel language for NeuronCores. Everything above that is its own: a scheduler with continuous batching and chunked prefill in which decodes never pause for prefill, a paged KV cache with FP8 pages, a radix prefix cache (which also works for linear-attention models, by checkpointing their recurrent state), tensor parallelism, expert parallelism and DP attention, NKI kernels for mixture-of-experts, linear attention and sparse attention, and an OpenAI-compatible server with streaming and tool-call and reasoning parsers. [FEATURES.md](<https://github.com/foxl-ai/kiln/blob/main/FEATURES.md>) compares it feature by feature with vLLM v0.30.0 and SGLang v0.5.21.

The constraint that shapes all of it is that a NeuronCore runs graphs compiled ahead of time. There is no eager kernel launch to fall back on, so every step the scheduler forms is padded to a bucket of three sizes, sequences, tokens and KV pages per sequence, and every bucket is a graph that has to exist before serving starts. The page axis is what stops a decode step from reading the full context length worth of KV for every sequence. On the executor Kiln uses today, a graph call costs about 94 microseconds at steady state, a floor under every decode step. Compiling those graphs is slow enough that Kiln compiles them on CPU hosts and ships the results to the Trainium box, which is the compile farm described below.

The runtime, collectives and compiler are binaries AWS does not let anyone redistribute, so Kiln ships as source and runs in the PyTorch inference environment of the AWS Deep Learning AMI for Neuron, SDK 2.32. The files it adapts from vLLM, SGLang and transformers are listed in [THIRD\_PARTY\_NOTICES.md](<https://github.com/foxl-ai/kiln/blob/main/THIRD_PARTY_NOTICES.md>).

## Why build an inference engine for Trainium?

Because the cheapest Neuron capacity is the part AWS's own inference software is leaving. As Kiln's design notes record it, NxD Inference has been in maintenance mode since SDK 2.30 and was removed from the Deep Learning AMIs in 2.32; its successor, vllm-neuron 0.24, runs on Trn2 and Trn3 only and supports five model families; and PyTorch/XLA inference on trn1 and inf2 has been in maintenance mode since 2.31. None of the top 13 open-weight models on the Artificial Analysis Intelligence Index (v4.3, 2026-10-01) had official Neuron support.

The other reason is price. Spot prices in us-east-2 on 2026-10-02 were $2.15 an hour for a trn1.32xlarge with 512 GB of HBM, $14.29 for a trn2.48xlarge with 1.5 TB, and $15.48 for a g7e.48xlarge GPU instance with 768 GB. A trn1.32xlarge is 16 Trainium chips with two NeuronCores each, 32 cores of 16 GiB, and FP8 matrix math runs at the BF16 rate on this generation. On demand the same instance is $21.50 an hour, ten times its spot price, a fact that matters again near the end of this post.

The goal, set on 2026-10-03, was to serve GLM-5.3-Flash on Trainium at a lower cost per token than a SageMaker deployment of vLLM on 8 x H200, on requests of the same shape.

## How was Kiln measured against vLLM on H200?

**The model.** zai-org/GLM-5.3-Flash with its real weights: 45 layers, 34 of them KDA linear attention and 11 DeepSeek-style sparse attention (DSA) over a compressed latent cache, 42 mixture-of-experts layers with 288 routed experts (8 per token) and a shared expert, four hyper-connection streams in place of one residual, and FP8 weights with 128 x 128 block scales, about 328 GB in all. AWS's managed Neuron path does not support it.

**The workload.** Every request is 8,192 prompt tokens and 256 output tokens, in a closed loop at the stated concurrency, 128 requests per level. Kiln's prompts are random tokens. With 32 prompt tokens per output token, every level is bound by prefill, which is why most of the work below is prefill work.

**Kiln's configuration.** One trn1.32xlarge, tensor parallelism 32 across all 32 NeuronCores, DP attention 4, prefill of 4,096 tokens per step, 12 mixture-of-experts layers per graph, Neuron SDK 2.32, every graph loaded from the compile farm and none compiled on the device. DP attention splits the 32 ranks into four groups of eight that each serve their own requests in the attention layers, while the expert layers run over every group's rows. It exists because this model's compressed attention cache is not split across heads, so under plain tensor parallelism every rank would hold every sequence's cache, about 15 GB per rank at 128 sequences of 8,448 tokens, and four groups divide that by four.

**The reference.** vLLM's purpose-built `vllm/vllm-openai:glm53-flash` image on a SageMaker real-time endpoint, an ml.p5en.48xlarge with 8 x H200, TP 4 / DP 2 / EP 8, measured with SageMaker's generative AI benchmarking on the same 8,192 / 256 streaming shape. The reference's prompt content is not recorded in Kiln's notes. SageMaker hosts that instance at $72.795 an hour.

**The pricing basis.** Like for like, spot against spot. A SageMaker endpoint has no spot price, so the reference's measured throughput is priced at the EC2 spot band of the same instance type, p5en.48xlarge at $28.77 to $30.29 an hour (us-east-2, read 2026-10-03), and Kiln at trn1.32xlarge spot. On-demand against on-demand is reported separately below, because Kiln does not win it.

**Correctness.** No throughput change counts until it passes two real-weight checks on the device: the mean log-probability of four short sentences, about -2.07, and a 3,071-token slice of wikitext-2 whose mean has to stay within 0.01 of -0.551. The four-sentence mean is oversensitive (one sentence, "Water boils", sits on a knife-edge and has moved by a nat from a few subnormal weight changes), so the wikitext slice is the one that measures quality.

## What does it cost per token?

![Figure 1, cost per 1M output tokens by concurrency, spot prices on both sides. At 16 concurrent requests vLLM on 8 x H200 costs $25.7-27.1 and Kiln on one trn1.32xlarge $6.84, 73-75% lower. At 32, $9.49-9.99 against $5.32, 44-47% lower. At 64, $5.48-5.77 against $4.86, 11-16% lower, and $4.36 with Kiln's opt-in decode kernels. At 128, vLLM costs $3.39-3.57; Kiln does not fit a trn1.32xlarge there, and on a trn2.48xlarge it costs $22.07.](<https://foxl.ai/blog/introducing-kiln-cost.png>)

| Concurrency | Kiln out tok/s | Kiln $ / 1M out | Kiln first token / between tokens, p50 | vLLM out tok/s | vLLM $ / 1M out | vLLM first token / between tokens, p50 |
| --- | --- | --- | --- | --- | --- | --- |
| 16 | 87.3 | **$6.84** | 8.7 s / 149 ms | 311 | $25.7-27.1 | 1,428 ms / 20.1 ms |
| 32 | 112.3 | **$5.32** | 6.5 s / 254 ms | 842 | $9.49-9.99 | 979 ms / 25.0 ms |
| 64 | 122.9 | **$4.86** | 6.9 s / 452 ms | 1,458 | $5.48-5.77 | 634 ms / 32.3 ms |
| 64, opt-in decode kernels | 137.1 | **$4.36** | 6.5 s / 410 ms |  |  |  |
| 64, opt-in mixed batches | 131.8 | $4.53 | 6.3 s / 419 ms |  |  |  |
| 128 | does not fit |  |  | 2,359 | $3.39-3.57 | 415 ms / 43.4 ms |

The H200 box produces 11.9 times Kiln's output tokens a second at concurrency 64, and at spot it costs 13.4 to 14.1 times as much an hour. Kiln's advantage is the difference between those two ratios, and it narrows as concurrency rises, because vLLM's throughput keeps climbing with the batch while Kiln on trn1 levels off: 87.3, 112.3 and 122.9 tokens a second at 16, 32 and 64, against 311, 842 and 1,458. At concurrency 128, where vLLM is at its cheapest, Kiln does not fit in trn1's memory at all.

The two opt-in rows are measured, and they are off by default for a reason given below. The combination of mixed batches with the decode kernels was not measured.

### Against model providers' prices

A provider billing GLM-5.3-Flash at $0.15 per million input tokens and $0.50 per million output tokens charges $0.001357 for one 8,192 / 256 request. Kiln's cost for the same request at concurrency 64 is $0.001244 with its defaults, 8% below that list price, and $0.001115 with the decode kernels, 18% below it. Discounted providers at $0.075 and $0.25 charge $0.0006784, which is still 1.6 to 1.8 times cheaper than Kiln's cost. These compare Kiln's instance cost with a provider's price, which carries everything else a provider pays for, so they say what self-hosting on trn1 spot costs against buying tokens, not that Kiln is cheaper to run than a provider's own fleet.

Prompt caching lowers Kiln's cost per request, and it lowers a provider's bill too. On an earlier tree of 2026-10-04, with 75% of every prompt one of four shared prefixes and the cache warm, concurrency 64 served 231.6 output tokens a second at $0.000660 per request, 58% below the same requests uncached. A provider bills cached input tokens at a discount ($0.03 per million at list), and against that bill Kiln's cost was 1.065 times the list price. The reason is the output token: split by device time, a cached input token costs Kiln about nothing and an uncached one about the list's $0.15 per million, but an output token costs $1.42 to $1.97 per million against a list price of $0.50. With most of the input cached the request is bound by decode, and the decode step is where trn1 is weakest.

## What moved the number from 19.4 to 122.9 tokens a second?

![Figure 2, output tokens per second at concurrency 64 on one trn1.32xlarge, one bar per change: the first 8192 / 256 sweep 19.4, the MoE prefill kernel and elementwise hyper-connections 29.8, 12 MoE layers per graph with prefill 2048 at 54.4, sequence-parallel prefill streams 67.7, the DSA and KDA kernels with prefill 4096 and KV sized for every sequence 88.3, reduce-scatter with the pool-key cache and tile skip 94.9, expert parallelism 106.9, per-tile FP8 dequantization 114.9, collectives over the 8-rank group 123.1, and the opt-in decode kernels 137.1. A shaded band at 103.5-109.0 marks where Kiln matches the cost of vLLM on 8 x H200 at spot; expert parallelism is the first step inside it.](<https://foxl.ai/blog/introducing-kiln-progression.png>)

The first end-to-end run of this workload at concurrency 64 produced 19.4 output tokens a second. The final default run produced 122.9, which is 6.3 times as many on the same instance type. Each step in the table was measured on the same box against the step before it.

| Change | Out tok/s at concurrency 64 |
| --- | --- |
| first 8192 / 256 sweep at concurrency 64 | 19.4 |
| prefill MoE kernel on real FP8 scales, elementwise hyper-connections | 29.8 |
| 12 MoE layers per graph, prefill 2048 per step | 54.4 |
| sequence-parallel prefill streams | 67.7 |
| DSA top-k and scoring kernels, KDA kernel, prefill 4096, KV sized for every sequence | 88.3 |
| reduce-scatter prefill, separate FP8 pool-key cache, MoE tile skip | 94.9 |
| expert parallelism | 106.9 |
| per-tile FP8 dequantization in the expert-parallel kernel | 114.9 |
| attention collectives over the 8-rank DP group | 123.1 (the final default run: 122.9) |
| opt-in KDA and sparse-DSA decode kernels, sequence-parallel decode streams | 137.1 |

The rest of this section explains the steps that taught us something about the hardware. Each one names the section of [docs/neuron-notes.md](<https://github.com/foxl-ai/kiln/blob/main/docs/neuron-notes.md>) that holds its measurements.

### Prefill was the workload

The first sweep, at concurrency 8, processed about 460 prompt tokens a second on the trn1.32xlarge. The GPU reference processed about 76,000 a second at concurrency 128. Decode waited behind those prefill steps, which is why its time between tokens was 447 ms. So the order of work was set by prefill: a prefill kernel for the experts, bigger prefill chunks, and anything that took time out of a prefill layer.

### A collective costs per graph, so the graphs got big

On trn1 a graph that holds a collective across chips pays a fixed cost every time it executes, and further collectives in the same graph are nearly free. At 32 ranks a graph with one all-reduce of a small hidden state took 5.1 ms chained, and one with twelve took 5.0 ms. The cost depends on how many chips the collective crosses. Inside one 32-rank world, an all-reduce over the two ranks of one chip took 0.15 ms, and over eight ranks on four chips 2.55 ms ("Collectives across chips: a fixed cost per graph EXECUTION"). One graph per layer would therefore pay that cost 45 times a step. Kiln groups 12 mixture-of-experts layers into one graph, which only fits under neuronx-cc's limit of 5 million instructions because the expert computation is one NKI kernel, a grouped GEMM that applies each expert once to all of its tokens. At concurrency 32, 12-layer prefill graphs ran at 38.4 tokens a second against 26.8 for two-layer graphs, 43% more.

### The hyper-connection arithmetic was what spilled

GLM-5.3-Flash's layer graphs at prefill cost 3 to 8 times the sum of their parts. The cause was not the kernels: a null NKI kernel that does no work inflated one layer to 81.6 ms exactly as the real kernel did. It was the four hyper-connection streams, written as in the reference implementation, an fp32 copy of a \[T, 16,384\] tensor and thousands of tiny per-token \[4, 4\] matrix products. neuronx-cc's allocator spilled 43% of the graph's 4,371 tensors and generated 26.7 GB of spill traffic. The same algebra written as multiply-adds over \[T, 4,096\] rows took layer 0 from 10.7 ms to 3.79 ms with the linear-attention kernel in it ("The layer-graph spill was the hyper-connection arithmetic").

Rewriting it raised a numerics problem that only showed on the device. The reference rounds the streams to bf16 at fixed points, and inside the 12-layer graphs neuronx-cc did not keep an explicit fp32 to bf16 to fp32 round trip between elementwise operations, so the device's results drifted from the CPU's. The fix rounds with arithmetic the compiler cannot fold: a Veltkamp split in fp32, c = 65537x and hi = c - (c - x), which gives exactly the round-to-nearest-even bf16 value (checked on 100,000 values) for about 1 ms per layer. With it, the wikitext slice scores -0.5515 on the device, against -0.548 for the reference form on both the device and the CPU.

### Sequence-parallel streams, and KV sized for every sequence

Between prefill layers each of the 32 ranks now keeps only 1/32 of the hyper-connection stream rows, and each block gathers them back as it needs them. That was worth 23.5% at concurrency 32 (46.3 to 57.2) and 24% at concurrency 64 (54.4 to 67.7). The next lesson was about memory rather than kernels. At an 8,448-token context, 16 sequences per DP group need 4,224 KV pages, and the 1.0 GB FP8 KV pool of the earlier runs held 3,972, so the scheduler could not keep every request in flight. A run that is short of KV reads like a kernel regression. Sizing the pool at 1.25 GB gave 16% more throughput, 1.5 GB gave 88.3 tokens a second, and 1.75 GB gave exactly the same 88.3, which showed the pool was no longer the limit.

### Expert parallelism without an all-to-all

Under tensor parallelism each rank holds 64 intermediate rows of all 288 experts. Under expert parallelism each of the 32 ranks holds 9 whole experts and computes only the token-expert pairs routed to them. The usual cost of expert parallelism is a dispatch all-to-all before the experts and a combine after them. Kiln has neither, because under DP attention every rank already holds every row of the batch at the expert layer, and the block's output is all-reduced anyway for the shared expert, which stays tensor-parallel. So the layout moves no extra byte between ranks. One expert block at 4,096 rows took 16.03 ms under tensor parallelism and 5.49 ms under expert parallelism, and the end-to-end gain at concurrency 64 was 12.6%, 94.9 to 106.9 ("Expert parallelism for GLM-5.3-Flash's routed experts"). That step is the first one inside the H200's cost band in Figure 2.

The next step was in how the FP8 weights are dequantized. The loader now fits each 128 x 128 block so that every kernel tile carries one scale, and each tile is dequantized by one instruction, alternating between the vector and the scalar engine. On the busiest rank, with the routing the benchmark's random-token prompts produce, that took layer 20's expert kernel from 8.03 to 5.61 ms, and the end-to-end number from 106.9 to 114.9.

### The token mixers' collectives inside their group

An attention layer under DP attention reads only its own group's rows, so its gather and reduce-scatter can run over the 8 ranks of its group, on 4 chips, instead of the whole 32-rank world. That took the 4,096-token prefill call at concurrency 32 from 0.665 to 0.593 s and the concurrency-64 number from 114.9 to 123.1, 7.1% ("The token mixers' collectives inside their attention group"). The expert layers cannot do the same, because the experts need every row.

After this step the prefill call breaks down as follows ("Where the 4096-token prefill call goes after the group collectives"): the 42 expert blocks are about 71% of it, the 11 sparse-attention blocks about 15%, the 34 linear-attention blocks about 10%, and fixed per-graph costs about 2%. Expert load balance and the expert block's three world collectives are what is left to work on.

### The decode kernels, and why they are opt-in

Three changes target the decode step: KDA decode as one NKI kernel that updates each sequence's recurrent state in place, sparse-attention decode that gathers only the 512 selected pools instead of masking all 8,448 keys (5.73 ms down to 0.73 ms at 16 rows), and sequence-parallel streams in decode as in prefill. Together they take the decode step at 64 rows per group from 456.5 to 233.7 ms ("Where a GLM-5.3-Flash decode step goes, from a device profile, and three decode kernels"). At concurrency 64 that is 137.1 tokens a second, $4.36 per million. The decode-path log-likelihood did not move: +0.0003 nats per token with the kernels, against a run-to-run floor of +0.0000.

They stay off by default because they were measured on three configurations and not on the others (concurrency 16, trn2, mixed batches, speculative decoding), and turning them on changes the cache key of every decode graph, so making them the default needs those configurations compiled and measured first.

## What did we measure that did not pay?

The notes keep the ideas that did not work beside the ones that did, with the same level of measurement. Five of them are worth knowing before trying them on Trainium.

- **An all-to-all in place of zero-padded all-reduce gathers.** Sending every rank's \[128, 4,096\] rows to every rank with an all-to-all was exact, and took 34.7 to 39.2 ms against 5.2 to 6.8 ms for the zero-padded all-reduce it would replace, about six times slower on trn1. The gathers stay all-reduces, also because a cached all-gather graph breaks when it is loaded in a later process.
- **DeepSeek-style fine-grained FP8 GEMM.** Applying the scales inside the matrix multiply would cost 4 to 8 times the vector-engine work of dequantizing each tile once, and a NeuronCore has one vector engine.
- **Static hot-expert placement.** Per layer, the 32 most-loaded experts on wikitext and on the benchmark's random tokens overlap in 13% of their members on average, and the top 8 in 1.8%. A placement tuned on either input is an artifact for the other, and a hybrid tuned on the batch's own hot set was no faster than plain expert parallelism.
- **Multi-token-prediction speculative decoding.** GLM-5.3-Flash's own MTP head drafted well on the benchmark's prompts: 96.7% of drafts accepted, 1.97 tokens per verify at concurrency 64. Measured separately, acceptance was 75% on wikitext and 88% on chat. It was still 4.6% slower end to end, 106.9 to 102.0, and slower at every level ("MTP speculative decoding on GLM-5.3-Flash at serving scale"). With 32 prompt tokens per output token every step that carries a prefill also pays the verify's extra rows and a 4,096-row draft pass, and the synchronous speculative step gives up the overlap scheduling the baseline has. It should pay on decode-heavy traffic once drafting runs on the device.
- **Expert-parallel decode.** Its floor is set by the bytes it has to read. A NeuronCore moved 224 to 230 GB/s from HBM in every variant tried, so reading a rank's 9 whole experts takes about 1.0 ms, while tensor parallelism reads only the slices of the experts a batch touches. An expert-parallel decode step therefore cannot match a tensor-parallel one at 8 rows per group whatever the kernel, and measured 24% slower there. Expert parallelism still wins overall from 8 rows per group, because its prefill gain is larger (98.9 against 90.8 tokens a second at concurrency 32), so Kiln turns it on only from there: at concurrency 16, with 4 rows per group, it measured 80.5 against 82.9.

## What could only a device show?

Most of Kiln is tested on CPU, where the same model code runs as the correctness reference. These bugs did not exist there.

**Padded rows that raced for one memory slot.** Two of six cold reruns of the same requests, with the batch composition held fixed, differed from the first run after output token 16. Under DP attention a decode call carries 8 rows per group, and a group with one request has 7 rows of padding. Every padded row wrote its KV into slot 0 of the null page and read that same slot back as its one visible key, in the same 12-layer graph: seven writers and seven readers of one address, whose order the program does not fix. A padded row's output is discarded, but it shares the call with the real row, most likely through the expert kernel's routing plan, and it changed the real request's log-probabilities by up to 0.13 nats. Replaying one decode call 1,000 times showed the padded rows of one group differing in 999 of them, while one-layer graphs happened to be stable. The fix gives padded rows the other slots of the null page and never writes slot 0 after start-up. No graph changed. On the final graphs, 1,000 replays of a decode call differed in none, synchronously or with four calls in flight ("Padded rows wrote the slot they read: why a cold rerun differed").

**A copy from the host does not wait for queued graphs.** The same investigation found that an eager copy into device memory is not ordered behind graph calls already launched on the core. A probe launched a decode call, then copied data into the slot that call writes, then read it back: the slot held the decode's write in 200 of 200 trials. The engine makes such copies when it restores a cached prefix from host memory while the previous step is still running, and a restored page could be one that step was still writing. Restores now first read back the output of the last call launched on the rank; the same probe then failed 0 of 200 times.

**A kernel the simulator passed.** Four accumulation groups written into four 128-column slices of one PSUM tile, read after the fourth, lost terms on the device while `nki.simulate` was exact. One block per PSUM tile, read right after its group, fixed it. The notes now advise comparing a kernel's device output with the simulator's block by block.

**Three compiler traps.** A comparison against a Python float literal can lower to f64, which neuronx-cc rejects for the whole graph. Integer division by a constant is computed through an fp32 reciprocal and can be off by one: gpt-oss-120b at tensor parallelism 32 located its vocabulary shard as `vocab_start // 6284` and returned top tokens from one shard too low. And because the compile cache key hashes the traced FX graph text, in which Dynamo names nodes after the first Python local a value is stored in, a refactor that only moves code into a helper can change the key of every graph that traces it while the HLO stays identical, and every cached graph then misses.

## Why compile on CPUs?

One 12-layer prefill graph of GLM-5.3-Flash took 45 minutes to compile on the trn1.32xlarge, and the compiler needs up to 174.9 GB of host memory for it. A device box compiles one graph at a time, and while it does, all 32 NeuronCores sit idle at the instance's hourly price. A compile, though, is a command-line call on an HLO file, and a CPU host produces the same artifact: the farm's compiled graph and the device's differ only in a recorded output path and debug files, 66 of 66 archive members otherwise identical ("Compile farm: neuronx-cc on CPU hosts, and capturing graphs without a NeuronCore").

So Kiln captures every graph a configuration will trace without a NeuronCore, one process per tensor-parallel rank on PyTorch's meta device with shapes read from the checkpoint's headers (all 32 ranks of GLM-5.3-Flash in about three minutes), compiles them on CPU instances, and stores them where a device box fetches them. On spot, 24 decode graphs compiling at once on a c8i.48xlarge come to 68 graphs an hour per dollar, against 3.4 for a trn1.32xlarge compiling them one at a time. The farm also finds failures sooner: a configuration whose prefill graphs exceeded the instruction limit failed on the farm 4 minutes after capture, where the device would have reached it about an hour into its warm-up. Every final run above loaded all of its graphs from the farm, with no compile on the device.

Capture had two traps of its own. The process group has to exist before the Neuron PyTorch layer is imported, or two operators are recorded in a different spelling and hash to a different key. And the Neuron collective lowerings have to be registered, or torch\_xla lowers the all-reduce itself into a different program under the very key the device computes, because the key hashes the FX graph and not the HLO. A diff of the two HLO files caught it before a wrong graph was ever loaded.

## What does Kiln not win?

- **Latency.** At concurrency 64 the first token takes 6.9 s at the median and 79.7 s at the 90th percentile, and each next token about 452 ms, against 634 ms and 32.3 ms for vLLM on H200. The win is cost per token, for throughput workloads whose requests can wait seconds for their first token; interactive chat is not what this configuration is for.
- **Concurrency 128.** It does not fit in trn1's 16 GiB per NeuronCore, and that is where vLLM is cheapest, at $3.39-3.57 per million output tokens.
- **trn2.** A trn2.48xlarge reaches 189.9 output tokens a second at concurrency 128, which at its spot price of $15.09 an hour is $22.07 per million, against $3.39-3.57 for the H200 at the same concurrency. It has about 3.5 times trn1's HBM bandwidth and compute per box at about 7 times the spot price, so today its case is capacity per box, not cost.
- **On-demand pricing.** At trn1.32xlarge's on-demand $21.50 an hour, concurrency 64 costs $48.59 per million output tokens, against $13.87 for vLLM at SageMaker's $72.795. No level is won on demand.
- **Discounted providers.** As above, they are 1.6 to 1.8 times cheaper than Kiln's cost per request.
- **The decode step's own cost.** With the default kernels a decode step takes 5 to 10 times the time it would need to read its weights, KV and state once from HBM, and bigger batches alone level off near $1.1 per million output tokens decode-only. That gap is the largest lever left on trn1, and the reason the output token, not the input, is what keeps Kiln above the list price on cached workloads.
- **Breadth.** This is one model measured to this depth. What runs on real weights, and what is still missing for the other top open models, is in [docs/model-coverage.md](<https://github.com/foxl-ai/kiln/blob/main/docs/model-coverage.md>).

## How do I read or run it?

The code is at [github.com/foxl-ai/kiln](<https://github.com/foxl-ai/kiln>) under Apache-2.0. The [README](<https://github.com/foxl-ai/kiln/blob/main/README.md>) has the results and a quickstart for a small model on one NeuronCore; [DESIGN.md](<https://github.com/foxl-ai/kiln/blob/main/DESIGN.md>) explains the layers and why each technique from vLLM and SGLang takes the shape it does on a static-shape accelerator; [docs/price-performance.md](<https://github.com/foxl-ai/kiln/blob/main/docs/price-performance.md>) holds every run with its configuration and the full iteration history, and its "How to resume" section gives the commands for the configurations above.

```
git clone https://github.com/foxl-ai/kiln && cd kiln
source /opt/aws_neuronx_venv_pytorch_inference_vllm_0_24_0_1_1_0/bin/activate
export PYTHONPATH=$PWD

# A small model on one NeuronCore; the first request compiles its graphs into the local cache
python -m kiln --model Qwen/Qwen3-0.6B --device neuron --port 8000
```

GLM-5.3-Flash uses all 32 NeuronCores of a trn1.32xlarge, and its graphs take hours to compile on the device, so compile them on a CPU host first with `tools/compile_farm.py` and serve from the cache. The benchmark is `bench/serve_sweep.py`, and the tests run on CPU.

We will keep the notes in the same shape as the code moves: a measurement for every claim, and the ideas that lost kept next to the ones that won. If you run Kiln on Trainium and get a number that disagrees with ours, the notes say how ours were taken, and an issue on the repository is the place to tell us. For the measurement habit behind this, see [A citation is not a measurement](<https://foxl.ai/blog/a-citation-is-not-a-measurement>).

## References and further reading

1. [Kiln on GitHub (Apache-2.0)](<https://github.com/foxl-ai/kiln>) - Reference
2. [Every measurement: docs/price-performance.md](<https://github.com/foxl-ai/kiln/blob/main/docs/price-performance.md>) - Reference
3. [Engineering notes: docs/neuron-notes.md](<https://github.com/foxl-ai/kiln/blob/main/docs/neuron-notes.md>) - Reference
4. [Kiln design](<https://github.com/foxl-ai/kiln/blob/main/DESIGN.md>) - Documentation