GPT-6 Astra Ultrafast: Same Model, a Third of the Time
GPT-6 Astra on Amazon Bedrock's Ultrafast tier, side by side with Standard in Foxl: one prompt sent to two windows at the same moment, sixteen times. Ultrafast finished first in all sixteen, at a median of 7.1 seconds against 20.9, and in direct API calls it streamed past the 300 tokens a second AWS quotes. How it was measured, why the multiple is three and not six, what the switch does and does not do, and what it costs.
On this page
The video above is two Foxl windows on the same Mac, recorded at Retina resolution. Both are set to GPT-6 Astra on Amazon Bedrock, through the same AWS account. The right one has the Ultrafast switch turned on, and you watch it being clicked. Then the same prompt goes into both and is sent to both at the same moment: write a TypeScript LRU cache with per-entry expiry, document every method, and test it.
The right window finished in 6.5 seconds. The left one took 20.5. Nothing between the send and the last token is sped up, slowed down or cut; the two timers are each window's own time from send to the last token.
One run is an anecdote, so we ran the same pair 16 times. Ultrafast finished first in 16 of 16, at a median of 7.1 s against 20.9 s: the same model, answers of the same length, in about a third of the time. Called directly, without Foxl in between, it streamed at more than 300 tokens a second in 11 of 12 runs.
Highlights
| Median of 16 paired runs in Foxl | Standard | Ultrafast |
|---|---|---|
| Send to the last token | 20.9 s | 7.1 s (2.9x sooner) |
| Send to the first text on screen | 2.75 s | 1.61 s |
| Generation speed, first text to last token | 117 tokens/s | 381 tokens/s |
| Output tokens per answer | 2,168 | 2,093 |
| Pairs Ultrafast finished first | 16 of 16; per pair 1.9x to 4.2x sooner, median 3.0x | |
| Direct API, six routes (12 pairs) | 112 to 116 tokens/s | median 406 tokens/s, 11 of 12 above 300 |
| Price of this turn at the US rates, nothing cached | about $0.21 | about $1.25 |
What Ultrafast is
On Bedrock, Ultrafast is a service tier of GPT-6 Astra rather than a separate model: the same model IDs, with one field on the request. The model card puts it in two sentences: “OpenAI reports API speeds up to six times faster than Standard. To request it with the Responses API, set "service_tier": "ultrafast".” AWS announced it on September 30, and Foxl v0.7.32 shipped it the same day.
It is available on bedrock-runtime through the US geographic and the global cross-Region profiles (us.openai.gpt-6-astra and global.openai.gpt-6-astra), and on bedrock-mantle in us-east-1. Its prices are, in the card's words, “six times the corresponding Standard prices”: $66 per million input tokens and $330 per million output tokens on the US profile, against $11 and $55.
Because the model is the same, Foxl does not give it a second GPT-6 Astra row in the picker. It adds a switch to the row you already chose.

Turning it on
- In Foxl Desktop, connect your AWS account under Settings > Model & provider, with Bedrock access to GPT-6 Astra.
- Open the model selector next to the message box and pick GPT-6 Astra in the AWS group.
- Turn on Ultrafast at the bottom of the selector. The lightning mark appears on the model button.
Four things it deliberately does not do:
- It is off by default, and it stays where you left it.
- It applies to the messages you send, and nothing else. Conversation titles, suggestions and scheduled runs reuse the same model but stay on Standard, so a background job never bills at six times the rate without you watching it. In every run below, the title calls that Foxl made alongside the chat went out without the tier.
- It is not available on Foxl credits. The Bedrock endpoint that Foxl credits use for GPT-6 Astra is regional Mantle in us-west-2, which the model card excludes (“Ultrafast does not support regional Mantle access in us-west-2”), and asked anyway it answers HTTP 400 with
Supported values are: 'auto', 'default', 'flex', and 'priority'. So the switch only appears where the request goes to your own account. - It does not hide the cost. Usage figures in Foxl count Ultrafast messages at six times the Standard rate, which is what AWS bills.
How we measured it
A race between two windows means something only if the tier is the only difference, so the setup was built to hold everything else equal and to record that it did.
- Two desktop servers from one build. In the desktop app the tier is a setting of the local server, so each window got its own server, with its own data directory, built from the same commit. One had Ultrafast on and one had it off, set through the same endpoint the switch calls, or in the filmed pairs by clicking the switch itself.
- One prompt, sent at once. The prompt above went into both message boxes and Enter was pressed in both together. In the seven filmed pairs, where both windows were timed on one clock, the two requests left between 0.4 and 15 milliseconds apart. Each window measured its own turn from the stream the app itself reads: the moment the request left the page, the first text, and the last token.
- The tier was checked on the wire. A 200 from Bedrock cannot tell “served on Ultrafast” from “accepted and ignored”, but the reply carries its own
service_tier. A tap on each server's outbound requests, after signing, recorded both sides. All 16 Ultrafast chat calls sent"ultrafast"and Bedrock echoed"ultrafast"on every one; all 16 Standard calls sent no tier and came back"default". Every call returned 200, and no turn reported an error. - Everything else held still. Region us-west-2 through
us.openai.gpt-6-astraonbedrock-runtime(the Responses API), reasoning effort Low on both, and the same input on both sides of every pair: 8,534 to 8,747 tokens, Foxl's system prompt and tools plus the prompt. Pairs 1 to 15 ran on 2026-10-01 between 06:44 and 07:07 UTC, and pair 16, the one in the video, at 08:09.
One disclosure about the answers. The servers were fresh installs, so in the first 14 pairs both replies opened with Foxl's one-line first-run greeting before the code. It was the same on both sides and a few dozen tokens out of about two thousand; for the last two, including the one in the video, first-run setup was completed beforehand and both replies start with the answer.
Results

Every pair went the same way. The spread is worth reading as well as the medians: Ultrafast's own runs ranged from 6.1 to 11.2 seconds, its slowest close to twice its fastest, where Standard stayed between 19.2 and 26.4. Even so, Ultrafast's slowest run finished 8.0 seconds before Standard's fastest. Per pair the advantage ran from 1.9x to 4.2x.

Splitting the turn shows where the time comes from. The wait for the first text shrinks by about a second, from 2.75 s to 1.61 s. The large difference is generation: 117 tokens a second on Standard against 381 on Ultrafast, a median of 3.3x per pair and up to 4.4x.
It also says what kind of turn gains the most. The 2,168-token answer here is mostly generation, and saved 13.8 seconds at the median, about two thirds of its length. By the same arithmetic a one-line answer is mostly wait and would save about a second; we did not measure short answers.
Up to six times, up to 300 tokens a second
The AWS announcement gives two figures: “According to OpenAI, Ultrafast delivers up to 6x faster inference in the API, with up to 300 tokens per second.” We measured both outside the app, so that no Foxl code could sit between the model and the clock. 12 more pairs went straight to the Bedrock API, two on each of six routes: bedrock-runtime through the US profile in us-west-2, us-east-1 and us-east-2, through the global profile in us-west-2 and ap-northeast-2, and bedrock-mantle in us-east-1. Same prompt, both tiers sent together, and every reply echoed the tier it was sent.

Ultrafast streamed above 300 tokens a second in 11 of 12 runs, at a median of 406 tokens a second and a peak of 436. The route made no difference to speak of.
The six-times figure did not appear, because Standard was faster than that figure implies. It ran between 112 and 116 tokens a second on every route, so the multiple came out at 2.2x to 3.9x per pair, median 3.6x. Six times faster at a top speed of 300 tokens a second implies a Standard baseline near 50, and no route we measured had a Standard that slow. Raising the reasoning effort to High did not move it either: two more pairs came out at 3.7x and 3.3x on throughput. So the result we measured on Bedrock today is past 300 tokens a second, and about three times sooner in a real app, rather than six times.
What it costs
Priced from the usage Bedrock returned, at the US profile's short-context rates, the video's turn (8,534 input tokens, none cached, and around 2,093 output tokens) comes to about $0.21 on Standard and about $1.25 on Ultrafast. Put the other way, Ultrafast bought 13.8 seconds on this answer for about $1.04 more.
Repeated turns came out cheaper, because the long prefix repeats. In 13 of the 16 pairs Bedrock reported 8,745 of the 8,747 input tokens as cached, on both tiers alike; at the cache-read rates ($1.10 and $6.60 per million) those turns come to about $0.13 and $0.75. Either way the output is most of the bill and the ratio stays six to one. The model card lists implicit caching for bedrock-mantle, not for the bedrock-runtime route these turns used, so treat the cached figures as what the replies reported rather than as a promise about your invoice. Above 272,000 input tokens both tiers move to their long-context prices, and the six-times relationship holds there too.
That makes it worth using on turns where you are waiting for the answer, such as reading code as it arrives or iterating on an answer with the model. For a long task you leave running, the wait costs nothing and Standard is the better buy, which is why Foxl never applies the tier on its own.
About the video
It is the last of the 16 pairs, recorded live in the light theme at twice the page's pixel density, the same density as a Retina display. Each window was captured from the browser's own paint timestamps and both were put on one clock, so a frame shows the same instant in both, and in the full-size video every window pixel is a captured pixel. The windows were driven by a script rather than by hand: the cursor is drawn where the script moved the pointer, the typing is the script's, and the switch click is a real click in the page that turned the tier on for that window. Nothing between send and the last token was trimmed or re-timed. The finish line also shows each window's generation speed, computed from the usage that came back with its reply.
The take was not chosen for its result. Earlier takes were set aside for problems with the setup (a cluttered sidebar, a capture format that dropped frames, the first-run greeting), and a finished dark-theme cut was replaced by this light, Retina one; all of their timings are among the 16 pairs. The rule for the final one was to use the first take with the final setup, whatever it showed. It came out at 3.1x, close to the 3.0x median of all 16 pairs. Choosing the fastest pair, at 4.2x, would have overstated the gain a typical turn gets.
Limits of this measurement
- One prompt and one kind of answer: about two thousand tokens of TypeScript. Shorter answers gain less, as above.
- One morning, 06:44 to 08:09 UTC, and the app pairs in one region (us-west-2, through the US profile). Capacity varies, and a launch-week tier may change.
- Speed only. Both tiers run the same model by AWS's description, and all 32 app answers contained a complete cache class and a Vitest suite, but we did not run the tests or grade the answers against each other.
- The app times include Foxl's own server and the stream to the window, which is what you experience, and is the same on both sides. The direct calls count text tokens only; the app figures count every output token.
Ultrafast is in Foxl Desktop from v0.7.32, for GPT-6 Astra on your own AWS account.
References and further reading
- AWS: GPT-6 Astra now supports UltraFast mode on Amazon BedrockReference
- GPT-6 Astra model card (Amazon Bedrock)Documentation
- Amazon Bedrock service tiersDocumentation
- Foxl v0.7.32 changelogRelease