One Request, One Answer: Fixing Subagent Delegation
In 30 days of one real Foxl install, 51 of the 71 delegations that ended in a synthesis were followed by two long replies: one from the turn that delegated, one from the subagents' results. Measured on real models before and after this release: GPT-6 Astra stops doing the delegated work itself and stops repeating its answer (5 of 12 delegations to none) at half the cost per delegated request, forked subagents stop refusing their own task as a prompt injection, and a subagent that reports without using a tool is asked once to do the work. The prompt cache hit rate stays where it was.

On this page
A request for a roundup of recent AI news went to GPT-6 Astra in Foxl Desktop. The agent handed the research to a subagent, which is what Foxl's instructions ask for broad research. Then it carried on researching the same question itself and wrote a full answer, and a second full answer, with two extra items, appeared under the first. The person got two answers to one question, and the research behind them was paid for twice.
That chat was not unusual. On the install it came from, 72% of the delegations in the last 30 days that ended in a synthesis were followed by two long replies. This post covers why that happens, what changed in this release, and how we measured the change. The same measurements turned up two problems inside the subagents themselves: some refused their own task as a prompt injection, and some reported results they had never computed.
Highlights
| Measured on 2026-10-02 | Before | This release |
|---|---|---|
| GPT-6 Astra delegations where the chat repeated its answer | 5 of 12 | 0 of 12 |
| GPT-6 Astra tool calls after delegating, mean | 9.1 | 0.1 |
| Cost of a delegated GPT-6 Astra request, mean | $1.04 | $0.52 |
| When GPT-6 Astra delegates, against Foxl's policy | 12 of 12 it should, 0 of 16 it should not | unchanged |
| Forked subagents that refused their task as an injection (Claude Haiku 4.5, old hand-off, 97-second task) | 5 of 20 | 0 of 20 |
| Subagents that reported a result without calling a tool (Haiku 4.5) | 1 of 30 | 0 of 30 |
| Execution benchmark: runs with every real result in one answer | 29 of 30 | 30 of 30 |
| Execution benchmark: median time to the final answer | 23.7 s | 17.4 s |
| Execution benchmark: prompt cache hit rate | 95.5% | 95.7% |
How delegation works in Foxl
The agent you talk to in Foxl Desktop, which we will call the parent, can start a subagent with the sessions_spawn tool. A subagent is a separate agent loop with its own context and toolset, and it runs while the parent's turn ends and the chat stays usable. Up to five run at once. Some are named definitions with a clean context, such as the researcher. Others are forks: they start as a copy of the conversation, under the parent's own system prompt, so the model provider can serve that long shared prefix from its prompt cache instead of processing it again.
When every subagent in a batch has finished, Foxl puts their reports into one message and starts a new turn with it, the synthesis. Its reply appears under a divider in the chat, and it is meant to be the answer to the request. So a delegated request has two turns that can answer it: the turn that delegated, and the synthesis. The design expects only the synthesis to answer.
How often it goes wrong
To see what delegation looks like in daily use, we read the database of one real install, in aggregate and read-only. Over 30 days it holds 2,405 assistant turns in conversations a person had. In 107 of them, 4.45%, the parent delegated.

Most of it comes from one model: 76 of 107 delegations were GPT-6 Astra's, 8.9% of its turns, and another 22 are in turns with no recorded model. Claude Opus 5.5, Foxl's default, delegated in 0.3% of its turns (1 of 397). Every model saw the same instructions, so whether to delegate is mostly the model's own choice. The controlled study below found the same split.
Of the 71 delegations in those 30 days that were followed by a synthesis, 51 were followed by two long replies in the chat: the turn that delegated had written a reply of at least 400 characters, and the synthesis wrote another. The synthesis started a median of about 7 minutes after the delegating turn began.
One request on the time axis

The desktop log of the reported chat shows the sequence. The subagent finished its research in 81 seconds. The parent, meanwhile, kept researching the same question for 7.2 minutes, 58 model calls in all, and then wrote its answer. When the synthesis started, it had the subagent's report and nothing telling it that the chat already held an answer, so it wrote the whole answer again from a prompt of 149,857 tokens. None of those tokens came from the prompt cache, because the synthesis started while the parent's turn was still unwinding and had to rebuild the agent from the database; a separate fix in the same release makes the synthesis wait for that turn to finish.
The instructions already said “Don't duplicate work you delegated.” They did not say what to do instead, and to a model that has just started a subagent and still has the user's question in front of it, continuing to work looks like the helpful move. In the controlled runs, GPT-6 Astra kept working after delegating in 8 of 12 delegations, with 9.1 tool calls on average.
What changed in the parent
There are two changes.
- The parent is told how to end the turn. The rule now reads: “After you delegate, do not do the delegated work yourself and do not write the final answer in that turn: work only on parts you did not delegate, then end the turn with one short sentence saying what is running. The answer is written once, when the results arrive.”
- The synthesis knows when an answer exists. If the turn that delegated still wrote a full answer, 400 characters or more, the synthesis gets a different instruction: “Your previous reply above already answered the user, so do not write that answer again.” It is asked for only what the results add or change, or one sentence saying they change nothing.
On the new tree GPT-6 Astra never answered before its subagents returned, so the second change did not come into play for it. It fired once in the controlled runs, for Claude Haiku 4.5, whose synthesis added 435 characters under a 6,196-character answer instead of writing a second copy. That answer was wrong, and the addendum did not fix it.
How we measured it
Every number from here on comes from a real Foxl desktop server running a real model on Amazon Bedrock, on main at the point where this work branched and on this release's tree. Two intermediate versions of the branch were measured as well, and they are named where they appear. Nothing was simulated, and no model output was replayed.
- Isolated servers. Each harness run got its own server and data directory, one server at a time, and each trial was a new conversation, up to three at once. The delegation study also gave every run a seeded workspace, a scratch home directory, so a model that guessed a path could not reach a real one, and a completed first-run setup, so no reply opened with the onboarding greeting.
- Results that cannot be guessed. In the execution benchmark every subagent's task is a shell command that hashes a random word. A model cannot know that value without running the command, so a correct final answer proves a real tool result travelled from the subagent into the chat.
- Requests a person would write. The delegation study used 14 natural requests, written in Korean, on a seeded workspace with a checkable fact behind each. Six are tasks Foxl's own policy says to delegate (multi-file work, review, broad research), and eight are tasks it says to do directly (a single command, a small question).
- Three models. GPT-6 Astra, the model in the reported chat; Claude Opus 5.5, Foxl's default; and Claude Haiku 4.5, which is cheap enough to run 30 times per cell. All ran on the Standard tier.
- Usage from the app's own records. Tokens, cache reads and cost are what Foxl recorded for each call, with a subagent's calls billed to the chat that asked for it. Foxl prices GPT-6 Astra at its global-profile rate while these calls went through the US profile, so the GPT-6 Astra costs below read about 10% low; the before-and-after ratio is unaffected.
Results: deciding to delegate, and what happens after
| 14 requests: GPT-6 Astra and Haiku 4.5 ran each twice, Opus 5.5 once | GPT-6 Astra, before | GPT-6 Astra, now | Haiku 4.5, before | Haiku 4.5, now | Opus 5.5, before and now |
|---|---|---|---|---|---|
| Delegated when the policy says to | 12 of 12 | 12 of 12 | 2 of 12 | 5 of 12 | 0 of 6 |
| Delegated when it says not to | 0 of 16 | 0 of 16 | 0 of 16 | 0 of 16 | 0 of 8 |
| Answers correct | 27 of 28 | 28 of 28 | 25 of 28 | 26 of 28 | 14 of 14, 14 of 14 |
| Repeated the answer, of delegations | 5 of 12 | 0 of 12 | 0 of 2 | 0 of 5 | none delegated |
| Parent kept working after delegating | 8 of 12 | 1 of 12 | 0 of 2 | 1 of 5 | |
| Delegated request: median time, mean cost | 46.7 s, $1.04 | 41.7 s, $0.52 | 36.9 s, $0.09 | 29.4 s, $0.12 | |
| Direct request: median time, mean cost | 2.9 s, $0.090 | 2.6 s, $0.061 | 3.8 s, $0.023 | 3.5 s, $0.010 | 7.4 s and 6.4 s, $0.038 |
GPT-6 Astra follows the policy exactly on both trees: it delegated every task the policy marks for delegation and none of the others, before and after. The difference is what it did after delegating. Before, it kept working in 8 of 12 delegations and repeated its answer in 5. Now it waited in 11 of 12, never repeated, and was right in all 12. A delegated request costs about half as much, because the research is done once.
The other two models show how much of this is the model. Claude Haiku 4.5 delegated fewer than half of the tasks the policy marks, a few more on the new tree. Claude Opus 5.5 delegated none of them on either tree, and still answered all 14 requests correctly, at a median of 6 to 7 seconds and about 4 cents each. On tasks of this size, the policy's list of work to delegate is broader than Claude Opus 5.5 needs.
The subagents had problems of their own
The same runs found two failures inside the subagents. Both were measured with Claude Haiku 4.5, which is cheap enough to run 20 to 30 times per version.
A task that looks like a prompt injection
A fork runs under the parent's system prompt, unchanged, because that prompt is the prefix the cache can reuse, and that prompt describes the main agent. Its own task then arrives as the last message in the user's position, and on main that message opened by telling the model to stop and that it was not the main agent. That is the pattern of a prompt injection: text in the user turn trying to change the model's role. A model trained to resist that will sometimes refuse, and on main at least 5 of 20 forks on a short task said so and declined to do it.
Longer tasks fared no better: with the same opening, 5 of 20 refused a task that runs a 97-second command. The fix explains the hand-off in the system prompt, the place the model already trusts. A short paragraph now says that some subagents run as a copy of the conversation, that Foxl itself adds a last message beginning with a marked block, and that the block is written by Foxl, not by the user, a page or a file. The block itself no longer shouts, and says in plain words what it is. The parent and every fork share the system prompt, so the paragraph sits inside the cached prefix for both, and the fork's prefix stays byte for byte identical to the parent's. With both changes, 0 of 20 refused the same task, and 0 of 20 refused tasks of 10 to 14 seconds.
Reporting work that was never done

This one took four rounds of 30 runs to understand. On main, 1 of 30 subagents reported a result without running anything. About half of them first tried to start subagents of their own, were refused with an instruction to do the task directly, and then did it, so the refusal seems to have worked as a reminder. The first version of this release stopped those pointless attempts, and invented reports went to 3 of 30, each with a value that could not be right.
A hand-off that stressed that the work itself was not yet done produced 6 of 30, no better. The change we kept is a check in the agent loop: when a subagent finishes without having called a single tool, it is asked once to do the task with its tools and report only what they return, or to repeat its report if the task genuinely needs no tool. That brought it to 0 of 30, with all 30 runs right. Against main, 1 of 30 to 0 is within noise; what the check adds is that a report built on no tool result is challenged once before it counts. It runs at most once per subagent and never for one that used a tool, so it costs nothing when the work was done.
Speed, cost and the prompt cache
With both subagent fixes in place we ran an execution benchmark: five workloads, six runs each, on Claude Haiku 4.5. One subagent; three in parallel; a natural request that needs two; the person asking something else while two run; and a subagent that runs for a while. Every subagent computes a value that cannot be guessed, and a run counts as a success only if the final answer contains every real value and is written once.
| 30 runs each | Before | This release |
|---|---|---|
| Runs with every real result, in one answer | 29 of 30 | 30 of 30 |
| Subagents that reported the right value | 53 of 54 | 54 of 54 |
| Time to the final answer, median and 90th percentile | 23.7 s, 35.6 s | 17.4 s, 22.2 s |
| Subagent run time, median | 17.2 s | 8.7 s |
| Subagents' attempts to start or stop other subagents | 105 | 9 |
| Model calls, all runs | 320 | 267 |
| Prompt and output tokens per run | 146,480, 1,762 | 124,944, 1,160 |
| Prompt cache hit rate, all calls and subagents only | 95.5%, 96.9% | 95.7%, 96.4% |
| Cost per run | $0.0310 | $0.0245 (21% less) |
| Answer in the language of the request (Korean) | 16 of 30 | 29 of 30 |
Most of the time saved is inside the subagents. On main, 36 of 54 of them spent their first steps trying to start subagents of their own, 105 attempts in all, each refused before the real work began. Removing that halved the median subagent run and cut the slow tail of the whole request from 35.6 to 22.2 seconds. Fewer calls and shorter outputs account for the lower cost.
The prompt cache hit rate did not move, which this release was designed to keep. A fork is only cheap because its prefix is identical to the parent's, so anything this release added for subagents had to be constant text that both sides carry. A probe that compares the fork's request prefix with the parent's, byte for byte, passes on the new tree, and the hit rate stayed at 95.7% overall and 96.4% for subagents.
The synthesis message is the turn's only input, and Foxl writes it in English. On main it named no language, so about half the answers to Korean requests came back in English. It now asks for a reply in the language the user has been writing in.
Two more defects the runs found
- A message sent during a synthesis cancelled it. If the person wrote something just as the subagents finished, the new turn stopped the synthesis, so the results never reached the chat; with this release's divider, it would also have left a divider with no answer under it. The new turn now waits for a running synthesis to finish, then runs.
- A resumed subagent could get more tools than it started with. This one was in the release's own new resume feature and was caught before release: a code-reviewer started with its 4 tools came back with 25 when it was resumed. A resume now keeps the original run's definition.
The work also went through a matrix of 100 failure cases, from a chat deleted while its subagents run to a message sent as the last one finishes, and three adversarial reviews. The fixes from those are listed in the v0.8.0 changelog.
What else changed for subagents
- A bar on top of the message box shows the subagents running for this chat, what each is doing and for how long, and can stop one or all of them. Their tool calls no longer appear in the chat itself.
- You can send a message to a subagent from its own screen. A running one reads it right after the step it is on; a finished or stopped one picks up where it left off, with everything it already did, and its new answer comes back to the chat.
- When you stop some subagents and others finish, the synthesis says which ones were stopped instead of promising their results.
- The divider above a synthesis says how many subagents it covers.
What we did not change
- The delegation policy. It is broader than Claude Opus 5.5 needs on small tasks, and Opus mostly ignores it to good effect. Narrowing it would change GPT-6 Astra's behaviour too, which now follows it well, so that is a product decision, and the numbers above are the input to it.
- How often Claude Haiku 4.5 delegates. It went from 2 of 12 to 5 of 12, which is within the noise of 12 runs, and when it did delegate on the new tree it was right in 3 of 5. One of the two misses is the run above where the parent answered first.
Limits of this measurement
- The 30-day figures come from one install, used heavily by one person. They show the failure is real and frequent there, not how often it happens across all users.
- In the 30-day data, “two long replies” means only that the delegating turn and the synthesis both wrote at least 400 characters; those cases were not checked for overlap. The controlled runs use a stricter test, a second answer at least half as long as a first that was already written.
- Model output varies from run to run. Most cells are 12 to 30 runs and the Opus ones 6 to 14, so a difference of one or two is noise. The changes this post leans on, GPT-6 Astra's repeats (5 of 12 to 0) and the refusals (5 of 20 to 0), are larger than that; the invented results against
mainand Haiku's delegation rate are not. - The seeded tasks are small. On a real research request the saving from not doing the work twice grows with the work: the reported chat spent 7.2 minutes duplicating a subagent that took 81 seconds.
- Claude Haiku 4.5 is not the default model. It was used for the subagent measurements because it is cheap enough to run 30 times per cell.
These changes are in Foxl Desktop v0.8.0.
References and further reading
- Sub-agents (Foxl docs)Documentation
- Foxl changelogRelease