Three tasks, six conversations
What one ambiguous handoff taught me about models, tools, prompt portability and the human gate
With contributions from OpenAI GPT-5.6 Sol (Codex), Anthropic Fable 5 (High effort, Claude Code) and Anthropic Opus 5 (Max, Claude Code).
At 9.35 this morning, I tried to start three pieces of archive research.
At 9.37, I had six conversations open.
The work was part of the Book Bench, the workspace I am using to turn a concentrated period of working with agentic AI, from the creation of this website through the work that followed, into a book. The Book Bench corpus contains 70 archived sessions. Across those sessions, its index counts 2,118 user messages, 20,868 assistant messages, 18,990 thinking blocks, 29,118 tool calls, 29,048 tool results and 161 attachments: more than 100,000 recorded items.
The wider evidence environment no longer ends with that original survey population. A separate post-cutoff DELTA addendum extends the preserved record through 11 August. Dates therefore belong to particular evidence windows; the archive should not be described as if everything stopped in July.
The Microsoft Excel master workbook is the selective integration layer. A structured workflow has so far systematically interrogated, extracted and integrated 860 evidence records from 25 archive sessions. It also holds 55 narrative timeline anchors, ten developing idea lineages and 40 provisional scene units. Each evidence row keeps its archive, event locator, time, source, model provenance, exact quotation or observation and a plain-English account of why it might matter.
On 12 August, I revisited the unexplored archive sessions. The purpose was not to extract everything indiscriminately. It was to identify the next representative sample worth exploring for the rich narrative detail that might genuinely help readers and audiences understand what developed, changed or failed. The survey produced ten deep-research priorities.
On this morning’s research board, Claude Fable 5 at High effort was handling one of those archived sessions, A049. Three separate GPT-5.6 Terra conversations, also at High reasoning, were intended to cover the other nine priorities in three groups of three.
This was not casual prompting. It was an attempt to trace how ideas developed across thousands of recorded actions without allowing a fluent summary to become accepted history merely because it sounded right.
The models and the tools
The distinction between a model and the tool around it is important here.
OpenAI’s GPT-5.6 Sol, running at High reasoning in Codex, helped prepare the research prompts. The three intended archive workers were OpenAI GPT-5.6 Terra at High reasoning in the Codex desktop environment. The parallel A049 worker was Anthropic’s Claude Fable 5 at High effort in Claude Code.
Model allocation was not based on task fit alone. Usage availability mattered too. OpenAI capacity was effectively fully available, while only about a tenth of Fable’s allowance remained before its reset; the contemporaneous instruction recorded 11 per cent. The practical question was therefore which model was best suited to each bounded task, and which available allowance should be spent before it expired.
The model generates language and decisions. The interface or agentic tool decides what actions phrases such as fresh task, delegate or workflow can trigger. A prompt does not meet a model in isolation. It meets a model inside an operating environment.
That is why naming only the model would hide part of what happened.
There was also an established operating pattern behind the handoff. For more than a month, I had regularly asked Codex to draft a prompt, reviewed and refined it, then pasted it into the conversation I had chosen. That pattern was visible within this conversation itself. Asking Codex to prepare a prompt did not normally mean asking Codex to operate the destination window.
Three working conversations became six
I opened three Codex conversations. Each had a different archive group and a separate output directory. Each received the same basic opening:
Use GPT-5.6 Terra at High reasoning in a fresh task with no subagents.
I expected each Terra conversation to do its assigned research directly.
Instead, each conversation became a launcher. Each launcher created another Terra conversation to do the work. Three intended research conversations became three parent conversations and three child workers.
The timings are unusually clear. The first parent was created at 09:35:01 BST and its child at 09:35:19. The second parent followed at 09:36:39 and its child at 09:37:19. The third parent began at 09:37:13 and its child at 09:37:36. Within two minutes and 35 seconds of the first conversation opening, all six existed.
The first pause instruction reached a worker at 09:39:47. The remaining affected work was paused or interrupted by 09:42:15. The execution incident itself lasted seven minutes and 14 seconds. Understanding it, correcting it and rebuilding confidence took much longer.
This was not six-desk research
It would be tempting to describe this as an overenthusiastic multi-agent system. That would make it sound more sophisticated and less useful than it was.
The three additional conversations were not independent research desks. Each was a worker created by a launcher shell. The parent did not provide a separate evidential judgement on the child’s work. It mainly launched or monitored it.
That distinction matters to Two desks, one gate, the method I have been developing while building this site and now while building the book.
A second desk is not another window.
It is not another copy of the same model working beneath the first. It is not a conversation count. It is a separate judgement with a defined role, a separate evidential responsibility and no authority to pass its own work through the final gate.
The archive work remained staged. Nothing from the interrupted conversations had entered the canonical evidence ledger, the master workbook or the manuscript. That boundary turned an operational failure into a contained incident rather than a corruption of the book’s record.
Why prompts move between my desks
My normal way of working is to ask one model to help prepare a prompt for another.
Sometimes the writer and recipient are in the same model family: Sol may help me brief Terra. Sometimes the handoff crosses model families and vendors: Codex may help me brief Claude, or Claude may help me brief Codex. I review and amend the text, choose the destination window and paste it myself.
That practice is part of the two-desk method, not an incidental convenience. A second model can expose assumptions in a brief before the implementing model acts on them. A cross-family handoff can also show whether instructions that appear ordinary to a person carry different operational meanings in different agentic tools.
The existing learning session A second opinion explains why disagreement between models is useful evidence rather than a contest. This incident adds another layer: different routing behaviour is itself contemporaneous evidence. It tells us something about the environment in which the work occurred, provided we record the model, effort, interface, prompt and outcome rather than generalising from one run.
Why the wording seemed ordinary
Claude Fable 5 had received materially similar wording that morning:
Use Claude Fable 5 at High effort in a fresh task with no subagents or workflows.
In the observed Fable run, the model worked directly inside the selected conversation. It began reading its governing files and did not visibly create another task, subagent or workflow.
In the observed Codex runs, in a fresh task became an operational instruction to create another conversation.
That does not establish a universal difference between Anthropic and OpenAI models. Nor does it tell us that either behaviour will occur every time. It establishes that materially similar natural-language handoffs produced different routing outcomes in these specific model-and-tool combinations.
It also explains the human side of the incident. I already had a working precedent for reading fresh task as a description of the conversation I had opened, not as an instruction to create another one.
Prompt portability cannot be assumed merely because the English looks portable.
Stop first, interpret second
We stopped all six conversations before deciding what their existence meant.
Then GPT-5.6 Sol and I reconstructed the topology from the native Codex records. Three conversations had been opened by me. Three had been created by Codex. Each child record linked back to its parent. The records also showed that the parent conversations had not independently performed the archive extraction.
That prevented a second mistake: treating every visible conversation as a separate writer and assuming the research itself had been duplicated six times.
The root cause was simpler.
The prompt had combined two audiences.
The model and effort selection, and the idea of using a fresh conversation, were instructions for me. The archive assignment was for the receiving model. By placing both inside one copy-and-paste block, the prompt turned my setup instruction into the receiving model’s execution instruction.
I pasted the text I had been given to paste. Codex acted on in a fresh task by creating one.
The words with no subagents did not cancel the positive instruction. They restricted one form of delegation while leaving the task-creation command intact.
The root cause analysis
We wrote a formal RCA while the native records were still available.
The primary failure was in prompt design. Model-selection and destination guidance intended for me had been placed inside the payload intended for Terra. The paste block was unsafe when copied exactly as presented.
Several things made the incident costlier:
- all three prompts were started before one had been tested;
- initial diagnosis looked at the wrong monitoring surface;
- the recovery briefly repeated the same uncertainty about who was authorised to create conversations;
no subagentslooked like a safety boundary but did not prohibit the action the preceding words requested;- the same working phrase had already behaved differently in another tool.
The impact was three unnecessary launcher conversations and almost an hour of interruption, diagnosis and recovery. Some staged archive work was substantially complete, some partial and some not begun. The six conversations are preserved together as an incident set. Anything they wrote will be inventoried and isolated before it is considered again.
The complete internal RCA retains the conversation identifiers, source paths and native links. A public-safe version now sits in the site’s gated Caught at the Desk record. It remains behind the desk login because the current desk is a working instrument, not yet the public-facing tool planned for a later rebuild. This public Dispatch retains only the archive-session label A049 because the later validation follows that same recorded session.
The two-file proposal was not mine
Codex’s first corrective suggestion was to split the operator instructions and the paste-only prompt into two files. It was a plausible technical fix: each audience would get its own file. But it did not fit how I work.
For a long assignment, the full governing prompt can already wait in the shared workspace. What I need at the handoff is one brief launcher I can read, amend and operate, with an unmistakable boundary between my instructions and the text I send. Splitting that launcher would add another choice, another opportunity to copy the wrong thing and another separation between the instructions I needed and the text I was reviewing.
I rejected the two-file design. It removed one ambiguity by creating a practical risk. The better correction kept one launcher, two audiences and a hard boundary.
One file, two audiences, one hard boundary
The correction keeps the whole brief launcher in one file, but gives each audience its own unmistakable region. The full governing prompt remains in the shared workspace for the model to read directly.
The upper region is labelled for me:
# KISH INSTRUCTIONS - DELETE THIS SECTION BEFORE SENDING
- Destination: an existing conversation selected by Kish
- Model: [model]
- Effort: [effort]
- Tool or interface: [where the model is running]
- Working root: [project]
- Permissions: [read/write boundary]
- Timing: test this prompt before starting the others
- Before sending: delete this entire section and both divider lines
Then comes the boundary:
================ DELETE EVERYTHING ABOVE THIS LINE ================
================ ACTUAL PROMPT STARTS BELOW THIS LINE ==============
The receiving model sees only what follows:
EXECUTE IN THIS CURRENT CONVERSATION ONLY - DO NOT CREATE ANOTHER TASK.
Do not create, fork, delegate, launch or message another task, thread,
conversation, subagent or workflow.
[one-sentence assignment summary]
Work in [project root].
Read [full governing-prompt path] completely, adopt it as governing and
execute only that assignment.
[critical source/write boundary and stop condition]
The upper section may tell me to open a new conversation. The lower section must never tell the receiving model to do so.
The method complements How do I write a good first prompt? and The prompting primer. Those sessions teach the brief: goal, reuse, scope, preserve, verify and output. The advanced session How do I hand a prompt to another AI? teaches the handoff itself: which words belong to the operator, which belong to the model and what the surrounding tool may interpret as an action.
The method also reflects current OpenAI guidance to state instructions once and define autonomy and approval boundaries clearly. That guidance supports the correction. It does not prove why these six conversations behaved as they did. The native records do that.
The workflow around the prompt
The prompt is only the handoff. The working method has six safeguards:
- Prepare and operate. One launcher names the model, effort and tool, then separates my instructions from the prompt I will paste. I choose the destination and send it myself. Asking a desk to prepare a prompt does not authorise that desk to operate the destination.
- Preflight. The first conversation starts alone. It should read the source or begin the work directly. If it proposes another task, the rollout stops.
- Record. Keep the model, effort, interface, conversation, time, prompt and routing outcome. Without them, comparison becomes anecdote.
- Separate. Each concurrent assignment has one writer and one output directory. More conversations do not mean more independent desks.
- Review. The first desk’s packet remains staged while a separate model checks its evidence and conclusions against the source.
- Gate. I decide what may enter the canonical evidence ledger, workbook, manuscript or public site. Neither desk promotes its own work.
That is the practical version of Two desks, one accountable editor. Separation exists to create challenge and preserve responsibility, not to decorate the workflow with more windows.
The rule was used before the morning ended
A rule written after a failure is not yet evidence that the rule works.
The first saved A049 output from Claude Fable 5 appeared at 09:38:41 BST, and its required checkpoint was saved at 09:40:56. The native first-message time is not exposed in the preserved Book Bench artefacts, so I will not reconstruct it from memory.
At 10:10, the one-file handoff became the workspace standard. At 10:16:16, the corrected continuation prompt was saved. I deleted the Kish Instructions region and pasted the actual prompt into the existing Fable conversation.
Fable resumed in the same recorded research session. At 10:23:24, its completion report was saved. It had completed the staged A049 packet, mechanically validated 16 research rows and 15 quotations, changed nothing outside its assignment directory, used no subagents or workflows and stopped at the human-review gate.
That is one successful reuse, not proof of universal reliability. But it means the correction moved beyond a proposal.
The morning contained the whole sequence:
accidental discovery -> recognition -> formalisation -> rule -> repeated use
The Book Bench became part of its own archive
The Book Bench exists to revisit a concentrated period of working with agentic AI and turn the preserved record into learning for a book.
One of the structures emerging from that work is a way of tracing ideas. Not every idea appears fully formed. Some begin as an accidental discovery. They are recognised, formalised into a method, turned into a rule and eventually repeated until they become practice.
This morning, while trying to recover those sequences from the archive, we generated a new one.
That is the part I want to preserve.
The lesson is not simply that a prompt contained the wrong phrase. It is that an operational mistake becomes useful only when the record survives long enough to distinguish what happened from what everyone thought had happened.
Three tasks became six conversations. The gate kept six conversations from becoming accepted evidence. The RCA turned confusion into a cause. A human correction turned the first proposed solution into a usable handoff. The next prompt turned that handoff into a rule. One successful reuse turned the rule into the beginning of a method.
The second desk is not another window. The gate is not a final click.
They are how the work remembers what it is allowed to claim.