Subagents

Independent work actually runs at the same time now.

delegate_task always described itself as parallel-friendly — call it more than once in a turn for independent research threads and they run at once. That wasn't true until now: multiple calls in the same turn ran one after another, on the same thread, with zero wall-clock benefit. They now run genuinely concurrently, each on its own worker thread, with the locking that makes that safe.

A pre-pass, not a rewritten dispatch loop

The tool's own description promised something the code didn't do

The problem

delegate_task told the model to split independent research into separate calls in the same turn for parallel execution. The actual dispatch loop ran every tool call strictly one after another on a single thread — two delegate_task calls got the exact same sequential wait as one call done twice, just with extra framing suggesting otherwise.

How Kryex solves it

Every delegate_task call in the current turn is now collected up front. When there are two or more, they run concurrently, each on its own worker thread, with results attributed back to the correct call once every thread finishes. Every other tool — edits, approvals, running a command — keeps its exact existing one-at-a-time order untouched; only delegate_task calls are pulled out and run together, and only when there are genuinely two or more of them.

Example

Asked to investigate three independent parts of a large codebase, the agent issues three delegate_task calls in one turn. All three start at once and run to completion in the time the slowest one takes — not the sum of all three.

The impact

A single delegate_task call — still the common case — is completely unaffected. Parallel delegation is now a real wall-clock win exactly when the model already believed it was one.

Before — one thread, one at a time
delegate_task #1
delegate_task #2
delegate_task #3
Three independent research threads, three sequential waits.
Now — genuinely concurrent
delegate_task #1
delegate_task #2
delegate_task #3
Same three threads, one wall-clock wait — each on its own worker thread.

Two roles, one tool set, different deliverables

A subagent is spawned with a role that shapes what its final answer looks like, not what it's capable of doing — both roles below share the identical, read-only tool set.

explorer

Investigates and reports back findings — files, patterns, how something currently works. Read-only tools: read_file, list_files, search_code, directory_tree.

planner

Same read-only tool set as explorer, but a different deliverable: a concrete, actionable plan — a one-line summary, ordered steps naming real files and what changes at each, and open risks for the lead agent to resolve. It cannot edit, run commands, delegate further, or ask the user anything.

Use planner when the deliverable isn't just facts but a concrete implementation strategy — and the investigation needed to produce it would otherwise burn a lot of the lead agent's own context on files it will never touch again.

A planner's plan doesn't stay buried in tool output

A returned plan is just text in the lead agent's context until it becomes a real, visible plan — so the lead is instructed to turn a planner's result directly into a submit_plan call as its very next action, and a mechanical backstop catches the exact case where that doesn't happen and forces one more turn to fix it, rather than letting a real plan silently render as ordinary chat text.

What running concurrently required making safe

Before this fix, every tool call ran one at a time, so nothing in a session's shared state was ever genuinely raced. Once multiple subagents can run on separate threads at once, three things needed real protection:

The event stream stays intact

Every update from every subagent, running or not, is serialized through one lock before it's written to the live conversation stream — two subagents finishing at the same instant can never corrupt each other's output.

The spawn cap can't be slipped past

Checking and incrementing how many subagents have been spawned this turn happens as one atomic step, so two subagents starting at the same moment can't both slip through under a cap meant to stop them.

Two pending approvals never collide

Only one approval can be waiting at a time. If two concurrently-running subagents both need human sign-off in the same turn, they take turns through this lock instead of one silently overwriting the other's pending request.

All three only ever come into play when two or more subagents are genuinely running at once — the common case of a single subagent never touches them.

Room to actually use it

A cap on how many subagents can spawn in one turn matters more once they can genuinely run in parallel — a conservative cap made sense when a higher number only meant more sequential waiting for no benefit. That cap has been raised from 3 to 5, matching the upper end of what published multi-agent research systems observe as the useful range for concurrent subagents on one query. The limit on how many tool calls a single subagent can make before finishing was separately raised as well, once a real, repo-scale investigation showed the old limit wasn't enough room to actually finish one.

The lead agent is actively steered toward this: independent research threads are split into separate delegate_task calls in the same turn specifically so the concurrency above actually gets used, and toward planner over explorer whenever the ask is for a plan or a change strategy, not just findings. See Skills and Memory for the other two pieces of how Kryex Code reasons across a large codebase without holding all of it in context at once.

CtrlI