RUNTIME / AGENT LOOP
A turn, from prompt to finish.
NAH keeps the coding-agent cycle in ordinary code you can inspect: prepare context, stream one model response, execute its tool calls, record the result, then decide whether to continue.
Two layers work together
The core package supplies runAgent, the prompt builder, coding tools, and session primitives. It accepts an AI SDK v5 language model and does not pick a provider or own a terminal UI.
The CLI assembles the application around that core: it resolves credentials and a model, discovers project instruction files, adds task tracking, memory and optional delegation, selects local or sandboxed tools, applies a permission policy, and persists the conversation.
Build on the core SDK →What happens during a turn
- Prepare the request. The prompt, prior session messages, the system prompt, discovered
AGENTS.md/CLAUDE.mdcontext, and — when memory is enabled — the memory index for this turn. - Start a step. One AI SDK streaming call with the current transcript and tool map. One step is one model round-trip, including any tool calls in that response.
- Stream and run tools. Text deltas and tool calls/results become typed events as they arrive. Tools enforce schemas, output caps, workspace boundaries, and read-before-write for existing files.
- Record usage and check stop conditions. The assistant response joins the transcript. The loop stops if there were no tool calls, a budget was reached, the response hit its output limit, or cancellation occurred.
- Continue with compacted context when needed. Older messages are summarized or truncated before the next step when the conversation itself gets large.
- Persist. The transcript is appended to the session file and, when memory is on, the turn is extracted and the snapshot saved.
prompt + prior messages + system + tools + memory index
→ streamed model step
→ text / tool calls / tool results
→ append response and account usage
├─ no tool calls or a stop condition → finish
└─ tools remain and budget allows → compact → next stepTwo different budgets
maxTokens is a spend budget and accumulates across the whole run, so it stops a run that is over-consuming even when the context is small. It defaults to 400,000; set it to 0 to disable the cap. The loop also defaults to 32 steps and 8,192 generated tokens per response.
compactAtTokens is a context budget, and the two are deliberately different. Because every step re-sends the whole transcript, cumulative usage roughly multiplies the real context size — triggering compaction off it compacted healthy runs and cost the agent its working memory mid-task. Compaction instead compares the size of the most recent request, taken from the provider's own reported input count, against the threshold. It runs only after a step that requested tools, and only when there is a middle section to summarize.
Steering a running turn
A run is steerable from the host or the CLI. A message sent mid-run is appended to the transcript at the next step boundary — after the current step's tools settle, before the next model request — so it is never injected into a request that is already streaming.
const run = runAgent(options);
run.steer("actually use TypeScript, not JavaScript");
run.followUp("then update the changelog"); // only if the run would otherwise finish
run.interrupt(); // abort now
run.pending(); // { steer: [], followUp: [] }
Steers jump ahead of follow-ups, and each keeps its order. A steer never aborts the in-flight model call — interrupt() is the separate action for that. Two details matter for reliability: a pending message prevents the run from ending (otherwise a message typed while the model was writing its final answer would be silently discarded), and a delivered message grants a fresh step window so maxSteps cannot drop something a human deliberately sent.
Steers reach the model as ordinary user messages, indistinguishable from any other turn. Nothing marks them.
Stopping a turn that will not stop
Tool execution receives the run's abort signal, so interrupting reaches a running command instead of merely looking like it did. A shell command that reads stdin gets EOF immediately rather than blocking until its timeout, and a timed-out command is killed as a process group so no grandchildren survive.
Ctrl-C stop the running turn and say so
Ctrl-C press again to exit the TUI
Ctrl-D exit on an empty promptTools and permissions
The standard tools are read, list, grep, glob, edit, write, and bash. Read-only tools are never approval-gated. Edit, write, and bash pass through the configured policy; the SDK allows them if no callback is supplied.
In an interactive session NAH asks before mutations. Answer y once, a to allow that whole tool for the session, or A for only that exact call. Scoping “always” to the exact command string meant a fresh prompt for every slightly different invocation, which read as the choice not being remembered.
readonly denies mutations and yolo allows them. Print and JSON modes default to yolo because they cannot pause for a prompt, so pass --permissions readonly for a read-only one-shot run.
Observe and use the result
runAgent returns an async event stream and a result promise. The stream reports run-start, step boundaries, text deltas, tool activity, mid-turn user messages, compaction, and finish/error events. The result contains final text, stop reason, usage, transcript, step count and compaction count.
const run = runAgent(options);
for await (const event of run.events) {
renderProgress(event);
}
const result = await run.result;
console.log(result.reason, result.steps, result.usage);The CLI adds session controls around a turn: continue or branch saved JSONL history, inspect usage and the task plan, view diffs, and undo supported workspace changes.
CLI modes, sessions, and the TUI →