Serving
The fisk serve command hosts an agent behind the endpoints its configuration enables. It runs until interrupted. The
agent is the one fisk run drives: tool set, prompt, model and harness settings come from the same
configuration file.
Queued jobs takes each job off a work queue, runs it, and stores the answer for the submitter to read later. Answering prompts takes a prompt from another agent and streams the run back while that agent waits. Serving tools runs one tool for another agent and starts no agent loop.
Note
At least one endpoint must be enabled. fisk serve exits with an error when the configuration enables none.
Starting a worker
A minimal configuration has the application path, the tool selection and one channel. Here that channel is queued jobs:
The startup banner lists the endpoints it started and the settings every run uses:
The banner adds an Agent Context line when the agent’s nats_context differs from the queue’s.
Each endpoint prints its own section below. It shows the addresses that endpoint answers on and the limits it uses. Answering prompts and serving tools show examples.
Shared resources
fisk serve builds the model provider, the session store, the memory store, the knowledge index and the NATS
connection once at startup. Every run shares them. A missing stream or bucket fails the process immediately instead of
failing whatever job happens to arrive first.
Warning
A worker whose storage does not exist fails at startup. Under a supervisor that restarts on failure it crash-loops.
When the knowledge index does not exist at startup, each run opens the index for itself, so an index built after the worker started is visible to later runs.
Concurrency
Each channel limits its own runs, so a process serving two channels at two runs each is running four.
--workers sets how many queued jobs run at once, overriding expose.agent.jobs.workers:
--workers affects the queued-jobs channel only. The prompts channel takes its count from
expose.agent.a2a.prompts.workers and refuses a caller when every slot is busy.
A work queue has a concurrency setting of its own that limits every worker on it together. Setting workers above what
the queue allows leaves slots idle rather than raising throughput.
Timeouts
harness.tool_timeout limits a single tool call, in fisk serve and fisk run alike. The default is five minutes.
0s removes the limit, for commands that run for hours.
Note
--workers overrides the configuration file. harness.tool_timeout in the file overrides the built-in default.
The timeout stops a command and its process group. It does not stop an in-process handler that ignores its context.
Where tools run
Command tools run in the worker’s own working directory unless --work-dir names another. It must be an absolute path
that already exists.
Every run shares it. Set the worker count to 1 when a tool writes local state that concurrent runs would corrupt.
Note
A CLI that reads a context or a profile of its own does not inherit the agent’s nats_context, so pass the selection
explicitly where it matters.
Shutdown
On the first interrupt the worker drains. It takes no new work, and runs in flight continue to their next resumable point:
A second interrupt stops the worker at once. The queue redelivers any queued job still running, and the redelivery resumes from the journal. The prompts channel answers its callers with a failure instead.
A drain stops every endpoint, so a worker also serving tools stops answering peers at the same point. A worker with no channel has nothing to resume:
Sessions
Each channel journals its runs, so an interrupted run resumes instead of making the same model calls a second time.
Sessions need a store every worker can read. On one machine the default file backend is enough. Across machines,
configure a shared harness.sessions backend. Without one, a job redelivered to a different worker cannot read the
journal and starts again.
When two workers reach the same journal, the second one claims it. The claim is written before the run starts, and the first worker sees it before its next tool call and stops. Only a tool already running can execute twice.
Settings a channel run ignores
The following settings narrow the MCP and a2a tool endpoints, not a run served over a channel:
expose.agent.toolsselects what is served over MCP and a2a. A channel runs the whole agent loop, so it uses the agent’s ownincludeandexcludeinstead- the waiver that lets a tool-serving configuration omit
identity,system_promptandllm.modeldoes not apply to a channel, since a run needs all three
Safety
A served run is a full agent loop driven by caller-supplied prompt text, running every tool the configuration allows. The channel’s own admission check is therefore the only access control: queue publish permission for queued jobs, NATS publish permission for prompts.
harness.tool_timeout limits each tool call and llm.budget limits the run. A caller may lower the budget, never
raise it. The tool safety rules described in the Reference hold here as everywhere else:
commands run as an argument vector rather than through a shell, each argument is checked against the command’s schema,
and credentials are stripped from tool environments.