Manager Socket Protocol#
The manager exposes a long-lived WebSocket for driving executions. It replaces
the HTTP run flow (POST /genvm/run + poll + DELETE), which remains
available for one release train as a deprecated adapter over the same core
(see Manager API). The admin HTTP endpoints are unaffected.
Clients connect by upgrading GET /ws on the manager’s HTTP listener, so the
protocol is reachable wherever that listener is: over the unix socket when the
manager runs with --socket, over TCP when it runs with --host/--port.
There is no separate address to configure.
No authentication: the manager assumes a trusted deployment (unix socket or loopback); do not expose this listener beyond the machine boundary.
Framing#
Every protocol message is one binary WebSocket message; WebSocket already delimits messages, so there is no length prefix. Header integers are big-endian. Payloads are calldata-encoded (see Calldata Encoding).
message := method_id:u16_be request_id:u64_be payload
payload := calldata(value)
A message is therefore payload length + 10 bytes. Note the byte order
differs from the executor host protocol, which is little-endian
(Host Loop Pseudocode).
request_id != 0: a request. Exactly one reply message follows, with the samemethod_idandrequest_idechoed (or anerrormessage with therequest_idechoed).request_id == 0: a notification; no reply. Only the manager sends notifications. A client message withrequest_id == 0is answered with anerror(bad_request_id,request_id == 0).
Request ids are chosen by the client and scoped to the connection; the manager never initiates requests, so there is no id-space split.
Unknown method ids, undecodable payloads, messages shorter than the 10-byte
header, and text messages each produce an error message without closing the
connection.
A message larger than the manager’s max_message_bytes is refused by the
WebSocket layer before the manager sees it. The cap is enforced while reading,
and reports as a read failure rather than a close handshake, so the connection
is dropped without a close frame – a client observes an abnormal closure
(1006), not 1009, and no error payload is sent. Clients that page artifacts
within the documented chunk cap never approach the cap.
Requests and notifications are externally-tagged: a single-key map whose key
names the variant. Replies, including errors, are untagged maps.
The method_id header is authoritative for routing
Method ids#
Generated from crates/modules-interfaces/codegen/data/manager-api.json
(Rust, Python and this page share one source).
The generated constants are rendered in Constants; see methods for the method id enum. Directions and request kinds are described in each method section below.
Connection hello#
Immediately after accepting a connection the manager sends a hello
notification:
{ "hello": { "boot_id": u64, "protocol_major": u32 } }
boot_id is random, generated once per manager process start. GenVM ids
restart from 1 with the process, so the durable identity of a run is the pair
(boot_id, genvm_id). Clients MUST remember the boot_id and pass it to
attach; a mismatch after a manager restart is surfaced as
boot_id_mismatch instead of silently binding to an unrelated, reused id.
run#
Request payload: the same logical structure as the deprecated
POST /genvm/run body (GenvmRunRequest in Manager API), plus:
host_genvm_id(string, optional) – client correlation token, echoed in events for this run. Also an idempotency key: arunrepeated with the same token before the retention TTL expires returns the id already allocated for it instead of starting a second execution. A run that ended withfailed_to_startholds its token too, soackit before retrying with the same token.host_hello_data(array of bytes, optional, default[]) – indexed by host connection index; the executor writes entry i verbatim to host i on connect, before the first method byte (see Host Loop Pseudocode). The manager rejects a non-empty entry for a host connection it owns itself (currently index 1, theconsume_resultsocketpair).hook_cross_contract_calls(bool, optional, defaultfalse) – whether this host wants to be asked where aCallContractruns. When false the manager answersresolve_call_contract_executoritself with a null reply, so every call stays in-process and the host need not implement that method. When true the question is routed to host 0 and the host may send the caller across a major boundary (see Host Loop Pseudocode). A nested run inherits the value from its parent.deadline(duration string, optional) – when set, overridesmax_execution_minutesas the strict deadline. The manager enforces it and pushes the terminal event; clients need no timeout timer of their own. Either way the deadline is capped at 24 hours: a longer one is silently shortened, not rejected. See Duration strings.unsafe_overrides(map, optional) – overrides that reach boundaries production traffic cannot. Each member states thedebug_modeit needs; with debugging disabled none of them apply, so consensus traffic always runs the manifest-resolved version with the executor’s own limits.reroute_to(string, optional) – run this version instead of the oneselectorresolves to. A plain string is an executor directory used as it stands; are:prefix makes the rest a regular expression matched against manifest version keys, and the newest match wins. Honored fromsafe.initial_recursion(u32, optional) – seeds the chain’s recursion budget, replacing the executor’s ownVM_RECURSION, so a boundary test need not spend one executor process per unit of budget. Honored fromunsafe.allow_two_workers(bool, optional) – overrides the manager config of the same name fromunsafe. The config defaults to true. In v0.3+, false makes deterministic execution await each submitted nondeterministic validation task before continuing; null keeps the configured value. The manager passes this in the execution input. v0.2 warns when false and keeps its existing scheduling. Nondeterministic calls remain allowed.
Response:
{ "genvm_id": u64 }
The id is allocated and returned immediately; validation, permit acquisition
and process spawn continue asynchronously and report through event
notifications. The requesting connection is subscribed to the run’s events
automatically.
A payload that fails to decode is answered with malformed_frame. An
idempotent retry can receive unknown_id if the retained run is removed
concurrently before subscription. A decoded request that fails a check, such
as a non-empty host_hello_data[1] or needing modules while they are stopped,
ends with failed_to_start
attach#
{ "attach": { "boot_id": u64, "genvm_id": u64 } }
Subscribes the connection to the run’s events and returns a snapshot – the
most recent lifecycle event for the run, in the same shape as an event
payload:
{ "snapshot": <event payload> }
Errors: boot_id_mismatch if boot_id is not the current process’s;
unknown_id if the run does not exist, was acked, or its retention TTL
expired. Disconnecting drops all of a connection’s subscriptions; it does not
affect the run (see below).
cancel#
{ "cancel": { "genvm_id": u64 } }
Requests termination. Response is an empty map; the outcome is reported through
event under the Lifecycle guarantees. Cancelling a run still queued on
permits aborts it before spawn (no permit is consumed) and ends in finished
with cause cancelled
ack#
{ "ack": { "genvm_id": u64 } }
Releases the retained result and state for a finished run. Response is an
empty map. An ack before a terminal event or result is refused with
not_finished and leaves the run fully usable. After ack (or after the
retention TTL expires), attach and get_artifact answer unknown_id.
Reads before ack are non-destructive and repeatable from any number of
connections.
get_artifact#
{ "get_artifact": { "genvm_id": u64, "field": str,
"offset": u64, "max_len": u32 } }
field is one of stdout, stderr, genvm_log. Response:
{ "total_len": u64, "data": bytes }
data is at most min(max_len, chunk_cap) bytes starting at offset
(chunk_cap is a server constant, 256 KiB); clients page until
offset + len(data) == total_len. genvm_log is served as JSON Lines
(one structured record per line). Artifact replies are sent on the same
connection through a low-priority writer queue, so bulk transfers cannot
starve lifecycle events.
Artifacts are available only for retained finished runs. Queued, running,
and retained failed_to_start runs answer unknown_id even though they
remain attachable. An invalid field on a retained finished run answers
malformed_frame
The finished event carries each artifact’s total size, so clients can skip
the calls entirely when the blobs are empty
event notifications#
Externally-tagged; every event carries genvm_id and, when the run was
started with one, host_genvm_id. Variants:
queuedThe run is allocated but no executor process exists yet; startup validation (module locks) or permit acquisition may still be pending. Non-terminal. Mainly seen as the
attachsnapshot of a run that has not spawned.attachreturns the current lifecycle state; later notifications follow the Lifecycle guarantees. Intermediate states may be coalesced, so a fast run can go straight fromqueuedto a terminal event{ "queued": { "genvm_id": u64, "host_genvm_id": str? } }startedThe executor process was spawned.
{ "started": { "genvm_id": u64, "host_genvm_id": str? } }failed_to_startTerminal. The run never reached a spawned executor: request validation, version resolution, or the spawn itself failed. Permits are released.
{ "failed_to_start": { "genvm_id": u64, "host_genvm_id": str?, "error": str } }An executor that was spawned and then died on its own – rejecting its arguments, say – reports
finishedwith that exit code, not this event. The two are distinguished by whether a process was created, which the manager knows exactly; “started, then died immediately” is not a state it can observe without guessing, and a timer that guesses it would both tax every healthy run and still race a slow failure.finishedTerminal; retries can replay its notification (see Lifecycle guarantees)
{ "finished": { "genvm_id": u64, "host_genvm_id": str?, "cause": str, # exited | cancelled | deadline | shutdown "exit_code": i64?, # null when killed before exit code known "consumed_result": bytes?, # what the executor sent via consume_result "metrics": map?, "finished_at": str, # RFC3339 "version_major": u32, "version_minor": u32, "artifact_sizes": { "stdout": u64, "stderr": u64, "genvm_log": u64 } } }An internal manager failure after the spawn also ends in
finishedwith a nullexit_code. Its cause isexitedunless a termination cause was already recorded
For a top-level run, consumed_result is one outer ResultCode byte
followed by a calldata-encoded ReportedResult map. Before retaining it, the
manager checks that:
The outer byte is a known result code and agrees with the map’s
kindThe map decodes completely
execution_hashandsmall_hashare each 32 bytes unless the result isInternalError
An invalid report is refused without an acknowledgement and is not published
as consumed_result. FatalVmError is also illegal at this boundary: a
debug manager asserts, while a release manager logs the executor violation and
rewrites both result-code locations to VmError before publication. Clients
therefore never receive a top-level FatalVmError
The manager does not decode reported leader_public_data. The bytes remain
opaque, executor-line-specific consensus proposals
Lifecycle guarantees:
Each run has 1 terminal state (
failed_to_startorfinished). Subscribed connections receive its notification while connected; retryingrunwith the same retained token can replay it. Manager shutdown may close connections before the terminal notification is deliveredDisconnect does not kill a run. Runs terminate only via
cancel, the deadline, or manager shutdown. A client may disconnect, reconnect, andattachby(boot_id, genvm_id)to recover the state and result.Results are retained until
ackor the retention TTL (manager configurationexecution_retention, a duration string, default5m).
Duration strings#
Every duration on this wire and in the manager configuration is a string: a decimal number followed immediately by a unit, with no space.
Unit suffix |
Meaning |
|---|---|
|
milliseconds |
|
seconds |
|
minutes |
|
hours |
Examples: "30s", "10.5m", "1h", "250ms". The fractional part
is optional and is resolved at millisecond granularity; a value that is not a
number followed by one of the units above is rejected as malformed_frame
(or as a configuration error at startup).
A bare number is not accepted – the unit is mandatory, so that a value can never be silently misread as the wrong scale.
Error messages#
method_id = error, request_id echoed from the failing request
(0 when no request id could be attributed):
{ "code": u8, "message": str }
See errors for the generated error code enum. Meanings:
internal is a handler failure with a diagnostic message;
malformed_frame is failed calldata decode, a bad payload shape, a message
shorter than the header, a non-binary message, or an invalid artifact field;
unknown_method is a method id absent from the generated table or unavailable
to client requests (hello, event, error); unknown_id means the run
never existed, was acked, expired, or has no finished result for get_artifact;
boot_id_mismatch is attach across a
manager restart; bad_request_id is a client request with
request_id == 0; and not_finished is ack before a terminal event
or result.
An oversized message has no error code: the connection is dropped, as described under Framing.
Scope note#
Multiple connections attaching to one run observe it through the manager
channel only. Host connection 0 – the socket the executor dials and speaks
the host protocol on (Host Loop Pseudocode) – is owned by whichever address the
run request supplied; attaching does not share it. host_hello_data is
the building block for future shared-host schemes.