persistent-ai

@persistent-ai/fireflow-execution-worker (0.39.0)

Published 2026-10-07 23:38:00 +00:00 by ak

Installation

@persistent-ai:registry=
npm install @persistent-ai/fireflow-execution-worker@0.39.0
"@persistent-ai/fireflow-execution-worker": "0.39.0"

About this package

fireflow-execution-worker

The DBOS runtime that executes flows. It dequeues runs from the fireflow-executions queue (and the background, sandbox and terminal queues) in the executions database (DATABASE_URL_EXECUTIONS) and runs them as durable workflows. It serves no API; fireflow-execution-api and Flame Chorus only enqueue. Run as many replicas as the load needs: each one registers the same queues and claims work from PostgreSQL.

src/main.ts starts it; the DBOS configuration itself is read in packages/fireflow-executor/server/utils/config.ts (config.dbos) and applied in packages/fireflow-executor/server/dbos/config.ts (initializeDBOS) and packages/fireflow-executor/server/dbos/queue.ts.

Throughput and replica settings

Defaults are for many concurrent agent runs per worker. Option names are those of @dbos-inc/dbos-sdk 4.27.6 (DBOSConfig in dist/src/dbos-executor.d.ts, QueueParameters in dist/src/wfqueue.d.ts).

env default sets what it does
DBOS_WORKER_CONCURRENCY 50 workerConcurrency of fireflow-executions (half of it, at least 1, for fireflow-executions-background) Most runs one worker process executes at once, and so the most one queue poll claims.
DBOS_QUEUE_CONCURRENCY 0 (no cap) globalConcurrency of both execution queues Most runs executing at once across all workers. 0 removes the cap: the claim runs READ COMMITTED with SKIP LOCKED and no COUNT(*), so replicas scale without contending on one lock; a positive number switches the claim to REPEATABLE READ with FOR UPDATE NOWAIT.
DBOS_QUEUE_POLLING_INTERVAL_MS 50 minPollingIntervalMs of fireflow-executions How often each worker polls for new runs, so how long an idle worker takes to pick one up — and how long a slot freed by a finished run waits for its next run, because DBOS 4.27.6 wakes the poll only after a poll (dist/src/wfqueue.js, runQueuePoll). Each poll is one claim transaction of four statements per worker. Measured 2026-09-23 on an idle worker with one queue: 25.3 statements/s (8.1 commits/s) at 200 ms, 74.4 statements/s (21.2 commits/s) at 50 ms.
DBOS_SYSTEM_DATABASE_POOL_SIZE 30 systemDatabasePoolSize Connections to the DBOS system database per process. A claim starts every claimed run in the same tick and each needs a connection for its first step, so the pool is what meters them; half of it is reserved for polling reads. Budget replicas × this against the server's max_connections.
FIREFLOW_WORKER_ID (alias DBOS__VMID) unset (local) executorID DBOS stamps this id on every run the process claims and, at launch, recovers only the PENDING runs stamped with it. Give each replica its own id that survives that replica's restarts, such as a StatefulSet pod name. Replicas sharing one id re-run each other's in-flight runs when any of them restarts; a replica that comes back under a new id never recovers what it left PENDING. Leave unset only for a single worker; unset, the worker logs a warning at start.
DBOS_RUN_MIGRATIONS true runMigrations false makes launch verify the DBOS system schema instead of migrating it, and fail if it is behind. See the upgrade procedure below.
DBOS_SHUTDOWN_DRAIN_TIMEOUT_MS 25000 workflowCompletionTimeoutMS of DBOS.shutdown() How long SIGTERM waits for running workflows. Keep it below the orchestrator's termination grace period; what is still running is recovered at the next launch under the same executor id.
DBOS_LOG_LEVEL info logLevel The DBOS runtime's own log level, which filters every DBOS.logger call. The per-run lines (run started, provisioned, flow cache hit, run completed) are debug, so a busy worker prints none of them at info.
FIREFLOW_PROVISION_CACHE_TTL_MS 10000 per-worker caches in dbos/workflows/provision-cache.ts How long a worker keeps what it resolved before a Chorus run: the caller's FireFlow user and the flow's pinned coordinate (branch tip, path, owner). A run inside the window makes no editor-database or lakeFS call before its flow starts; a moved branch, a removed workspace or a changed owner reaches new runs after it. 0 turns the caches off.
NODE_OPTIONS --max-semi-space-size=32 (set in the execution-worker stage of the root Dockerfile) V8's young-generation size Each semi-space may grow to 32 MB, twice Node's default, so the short-lived objects every run allocates (node events, step outputs) die young instead of being promoted. A deployment that sets its own NODE_OPTIONS replaces the image's value and should carry the flag.

Other variables the worker reads (DBOS_APPLICATION_NAME, DBOS_ADMIN_*, the sandbox and terminal queue sizes) are listed with comments in .env.example.

Every DBOS process sharing a system database must use the same DBOS_APPLICATION_NAME. Since dbos-sdk 4.26 the queue claim, recovery and application versions are scoped by application name (dist/src/system_database.js, createApplicationVersion): the first named application to launch claims the pinned version 1.0.0, and a process with a different name then fails at launch. Runs enqueued with no name (fireflow-execution-api's DBOSClient, Flame Chorus) are claimed by any worker.

Keep DBOS's useListenNotify on (the default; nothing here sets it). Every live stream reader — FireFlow's own and Flame Chorus's — wakes on the notification DBOS sends on dbos_streams_channel after a stream write commits, and DBOS sends it only from a launched runtime with LISTEN/NOTIFY on (dist/src/system_database.js, #signalNotification). Off, no reader learns of a new value until the run ends. The run's end is announced by FireFlow's own trigger on ff_workflow_terminal, which the worker and the API install at start (PostgreSQLMigrations.ts).

This repository carries no Helm chart or other cluster manifest for the worker (searched git ls-files for helm, chart, k8s, kustom and contour on 2026-09-23); the settings above are set in whichever deployment repository runs it.

Upgrading @dbos-inc/dbos-sdk: the system schema migration

A new SDK version can add migrations to the dbos schema, and DBOS.launch() applies them. Moving from 4.25.14 to 4.27.6 takes the schema from version 69 to 108. Migrations 100 to 104 run ALTER TABLE … ADD COLUMN on dbos.workflow_status and dbos.operation_outputs: no table rewrite, but each takes an ACCESS EXCLUSIVE lock, waits behind every open transaction on the table and blocks every writer until it commits. The migration runner sets no lock_timeout (dist/src/sysdb_migrations/migration_runner.js). Launches of FireFlow processes are serialised by an advisory lock (initializeDBOS), so replicas do not race each other, but their running workflows and incoming enqueues do contend with the migration.

  1. Scale every DBOS process on the executions database to zero: all worker replicas. Pause Flame Chorus enqueues if a stall of the enqueue path is not acceptable; a blocked enqueue waits, it does not fail.

  2. Run the migration once, with a lock timeout so it fails fast instead of queueing behind a long reader:

    PGOPTIONS='-c lock_timeout=5s' pnpm --filter @persistent-ai/fireflow-executor exec dbos schema "$DATABASE_URL_EXECUTIONS"
    

    On a timeout, find the blocking session and run it again. dbos schema --print-migrations 70 prints the SQL from migration 70 on, without connecting, for a database administrator to apply instead. Starting one worker with DBOS_RUN_MIGRATIONS=true does the same thing without the lock timeout.

  3. Start the replicas with DBOS_RUN_MIGRATIONS=false. Each verifies the schema at launch and refuses to start against an unmigrated database.

Dependencies

Dependencies

ID Version
@mixmark-io/domino ^2.2.0
@modelcontextprotocol/sdk ^1.30.0
@opentelemetry/api ^1.9.1
@opentelemetry/context-async-hooks ^2.11.0
@opentelemetry/core ^2.11.0
@opentelemetry/exporter-logs-otlp-proto ^0.222.0
@opentelemetry/exporter-trace-otlp-proto ^0.222.0
@opentelemetry/resources ^2.11.0
@opentelemetry/sdk-logs ^0.222.0
@opentelemetry/sdk-trace ^2.11.0
@opentelemetry/sdk-trace-base ^2.11.0
@persistent-ai/fireflow-executor 0.39.0
@persistent-ai/fireflow-metrics 0.39.0
@persistent-ai/fireflow-nodes 0.39.0
@persistent-ai/fireflow-sandbox 0.39.0
@persistent-ai/fireflow-telemetry 0.39.0
@persistent-ai/fireflow-trpc 0.39.0
@persistent-ai/fireflow-types 0.39.0
@persistent-ai/persistentai-api 0.39.0
@trpc/server ^11.18.0
bigint-crypto-utils ^3.3.0
cors ^2.8.6
dotenv ^17.4.2
drizzle-orm ^0.45.2
esm ^3.2.25
grammy ^1.46.0
kafkajs ^2.2.4
nanoid ^6.0.1
nanoid-dictionary ^5.0.0
pg ^8.23.0
pg-listen ^1.7.0
superjson ^2.2.6
winston ^3.19.0
winston-transport ^4.9.0
ws ^8.21.3
xxhashjs ^0.2.2
yaml ^2.9.0

Development Dependencies

ID Version
@persistent-ai/fireflow-overcast-aztec 0.39.0
@persistent-ai/typescript-config 0.39.0
@types/cors ^2.8.19
@types/node ^26.6.2
@types/ws ^8.18.1
vite-plugin-top-level-await ^1.6.0
Details
npm
2026-10-07 23:38:00 +00:00
0
BUSL-1.1
6.4 MiB
Assets (1)
Versions (26) View all
0.40.5 2026-10-08
0.39.1 2026-10-08
0.39.0 2026-10-07
0.38.1 2026-09-30
0.37.1 2026-09-29