Most AI applications stop at calling an API. The hard part starts after that: batching thousands of jobs without blowing through rate limits, surviving a worker crash mid-submission, recovering when a provider key runs out of quota, and proving that the output you stored actually matches the schema you asked for.
Forge is my answer to that operational layer. It runs LLM batch and realtime jobs across OpenAI and Anthropic with two Node daemons — Anvil (the builder) and Hammer (the executor) — that share nothing but a Postgres database. No queue service, no IPC, no shared process. Every coordination problem is solved inside the database.
Postgres as the coordination spine
The system revolves around one handoff table. An upstream producer seeds pending rows; Anvil claims them and shapes them into provider-compatible batches; Hammer advances those batches through a fleet of worker loops and provider endpoints; terminal outcomes write back as done or failed.
The conventional move here is a queue: Redis, SQS, RabbitMQ. Each adds a second source of truth, and the moment job state lives in two places you inherit a class of bugs about keeping them consistent. Forge takes the other path: if Postgres already holds the state, let Postgres be the coordination layer.
Claiming work without stepping on itself
Anvil claims eligible jobs in a way that lets several copies run at once without two of them grabbing the same row, and opening a batch is idempotent — running it twice can't create a duplicate. Workers claim batches by advancing their status.
That combination has a property worth spelling out: every daemon and every worker is safe to run N times. Two Anvil replicas racing for the same rows simply skip past each other. A worker that crashes after claiming but before finishing leaves a row that someone else will reclaim. Horizontal scaling is not a feature bolted on later — it falls out of the claiming discipline.
Leases, not timestamps
Crash recovery keys on an explicit lease column, not a last-updated timestamp. The distinction matters: "this row hasn't been touched in five minutes" is ambiguous (slow worker? dead worker?), while "this lease expired" is a contract. Long-running realtime executions heartbeat their lease; terminal states clear it; a reaper worker reclaims anything whose lease has lapsed and either releases it for retry or fails it as an orphan.
Side effects that survive a crash
When a job completes, two things must happen: the database flips state, and the outside world hears about it — a tag stamped back into the upstream system, a notification fired. If those are two separate operations, a crash between them leaves the database and the world disagreeing.
Forge records the side effect in the same database write as the state change, then lets a dedicated worker deliver it out-of-band. A crash can lose timeliness, never consistency.
Batch and realtime in one model
Batch-mode jobs walk the long lane: created, submitting, submitted, polling, ready to collect, collecting, completed. Realtime jobs skip the provider batch path entirely: realtime, executing, completed. Same tables, same lease discipline, same delivery guarantees — one architecture, two latency profiles.
Provider ports
OpenAI and Anthropic ship in-tree behind a provider registry that resolves per batch. The differences are real — OpenAI wants JSONL file uploads, Anthropic takes inline batches, and a JSON schema that works on one can fail on the other — so provider quirks live in adapters and get validated up front, while the core execution model stays provider-agnostic.
Quota handling closes the loop: when a provider key exhausts its quota, it pauses; a recovery worker probes paused keys with tiny no-spend requests and unpauses them on refill. A control loop, fully instrumented with Prometheus metrics, instead of a human noticing failures in a dashboard.
Fail-closed validation
Every provider success line is revalidated against the originating processor's response schema before Forge accepts it. On an unresolvable schema, the line is reclassified as invalid output and routed through retry caps — a transient error gets a retry; a persistent mismatch terminates as failed rather than silently storing an unvalidated payload. With LLM outputs, "the provider said 200" is not the same as "the data is what you asked for."
The prompt platform behind this treats prompts as versioned, tagged artifacts — production, canary, draft — with strict templating, schema validation on both input and response, and promotion surfaces with preview. Prompts are deployable artifacts, not strings in code.
Lessons learned
- One source of truth beats two fast ones. Every coordination feature got simpler because state had exactly one home.
- Idempotency is the design test. If a step can't safely run twice, the claiming discipline is wrong.
- Leases make failure explicit. Encoding "who owns this and until when" turns crash recovery into a simple check.
- Validate at the trust boundary. Provider responses are external input; schema-check them before they become your data.
- Make the operator a first-class user. Health endpoints, metrics, cost rollups, and a live architecture view turned debugging sessions into glances.
Postgres-as-coordination-spine is not the answer for every scale, but for LLM job orchestration — where jobs are heavy, throughput is provider-bounded, and correctness matters more than microseconds — it removes an entire layer of infrastructure and the failure modes that come with it.