Observation before retry
After an ambiguous transport failure, Corrobo asks the authoritative system what is true before deciding whether another attempt is safe. observe() may only read, so recovery can never create the effect.
TypeScript library for safe retries of external side effects: when a timeout or crash leaves a write's outcome unknown, it checks the real system before acting again.
State-changing API calls can time out after the mutation already happened. Treating every caught exception as a failed write, or retrying blindly, can create duplicate refunds, duplicate messages, or unsafe retries.
Idempotency keys and durable workflows (Temporal, DBOS, Inngest) help, but neither answers the question directly: a step that crashes after its external write still runs again, and many APIs have no keys or expire them.
A small TypeScript library that wraps one consequential call in an effect contract (execute, observe, reconcile, recover). It records the attempt before calling out, then asks the authoritative system what actually happened, so a transport error is treated as evidence rather than business truth.
PostgresStore for crash-safe, concurrent operation records (advisory locks, version-checked writes, database-clock time windows) and InMemoryStore for tests.
Controls for agent runtimes that run before anything executes: an authorize() gate for human review, attributed review decisions tied to one review's token and the recorded intent, approval expiry, and a revalidate() hook that re-checks the world before every attempt.
A conformance harness (corrobo/testing) that runs a user's own contract against a fake provider through lost responses, late landings, crashes, and races, and counts real effects on the fake.
Operation identity → durable reservation → authorize / revalidate → execute side effect → authoritative observe → reconcile evidence state → settle in-flight window → derive recovery disposition
After an ambiguous transport failure, Corrobo asks the authoritative system what is true before deciding whether another attempt is safe. observe() may only read, so recovery can never create the effect.
APPLIED, NOT_APPLIED, CONFLICTED, PENDING, and UNKNOWN describe the evidence; COMPLETE, RETRY, REPLAN, REVIEW, and INVESTIGATE describe what is safe to do next.
A timed-out request isn't undone; it can still land after the first check. A NOT_APPLIED after a failed call only becomes RETRY once the contract's declared in-flight window has passed, and each attempt records the window it was sent under so a later deploy can't shorten it.
runEffect() takes no review decision, so the code making attempts can't approve them. reviewEffect() records who approved what, bound to one review's token, and never executes. That is an API boundary that enables capability separation, not a privilege boundary, and the docs say so.
Attempts are reserved before execution, same-identity callers are serialized with advisory locks, and version-checked writes keep a caller that lost its lock from overwriting anything.
Every row of the failure matrix (lost responses, late landings, crashes at each boundary, races, lost locks, review, identity misuse) cites the test that proves it, and CI fails if a cited test disappears. A model-based test drives random sequences of runs, review decisions, crashes, clock jumps, and rolling upgrades against an independent oracle, in memory and against Postgres; a mutation check reintroduces 35 known bugs and the model catches every one. Runnable examples include a real HTTP ledger, Stripe refunds with idempotency keys, and a DBOS workflow step killed between its write and checkpoint (naive retry: 2 credits; Corrobo: 1).