Engineering evidence · 9 October 2026

Code, tests and
operational records

The engineering behind a platform used by customer businesses with approximately $32M in combined sales since 2022. Explore a TypeScript sample, validation and recovery decisions, with operational records available for technical review

01 · Reviewable source

Small enough to read.
Grounded in a real client

ClaudeCode Octo is my separate TypeScript, React and Electron desktop/web agent client. Its operating contract covers isolated workspaces, local and remote execution, a shared lifecycle interface, and safe handling of retries and interrupted sessions

25 / 25

Standalone regression cases passed

Ported from the project’s existing tests to Node’s built-in test runner. The source modules compiled with strict TypeScript. The original suites also passed: 12 request-gate assertions and 13 external-link assertions

2 modules

Executable logic preserved

The sample contains the real request-gate and URL-policy logic. Private incident identifiers were removed from comments. Original paths and SHA-256 fingerprints are included; the test harness was prepared for this portfolio

Run locally
npm ci
npm test

Node.js 20.19+ · TypeScript 5.9.3 · no runtime dependencies · no model calls or network requests in the tests

One request, one decision

A transport retry should attach to its existing execution. A new message arriving during a live turn should steer that turn. A recently completed request can replay its result. Keeping this decision pure makes the concurrency policy directly testable

SituationDecision
Same request ID, live turn and registered completionJoin
Same request ID, completed within ten minutesReplay
Different message while a turn is liveInject
Idle engine, no retained matchRun

The integration keeps side effects outside the pure gate and guards cleanup by completion ownership. The extract alone is not a lock or a distributed exactly-once guarantee. Result retention expires after ten minutes; the policy trusts internal state and does not survive a process restart by itself

Validate navigation at the boundary

The second module resolves relative URLs and permits HTTP, HTTPS and mailto. Script, data, local-file and malformed absolute URLs are rejected. The reload predicate distinguishes the exact current document from navigation to another route

This is a protocol policy, not a destination allowlist or server-side request-forgery defense. Its documented scope matters as much as its passing tests

Read askGate.ts
// Extracted from ClaudeCode Octo on 9 October 2026; implementation unchanged.
export interface AskGateInput {
  /** Session has a live turn RIGHT NOW per the engine trackers
   *  (activeStreams / activeRemote / activeCodex) — ground truth, not the gate's
   *  own bookkeeping, so a stale gate entry can never wedge a session. */
  isLive: boolean
  /** askId of the registered in-flight ask for this session, if any. */
  inFlightAskId?: string | null
  /** There IS a registered in-flight completion to join. */
  hasInFlight: boolean
  /** askId of the most recently COMPLETED ask for this session, if any. */
  recentAskId?: string | null
  /** When that recent ask completed (ms epoch). */
  recentAt?: number
  /** The incoming ask's client-generated id (absent on legacy clients and
   *  engine-internal dispatches). */
  askId?: string | null
  now: number
}

export type AskGateDecision = 'run' | 'join' | 'replay' | 'inject'

/** How long a completed askId still swallows a late transport retry. Chromium
 *  retries fire immediately on connection failure, so minutes of slack is ample;
 *  10 min also covers a sleep/wake gap. */
export const RECENT_ASK_TTL_MS = 10 * 60_000

export function decideAsk(i: AskGateInput): AskGateDecision {
  if (i.isLive) {
    // A turn is genuinely running. The ONLY ask that may attach to it is its own
    // transport-level retry (identical askId). Everything else — different
    // askId, or no askId at all — is a different message and must be steered in,
    // never run as a second concurrent query.
    if (i.askId && i.hasInFlight && i.inFlightAskId === i.askId) return 'join'
    return 'inject'
  }
  // No live turn. A matching recently-completed askId means this is a late
  // retry of a turn that already ran to completion — do NOT run it again.
  if (
    i.askId &&
    i.recentAskId === i.askId &&
    typeof i.recentAt === 'number' &&
    i.now - i.recentAt < RECENT_ASK_TTL_MS
  ) {
    return 'replay'
  }
  return 'run'
}
Read externalLinks.ts
// Extracted from ClaudeCode Octo on 9 October 2026; implementation unchanged.
const EXTERNAL_PROTOCOLS = new Set(['http:', 'https:', 'mailto:'])

/**
 * Resolve an anchor/navigation target to a URL that may be handed to the
 * operating system. Relative links are resolved against the current document,
 * which matters for the served thin-client UI: `/docs` is still a web link and
 * must not replace the Electron app.
 */
export function resolveExternalUrl(href: string, baseUrl?: string): string | null {
  try {
    const url = baseUrl ? new URL(href, baseUrl) : new URL(href)
    return EXTERNAL_PROTOCOLS.has(url.protocol.toLowerCase()) ? url.href : null
  } catch {
    return null
  }
}

/**
 * Electron emits will-navigate for renderer-initiated top-level loads, but not
 * for BrowserWindow.loadURL/loadFile. The only renderer navigation we retain is
 * an exact reload of the current document (needed by dev/HMR and normal reload).
 */
export function isExactDocumentReload(targetUrl: string, currentUrl: string): boolean {
  try {
    return new URL(targetUrl).href === new URL(currentUrl).href
  } catch {
    return targetUrl === currentUrl && targetUrl.length > 0
  }
}

A separate accepted-workspace bootstrap regression also passed. These focused checks do not replace the full application suite or a live recovery benchmark. The historical OctoFunnel checks are documented separately in the technical audit

02 · Feedback and recovery

An attempt is one step
in a customer workflow

My commercial proof is customers selling through the platform: approximately $32M in combined customer sales since 2022, including the approximately $3.75M customer case. The customer case documents the business workflow and purchase records behind that example

Tool request → validation feedback → corrected request → result

The built-in agent adds ordinary tool errors to the conversation and can use that feedback in its next model round. A rejected write can therefore be followed by a corrected write within the same customer task. Validation also protects the workspace by refusing invalid changes

The loop has tool budgets, cancellation and terminal error handling. Some failures need another turn or intervention. With MCP, the connected agent controls its own recovery. Attempt-level error counts help diagnose friction and cost; they do not by themselves describe the final task outcome or customers’ sales

Source reviewed on 9 October 2026: apps/server/src/agent/loop.ts, snapshot c347235982d1. Tool results return to the model; thrown exceptions and cancellation can end the turn. Model choice can affect behavior, but this review did not compare outcomes by model price. The commercial figures span the platform’s history and are not attributed specifically to AI

Inspect the operational records and measurement definitions

Window: 8 September–7 October 2026 UTC. Aggregate queries ran against live PostgreSQL on 9 October in a repeatable-read, read-only transaction. Only aggregate results were exported

Chat lifecycle stateTurnsStop requestedMedian lifecycle time
completed3,509No72.85 s
completed380Yes37.24 s
failed58No71.34 s
interrupted6No144.38 s
interrupted1Yes223.45 s
Total recorded turns3,954One turn is one durable reply lifecycle

Of 3,889 turns marked completed, 3,529 retained a nonempty final-answer record. The completed total includes 380 turns with a stop request. A further 58 turns failed and 7 were interrupted. These are distinct operational states, not customer acceptance labels

Timing

Completed turns without a stop flag had a median lifecycle of 72.85 seconds and a 95th percentile of 445.08 seconds. This measures creation-to-completion timestamps, including the turn’s orchestration. It does not measure human time saved or the time to finish an entire customer project

Cost attribution

The retained AI-chat wallet ledger contains 24,743 debit entries totaling 5,302.056188 platform credits. Billing exists, but these records do not supply a reconciled cost per accepted task. External MCP clients also pay for their own models outside this ledger

Generation jobs

The same window contains 18 funnel-generation job records: 16 complete and 2 failed. The completed jobs record 522.628362 credits spent. Their 84.70-second median update-time interval is a lifecycle proxy, not an end-to-end funnel-creation benchmark

Interventions and acceptance

Stop flags, failures and automatic recovery fields are observable. Human review, corrective edits, abandonment and explicit customer acceptance are not reliably classified by the reviewed stores. I do not treat a technical completion as an accepted customer outcome

The provisional cohort excludes configured operator accounts and the earlier test-like email filter. It is not a verified paying-customer roster and can contain staff-assisted work. Retention and deletion affect coverage. Source semantics were reviewed at snapshot 3b91d462f; historical deployed versions across the window were not reconstructed

Optional evaluation of agent task quality

For claims specifically about agent accuracy, time saved or cost per completed task: a stable task ID linking all attempts, model/tool charges and final artifacts; an explicit accepted/rejected/abandoned outcome; recorded interventions and review time; and a predefined acceptance rubric. A prospective comparison should retain failed attempts and separate active human time from wall-clock time. This would extend the engineering evidence; it is not a prerequisite for presenting the existing customer-sales results

03 · Decisions and trade-offs

The engineering reasoning

Shared data for agent and UI

OctoFunnel exposes the same canonical business resources through a VFS and a visual editor. This avoids synchronizing an independent agent copy. The cost is maintaining schema, permissions and projections together whenever a feature changes

Structured tools over interface clicks

A compact file-like tool surface lets agents inspect and change structured resources without depending on screen coordinates. It needs explicit validation and contextual instructions; tool success alone cannot establish whether the resulting customer workflow is good

Durable turns and guarded writes

Persisted lifecycle rows support reconnects and interruption handling. Compare-and-set completion prevents a stale owner from overwriting an interruption decision. Revision checks and checkpoints protect changes, with storage and recovery complexity as the trade-off

Isolated agent workspaces

The separate client keeps workspace state and credentials isolated while sharing the rendering contract. Its bootstrap regression checks that accepted helper code and matching dependencies are used, even when a shared checkout has unrelated edits

The sample and focused tests make selected implementation decisions reviewable. The linked customer case shows how the platform supports an operating business

Discuss the implementation

I can walk through the source, the failure modes behind these policies and the limits of the measurements

Get in touch Read my résumé