Architecture & reliability
Core services are understood, but customer-flow reliability is not measured end to end.
Public example / fictional company
A complete fictional example showing how payment, reliability, delivery, security, organization, AI, and diligence evidence becomes a ranked 90-day plan.
Recommended decision
Proceed with one corridor only after protecting transaction-state invariants, installing business-flow reliability signals, and bounding the AI support workflow. Defer a broad platform rewrite.
Context and trigger
A Seed-stage cross-border payments company is preparing to add two corridors while introducing an AI-assisted support workflow. Leadership needs to decide which controls must precede expansion and which platform work can wait.
Workflows under review
Payment workflow
Customer request → payment service → provider → webhook → ledger state → reconciliation → settlement/refund
AI support workflow
Support question → retrieval → policy context → model answer → confidence/control gate → human escalation
Readiness score
Equal weighting across eight dimensions. Selected deep dives change evidence depth, not scoring weight.
Core services are understood, but customer-flow reliability is not measured end to end.
Retry and reconciliation behavior permits ambiguous transaction state.
Deployment is repeatable; cross-team dependencies remain difficult to forecast.
Technical symptoms are reviewed, but financial and customer consequences are not consistently classified.
Baseline controls operate; AI-specific access and prompt-injection boundaries remain untested.
Leadership coverage is credible, with fragmented decision rights across corridor expansion.
Evaluation, fallback, and permission controls are not sufficient for market expansion.
Evidence exists, but it is distributed across owners and cannot yet support a concise diligence narrative.
Prioritized risk register
F-01 / Payments & ledger controls
Confidence: High
Evidence inspected
Sequence review shows the provider request can succeed after the local timeout while the retry receives a second reference.
Consequence
Duplicate movement or delayed customer resolution can appear when the next corridor increases timeout variance.
Action
Define idempotency, state-transition, and reconciliation invariants before enabling the next corridor.
Owner
Payments lead
Completion evidence
Every timeout path resolves to one canonical financial state in an automated replay test.
F-02 / Payments & ledger controls
Confidence: High
Evidence inspected
Exception export contains provider mismatch, delayed settlement, and duplicate-reference cases under the same operational status.
Consequence
Material exceptions can age behind low-risk operational noise.
Action
Classify exceptions by financial exposure, assign aging targets, and publish a daily owner view.
Owner
Payments operations
Completion evidence
High-risk exceptions have an owner within one business hour and no unexplained aging beyond the agreed threshold.
F-03 / Architecture & reliability
Confidence: High
Evidence inspected
Dashboards report service health and latency but cannot quantify completion or correctness for a customer payment flow.
Consequence
Leadership cannot distinguish acceptable service health from customer-impacting financial degradation.
Action
Instrument business-flow SLIs and establish initial objectives with error-budget review.
Owner
Platform lead
Completion evidence
Weekly review reports successful and ambiguous outcomes for each critical payment stage.
F-04 / AI development practices
Confidence: Medium
Evidence inspected
Current evaluation emphasizes fluent answers from one market and does not score citation support or safe refusal.
Consequence
Expansion can produce confident but incorrect guidance for customers in a new corridor.
Action
Build market-specific regression sets, require supported citations, and measure escalation quality before launch.
Owner
AI product lead
Completion evidence
The release gate passes agreed retrieval, groundedness, refusal, and escalation thresholds for each active market.
F-05 / Security & SOC 2 readiness
Confidence: Medium
Evidence inspected
The workflow retrieves customer-provided text and internal content without an adversarial test suite or explicit action allowlist.
Consequence
Manipulated context could expose internal instructions or trigger unintended workflow behavior.
Action
Separate trusted instructions from retrieved content, apply least-privilege tool access, and add adversarial regression tests.
Owner
Security owner
Completion evidence
The workflow resists the agreed injection suite and cannot perform actions outside its documented allowlist.
F-06 / Technical diligence
Confidence: High
Evidence inspected
The assessment required repeated reconciliation of conflicting diagrams, roadmap notes, and incident ownership.
Consequence
Investor, partner, or audit preparation will consume leadership time and expose inconsistent answers.
Action
Create one maintained engineering evidence index with owners and review dates.
Owner
VP Engineering
Completion evidence
A diligence dry run answers the defined architecture, reliability, security, and roadmap questions from current evidence.
90-day plan
Days 0–30
Completion evidence
Replay tests, exception aging view, and documented AI allowlist.
Days 31–60
Completion evidence
Weekly flow review, evaluation report, and completed diligence dry run.
Days 61–90
Completion evidence
Launch decision record, ownership review, and approved next-quarter roadmap.
Decisions enabled
Work deliberately deferred
Boundaries and exclusions