A governed agentic platform across engineering, manufacturing, stock, procurement and finance, on an ERP with no API
1,184 commits · one engineer · Dec 2025 → Aug 2026
24 views · 83 AI tools · 75 agent capabilities · 10 SAP connectors · 455 test files
The mistake people make
A finance dashboard tells you the score.
An operating system tells you why the score is what it is —
and what you can still change.
The constraint
Our ERP cannot give us an API
SAP ECC 6.0 — no HANA, no Fiori, no OData
"Reporting" = run a t-code, export an ALV grid, email it
Enabling OData needs BASIS, change management and transport requests — 3 months minimum
The upgrade maths
S/4HANA migrations start at ~$250K for small companies, millions for mid-market.
Gartner: of 35,000 ECC customers, ~17,000 will still be on legacy past the 2027 deadline.
We are one of them.
The actual question
How do you get AI-native decision support out of an ERP that
cannot give you an API —
without a transformation budget,
and without ever being wrong about a number?
Where it landed
The platform today
1,184commits
1engineer + AI agents
24routed views
83AI tool definitions
75agent capabilities
10SAP connectors
12 / 34roles / permissions
455test files
11cosmos containers
11admin capabilities
~364klines of source
$ single digitsmonthly db spend
Built alongside a day job. The shape matters more than the scale, the next slide is the shape.
The spine technical
One calculation engine, three runtimes
financialModel.js
(pure function — no React, no I/O)
│
┌────────────┼────────────┐
▼ ▼ ▼
BROWSER API AI SERVICE
live server-side copilot
simulation calculation arithmetic
Written on day one, before any UI, backend or database. It survived every
refactor, three hosting migrations and a service extraction.
Why this matters for AI
The copilot does not estimate margin. It calls the same function
the dashboard renders.
There is exactly one definition of margin in this system — so the AI
and the screen can never disagree.
This is the half of the platform people don't expect.
Product engineering
The roadmap is a data source, not a slide, six lenses, live from Jira
Flow HealthWeekly throughput, aging WIP, stuck >90d, lead time — per project.
Is the run rate rising or falling against the trailing quarter?QualityBug ratio, hygiene signals, oldest open item. The same permission
model that gates revenue gates this.
Production InsightsThroughput and build state — the operational constraint
sitting behind every revenue plan.RepairsRepair jobs and backlog risk, drillable to individual job detail.
The order book is only real if you can build against it.
Not an inventory dashboard — a stock decision system.
"What can we ship now? Build now? What unblocks as supplier POs land?"
Procurement & supply
Three of the ten SAP connectors are procurement-side: purchase orders · PO lines · GRN receipts
Supplier PO InsightsSupplier price comparison, procurement scope analysis and
commitment trends — the inbound side of the build model.Unblock as POs landWhich open POs actually clear a build — and which are
landing on stock that isn't the current bottleneck.
Procurement and manufacturing are the same question asked from two ends. The platform answers both from one dataset.
Manufacturing · capacity
The Capacity Planner, the most operational surface in the platform
Contract at a glanceOrder register, delivery clock, penalty exposure, compliance.
"Every figure is computed from the order register — the summary sheet writes itself."Order RegisterThe source of truth the dashboard is derived from — not a
parallel set of numbers.
Contract management
The translation layer — client ↔ usClient item codes mapped to ours, with
mapping confidence. Unglamorous, and the reason penalty maths can be trusted.Reports & ExportThe monthly client report — generated from the register,
not re-keyed.
The bug that shows the rigour: the dashboard counted PO lines while labelling them
POs, 528 lines vs 499 distinct PO documents, and penalties were counted per PO. The KPI block was silently mixing two grains.
Procurement 3 of the 10 connectors
Supplier PO Insights — a suite, not a widget
Seven tabs — Overview · Inbound Runway · Suppliers · Savings · Supplier Trends · Material Costs · Exceptions —
over purchase orders, PO lines and GRN receipts.
Procurement
Where the supply side becomes a decision
Inbound RunwayWhat arrives, when — and whether it lands on the current bottleneck.Supplier TrendsPrice movement over time — how you catch a supplier walking a price up quietly.Material CostsProduct-material cost comparison across suppliers.
Month-end the number the business reports on
FinPack — a real management-accounts pack
Income StatementBalance SheetCash Flow
Forecast DriftThe view that earns the module.ReconciliationWhere the pack proves itself.+ month-by-monthM01…M12 navigation across the fiscal year.
FinPack is declared authoritative for PBT in the shared grounding —
so neither the copilot nor an agent may quote a PBT figure from anywhere else.
Act two
The financial spine
Plan · simulate · actuals · the official monthly pack
Plan and simulate
Budget PlanBottom-up customer and product planning, regional goals,
logistics markup configuration, and the budget lock lifecycle.Budget SimulatorDrag a margin or logistics slider — PBT, break-even and the
profitability timeline recompute client-side in under 16ms.
FinPack the number the business reports on
The official monthly management pack
The strictest surface in the platform gets its own tool family:
list_finpack_packs compare_finpack_packs
query_finpack_pack query_finpack_line
query_finpack_reconciliation
query_finpack_narrative_register
query_finpack_forecast_settlement
get_finpack_kpi_evolution
explain_finpack_drift ← the important one
Why explain_finpack_drift matters
At month-end the useful question is never "what is the number".
It is "what moved, and why".
A pack that can explain its own drift against the prior month is a different
artefact from a pack that merely states a total.
The surface the sales team actually lives in
Sales Command Center
One place to inspect an account portfolio — and to hold it to account
The organising idea
One filter bar re-scopes the entire page
ACCOUNT MANAGER · REGION GROUP · REGION · SPECIFIC CUSTOMER
[ Reset ]
│
└─ every KPI, chart, table and drilldown below
re-computes against the selected scope
Four hero KPIs — Total Portfolio Budget · Actual (Full Year) ·
Full-Year Variance · Open Order Book — then seven analytical sections, all scope-aware.
Why this matters organisationally
The GM and the account manager look at the same page, not two different
reports that disagree.
A portfolio review stops being "send me your numbers" and becomes
"let's both look at the same scope."
Full portfolio view. Filter bar, four hero KPIs, and Gap to Budget Analysis —
actual vs latest forecast vs budget target, with Final Attainment and Forecast Visibility bars.The same page, scoped to one account manager. Note the scope chip on every widget,
and "Drill down through your portfolio" with BACK navigation.
Portfolio inspection
Customer Health & Detailed Performance, the account manager's working list
Status badges — Behind Plan · At Risk · Exceeded · Unplanned — beside Planned Margin vs
Actual Margin, full-year budget, actual, three years of history, open order book and variance.
Filterable by plan status and customer type; exportable to Copy / CSV / Excel.
Lenses
The same portfolio, four ways, because the question changes the shape
FULL YEAR / QUARTERLY / 12 MONTHSAttainment reads differently over a quarter than a year.Net Sales / Gross Profit / QuantityA rep can hit revenue and miss margin. Both are visible.Product Mix by CategoryWhat the portfolio is actually made of — and how that shifted.
Sales Performance Explorer — a pivot builder over the active scope.
Here: variance by account manager, split by customer type, with linked account detail.
Drill-down
Customer Order Book InspectionFrom a portfolio number down to the individual open
order lines behind it — without leaving the page or changing tool.Scoped customer listThe same table, filtered to one manager's accounts —
their working list for the week.
Portfolio → customer → order line. Three levels, one surface, one permission model.
Every widget carries an Explain affordance opening this drawer:
what this shows · data sources (with the exact ETL defType) · how it is calculated — and
"Ask AI About This", which hands the widget's context straight to the copilot.
Promote to latest · bump the per-connector manifest ·
invalidate exactly the affected client caches.
If one row vanished
The run fails. The previous good data stays live.
Nobody sees a silently wrong number.
Silent data loss is the failure mode that destroys trust in a finance tool —
so it's designed out at the pipeline boundary, not checked for afterwards.
Cost engineering exectechnical
Separate freshness signalling from data transfer
Naive
Re-read all SAP data on every page load, for every user, forever —
for data that changes once a week.
What we do
Poll a tiny manifest every ~60s. Cache payloads for days.
Refetch only when a per-connector version moves — a weekly stock ETL
no longer invalidates the monthly FinPack cache.
Encrypted client cache
AES-GCM · per-surface keys · gzip-before-encrypt · operator-controlled key rotation ·
fail-closed if WebCrypto or the key fetch is unavailable.
No plaintext financial data reaches localStorage.
Result
Single-digit monthly Cosmos spend on a system carrying a decade of fiscal history.
Act four
The AI engine
Tools, not retrieval
This is not a document-search problem. It's a database-query problem.
Retrieval is probabilistic. Its failure mode is a confident answer derived from
approximately the right rows — indistinguishable from a correct one until someone checks.
Coverage
83 tools, routed by question and page
Tool group
Tools
Answers questions like
finance
67
plan vs actual, PBT gap, margin analysis, fiscal-year comparison
competitor dominance, coverage gaps, installed base
scenario
13
what-if simulation, services model, goal seek
stock_decision
5
available-to-build, orderbook coverage, material 360
production · contract · roadmap
4 ea
repair jobs, contract 360, roadmap delivery
methodology
3
"how exactly is this computed?" — the trust tools
Organised by business domain, not storage layout, so a user's question maps onto a tool without the model inventing a join.
Across the whole business
I drove it as an executive would. Prose blurred; structure and provenance are not.
Manufacturing"Which products can we build right now, and which single
component blocks the most builds?"Procurement"Which suppliers moved prices most, and where are we
over-committed on open POs?"
Across the whole business
Month-end"Explain the drift in the latest FinPack versus the prior
month, and what caused it."Contracts"What is our delivery penalty exposure, and which orders are
at risk of going late?"
What a good answer looks like
Structure, not fluency
Asked where gross margin was eroding, the copilot produced, in this order:
Provenance first. Named the retained FY26-W53 snapshot, noted the data was 8 days old,
and disclaimed it was not the final fiscal-year-close position.
Contradicted the premise. "Your overall gross margin is not eroding — but three pockets are."
Decomposed the erosion by product category with the rate change on each.
Line-level evidence — part numbers showing cost bookings against zero revenue.
Tabulated the customers driving the drag, ranked by GP impact.
Proposed a causal hypothesis — tender pricing on a volume ramp — flagged as a hypothesis.
None of that is prompt politeness. It falls out of tools returning typed, timestamped, scope-filtered
payloads. The model does synthesis, which LLMs are good at, over deterministic retrieval and arithmetic, which they are not.
Captured live, not staged. This slide is the one I'd put in front of a sceptical CTO.
The semantic layer technical
170 insight definitions — the platform's own dictionary
Every metric on every dashboard has a catalog entry, and the AI has tools to read it:
get_insight_definition, search_insights.
The instruction that matters
The grounding tells the model that catalog formulas and assumptions beat general
finance knowledge.
Asked "how is margin calculated here", it answers from this platform's definition —
not from what margin usually means.
Why this is the difference
An assistant that knows finance is a search engine with manners.
An assistant that knows your finance can be argued with — and corrected —
because its definitions are inspectable artefacts rather than model priors.
Human-in-the-loop technical
The clarification gate — deterministic, not prompted
Resolver outcome
Action
not resolved
clarify — stop and ask
resolved, confidence < 0.70
clarify — catches all-token guesses
resolved, 0.70 – 0.90
echo — proceed, but say "reading X as Y — correct me"
resolved, ≥ 0.90
answer — proceed silently
When ambiguous, the handler returns a slim payload, candidates only, no data.
The system prompt already said "if ambiguous, ASK", and the model still
guessed and proceeded. You cannot make the model ask by instructing it harder; you make answering
impossible without asking by withholding the data.
The point
Model choice is operational configuration,
not a deployment.
A new frontier model ships → an admin picks it from a dropdown that populated itself.
One degrades → the platform routes around it and recovers on its own.
Neither event needs an engineer.
Prompt architecture technical
One truth layer, two surfaces
The agent's system prompt is composed from the same exported blocks as the
chat prompt: company context · data-worlds dictionary · reference-resolution discipline ·
actuals-routing and methodology anchors.
Editing one shared block updates both surfaces at once.
Drift is structurally impossible, not a discipline problem.
Runtime-composed and never persisted: existing agent definitions inherit an improved
grounding with no document migration, and no customer-authored brief is ever rewritten.
Two of the eight agent conduct rules
NAMES, NOT IDS. Never surface raw SAP customer numbers or internal keys in
narrative. Resolve to friendly names — or say explicitly that you couldn't.
FRESHNESS IS PART OF THE TRUTH. State "Data as of <date>" from tool metadata,
and flag any source older than seven days as stale.
What's deliberately absent
Page awareness, UI context, follow-up framing — a scheduled report has no user to ask
and no page to look at.
Act five
The agent platform
Chat answers the question you thought to ask.
Agents answer the ones you should be asking.
Agents are typed Cosmos documents, not hardcoded prompts — browsable, cloneable, remixable.
Each card shows intent, the exact tool allowlist, cadence, next run and owner.
Runtime technical
Five decisions that make agents safe
1 · Quarantine the SDK
runtime.js is the only file importing the Agent SDK. Types may not leak into
stores, hooks or ai-core. Swapping runtimes = one week, one file.
2 · Lock the built-in tool surface
allowedTools is an approval gate, not a visibility gate — so we also populate
disallowedTools with every SDK built-in, and reject definitions that smuggle one in.
3 · Enforce policy in the handler
The SDK can block a tool call; it cannot filter a tool's output.
BU data policy, cross-BU block and opt-in ABAC live in the handlers.
4 · Extract citations structurally
A PostToolUse hook pulls citations out of tool returns into the run document —
never parsed from the model's prose.
5 · Validate before publishable
An independent oracle grades every generated report for grounding, scope and
publishability in strict JSON. Failures land in a Needs Review queue, not an inbox.
Plus
Per-BU and per-agent cost quotas, and an audit hook recording who created, ran,
changed or scheduled what.
Scheduled trigger · run duration · Succeeded + Validated badges ·
explicit BU-wide visibility scope · citation chips bound to date ranges · charts with a "how to read" note.
This ran at 08:00 without anyone pressing a button.
The invariant that made it shippable
V1-indistinguishability
Every agent published by Composer V2 must be byte-equivalent to one authored
through the old typed pipeline. A snapshot suite enforces it on every PR.
Which means the scheduler, the runtime, the policy hooks and the report format
never had to learn that Composer V2 exists.
A new authoring UX shipped with zero blast radius on the execution path.
How to get invariants like this
Twelve rounds of independent adversarial agent review are traced in the plan document
before a line was written.
The review didn't just find bugs — it deleted features. A parallel chat runtime,
a marker protocol, a separate MCP factory map: all proposed, all cut.
Act six
Governance
The part that took longest and demos worst
Eleven admin capabilities
"Enterprise grade" is mostly this, the unglamorous half
RBACRole → permission mapping is data, not code. No deploy to change it.AI Agents PolicyPer-BU data policy, cost quotas, permitted capability surface.AuditField-level deep diff. Denied writes are audited too.
AI Chat Logs"What did the AI tell people?" is an auditable question, under a
documented envelope contract and retention policy.ETL PipelinePer-connector run history, import status, reconciliation state.System HealthCosmos, ETL, AI probes, storage, and an Azure cost ledger — in-product.
Full set: User Governance · Global Settings · Business Config & Admin · Feedback Triage · RBAC ·
AI Agents Policy · Audit Logs · AI Chat Logs · Operational Logs · ETL Pipeline · System Health
Testing technical
455 test files across five layers
Layer
What it protects
Domain model
The pure financial engine — margin, discount compounding, freight modes, break-even, projection
API contract
Request shape → response shape; permission gate → HTTP status. Named *.contract.test.js because they encode contracts other systems depend on
Frontend component
Not pixels — logic. Does the ETL tab appear for sap_exporter? Does the lock banner render when the FY is locked?
Authorization harness
Oracle vs. executor, golden + pairwise matrix, quarantine with expiry. Nightly at 02:00 UTC
AI service
ai-core contracts, extraction, and SSE event-shape validation — the stream format is a contract the frontend depends on
I did not write these because I'm disciplined. I wrote them because of the next slide.
Incidents became rules
Every contract clause has a scar behind it
What happened
What it became
An AI agent "simplified" a component and silently deleted 200 lines
NEVER OMIT CODE — in the agent contract, in capitals
One ESM import crossing a function-folder boundary crashed the entire Functions worker — not one endpoint, all of them
runtime-boundary.contract.test.js
Months later: contract code imported a workspace package that resolved in the monorepo but was absent from the deployed package
Byte-parity mirroring + guard extended to reject deploy-absent imports
Eight formula bugs in one day, each producing plausible-looking numbers
financialModel.js became a protected file
Different scripts used different partition-key paths for the same containers
One path (/pk) everywhere; one config file as sole source of truth
Source resolution is not package resolution. The exact deployment artifact must load
against only its own production dependency graph before it ships.
My favourite bug, because it is so quiet
findDue(): WHERE c.enabled = true
AgentDefinitionV1 has no enabled field.
The predicate never matched. No scheduled agent run could ever fire.
The weekly briefing was scheduled. The system reported no errors. Nothing happened.
There is no exception to catch when your filter silently matches nothing.
Supply chain technical
Freeze the deploy graph
Deploys install the exact graph CI validated — npm ci against a committed lockfile
Any version move shows up as a reviewable lockfile diff in a PR, not invisibly at deploy time
Four surfaces — web host, api, ai-service, ETL Python — are separate deploy units, never co-mingled
The guiding principles, verbatim
Pin what works; don't chase latest. Subtraction beats migration — removing an unused dependency shrinks the
audit surface with zero functional risk. One surface, one wave, one PR.
That is the core defence against slopsquatting and surprise upgrades.
Act eight
What I got wrong
A self-assessment that only lists strengths isn't one
June 2026 — formal AI/agentic maturity audit
Level 3 of 5
Dimension
Now
Target
Grounding & provenance
4
4
Tool use & capability management
4
4
Human-in-the-loop & clarification
4
4
Agent runtime & execution model
3
4
Model orchestration & reliability
3
4
Guardrails, safety & governance
3
4
Evaluation (agent output quality)
2
4
Observability & telemetry
2
4
Business context & memory
2
4
Multi-agent & durable workflows
2
3
How it was produced
A multi-agent workflow mapped 12 AI subsystems against ~60 primary sources on the
2025–26 state of the art — then adversarially verified every claimed gap back
against the codebase.
13 of 26 gap-claims turned out to be overstated. The code already had a
partial implementation, or the "gap" was a deliberate documented decision.
The sharpest criticism in it
A recurring "documented-but-not-enforced" risk. Several controls are described
in contracts but not mechanically verified.
Writing the rule is the easy half.
3–4 on the hard, differentiating dimensions. 2 on the operational hygiene the industry standardised in 2025.
Most enterprise teams are the mirror image: decent ops tooling bolted onto shallow grounding.
The backlog
What's next
Near — close the hygiene gap
· Agent-output evals. World-class contract tests, zero quality evals. Golden questions, scored grounding + citation coverage, in CI.
· OpenTelemetry gen_ai.* spans. Bespoke correlation logging works; it doesn't interop.
· Durable execution for the interactive chat loop.
· Turn declared controls into enforced ones.
Medium — product
· Composer V3. Invert the canvas: artefact as viewport, chat as side rail. Plus a lineage chip, plan-mode preview and per-section execution badges.
· Material Cost Service. A cost needs an auditable basis, effective date, coverage and evidence — not just a number.
· Stock beyond gross_buildable — reservation-aware free stock.
· A memory layer. There is none today.
Long — the big one
Intersection Intelligence. A map-first digital twin of every road junction in South Africa — signalised, roundabout, stop-controlled — with confidence scores and a proprietary overlay of deployed controller make/model.
The most ambitious thing on the roadmap was designed, from the first line,
to be switchable off.
Intersection Intelligence ships as a standalone Fastify service on its own App Service plan,
mirroring the AI service, because, in the owner's words:
"if it does not work, it should be switched off without breaking the core of the app."
A kill-switch ladder, its own operator console, and a proxy line in the host that can simply stop resolving.
…and then I retried it, and found a real bug
The trust machinery worked. The engineering had a hole.
What happened
I re-ran the Jira question for this edition. The tools timed out again — three
attempts out of three.
This time it degraded better: it fell back to the Roadmap page KPIs (same live Jira data,
already in Cosmos), gave a real quantified answer on throughput and WIP, and refused only
the part it couldn't source.
"22 items have been stuck for more than 90 days. I attempted to pull the ranked
list but the Jira query timed out this turn — I don't have the individual items,
so I won't guess at which ones are oldest."
Root cause
TOOL_TIMEOUT_MS = 5000 is the default tool budget.
Production tools were recognised as slow and given a 30s override.
No Jira tool was ever added to that map — so a cross-internet call to
Atlassian Cloud gets the same five seconds as an in-region Cosmos point read.
The fix isn't a bigger number
Derive the timeout from a declared executionClass on each tool, and add a contract
test asserting every tool making an outbound non-Azure call is declared
external-api.
That turns a silent 5s default into a build failure.
A pattern worth naming
Beware the opt-in registry.
Two of this platform's real bugs are the same shape:
The missed model tier
A model matching no tier pattern lands in other — and auto mode never selects it.
"…this is exactly how claude-fable-5 was missed."
The Jira timeout
A tool absent from the override map silently inherits a 5s budget meant for
in-region reads.
Derive from a declared property; test the derivation.
A map that new items must be added to will eventually not be.
Act ten
What building this actually taught me
Evidence, not adjectives
The claim
Not "I know about AI agents"
I have made every architectural decision that determines whether an
enterprise agent platform is trustworthy, and I can show you the specific incident that taught me each one.
What a CV usually claims
"Led AI initiatives" · "delivered a GenAI pilot" · "familiar with RAG and agents".
Exposure. Cheap in 2026.
What this is instead
A production system with named failure modes, a published maturity score, and a
bug report I wrote about my own platform three slides ago.
Competency map exec
Ten things this platform is evidence of
Competency
The evidence
AI strategy under constraint
Bridge around an un-migratable ERP instead of a seven-figure programme; single-digit monthly DB cost
Knowing when not to use the fashionable pattern
Refused RAG for numbers; cut a sub-agent swarm, a custom DSL, a parallel data plane and a second chat runtime
Grounding & the economics of trust
Structural citation extraction; provenance + confidence + model on every answer; it refuses when tools fail
Agentic architecture & vendor risk
One file imports the SDK; swap cost stated as ~1 week; four SDK gotchas known from being cut by them
Model portfolio management
Auto-discovery, tier policy, health circuit breaker with demotion and recovery — one resolver for the whole platform
Governance that is enforced, not declared
Policy in tool handlers, not prompts; built-ins denied by construction; quotas, audit, validation oracle
Deterministic / probabilistic decomposition
Ranking, thresholds, scope and identity are code; identity resolution shipped before the agent feature
Evaluating your own programme honestly
Level 3/5 published with the 2s visible; 13 of 26 gap-claims adversarially disproved
AI-augmented delivery as an operating model
1,184 commits solo; leverage came from contracts, not prompting; two-model pairing with human arbitration
Product judgement in an AI surface
The LLM has no publish tool; V1-indistinguishability let a new UX ship with zero blast radius
The one-paragraph version
I built and operate a production agentic platform over a legacy ERP: 83 typed tools,
75 agent capabilities, a scheduled agent framework on the Claude Agent SDK with
policy enforced at the tool-handler boundary, structural citation extraction, an
independent validation oracle, per-tenant cost quotas, model auto-discovery with a health
circuit breaker, and a three-layer authorization model differentially tested by a decision
oracle every night.
I can tell you which of those decisions were load-bearing, which were fashion, and which one I
would sequence differently, because I have the incident behind each.
For anyone hiring in this space
The competency map
Self-taught, self-rated, and deliberately not all fives
Competency map 1 = aware · 3 = working · 5 = shipped & operating in production
Where I actually am
Core agentic engineering
Prompt & system-prompt architecture5
Tool / function-calling design5
Conversational chat engine (multi-turn, SSE)5
Agent harness & orchestration5
Claude Agent SDK (production)5
MCP — Model Context Protocol4
Multi-agent / durable workflows2
Trust, safety & operations
Grounding, citations & provenance5
AI governance & guardrails4
Security & authorization for AI5
Model routing & multi-model ops4
Cost control / FinOps for AI4
Evals & output-quality measurement2
AI observability (OTel gen_ai.*)2
Platform, data & delivery
Data engineering for AI (ETL, semantics)5
AI product design & explainability UX5
AI-augmented software delivery5
Cloud architecture (Azure, serverless)4
RAG / vector retrieval2
Memory / long-horizon context2
Fine-tuning / model training1
Every rating above is defended by a named artefact in this deck. The next slide shows the evidence for the fives —
and why three of these are deliberately low.
The evidence behind the fives
Each rating maps to something in this deck
Competency
What backs it
Prompt & system-prompt architecture
"One truth layer, two surfaces" — chat and agent prompts composed from shared blocks so drift is structurally impossible; 8 non-negotiable agent conduct rules
Built from scratch: multi-turn, SSE streaming with a frozen event contract, 3-layer cache, abort/retry, deterministic clarification gate
Agent harness & orchestration
Scheduled agents in production: definitions as typed documents, scheduler, run store, citations, quotas, validation oracle, Needs-Review queue
Claude Agent SDK
One-file adapter boundary; disallowedTools defence in depth; PreToolUse/PostToolUse/Stop hooks; four SDK gotchas known from being cut by them
Grounding & provenance
Structural citation extraction from tool returns; snapshot-dated citation chips; provenance + confidence + model on every answer; it refuses when tools fail
Security & authorization for AI
3-layer authz enforced in tool handlers, not prompts; 12 roles / 34 permissions; decision oracle differentially testing it nightly
Data engineering for AI
10 SAP connectors with no-loss reconciliation; 170-entry insight catalog as a semantic layer the model reads
AI product design & explainability
Insight Explainer drawer (what / sources / how-calculated / "Ask AI About This"); confidence chips; clarification UX
AI-augmented software delivery
1,184 commits solo in 8 months; contracts-not-prompts operating model; two-model pairing with human arbitration
And the low ones
Why three of those are twos — and one is a one
Evals & output-quality measurement — 2
World-class contract tests; until recently zero quality evals. I built the hard part
first and left the layer the industry standardised in 2025.
It is the top item on my backlog, and the first thing I'd build at a new company.
AI observability — 2
Solid bespoke correlation logging; no OpenTelemetry gen_ai.* spans, so it doesn't
interoperate with anything.
RAG / vector retrieval — 2
I deliberately refused RAG for structured financial data and can defend that
decision in detail — but I won't claim depth I don't have. If your problem is
genuinely document search, hire someone with a 5 here.
Fine-tuning / model training — 1
I have not trained or fine-tuned a model. My leverage has been architecture around
frontier models, not producing them.
A matrix with no low scores is a marketing document.
These four are exactly what I'd want to be hired to go and close.
If you're hiring
Everything on the previous slides is self-taught,
built in production, against real money, and documented well enough that you can audit the claims.
Best fit
Head of AI · VP AI Engineering · AI Platform / Applied AI Architect ·
Chief AI Officer (mid-market) · CTO (mid-market)
What I'd bring in month one
A benchmark of your AI estate against the current state of the art, with an
evidence-cited, prioritised backlog — the same exercise I ran on my own platform,
which disproved 13 of its own 26 findings.
The two interview questions I'd ask myself
"What would you do differently?"
Sequence evals before the agent framework, not after. I have world-class
contract tests and, until recently, zero output-quality evals — I built the hard
part first and left the hygiene layer the industry standardised in 2025.
And identity resolution even earlier than I did.
"What is the risk in what you built?"
Single-instance assumptions in the model catalog and health breaker · no memory layer ·
a hand-rolled, non-durable interactive chat loop · and the
"documented-but-not-enforced" gap.
Every one is on the backlog with a named fix — and I'd rather be asked about them than
have them found.
What actually changed
Outcomes
Speed of answer
Emailed spreadsheets → a live operating picture.
SAP export to dashboard in ~30 seconds.
Speed of decision
Budget revisions were a quarterly
event. Now a director drags a slider and sees PBT move in under 16ms.
Breadth
One question surface across engineering,
manufacturing, stock, procurement, contracts and finance.
Questions nobody had time to ask
Scheduled agents produce cited briefings
at 08:00 — margin drift, customer health, debtor exposure — without anyone requesting them.
Cost
Single-digit monthly database spend.
Built for a fraction of one SAP consulting engagement.
Trust
The system says "I don't know" when its tools
fail — which is why people believe it when it doesn't.
If you take eight things
Lessons
Build the calculation engine first, as a pure function, and test it before you write a component.
Give agents tools, not documents. For structured financial data a typed tool over a partition-targeted query wins on every axis.
Enforce policy where the data is, not where the prompt is. Anything enforced only in a system prompt is a suggestion.
Keep the deterministic layer deterministic. Ranking, thresholds, scope and identity are code. Let the model do the writing.
Sequence the boring dependency first. Identity before Movement Pack; Movement Pack before agents.
Prove conservation before you publish. No-loss reconciliation at the boundary beats reconciliation reports afterwards.
A silent no-op is worse than a crash. Assert the observable outcome, not that your code ran.
Write contracts, not prompts. Every rule came from a specific incident — which is exactly why agents follow them.
Method note — optional, drop if short on time
How these screenshots were made
Every image is the live production system with real data on screen. A redaction
layer is injected before capture; capture loops, redact, verify, repeat, and writes a PNG only once the page verifies clean.
It failed eight times first
1. The checker shared the redactor's blind spot — both scanned leaf elements, so prose in
<li><strong>Facts:</strong>…</li> passed through while the check said "clean". 2. Pattern matching was the wrong model — an answer named five customers in prose with
no digits, so nothing matched. 3. I fixed the leaf-only bug, then immediately rebuilt it in the new pass. 4–6. Form values aren't text nodes · names that exist only in the page's own data ·
tenant identifiers that look like ordinary product codes. 7. The 25-character prose threshold never described the risk — a 17-character Jira title
in a model-rendered table slipped under it. 8.Verified clean, captured dirty. A clipped capture resizes the surface, Recharts re-renders its axes,
and real currency was painted into the PNG after the check passed. The run logged $=0. The file leaked.
The rule that worked
You cannot enumerate what sensitive text looks like.
So AI output became sensitive by default — and, after failure 8,
the assertion straddles the capture: shoot, re-verify, re-shoot if the page moved,
so the number reported describes the file on disk rather than the DOM a moment earlier.
A test that agrees with your bug is worse than no test —
it converts an unknown risk into false confidence.
Same class of error as the silent scheduler no-op. My redactor and my verifier were
written by the same mind, from the same mental model, on the same afternoon, so they failed together.
The first ten minutes of every conversation
FAQ
Four decisions people push back on — and the answer to each.
LangChain · one model provider · a document database · no evals
FAQ 1 most common question
"Why didn't you just use LangChain?"
Product
What it gives you
What I did instead
Verdict
LangChain (OSS)
Model / tool / prompt abstractions across providers
Anthropic Messages API + frozen AiProviderTransport
Primary reason, stated first: the point was to learn this from the ground up.
A framework teaches you its opinion about a tool loop, not what a tool loop is.
The research note that locked this in: the SDK "saves us from rebuilding the runtime / orchestration layer" —
v0 became days of work, not weeks.
What no framework could do for me
ABAC enforced inside tool handlers, before a row reaches the model ·
the Movement Pack so the model narrates rather than calculates ·
a 170-entry governed insight catalog ·
SAP reconciliation that refuses to publish on row-count drift.
A framework would have abstracted the tractable part of the problem and left the difficult part untouched.
Frugality is a real constraint, not a flourish. ~12 internal users on a shoestring.
LangSmith Plus is $39/seat/month + trace overages; a typical LangSmith + LangGraph Platform setup lands
$175–375/month before model spend.
Where I change my mind, and it's the same place twice. My two lowest scores are
evals 2/5 and observability 2/5. That is exactly LangSmith's core competence.
The move is to close a named gap rather than adopt a framework wholesale: adopt an eval/trace layer against interfaces that are
already frozen and contract-tested.
FAQ 2 the lock-in question
"Why go all-in on Claude? Where's the router?"
There is a seam — not a router
AiProviderTransport is a frozen, asserted contract: the orchestration layer talks to
a provider, not to Anthropic. providerCapabilities.js says so in its own header — adding OpenAI or Gemini
"would add sibling files with the same classify() / resolveThinkingConfig() contract."
One implementation exists. The audit scores this 3/5 and names it:
"sophisticated single-provider routing; no cross-provider failover."
What I do have inside the bet
Model auto-discovery every 12h against /v1/models · tier classification ·
a health circuit breaker that demotes and recovers automatically.
Failover exists — within the family, not across vendors.
This is a resourcing decision, not an architectural belief.
A provider-abstraction layer is a tax paid daily on every feature, against a risk that hasn't materialised.
One person, running an entire business: you pick a platform and move.
So I went vertically integrated, not just the LLM API, but the
Agent SDK and the MCP tool surface. That verticality is why the agent framework took days.
If Anthropic goes down, the AI degrades. The platform doesn't, every number comes from Cosmos.
The model only narrates.
Revisit triggers: sustained availability problems · a pricing shift · a capability that exists only elsewhere.
The seam is there so that day is an implementation, not a rewrite.
FAQ 3 the architect's question
"Why Cosmos DB and not a relational database?"
It began as a cost decision on a shoestring proof of concept, and stayed for engineering reasons.
Schema evolution without migrations. Fiscal-year structures, contract shapes and connector outputs changed repeatedly across eight months. Not one needed a migration window.
The ETL has no impedance mismatch. SAP exports arrive as records and land as documents. No ORM, no staging schema. That is most of why the ETL is ten connectors and not a data-engineering department.
The partition key is the security boundary. Everything partitions on /pk carrying BU + fiscal year, so hasBuScope is enforced in the storage layout — not just a WHERE clause someone can forget.
Operational features I'd otherwise build. TTL gives one-year audit retention as a field, not a cron job. ETags give optimistic concurrency without a locking strategy.
One dependency, one region, an emulator.npm run dev:up and there is no second database technology to run, secure, back up or patch.
The counterpoint
Cross-container joins and aggregations are my problem — hand-written and app-side.
A relational database would make ad-hoc analytical queries dramatically easier.
Which is exactly why the AI reads a governed semantic layer instead of writing SQL.
The database choice and "tools, not text-to-SQL" are the same decision seen twice.
FAQ 4 — rapid fire the uncomfortable ones
Answers I don't get to soften
Question
The answer
Tools instead of text-to-SQL?
A generated query cannot be authorized. All 83 tools enforce ABAC inside the handler before a row reaches the model, and return the citations the hooks accumulate. Text-to-SQL inverts that — the model composes the access path and you audit afterwards.
How do you stop it inventing numbers?
Structurally. Deterministic handlers emit citations and datasets; PostToolUse hooks accumulate them; the model is not the source of provenance. Honest caveat: the ClaimV1 refine gate currently passes trivially — a guarantee that always passes is not a guarantee. It's documented and on the near list.
Weakest part of the platform?
Durable execution. The dispatch queue has no consumer, so a restart orphans in-flight runs; the scheduler's 5-min stale lease is shorter than the 15-min run timeout — a latent double-fire. A documented Phase 4 deferral, and the one place LangGraph's checkpointing would have bought me something real.
No evals? In a financial app?
Correct, and it tops the backlog. What exists is the opposite trade: ~1,100+ deterministic contract tests, frozen schemas, a named regression test per incident. Every model call in every test is a deterministic fake — strong guarantees about behaviour, none about answer quality. The probabilistic layer goes on top of the deterministic gate.
What would you do differently?
OpenTelemetry from day one — retrofitting gen_ai.* spans over a bespoke scheme is pure rework. Evals before feature twenty. And enforce before declaring: the sharpest audit finding was controls that were persisted yet inert — an admin sets restrictedFields and believes it works.
Documentation describing runtime behaviour you haven't written is worse than none —
for the same reason a test that agrees with your bug is worse than no test.
Thank you
Contracts, Not Prompts
The bottleneck was never typing speed.
It was clarity of specification.