Syntell BizOps Intelligence

The Operating System
of a Business

A governed agentic platform across engineering, manufacturing,
stock, procurement and finance, on an ERP with no API

1,184 commits · one engineer · Dec 2025 → Aug 2026
24 views · 83 AI tools · 75 agent capabilities · 10 SAP connectors · 455 test files

The mistake people make

A finance dashboard tells you the score.

An operating system tells you why the score is what it is — and what you can still change.

The constraint

Our ERP cannot give us an API

  • SAP ECC 6.0 — no HANA, no Fiori, no OData
  • "Reporting" = run a t-code, export an ALV grid, email it
  • Enabling OData needs BASIS, change management and transport requests — 3 months minimum

The upgrade maths

S/4HANA migrations start at ~$250K for small companies, millions for mid-market.

Gartner: of 35,000 ECC customers, ~17,000 will still be on legacy past the 2027 deadline.

We are one of them.
The actual question

How do you get AI-native decision support out of an ERP that cannot give you an API — without a transformation budget,
and without ever being wrong about a number?

Where it landed

The platform today

1,184commits
1engineer + AI agents
24routed views
83AI tool definitions
75agent capabilities
10SAP connectors
12 / 34roles / permissions
455test files
11cosmos containers
11admin capabilities
~364klines of source
$ single digitsmonthly db spend

Built alongside a day job. The shape matters more than the scale, the next slide is the shape.

Platform stack
The spine technical

One calculation engine, three runtimes

financialModel.js (pure function — no React, no I/O) │ ┌────────────┼────────────┐ ▼ ▼ ▼ BROWSER API AI SERVICE live server-side copilot simulation calculation arithmetic

Written on day one, before any UI, backend or database. It survived every refactor, three hosting migrations and a service extraction.

Why this matters for AI

The copilot does not estimate margin. It calls the same function the dashboard renders.

There is exactly one definition of margin in this system — so the AI and the screen can never disagree.
Act one

The operating model

Product engineering · manufacturing · stock · procurement · contracts

This is the half of the platform people don't expect.

Product engineering

The roadmap is a data source, not a slide, six lenses, live from Jira

Flow health
Flow HealthWeekly throughput, aging WIP, stuck >90d, lead time — per project. Is the run rate rising or falling against the trailing quarter?
Quality
QualityBug ratio, hygiene signals, oldest open item. The same permission model that gates revenue gates this.
Manufacturing production

Overview · BuildState · Repairs · Capacity Planner

Production overview
Production InsightsThroughput and build state — the operational constraint sitting behind every revenue plan.
Repairs
RepairsRepair jobs and backlog risk, drillable to individual job detail.

The order book is only real if you can build against it.

Stock Management Insights
Not an inventory dashboard — a stock decision system. "What can we ship now? Build now? What unblocks as supplier POs land?"
Procurement & supply

Three of the ten SAP connectors are procurement-side: purchase orders · PO lines · GRN receipts

Supplier PO Insights
Supplier PO InsightsSupplier price comparison, procurement scope analysis and commitment trends — the inbound side of the build model.
Unblock as POs land
Unblock as POs landWhich open POs actually clear a build — and which are landing on stock that isn't the current bottleneck.

Procurement and manufacturing are the same question asked from two ends. The platform answers both from one dataset.

Manufacturing · capacity

The Capacity Planner, the most operational surface in the platform

Capacity Planner
Contract management

Multi-year client contracts, deliberately fiscal-year-agnostic

Contract dashboard
Contract at a glanceOrder register, delivery clock, penalty exposure, compliance. "Every figure is computed from the order register — the summary sheet writes itself."
Order register
Order RegisterThe source of truth the dashboard is derived from — not a parallel set of numbers.
Contract management
Contract master
The translation layer — client ↔ usClient item codes mapped to ours, with mapping confidence. Unglamorous, and the reason penalty maths can be trusted.
Reports and export
Reports & ExportThe monthly client report — generated from the register, not re-keyed.

The bug that shows the rigour: the dashboard counted PO lines while labelling them POs, 528 lines vs 499 distinct PO documents, and penalties were counted per PO. The KPI block was silently mixing two grains.

Procurement 3 of the 10 connectors

Supplier PO Insights — a suite, not a widget

Supplier PO Insights

Seven tabs — Overview · Inbound Runway · Suppliers · Savings · Supplier Trends · Material Costs · Exceptions — over purchase orders, PO lines and GRN receipts.

Procurement

Where the supply side becomes a decision

Inbound runway
Inbound RunwayWhat arrives, when — and whether it lands on the current bottleneck.
Supplier trends
Supplier TrendsPrice movement over time — how you catch a supplier walking a price up quietly.
Material costs
Material CostsProduct-material cost comparison across suppliers.
Month-end the number the business reports on

FinPack — a real management-accounts pack

Income statement
Income Statement
Balance sheet
Balance Sheet
Cash flow
Cash Flow
Forecast drift
Forecast DriftThe view that earns the module.
Reconciliation
ReconciliationWhere the pack proves itself.
Suppliers
+ month-by-monthM01…M12 navigation across the fiscal year.

FinPack is declared authoritative for PBT in the shared grounding — so neither the copilot nor an agent may quote a PBT figure from anywhere else.

Act two

The financial spine

Plan · simulate · actuals · the official monthly pack

Plan and simulate
Budget plan
Budget PlanBottom-up customer and product planning, regional goals, logistics markup configuration, and the budget lock lifecycle.
Budget simulator
Budget SimulatorDrag a margin or logistics slider — PBT, break-even and the profitability timeline recompute client-side in under 16ms.
FinPack the number the business reports on

The official monthly management pack

The strictest surface in the platform gets its own tool family:

list_finpack_packs compare_finpack_packs query_finpack_pack query_finpack_line query_finpack_reconciliation query_finpack_narrative_register query_finpack_forecast_settlement get_finpack_kpi_evolution explain_finpack_drift ← the important one

Why explain_finpack_drift matters

At month-end the useful question is never "what is the number".

It is "what moved, and why".

A pack that can explain its own drift against the prior month is a different artefact from a pack that merely states a total.
The surface the sales team actually lives in

Sales Command Center

One place to inspect an account portfolio —
and to hold it to account

The organising idea

One filter bar re-scopes the entire page

ACCOUNT MANAGER · REGION GROUP · REGION · SPECIFIC CUSTOMER [ Reset ] │ └─ every KPI, chart, table and drilldown below re-computes against the selected scope

Four hero KPIs — Total Portfolio Budget · Actual (Full Year) · Full-Year Variance · Open Order Book — then seven analytical sections, all scope-aware.

Why this matters organisationally

The GM and the account manager look at the same page, not two different reports that disagree.

A portfolio review stops being "send me your numbers" and becomes "let's both look at the same scope."
Sales Command Center — full portfolio
Full portfolio view. Filter bar, four hero KPIs, and Gap to Budget Analysis — actual vs latest forecast vs budget target, with Final Attainment and Forecast Visibility bars.
Single account manager portfolio
The same page, scoped to one account manager. Note the scope chip on every widget, and "Drill down through your portfolio" with BACK navigation.
Portfolio inspection

Customer Health & Detailed Performance, the account manager's working list

Customer health table

Status badges — Behind Plan · At Risk · Exceeded · Unplanned — beside Planned Margin vs Actual Margin, full-year budget, actual, three years of history, open order book and variance. Filterable by plan status and customer type; exportable to Copy / CSV / Excel.

Lenses

The same portfolio, four ways, because the question changes the shape

Quarterly
FULL YEAR / QUARTERLY / 12 MONTHSAttainment reads differently over a quarter than a year.
Gross profit lens
Net Sales / Gross Profit / QuantityA rep can hit revenue and miss margin. Both are visible.
Product mix
Product Mix by CategoryWhat the portfolio is actually made of — and how that shifted.
Sales Performance Explorer grouped by account manager
Sales Performance Explorer — a pivot builder over the active scope. Here: variance by account manager, split by customer type, with linked account detail.
Drill-down
Customer order book inspection
Customer Order Book InspectionFrom a portfolio number down to the individual open order lines behind it — without leaving the page or changing tool.
Account manager customer list
Scoped customer listThe same table, filtered to one manager's accounts — their working list for the week.

Portfolio → customer → order line. Three levels, one surface, one permission model.

Insight explainer drawer
Every widget carries an Explain affordance opening this drawer: what this shows · data sources (with the exact ETL defType) · how it is calculated — and "Ask AI About This", which hands the widget's context straight to the copilot.
Act three

The SAP bridge

Start with the export, not the API

SAP connector pattern
The gate technical

No-loss reconciliation


importStatus            === 'SUCCESS'
&& latestPointerUpdated === true
&& reconciliation.isBalanced === true
//  sourceRows === includedRows + excludedRows
  

If it balances

Promote to latest · bump the per-connector manifest · invalidate exactly the affected client caches.

If one row vanished

The run fails. The previous good data stays live. Nobody sees a silently wrong number.

Silent data loss is the failure mode that destroys trust in a finance tool — so it's designed out at the pipeline boundary, not checked for afterwards.

Cost engineering exec technical

Separate freshness signalling from data transfer

Naive

Re-read all SAP data on every page load, for every user, forever — for data that changes once a week.

What we do

Poll a tiny manifest every ~60s. Cache payloads for days. Refetch only when a per-connector version moves — a weekly stock ETL no longer invalidates the monthly FinPack cache.

Encrypted client cache

AES-GCM · per-surface keys · gzip-before-encrypt · operator-controlled key rotation · fail-closed if WebCrypto or the key fetch is unavailable.

No plaintext financial data reaches localStorage.

Result

Single-digit monthly Cosmos spend on a system carrying a decade of fiscal history.
Act four

The AI engine

Tools, not retrieval

This is not a document-search problem.
It's a database-query problem.

Retrieval is probabilistic. Its failure mode is a confident answer derived from approximately the right rows —
indistinguishable from a correct one until someone checks.

AI engine core
Coverage

83 tools, routed by question and page

Tool groupToolsAnswers questions like
finance67plan vs actual, PBT gap, margin analysis, fiscal-year comparison
sap66debtors risk, delivery performance, customer concentration, YoY
engineering16Jira throughput, stuck items, project health
market16competitor dominance, coverage gaps, installed base
scenario13what-if simulation, services model, goal seek
stock_decision5available-to-build, orderbook coverage, material 360
production · contract · roadmap4 earepair jobs, contract 360, roadmap delivery
methodology3"how exactly is this computed?" — the trust tools

Organised by business domain, not storage layout, so a user's question maps onto a tool without the model inventing a join.

Across the whole business

I drove it as an executive would. Prose blurred; structure and provenance are not.

Stock buildability
Manufacturing"Which products can we build right now, and which single component blocks the most builds?"
Procurement
Procurement"Which suppliers moved prices most, and where are we over-committed on open POs?"
Across the whole business
FinPack drift
Month-end"Explain the drift in the latest FinPack versus the prior month, and what caused it."
Contract penalties
Contracts"What is our delivery penalty exposure, and which orders are at risk of going late?"
What a good answer looks like

Structure, not fluency

Asked where gross margin was eroding, the copilot produced, in this order:

  1. Provenance first. Named the retained FY26-W53 snapshot, noted the data was 8 days old, and disclaimed it was not the final fiscal-year-close position.
  2. Contradicted the premise. "Your overall gross margin is not eroding — but three pockets are."
  3. Decomposed the erosion by product category with the rate change on each.
  4. Line-level evidence — part numbers showing cost bookings against zero revenue.
  5. Tabulated the customers driving the drag, ranked by GP impact.
  6. Proposed a causal hypothesis — tender pricing on a volume ramp — flagged as a hypothesis.

None of that is prompt politeness. It falls out of tools returning typed, timestamped, scope-filtered payloads. The model does synthesis, which LLMs are good at, over deterministic retrieval and arithmetic, which they are not.

Honest failure
Captured live, not staged. This slide is the one I'd put in front of a sceptical CTO.
The semantic layer technical

170 insight definitions — the platform's own dictionary

Every metric on every dashboard has a catalog entry, and the AI has tools to read it: get_insight_definition, search_insights.

The instruction that matters

The grounding tells the model that catalog formulas and assumptions beat general finance knowledge.

Asked "how is margin calculated here", it answers from this platform's definition — not from what margin usually means.

Why this is the difference

An assistant that knows finance is a search engine with manners.

An assistant that knows your finance can be argued with — and corrected — because its definitions are inspectable artefacts rather than model priors.
Human-in-the-loop technical

The clarification gate — deterministic, not prompted

Resolver outcomeAction
not resolvedclarify — stop and ask
resolved, confidence < 0.70clarify — catches all-token guesses
resolved, 0.70 – 0.90echo — proceed, but say "reading X as Y — correct me"
resolved, ≥ 0.90answer — proceed silently

When ambiguous, the handler returns a slim payload, candidates only, no data.

The system prompt already said "if ambiguous, ASK", and the model still guessed and proceeded. You cannot make the model ask by instructing it harder; you make answering impossible without asking by withholding the data.

Model discovery and routing
The point

Model choice is operational configuration, not a deployment.

A new frontier model ships → an admin picks it from a dropdown that populated itself.

One degrades → the platform routes around it and recovers on its own.

Neither event needs an engineer.

Claude Agent SDK anatomy
Prompt architecture technical

One truth layer, two surfaces

The agent's system prompt is composed from the same exported blocks as the chat prompt: company context · data-worlds dictionary · reference-resolution discipline · actuals-routing and methodology anchors.

Editing one shared block updates both surfaces at once. Drift is structurally impossible, not a discipline problem.

Runtime-composed and never persisted: existing agent definitions inherit an improved grounding with no document migration, and no customer-authored brief is ever rewritten.

Two of the eight agent conduct rules

NAMES, NOT IDS. Never surface raw SAP customer numbers or internal keys in narrative. Resolve to friendly names — or say explicitly that you couldn't.

FRESHNESS IS PART OF THE TRUTH. State "Data as of <date>" from tool metadata, and flag any source older than seven days as stale.

What's deliberately absent

Page awareness, UI context, follow-up framing — a scheduled report has no user to ask and no page to look at.
Act five

The agent platform

Chat answers the question you thought to ask.
Agents answer the ones you should be asking.

Agent gallery
Agents are typed Cosmos documents, not hardcoded prompts — browsable, cloneable, remixable. Each card shows intent, the exact tool allowlist, cadence, next run and owner.
Agent framework
Runtime technical

Five decisions that make agents safe

1 · Quarantine the SDK

runtime.js is the only file importing the Agent SDK. Types may not leak into stores, hooks or ai-core. Swapping runtimes = one week, one file.

2 · Lock the built-in tool surface

allowedTools is an approval gate, not a visibility gate — so we also populate disallowedTools with every SDK built-in, and reject definitions that smuggle one in.

3 · Enforce policy in the handler

The SDK can block a tool call; it cannot filter a tool's output. BU data policy, cross-BU block and opt-in ABAC live in the handlers.

4 · Extract citations structurally

A PostToolUse hook pulls citations out of tool returns into the run document — never parsed from the model's prose.

5 · Validate before publishable

An independent oracle grades every generated report for grounding, scope and publishability in strict JSON. Failures land in a Needs Review queue, not an inbox.

Plus

Per-BU and per-agent cost quotas, and an audit hook recording who created, ran, changed or scheduled what.
Filed agentic report
Scheduled trigger · run duration · Succeeded + Validated badges · explicit BU-wide visibility scope · citation chips bound to date ranges · charts with a "how to read" note.
This ran at 08:00 without anyone pressing a button.
Composer V2
The invariant that made it shippable

V1-indistinguishability

Every agent published by Composer V2 must be byte-equivalent to one authored through the old typed pipeline. A snapshot suite enforces it on every PR.

Which means the scheduler, the runtime, the policy hooks and the report format never had to learn that Composer V2 exists.

A new authoring UX shipped with zero blast radius on the execution path.

How to get invariants like this

Twelve rounds of independent adversarial agent review are traced in the plan document before a line was written.

The review didn't just find bugs — it deleted features. A parallel chat runtime, a marker protocol, a separate MCP factory map: all proposed, all cut.
Movement Pack
Act six

Governance

The part that took longest and demos worst

Authorization and decision oracle
Eleven admin capabilities

"Enterprise grade" is mostly this, the unglamorous half

RBAC
RBACRole → permission mapping is data, not code. No deploy to change it.
AI Agents Policy
AI Agents PolicyPer-BU data policy, cost quotas, permitted capability surface.
Audit logs
AuditField-level deep diff. Denied writes are audited too.
AI chat logs
AI Chat Logs"What did the AI tell people?" is an auditable question, under a documented envelope contract and retention policy.
ETL pipeline
ETL PipelinePer-connector run history, import status, reconciliation state.
System health
System HealthCosmos, ETL, AI probes, storage, and an Azure cost ledger — in-product.

Full set: User Governance · Global Settings · Business Config & Admin · Feedback Triage · RBAC · AI Agents Policy · Audit Logs · AI Chat Logs · Operational Logs · ETL Pipeline · System Health

Testing technical

455 test files across five layers

LayerWhat it protects
Domain modelThe pure financial engine — margin, discount compounding, freight modes, break-even, projection
API contractRequest shape → response shape; permission gate → HTTP status. Named *.contract.test.js because they encode contracts other systems depend on
Frontend componentNot pixels — logic. Does the ETL tab appear for sap_exporter? Does the lock banner render when the FY is locked?
Authorization harnessOracle vs. executor, golden + pairwise matrix, quarantine with expiry. Nightly at 02:00 UTC
AI serviceai-core contracts, extraction, and SSE event-shape validation — the stream format is a contract the frontend depends on

I did not write these because I'm disciplined. I wrote them because of the next slide.

Incidents became rules

Every contract clause has a scar behind it

What happenedWhat it became
An AI agent "simplified" a component and silently deleted 200 lines NEVER OMIT CODE — in the agent contract, in capitals
One ESM import crossing a function-folder boundary crashed the entire Functions worker — not one endpoint, all of them runtime-boundary.contract.test.js
Months later: contract code imported a workspace package that resolved in the monorepo but was absent from the deployed package Byte-parity mirroring + guard extended to reject deploy-absent imports
Eight formula bugs in one day, each producing plausible-looking numbers financialModel.js became a protected file
Different scripts used different partition-key paths for the same containers One path (/pk) everywhere; one config file as sole source of truth

Source resolution is not package resolution. The exact deployment artifact must load against only its own production dependency graph before it ships.

My favourite bug, because it is so quiet
findDue(): WHERE c.enabled = true

AgentDefinitionV1 has no enabled field.

The predicate never matched. No scheduled agent run could ever fire.

The weekly briefing was scheduled. The system reported no errors. Nothing happened.

There is no exception to catch when your filter silently matches nothing.

Supply chain technical

Freeze the deploy graph

  • Deploys install the exact graph CI validatednpm ci against a committed lockfile
  • Any version move shows up as a reviewable lockfile diff in a PR, not invisibly at deploy time
  • Four surfaces — web host, api, ai-service, ETL Python — are separate deploy units, never co-mingled

The guiding principles, verbatim

Pin what works; don't chase latest.
Subtraction beats migration — removing an unused dependency shrinks the audit surface with zero functional risk.
One surface, one wave, one PR.

That is the core defence against slopsquatting and surprise upgrades.

Act eight

What I got wrong

A self-assessment that only lists strengths isn't one

June 2026 — formal AI/agentic maturity audit

Level 3 of 5

DimensionNowTarget
Grounding & provenance44
Tool use & capability management44
Human-in-the-loop & clarification44
Agent runtime & execution model34
Model orchestration & reliability34
Guardrails, safety & governance34
Evaluation (agent output quality)24
Observability & telemetry24
Business context & memory24
Multi-agent & durable workflows23

How it was produced

A multi-agent workflow mapped 12 AI subsystems against ~60 primary sources on the 2025–26 state of the art — then adversarially verified every claimed gap back against the codebase.

13 of 26 gap-claims turned out to be overstated. The code already had a partial implementation, or the "gap" was a deliberate documented decision.

The sharpest criticism in it

A recurring "documented-but-not-enforced" risk. Several controls are described in contracts but not mechanically verified.

Writing the rule is the easy half.

3–4 on the hard, differentiating dimensions.
2 on the operational hygiene the industry standardised in 2025.

Most enterprise teams are the mirror image:
decent ops tooling bolted onto shallow grounding.

The backlog

What's next

Near — close the hygiene gap

· Agent-output evals. World-class contract tests, zero quality evals. Golden questions, scored grounding + citation coverage, in CI.
· OpenTelemetry gen_ai.* spans. Bespoke correlation logging works; it doesn't interop.
· Durable execution for the interactive chat loop.
· Turn declared controls into enforced ones.

Medium — product

· Composer V3. Invert the canvas: artefact as viewport, chat as side rail. Plus a lineage chip, plan-mode preview and per-section execution badges.
· Material Cost Service. A cost needs an auditable basis, effective date, coverage and evidence — not just a number.
· Stock beyond gross_buildable — reservation-aware free stock.
· A memory layer. There is none today.

Long — the big one

Intersection Intelligence. A map-first digital twin of every road junction in South Africa — signalised, roundabout, stop-controlled — with confidence scores and a proprietary overlay of deployed controller make/model.

Replaces municipality-level market analysis (aggregate, stale) with intersection-level reality (spatial, refreshable).

86 locked design decisions.
The design decision I'd want a CTO to notice

The most ambitious thing on the roadmap was designed, from the first line, to be switchable off.

Intersection Intelligence ships as a standalone Fastify service on its own App Service plan, mirroring the AI service, because, in the owner's words: "if it does not work, it should be switched off without breaking the core of the app."

A kill-switch ladder, its own operator console, and a proxy line in the host that can simply stop resolving.

…and then I retried it, and found a real bug

The trust machinery worked. The engineering had a hole.

What happened

I re-ran the Jira question for this edition. The tools timed out again — three attempts out of three.

This time it degraded better: it fell back to the Roadmap page KPIs (same live Jira data, already in Cosmos), gave a real quantified answer on throughput and WIP, and refused only the part it couldn't source.

"22 items have been stuck for more than 90 days. I attempted to pull the ranked list but the Jira query timed out this turn — I don't have the individual items, so I won't guess at which ones are oldest."

Root cause

TOOL_TIMEOUT_MS = 5000 is the default tool budget.

Production tools were recognised as slow and given a 30s override. No Jira tool was ever added to that map — so a cross-internet call to Atlassian Cloud gets the same five seconds as an in-region Cosmos point read.

The fix isn't a bigger number

Derive the timeout from a declared executionClass on each tool, and add a contract test asserting every tool making an outbound non-Azure call is declared external-api.

That turns a silent 5s default into a build failure.
A pattern worth naming

Beware the opt-in registry.

Two of this platform's real bugs are the same shape:

The missed model tier

A model matching no tier pattern lands in other — and auto mode never selects it. "…this is exactly how claude-fable-5 was missed."

The Jira timeout

A tool absent from the override map silently inherits a 5s budget meant for in-region reads.

Derive from a declared property; test the derivation. A map that new items must be added to will eventually not be.

Act ten

What building this
actually taught me

Evidence, not adjectives

The claim

Not "I know about AI agents"

I have made every architectural decision that determines whether an enterprise agent platform is trustworthy, and I can show you the specific incident that taught me each one.

What a CV usually claims

"Led AI initiatives" · "delivered a GenAI pilot" · "familiar with RAG and agents".

Exposure. Cheap in 2026.

What this is instead

A production system with named failure modes, a published maturity score, and a bug report I wrote about my own platform three slides ago.
Competency map exec

Ten things this platform is evidence of

CompetencyThe evidence
AI strategy under constraintBridge around an un-migratable ERP instead of a seven-figure programme; single-digit monthly DB cost
Knowing when not to use the fashionable patternRefused RAG for numbers; cut a sub-agent swarm, a custom DSL, a parallel data plane and a second chat runtime
Grounding & the economics of trustStructural citation extraction; provenance + confidence + model on every answer; it refuses when tools fail
Agentic architecture & vendor riskOne file imports the SDK; swap cost stated as ~1 week; four SDK gotchas known from being cut by them
Model portfolio managementAuto-discovery, tier policy, health circuit breaker with demotion and recovery — one resolver for the whole platform
Governance that is enforced, not declaredPolicy in tool handlers, not prompts; built-ins denied by construction; quotas, audit, validation oracle
Deterministic / probabilistic decompositionRanking, thresholds, scope and identity are code; identity resolution shipped before the agent feature
Evaluating your own programme honestlyLevel 3/5 published with the 2s visible; 13 of 26 gap-claims adversarially disproved
AI-augmented delivery as an operating model1,184 commits solo; leverage came from contracts, not prompting; two-model pairing with human arbitration
Product judgement in an AI surfaceThe LLM has no publish tool; V1-indistinguishability let a new UX ship with zero blast radius
The one-paragraph version

I built and operate a production agentic platform over a legacy ERP: 83 typed tools, 75 agent capabilities, a scheduled agent framework on the Claude Agent SDK with policy enforced at the tool-handler boundary, structural citation extraction, an independent validation oracle, per-tenant cost quotas, model auto-discovery with a health circuit breaker, and a three-layer authorization model differentially tested by a decision oracle every night.

I can tell you which of those decisions were load-bearing, which were fashion, and which one I would sequence differently, because I have the incident behind each.

For anyone hiring in this space

The competency map

Self-taught, self-rated, and deliberately not all fives

Competency map 1 = aware · 3 = working · 5 = shipped & operating in production

Where I actually am

Core agentic engineering
Prompt & system-prompt architecture5
Tool / function-calling design5
Conversational chat engine (multi-turn, SSE)5
Agent harness & orchestration5
Claude Agent SDK (production)5
MCP — Model Context Protocol4
Multi-agent / durable workflows2
Trust, safety & operations
Grounding, citations & provenance5
AI governance & guardrails4
Security & authorization for AI5
Model routing & multi-model ops4
Cost control / FinOps for AI4
Evals & output-quality measurement2
AI observability (OTel gen_ai.*)2
Platform, data & delivery
Data engineering for AI (ETL, semantics)5
AI product design & explainability UX5
AI-augmented software delivery5
Cloud architecture (Azure, serverless)4
RAG / vector retrieval2
Memory / long-horizon context2
Fine-tuning / model training1

Every rating above is defended by a named artefact in this deck. The next slide shows the evidence for the fives — and why three of these are deliberately low.

The evidence behind the fives

Each rating maps to something in this deck

CompetencyWhat backs it
Prompt & system-prompt architecture"One truth layer, two surfaces" — chat and agent prompts composed from shared blocks so drift is structurally impossible; 8 non-negotiable agent conduct rules
Tool / function-calling design83 typed tools across 13 groups, two-layer routing (chat profile / tool profile), hard budgets: 15 calls, 10 iterations, 300s ceiling
Conversational chat engineBuilt from scratch: multi-turn, SSE streaming with a frozen event contract, 3-layer cache, abort/retry, deterministic clarification gate
Agent harness & orchestrationScheduled agents in production: definitions as typed documents, scheduler, run store, citations, quotas, validation oracle, Needs-Review queue
Claude Agent SDKOne-file adapter boundary; disallowedTools defence in depth; PreToolUse/PostToolUse/Stop hooks; four SDK gotchas known from being cut by them
Grounding & provenanceStructural citation extraction from tool returns; snapshot-dated citation chips; provenance + confidence + model on every answer; it refuses when tools fail
Security & authorization for AI3-layer authz enforced in tool handlers, not prompts; 12 roles / 34 permissions; decision oracle differentially testing it nightly
Data engineering for AI10 SAP connectors with no-loss reconciliation; 170-entry insight catalog as a semantic layer the model reads
AI product design & explainabilityInsight Explainer drawer (what / sources / how-calculated / "Ask AI About This"); confidence chips; clarification UX
AI-augmented software delivery1,184 commits solo in 8 months; contracts-not-prompts operating model; two-model pairing with human arbitration
And the low ones

Why three of those are twos — and one is a one

Evals & output-quality measurement — 2

World-class contract tests; until recently zero quality evals. I built the hard part first and left the layer the industry standardised in 2025. It is the top item on my backlog, and the first thing I'd build at a new company.

AI observability — 2

Solid bespoke correlation logging; no OpenTelemetry gen_ai.* spans, so it doesn't interoperate with anything.

RAG / vector retrieval — 2

I deliberately refused RAG for structured financial data and can defend that decision in detail — but I won't claim depth I don't have. If your problem is genuinely document search, hire someone with a 5 here.

Fine-tuning / model training — 1

I have not trained or fine-tuned a model. My leverage has been architecture around frontier models, not producing them.

A matrix with no low scores is a marketing document. These four are exactly what I'd want to be hired to go and close.

If you're hiring

Everything on the previous slides is self-taught, built in production, against real money, and documented well enough that you can audit the claims.

Best fit

Head of AI · VP AI Engineering · AI Platform / Applied AI Architect · Chief AI Officer (mid-market) · CTO (mid-market)

What I'd bring in month one

A benchmark of your AI estate against the current state of the art, with an evidence-cited, prioritised backlog — the same exercise I ran on my own platform, which disproved 13 of its own 26 findings.
The two interview questions I'd ask myself

"What would you do differently?"

Sequence evals before the agent framework, not after. I have world-class contract tests and, until recently, zero output-quality evals — I built the hard part first and left the hygiene layer the industry standardised in 2025.

And identity resolution even earlier than I did.

"What is the risk in what you built?"

Single-instance assumptions in the model catalog and health breaker · no memory layer · a hand-rolled, non-durable interactive chat loop · and the "documented-but-not-enforced" gap.

Every one is on the backlog with a named fix — and I'd rather be asked about them than have them found.
What actually changed

Outcomes

Speed of answer

Emailed spreadsheets → a live operating picture. SAP export to dashboard in ~30 seconds.

Speed of decision

Budget revisions were a quarterly event. Now a director drags a slider and sees PBT move in under 16ms.

Breadth

One question surface across engineering, manufacturing, stock, procurement, contracts and finance.

Questions nobody had time to ask

Scheduled agents produce cited briefings at 08:00 — margin drift, customer health, debtor exposure — without anyone requesting them.

Cost

Single-digit monthly database spend. Built for a fraction of one SAP consulting engagement.

Trust

The system says "I don't know" when its tools fail — which is why people believe it when it doesn't.
If you take eight things

Lessons

  1. Build the calculation engine first, as a pure function, and test it before you write a component.
  2. Give agents tools, not documents. For structured financial data a typed tool over a partition-targeted query wins on every axis.
  3. Enforce policy where the data is, not where the prompt is. Anything enforced only in a system prompt is a suggestion.
  4. Keep the deterministic layer deterministic. Ranking, thresholds, scope and identity are code. Let the model do the writing.
  5. Sequence the boring dependency first. Identity before Movement Pack; Movement Pack before agents.
  6. Prove conservation before you publish. No-loss reconciliation at the boundary beats reconciliation reports afterwards.
  7. A silent no-op is worse than a crash. Assert the observable outcome, not that your code ran.
  8. Write contracts, not prompts. Every rule came from a specific incident — which is exactly why agents follow them.
Method note — optional, drop if short on time

How these screenshots were made

Every image is the live production system with real data on screen. A redaction layer is injected before capture; capture loops, redact, verify, repeat, and writes a PNG only once the page verifies clean.

It failed eight times first

1. The checker shared the redactor's blind spot — both scanned leaf elements, so prose in <li><strong>Facts:</strong>…</li> passed through while the check said "clean".
2. Pattern matching was the wrong model — an answer named five customers in prose with no digits, so nothing matched.
3. I fixed the leaf-only bug, then immediately rebuilt it in the new pass.
4–6. Form values aren't text nodes · names that exist only in the page's own data · tenant identifiers that look like ordinary product codes.
7. The 25-character prose threshold never described the risk — a 17-character Jira title in a model-rendered table slipped under it.
8. Verified clean, captured dirty. A clipped capture resizes the surface, Recharts re-renders its axes, and real currency was painted into the PNG after the check passed. The run logged $=0. The file leaked.

The rule that worked

You cannot enumerate what sensitive text looks like.

So AI output became sensitive by default — and, after failure 8, the assertion straddles the capture: shoot, re-verify, re-shoot if the page moved, so the number reported describes the file on disk rather than the DOM a moment earlier.

A test that agrees with your bug is worse than no test — it converts an unknown risk into false confidence.

Same class of error as the silent scheduler no-op. My redactor and my verifier were written by the same mind, from the same mental model, on the same afternoon, so they failed together.

The first ten minutes of every conversation

FAQ

Four decisions people push back on —
and the answer to each.

LangChain · one model provider · a document database · no evals

FAQ 1 most common question

"Why didn't you just use LangChain?"

ProductWhat it gives youWhat I did insteadVerdict
LangChain (OSS)Model / tool / prompt abstractions across providersAnthropic Messages API + frozen AiProviderTransportValue is portability. I chose one provider.
LangGraphGraph orchestration, checkpointing, durable executionClaude Agent SDK (agents) + own loop (chat)Competitive on the merits. Better durability than mine.
LangSmithTracing, datasets, trajectory evals, prompt versioningBespoke correlationId log · 1,100+ contract tests · zero output evalsStrongest case to adopt.
LangGraph PlatformManaged agent deployment and persistenceFastify on an App Service I already pay forOne region, twelve users.
LangMemLong-term agent memoryNo memory layer (scored 2/5)Real gap — not obviously a framework problem.

Primary reason, stated first: the point was to learn this from the ground up. A framework teaches you its opinion about a tool loop, not what a tool loop is.

FAQ 1 continued

The agent loop was never the hard part

What the SDK gave me

Tool loop · sub-agents · hooks · permissions · sessions.

The research note that locked this in: the SDK "saves us from rebuilding the runtime / orchestration layer" — v0 became days of work, not weeks.

What no framework could do for me

ABAC enforced inside tool handlers, before a row reaches the model · the Movement Pack so the model narrates rather than calculates · a 170-entry governed insight catalog · SAP reconciliation that refuses to publish on row-count drift.

A framework would have abstracted the tractable part of the problem and left the difficult part untouched.

Frugality is a real constraint, not a flourish. ~12 internal users on a shoestring. LangSmith Plus is $39/seat/month + trace overages; a typical LangSmith + LangGraph Platform setup lands $175–375/month before model spend.

Where I change my mind, and it's the same place twice. My two lowest scores are evals 2/5 and observability 2/5. That is exactly LangSmith's core competence. The move is to close a named gap rather than adopt a framework wholesale: adopt an eval/trace layer against interfaces that are already frozen and contract-tested.

FAQ 2 the lock-in question

"Why go all-in on Claude? Where's the router?"

There is a seam — not a router

AiProviderTransport is a frozen, asserted contract: the orchestration layer talks to a provider, not to Anthropic. providerCapabilities.js says so in its own header — adding OpenAI or Gemini "would add sibling files with the same classify() / resolveThinkingConfig() contract."

One implementation exists. The audit scores this 3/5 and names it: "sophisticated single-provider routing; no cross-provider failover."

What I do have inside the bet

Model auto-discovery every 12h against /v1/models · tier classification · a health circuit breaker that demotes and recovers automatically. Failover exists — within the family, not across vendors.

This is a resourcing decision, not an architectural belief. A provider-abstraction layer is a tax paid daily on every feature, against a risk that hasn't materialised. One person, running an entire business: you pick a platform and move.

So I went vertically integrated, not just the LLM API, but the Agent SDK and the MCP tool surface. That verticality is why the agent framework took days.

If Anthropic goes down, the AI degrades. The platform doesn't, every number comes from Cosmos. The model only narrates.

Revisit triggers: sustained availability problems · a pricing shift · a capability that exists only elsewhere. The seam is there so that day is an implementation, not a rewrite.

FAQ 3 the architect's question

"Why Cosmos DB and not a relational database?"

It began as a cost decision on a shoestring proof of concept, and stayed for engineering reasons.

  • Schema evolution without migrations. Fiscal-year structures, contract shapes and connector outputs changed repeatedly across eight months. Not one needed a migration window.
  • The ETL has no impedance mismatch. SAP exports arrive as records and land as documents. No ORM, no staging schema. That is most of why the ETL is ten connectors and not a data-engineering department.
  • The partition key is the security boundary. Everything partitions on /pk carrying BU + fiscal year, so hasBuScope is enforced in the storage layout — not just a WHERE clause someone can forget.
  • Operational features I'd otherwise build. TTL gives one-year audit retention as a field, not a cron job. ETags give optimistic concurrency without a locking strategy.
  • One dependency, one region, an emulator. npm run dev:up and there is no second database technology to run, secure, back up or patch.

The counterpoint

Cross-container joins and aggregations are my problem — hand-written and app-side. A relational database would make ad-hoc analytical queries dramatically easier.

Which is exactly why the AI reads a governed semantic layer instead of writing SQL. The database choice and "tools, not text-to-SQL" are the same decision seen twice.

FAQ 4 — rapid fire the uncomfortable ones

Answers I don't get to soften

QuestionThe answer
Tools instead of text-to-SQL?A generated query cannot be authorized. All 83 tools enforce ABAC inside the handler before a row reaches the model, and return the citations the hooks accumulate. Text-to-SQL inverts that — the model composes the access path and you audit afterwards.
How do you stop it inventing numbers?Structurally. Deterministic handlers emit citations and datasets; PostToolUse hooks accumulate them; the model is not the source of provenance. Honest caveat: the ClaimV1 refine gate currently passes trivially — a guarantee that always passes is not a guarantee. It's documented and on the near list.
Weakest part of the platform?Durable execution. The dispatch queue has no consumer, so a restart orphans in-flight runs; the scheduler's 5-min stale lease is shorter than the 15-min run timeout — a latent double-fire. A documented Phase 4 deferral, and the one place LangGraph's checkpointing would have bought me something real.
No evals? In a financial app?Correct, and it tops the backlog. What exists is the opposite trade: ~1,100+ deterministic contract tests, frozen schemas, a named regression test per incident. Every model call in every test is a deterministic fake — strong guarantees about behaviour, none about answer quality. The probabilistic layer goes on top of the deterministic gate.
What would you do differently?OpenTelemetry from day one — retrofitting gen_ai.* spans over a bespoke scheme is pure rework. Evals before feature twenty. And enforce before declaring: the sharpest audit finding was controls that were persisted yet inert — an admin sets restrictedFields and believes it works.

Documentation describing runtime behaviour you haven't written is worse than none — for the same reason a test that agrees with your bug is worse than no test.

Thank you

Contracts,
Not Prompts

The bottleneck was never typing speed.
It was clarity of specification.

Questions?