# Jason Lovell: At the frontier. In the code. Frontier AI research, solution design and engineering. I research what AI can do next, then build the systems and experiments to find out what holds up. Day role: Frontier AI strategy and enterprise innovation, PwC Intelligent Enterprise. Independent work, separate from my role at PwC. ## Studies Ten studies, each with its result and where its evidence stops. - [MCP Continuation Replay](https://jasonlovell.ai/work/mcp-continuation-replay): The action succeeded. The reply vanished. What happens next? Demonstrated: recovers, no duplicate. 24 Sep 2026. - [Scaffold Arena](https://jasonlovell.ai/work/scaffold-arena): How much of an agent’s performance comes from everything around the model? Status: instrument built. 14 Sep 2026. - [RealityBridge](https://jasonlovell.ai/work/realitybridge): A simulator says it worked. Would the real service agree? Finding: 3 of 16 diverged. 24 Sep 2026. - [ShadowSkillBench](https://jasonlovell.ai/work/shadowskillbench): When an agent learns a procedure, what else does it carry forward? Status: methods preview. 22 Sep 2026. - [Living Context](https://jasonlovell.ai/work/living-context): What should an agent revisit when the world changes? Demonstrated: refresh plans; learned extension inconclusive. 24 Sep 2026. - [Jev Decision Lab](https://jasonlovell.ai/work/jev-decision-lab): What changes when a model returns a typed judgment? Status: workbench. 17 Sep 2026. - [GhostTrace](https://jasonlovell.ai/work/ghosttrace): Does a behavioural signal survive repeated self-distillation? Finding: toy tier only. 21 Sep 2026. - [No Model Floor](https://jasonlovell.ai/work/no-model-floor): Does this benchmark even need a model? Finding: headroom 0. 21 Sep 2026. - [CascadeShift](https://jasonlovell.ai/work/cascadeshift): Are we measuring the model, or the configuration? Finding: interval includes 0. 21 Sep 2026. - [Nano interpretability](https://jasonlovell.ai/work/nano-interpretability): When is expensive interpretability worth its cost? Finding: wins only when distributed. 9 Jun 2026. ## Use-case design - [Agent-Ready Checkout Gateway](https://jasonlovell.ai/projects/agent-ready-checkout-gateway): Alpha reference gateway that lets shops accept consented orders from AI agents. Jun 2026. ## More - Builds by area: https://jasonlovell.ai/builds.md - Releases (specifications, a white paper, research software): https://jasonlovell.ai/releases.md - About and career: https://jasonlovell.ai/about.md - One-page summary: https://jasonlovell.ai/summary - For agents (MCP server, WebMCP tools, Markdown): https://jasonlovell.ai/agents - The studies, as music (two decks to mix them in the browser): https://jasonlovell.ai/listen # Jason Lovell Frontier AI research, solution design and engineering. I research what AI can do next, then build the systems and experiments to find out what holds up. Based in Austin, Texas. More than twenty years in emerging technology, leading teams in the UK and US. ## Today Frontier AI strategy and enterprise innovation, PwC Intelligent Enterprise (PwC US since September 2025). Independent work, separate from my role at PwC. ## How I build I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. ## What I care about now The space between an emerging capability and a redesigned enterprise. That is where evidence, judgment and thoughtful implementation matter just as much as the technology. ## The pattern I keep seeing The arc tends to repeat: loud enthusiasm, then a small number of teams quietly figuring out what is real and shipping it. ## Career - From 2004: **Mobile.** I started in mobile, when the consumer-tech frontier moved one new category at a time: operators first (O2, T-Mobile, EE and TalkTalk), then product and portfolio leadership at the device maker KAZAM. - **Connected devices and immersive computing.** IoT, wearables, tablets, connected home, smart audio, then VR and AR, including senior product management for VR, wearables and SmartThings at Samsung. Leadership teams across sectors, UK first, then the US. - 2016 to 2021: **Founding Captivate.** I founded Captivate, an applied-innovation consultancy for XR, AI and emerging technology, and ran it for five years. Clients included Synthesia, Gucci, NBC Universal and Lloyd's of London; for Synthesia I was VP of global partnerships in 2018 and 2019. That is where AI caught me: early generative models for synthetic media, alongside immersive work. In parallel, from 2017 to 2018, I led brand partnerships across EMEA for the VR company Jaunt. - 2017: **Pulled into AI.** AlphaGo, OpenAI's Dota work and AlphaGo Zero. I have read the field closely since. - 2019 to 2022: **Emerging technology and XR strategy, PwC UK.** I led strategy and client work on XR and spatial computing where it met AI and data, with a specialist team building training simulations and spatial analysis. - 2022 to 2024: **AI and emerging technology, PwC US.** An international assignment to PwC US in Austin. It began with immersive technology and grew into AI and generative AI: commercialising both for clients and co-developing solutions with technology alliance partners. - 2024 to 2025: **Technology business partner, PwC UK.** Back in the UK, the link between PwC's central technology team and its consulting teams: shaping AI and generative-AI propositions for clients, and bringing together the people to deliver them. - Sep 2025 to now: **Frontier AI strategy and enterprise innovation, PwC US.** A move to PwC US. First, agentic AI delivery for clients; now, with PwC Intelligent Enterprise, strategy, research and hands-on solution development: taking an idea from an early research question to a testable solution and a credible path to scale. - 2026: **Building in public.** Specifications from February 2026, an interpretability study in June, then nine more studies in September. ## Contact - Email: hello@jasonlovell.ai - Contact form: https://jasonlovell.ai/#contact (topics: a role or research collaboration, a question about a study, an AI idea to explore or prototype, a specification) - GitHub: https://github.com/jlov7 - LinkedIn: https://www.linkedin.com/in/jalovell/ - ORCID: https://orcid.org/0009-0001-6300-9155 # The action succeeded. The reply vanished. What happens next? MCP Continuation Replay: Reference implementation. Published 24 Sep 2026. Question: Does the state hold up? Scope: 1 synthetic tool · MCP SDK over stdio · SQLite record Page: https://jasonlovell.ai/work/mcp-continuation-replay Code: https://github.com/jlov7/mcp-continuation-replay A continuation and replay reference for the awkward moment when an agent cannot tell whether a tool finished its work. ## Demonstrated The reference makes completed work recoverable after a lost reply. A retry can recover the recorded result instead of blindly creating another issue. ## In practice A lost reply means unknown, not failed. Ask the system of record before retrying anything with side effects. ## Why it matters A missing reply leaves an agent with an awkward choice: retry and risk doing the work twice, or stop without knowing whether it finished. That uncertainty is the problem this project takes on. ## The design decision A retry reuses the original operation ID, and a status read never authorises a fresh write. The easier design gives each attempt a new ID. The repository keeps a deliberately unsafe control that does exactly that, and it creates a second issue. The price of refusing it: when the record cannot settle what happened, the outcome stays unknown instead of being guessed. ## What I built A SQLite-backed execution record and lost-reply recovery path, exercised through the MCP SDK over stdio with a synthetic issue-creation tool. ## What was recorded | Measure | Value | Note | | --- | --- | --- | | Interruption | After commit | The fixture crashes after the work has been recorded. | | Recovery | Status readback | The restarted caller uses the same operation identity. | | Database assertion | 1 physical issue | The test checks that recovery has not created a second issue. | ## Evidence - [Read the recovery test](https://github.com/jlov7/mcp-continuation-replay/blob/main/tests/test_next_recovery.py): Crash after commit, restart, status-only recovery and physical row assertions. - [Inspect the lost-reply demo](https://github.com/jlov7/mcp-continuation-replay/blob/main/scripts/demo_lost_reply.py): The frozen case-03 scenario and its retained wire transcripts and SQLite checks. ## Where the evidence stops This is one bounded tool and recovery scenario. It is not a production guarantee for arbitrary MCP tools or external side effects. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # How much of an agent’s performance comes from everything around the model? Scaffold Arena: Research workbench. Published 14 Sep 2026. Question: Can we trust the score? Scope: Beta workbench · synthetic fixtures and mock runs Page: https://jasonlovell.ai/work/scaffold-arena Code: https://github.com/jlov7/scaffold-arena A workbench and experimental protocol for studying orchestration, tools and control flow as part of the system being evaluated. ## Status The workbench makes scaffold choices explicit before any comparison. So far it runs on local synthetic fixtures and mock runs; there is no measured result yet. ## In practice Treat the scaffold as part of what you evaluate. Two models compared on different scaffolds is not a model comparison. ## What I built A research workbench that makes scaffold choices explicit, with a protocol for comparing quality, reliability, cost and latency. ## Where the evidence stops A beta research workbench, not a completed comparative live study or a validated ranking of agent scaffolds. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # A simulator says it worked. Would the real service agree? RealityBridge: Comparison harness. Published 24 Sep 2026. Question: Does the state hold up? Scope: 16 evaluation cases · 6 operations · pinned local Gitea Page: https://jasonlovell.ai/work/realitybridge Code: https://github.com/jlov7/realitybridge A harness that compares responses and resulting state against a pinned local reference service, and keeps the disagreements visible. ## Finding 3 of 16 new cases diverged from the pinned Gitea reference: two in the response, and one where both sides reported success but the resulting state differed. Checking the response alone would have missed it. ## In practice A mock can agree on the reply and still leave a different state behind. ## Why it matters An agent can learn to succeed in a simulator whose behaviour differs from the service it represents. A matching success response can hide that difference, so this project looks underneath it. ## The design decision The observation contract decides what counts as the same. Rows can be sorted, but duplicates are kept. Raw database IDs are not compared literally, and exact timestamps are dropped while their presence stays testable. A looser contract would agree more often and detect less, so I chose the stricter one and wrote it down. Changing it changes what the harness can see. ## What I built A bounded comparison harness for six operations against a pinned local Gitea reference, with recorded evidence and offline replay. ## What was recorded | Measure | Value | Note | | --- | --- | --- | | New evaluation cases | 3 of 16 diverged (differs) | E04, E06 and E09. The other 13 agreed. | | E06 responses | Agreed | Both paths reported success for the duplicate-label operation. | | Simulator state | Row overwritten (differs) | The baseline simulator replaced the duplicate label row. | | Reference state | Distinct rows | Pinned Gitea retained both. Response agreement missed this. | ## Evidence - [Inspect the recorded results](https://github.com/jlov7/realitybridge/blob/main/docs/research-completion-2026-09-22/RESULTS.md): The frozen local E06, E04 and E09 comparisons, including the state divergence. - [Read the comparison architecture](https://github.com/jlov7/realitybridge/blob/main/docs/ARCHITECTURE.md): Execution, reference readback, response comparison and the independent state oracle. ## Where the evidence stops The evidence covers the specified local operations. It does not establish platform-wide simulator equivalence or independent replication. The reference and evaluator code were AI-authored and have not had independent human review. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # When an agent learns a procedure, what else does it carry forward? ShadowSkillBench: Software & methods preview. Published 22 Sep 2026. Question: What carries over? Scope: Software and methods preview · synthetic procedures Page: https://jasonlovell.ai/work/shadowskillbench Code: https://github.com/jlov7/ShadowSkillBench A synthetic testbed for investigating learned procedures and the boundaries of authorised action. ## Status The release is the experimental machinery: a synthetic testbed for asking whether a learned procedure stays within authorised action. There is no confirmatory result yet. ## In practice Knowing how to do a task and being allowed to do it are separate questions. Test them separately. ## What I built Research software and a methods package for controlled experiments on learned procedures, with explicit separation between the testbed and any confirmatory result. ## Where the evidence stops A software and methods preview. No confirmatory provider study, model ranking or production finding is claimed. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # What should an agent revisit when the world changes? Living Context: Reference implementation. Published 24 Sep 2026. Question: Does the state hold up? Scope: In-memory, model-free reference Page: https://jasonlovell.ai/work/living-context Code: https://github.com/jlov7/living-context Revision-aware context snapshots, text retrieval and advisory refresh plans. ## Demonstrated Tracking revisions explicitly is enough to propose which context to refresh, with no model involved. The learned-context extension came out negative or inconclusive. ## In practice Stored context goes stale. Track what changed and plan what to re-read, instead of keeping everything. ## What I built An in-memory, model-free reference for tracking revisions and planning which context to refresh. ## Where the evidence stops An in-memory reference, not a deployed model-serving or replicated distributed-memory system. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # What changes when a model returns a typed judgment? Jev Decision Lab: Experimental workbench. Published 17 Sep 2026. Question: Where does judgment sit? Scope: Local workbench · authored business cases Page: https://jasonlovell.ai/work/jev-decision-lab Code: https://github.com/jlov7/jev-decision-lab A local lab for exploring typed answers, probabilities and policy in authored business cases. ## Status The lab puts a typed Jev judgment and its probability beside the policy code that acts on it, so each can be inspected on its own, using synthetic business cases. ## In practice Keep a model’s judgment and the policy that acts on it separate, so each can be inspected. ## What I built A teaching and experiment workbench around TypeSafe’s Jev judgment model, with synthetic cases and inspectable decision records. ## Where the evidence stops An experimental workbench. Authored cases and owner-recorded observations do not establish business ROI or production readiness. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # Does a behavioural signal survive repeated self-distillation? GhostTrace: Research software. Published 21 Sep 2026. Question: What carries over? Scope: Toy tier and local-LLM tier, reported separately Page: https://jasonlovell.ai/work/ghosttrace Code: https://github.com/jlov7/GhostTrace Controlled experiments on behavioural signal decay across recursive self-distillation. ## Finding In the toy tier the behavioural signal decays with a measurable half-life. The local-LLM tier came out negative, which marks where the result stops. ## In practice A behavioural result from a toy setting may not carry over to a real model. Test it where you deploy. ## What I built A toy-tier experiment and a separate local-LLM investigation, with distinct evidence boundaries. ## Where the evidence stops The toy finding is not a recursive LLM transfer law. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # Does this benchmark even need a model? No Model Floor: Offline diagnostics. Published 21 Sep 2026. Question: Can we trust the score? Scope: 328 configurations · 1,512 draw classes · offline Page: https://jasonlovell.ai/work/no-model-floor Code: https://github.com/jlov7/no-model-floor Diagnostics for benchmark floors and whether an adaptive system has room to improve. ## Finding Under the stated controls, a fixed action behind a guard reached the benchmark ceiling. That left no room for a model to show it could adapt. ## In practice If a fixed baseline already reaches the ceiling, the benchmark cannot credit a model. ## Why it matters Before crediting adaptation for a strong result, it helps to ask how far a fixed policy can get. If a simpler control reaches the ceiling, the benchmark has no room left to show an advantage. ## The design decision I removed the model before comparing models: fixed policies in its place, scaffold left in, then rescore. The usual order compares models first and credits the best one. Running the substitution first showed that the guards already encode the answer rule. I then checked all 328 settings in the control family rather than a sample, so the zero-headroom result covers the whole family. ## What I built An offline synthetic benchmark and controls for measuring the performance floor before attributing gains to adaptation. ## What was recorded | Measure | Value | Note | | --- | --- | --- | | Control coverage | 328 | Configurations in the finite synthetic family. | | Draw coverage | 1,512 | Draw classes covered by the result artifact. | | Ceiling shortfalls | 0 | Across the integer floors in this family, with one universally dominating configuration. | ## Evidence - [Inspect the result artifact](https://github.com/jlov7/no-model-floor/blob/main/evidence/theorem.json): The machine-readable finite-family result and its coverage. - [Read the exhaustive tests](https://github.com/jlov7/no-model-floor/blob/main/tests/test_theorem.py): Configuration and draw coverage, unique universal dominance, ceiling and guard assertions. ## Where the evidence stops A bounded synthetic diagnostic, not a general claim about adaptive systems or production safety. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # Are we measuring the model, or the configuration? CascadeShift: Synthetic case study. Published 21 Sep 2026. Question: Can we trust the score? Scope: Synthetic case study · offline measurement checks Page: https://jasonlovell.ai/work/cascadeshift Code: https://github.com/jlov7/CascadeShift-Research A synthetic study of configuration-aware tool use and the validity of the measurements around it. ## Finding The main comparison’s confidence interval includes zero, so the study claims no benefit and says so. ## In practice Hold configuration fixed, or measure it, before attributing a change in tool use to the model. ## What I built A research package examining configuration-aware tool use with offline measurement checks. ## Where the evidence stops Synthetic case-study evidence, not a proven general performance improvement. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # When is expensive interpretability worth its cost? Nano interpretability: Research software, three releases. Published 9 Jun 2026. Question: Can we trust the score? Scope: InterpBench, GPT-2-small, Gemma-2-2B and 9B · one Apple M4 Max Page: https://jasonlovell.ai/work/nano-interpretability Code: https://github.com/jlov7/nanoassembly Three small, calibrated studies of circuit and feature attribution, each graded against a known answer or the strongest cheap baseline rather than a convincing diagram. ## Finding Attribution beat the strongest gradient-free baseline only when the circuit was spread across token positions: on IOI in Gemma-2-2B by 15 to 45 points in 12 of 12 cells, with no meaningful advantage on single-token tasks. The same split held on GPT-2-small and Gemma-2-9B. Across 75 layer cells, gradient-selected circuits matched per-feature ablation to within a mean of 0.028. ## In practice Test attribution against the strongest cheap baseline first. It only earned its cost when the circuit was spread across token positions. ## Why it matters Interpretability work often ends with a circuit you are asked to trust. Exact per-feature ablation costs one forward pass per feature, so it matters whether it finds anything a cheaper score would miss, and whether the known answer it is graded against actually reproduces the behaviour. ## The design decision I grade every method against the strongest cheap baseline I can build, not a convenient weak one, and I measure the known answer instead of assuming it. The baseline decides the apparent result. Ranked by summed activation change, attribution beat the cheap score by 17 to 35 points on factual recall. Ranked by peak per-position change, almost all of that gap closed. The cost is smaller claims: when nanocircuits added a baseline built from a single forward pass, one of its two clean wins tied it exactly and dropped out. ## What I built nanocircuits grades circuit discovery on InterpBench transformers that contain a known circuit, against a leave-one-case-out structural baseline. nanofeatures and nanoassembly carry the same discipline to SAE features on GPT-2-small and Gemma-2-2B, with paired-bootstrap confidence intervals. Everything runs on one Apple M4 Max. ## What was recorded | Measure | Value | Note | | --- | --- | --- | | IOI, a distributed circuit | +15 to +45 pts | Attribution over the strongest gradient-free baseline in 12 of 12 cells. Every confidence interval excludes zero. | | Seven single-token tasks | No meaningful advantage | A pre-registered equivalence test (TOST, 5-point margin) over 63 cells: 10 small attribution wins of 6 points or less, 13 equivalent, 4 cheap-baseline wins, 36 inconclusive. | | Gradient vs exact ablation | Mean gap 0.028 | Faithfulness of gradient-selected and exact-selected circuits across 75 layer cells on GPT-2-small and Gemma-2-2B. | | Known circuits (InterpBench) | 1 of 7 cases | Two cases pass the baseline and faithful-circuit filters. Against a baseline built from a single forward pass, only case 11 still wins at node level. | ## Evidence - [Read the argument across all three](https://github.com/jlov7/nanoassembly/blob/main/when-is-interpretability-worth-it.md): when-is-interpretability-worth-it.md: the three studies as one argument. - [Read the method and exact numbers](https://github.com/jlov7/nanoassembly/blob/main/THESIS.md): nanoassembly THESIS.md: the 75-cell calibration, controls and limits. - [nanocircuits on Zenodo](https://doi.org/10.5281/zenodo.20611793): DOI 10.5281/zenodo.20611793. Circuit discovery graded against known circuits. - [nanofeatures on Zenodo](https://doi.org/10.5281/zenodo.20611795): DOI 10.5281/zenodo.20611795. When attribution beats a free baseline on SAE features. - [nanoassembly on Zenodo](https://doi.org/10.5281/zenodo.20611788): DOI 10.5281/zenodo.20611788. Multi-layer feature circuits and attribution cost. ## Where the evidence stops Small models and a finite task set. In nanocircuits two InterpBench cases plus IOI beat the structural baseline with a faithful ground-truth circuit; against a one-forward-pass behavioural baseline only case 11 and IOI still win. None of this is a claim about frontier-scale models. --- By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. # What should an agent be allowed to do with what it overhears? OverhearOps: Use-case design, Python. Repository created Sep 2025. Code: https://github.com/jlov7/OverhearOps Page: https://jasonlovell.ai/projects/overhearops An R&D system that overhears team chat and drafts dry-run fixes. ## Design choice It drafts but ships nothing by default: side effects run only in live mode with an approver's sign-off, and an uncertainty gate refuses to ship when confidence is low. ## In practice An agent that listens in on a team should draft, not act. Keep its side effects behind a dry-run default and a named approver. ## What I built - Detects actionable threads in Teams-shaped conversations and drafts plans, down to a PR diff and a Jira stub - Plan branches with a judge that picks one; replay regenerates the decisions and checks a stored hash - Dry-run by default: live shipping needs an approver's sign-off - OpenTelemetry traces, action graphs and a governance view with trace IDs and the replay hash ## Where it stops A personal R&D system built to run locally. By default it replays two Teams-shaped demo threads with deterministic model fixtures; the Microsoft Graph adapter needs tenant credentials, and live model calls need a provider key. Its guard blocked every attack in its own 8x5 prompt-injection suite, which says nothing about attacks outside that suite. --- By Jason Lovell. Independent work, separate from my role at PwC. # What would an underwriter need before trusting an agent's risk report? TerraRisk Agent: Use-case design, Python. Repository created Oct 2025. Code: https://github.com/jlov7/TerraRisk-Agent Page: https://jasonlovell.ai/projects/terrisk-agent A reference underwriting copilot for geospatial risk, with signed and traceable reports. ## Design choice Every decision the agent makes writes an action credential, a deny-by-default policy bundle sets what it may do and spend, and each report carries checksummed artefacts ready for keyless signing. ## In practice When an agent's analysis will inform a regulated decision, design the audit trail and the policy limits with the workflow, not after it. ## What I built - Natural-language questions decomposed into geospatial analysis steps - FEMA National Risk Index and BigQuery Earth Engine connectors - Action credentials for every decision, and OPA/Rego policies that deny by default, cap spend and keep PII at county level - Reports as PDF, GeoJSON and CSV with checksums and OpenTelemetry traces ## Where it stops A personal R&D reference implementation, not a product. It runs offline on synthetic fixtures by default; Google Earth AI stays stubbed until API access is available, and cloud mode needs your own BigQuery and Earth Engine access. Its evaluation harness (golden questions, ROUGE-L, join integrity) checks the reports' narrative and data joins, not underwriting accuracy. --- By Jason Lovell. Independent work, separate from my role at PwC. # What does a merchant need before an agent can buy? Agent-Ready Checkout Gateway: Use-case design, Python. Repository created Jun 2026. Code: https://github.com/jlov7/Agent-Ready-Checkout-Gateway Page: https://jasonlovell.ai/projects/agent-ready-checkout-gateway Alpha reference gateway that lets shops accept consented orders from AI agents. ## Design choice Consent goes into an append-only, hash-chained ledger before a policy hook can allow, deny or flag the order for review. ## In practice Settle what an agent may buy on its own, and what needs a person, while designing the checkout, not after an order ships. ## What I built - FastAPI gateway for intent, confirmation, authorisation and fulfilment - Append-only consent ledger with hash chaining - Policy hook for allow, deny, and review decisions - Receipt pipeline with detached provenance manifests ## Where it stops An alpha reference implementation (v0.1.0): a pattern demo, not a drop-in payment or PCI replacement. Payments go through a Stripe test-mode adapter, which does not simulate 3DS. The inventory and pricing MCP servers are mocks without full authentication, and receipts carry detached provenance manifests but no embedded C2PA signature. The default policy is a stub that allows everything and sends orders over $1,000 to review. A review decision is reported with its reasons in the authorisation response; the gateway does not pause the payment for it. --- By Jason Lovell. Independent work, separate from my role at PwC. # Builds Public repositories by area, newest first. Every one is Jason's own, built end to end. ## Use-case design - [Agent-Ready Checkout Gateway](https://jasonlovell.ai/projects/agent-ready-checkout-gateway): Alpha reference gateway that lets shops accept consented orders from AI agents. Jun 2026. - [TerraRisk Agent](https://jasonlovell.ai/projects/terrisk-agent): A reference underwriting copilot for geospatial risk, with signed and traceable reports. Oct 2025. - [OverhearOps](https://jasonlovell.ai/projects/overhearops): An R&D system that overhears team chat and drafts dry-run fixes. Sep 2025. ## Assurance and provenance - [Sentinel MCP](https://github.com/jlov7/Sentinel-MCP): R&D governance for MCP tools: every call can be authorised and replayed. Jun 2026. - [Agent HQ Guard](https://github.com/jlov7/Agent-HQ-Guard): A GitHub App that blocks agent merges until policy checks pass. Jun 2026. - [PAISL / Agent Boundary Assurance](https://github.com/jlov7/personal-ai-sovereignty-lab): A benchmark scaffold asking whether personal agents stay useful inside data boundaries. May 2026. - [Agent Assurance Case](https://github.com/jlov7/agent-assurance-case): Draft specification for signed agent-release evidence, with an offline reference verifier. May 2026. - [Damn Vulnerable Agent Asset Corpus (DVAAC)](https://github.com/jlov7/damn-vulnerable-agent-asset-corpus): Known-answer agent fixtures, vulnerable and clean, so scanner claims can be compared. May 2026. - [ProofPack](https://github.com/jlov7/ProofPack): Portable signed receipts that prove an agent run's record was not rewritten. Feb 2026. - [Runwright](https://github.com/jlov7/runwright): One pinned and scanned set of agent skills across AI coding tools. Feb 2026. - [AgentGate](https://github.com/jlov7/agentgate): An R&D containment gateway that can stop a misbehaving agent mid-run. Jan 2026. - [SKILLCHECK](https://github.com/jlov7/SKILLCHECK): A research-preview auditor for checking an Agent Skill before it is enabled. Oct 2025. ## Agent tooling - [SkillScope](https://github.com/jlov7/SkillScope): Research-only tracing of what an Agent Skill did and what it cost. Jun 2026. - [Switchboard](https://github.com/jlov7/Switchboard): An R&D sandbox: one approval and audit layer across three agent providers. Jun 2026. - [BranchLab](https://github.com/jlov7/branchlab): Replay an agent run locally, change one step and compare the outcomes. Mar 2026. - [Baton Studio](https://github.com/jlov7/baton-studio): Agent teams share a world model, passing a baton for high-stakes commits. Mar 2026. - [Meta Memory Studio](https://github.com/jlov7/meta-memory-studio): A control plane showing whether an agent's memory helps or hurts. Feb 2026. - [RLM-Lens](https://github.com/jlov7/rlm-lens): Query code and docs locally and see the path behind each answer. Feb 2026. - [Agent Director](https://jasonlovell.ai/projects/agent-director): A trace debugger that turns agent runs into repeatable eval cases. Jan 2026. ## Interpretability and model labs - [nanoIM](https://github.com/jlov7/nanoIM): Chat collapses time; nanoIM restores it. May 2026. - [nanoAWM](https://github.com/jlov7/nanoAWM): Predict an action's consequences before the agent acts, in a symbolic OS. May 2026. - [VoiceForge AI](https://github.com/jlov7/voiceforge-AI): Emotion-aware speech experiment: an LLM reads the tone before the voice speaks. Jun 2025. ## Evaluation - [ProofKern](https://github.com/jlov7/proofkern): Faster, until the baseline is fair: 4 winners against MLX, 0 against torch.compile. Jun 2026. - [MCP Interop BakeOff](https://github.com/jlov7/MCP-Interop-BakeOff): A portability lab: the same MCP tasks run across different agent runtimes. Jun 2026. - [SkillBench-PD](https://github.com/jlov7/SkillBench-PD): An R&D benchmark comparing full Agent Skill loading with progressive disclosure. Jun 2026. - [Adaptive Multi-Dimensional Monitoring (AMDM)](https://github.com/jlov7/AMDM): Near-real-time anomaly detection for agents, benchmarked on synthetic scenarios. Oct 2025. # Releases 12 releases carry a DOI or preprint identifier. A DOI records a release; it is not peer review. Full record on ORCID: https://orcid.org/0009-0001-6300-9155 - **Agent Boundary Assurance** (White paper, Jun 2026): An evidence discipline for what a local or enterprise agent accesses, remembers, transforms and sends. DOI [10.5281/zenodo.20815336](https://doi.org/10.5281/zenodo.20815336). - **Personal AI Sovereignty Lab (PAISL)** (Benchmark scaffold, Jun 2026): A local-first benchmark scaffold for testing whether a personal agent stays useful while respecting the user's data boundaries. DOI [10.5281/zenodo.20753227](https://doi.org/10.5281/zenodo.20753227). - **The nano interpretability sequence** (Research software, three releases, Jun 2026): Circuit discovery calibrated against ground truth, the per-element attribution-cost boundary on Gemma-2-2B and GPT-2, and multi-layer SAE feature circuits. DOI [10.5281/zenodo.20611788](https://doi.org/10.5281/zenodo.20611788). - **ProofKern** (Research software, Jun 2026): Verifier-first measurement for AI-generated GPU kernels. Four MLX-relative winners became zero cross-framework winners once torch.compile was added on the same GPU. [Code](https://github.com/jlov7/proofkern). - **nanoIM** (Research software, Jun 2026): A from-scratch lab for one question: if two interactions flatten to the same final chat transcript but require different next actions, what can a transcript-only model know? DOI [10.5281/zenodo.20492362](https://doi.org/10.5281/zenodo.20492362). - **nanoAWM** (Research software, May 2026): A tiny agent world model lab: compact models predict the consequences of an action before it runs, then plan against those predictions. DOI [10.5281/zenodo.20429920](https://doi.org/10.5281/zenodo.20429920). - **Agent Assurance Case** (Draft specification and verifier, May 2026): A portable, signed evidence object for agent release decisions. The verifier recomputes the verdict offline. DOI [10.5281/zenodo.20185170](https://doi.org/10.5281/zenodo.20185170). - **Damn Vulnerable Agent Asset Corpus** (Conformance corpus, May 2026): Deliberately vulnerable and deliberately clean agent release fixtures, with expected findings and Agent Assurance Case templates, for testing agent-asset assurance tools. DOI [10.5281/zenodo.20186918](https://doi.org/10.5281/zenodo.20186918). - **Tool Security Advisory** (Specification and preprint, Feb 2026): Machine-readable security advisories for MCP tools. DOI [10.36227/techrxiv.177155646.67434382/v1](https://doi.org/10.36227/techrxiv.177155646.67434382/v1). - **Skill Bundle Attestation** (Specification, Feb 2026): Deterministic bundle identity, content attestation and verification tooling for agent skills. DOI [10.5281/zenodo.18485741](https://doi.org/10.5281/zenodo.18485741). - **Tool Bill of Materials** (Specification, Feb 2026): Signed manifests that bind MCP server releases to immutable tool metadata, so a client can check what it is about to call. DOI [10.5281/zenodo.18458945](https://doi.org/10.5281/zenodo.18458945).