Skip to content

At the frontier. In the code.

Illustration: loose strands of light, what AI can do next, draw together at one point, design and build, into a single braid: what holds up. A few strands peel away in coral: what didn’t.

I’m Jason. I research what AI can do next, then build the systems and experiments to find out what holds up.

Day role Frontier AI strategy and enterprise innovation at PwC Intelligent Enterprise. The research here is my own, separate from that role.

Background 20+ years in emerging technology, UK and US

Based in Austin, Texas GitHub LinkedIn

Work

Built to find out.

The arc tends to repeat: loud enthusiasm, then a small number of teams quietly figuring out what is real and shipping it.

4 worked examples10 studies by question builds by area 12 citable releases

Both sides said success.
The state disagreed.

I built a harness that compares both the response and the resulting state against a real system: a pinned local Gitea service (an open-source Git server).

A mock can agree on the reply and still leave a different state behind.

Recorded pilot run.

Now check what each side stored.

  • responses differed: E04, E09
  • state differed: E06

E06 response: two create_label calls for p2-eval-dupe, colours 111111 then 222222. The simulator and Gitea both returned success with the same name and colour each time.

16 new evaluation cases. 13 agreed; 3 diverged. Read the recorded result

E06  responses agreed: both returned success

Recorded outcome of each new evaluation case
CaseCalls sent to bothResponseResulting state
E01create_issue, create_issue, get_issueagreedagreed
E02edit_issue, edit_issue, get_issueagreedagreed
E03add_label, add_label, get_issueagreedagreed
E04remove_label, remove_label, get_issuediffered: Gitea succeeded, the simulator returned 404agreed
E05create_labelagreedagreed
E06create_label, create_labelagreeddiffered: Gitea kept 2 label rows, the simulator kept 1
E07create_issue, get_issueagreedagreed
E08edit_issue, edit_issue, get_issueagreedagreed
E09edit_issue, add_label, get_issuediffered: Gitea succeeded, the simulator returned 404agreed
E10remove_labelagreedagreed
E11create_label, add_label, add_label, get_issueagreedagreed
E12edit_issue, remove_label, edit_issue, get_issueagreedagreed
E13create_issue, add_label, get_issueagreedagreed
E14create_label, create_label, add_label, get_issueagreedagreed
E15get_issue, get_issueagreedagreed
E16create_label, create_labelagreedagreed
Design choice
A strict observation contract decides what counts as the same: duplicates are kept, and raw database IDs are not compared literally.

A circuit diagram has to beat a cheap guess.

Test attribution against the strongest cheap baseline first. It only earned its cost when the circuit was spread across token positions.

Attribution minus the strongest cheap baseline, by taskDot-and-whisker chart, SAE-feature basis and raw-neuron control. On all seven single-token tasks, every estimate sits within 6.0 points of zero. On IOI, where the signal is distributed, the gap is +30.1 points (95% interval +16.0 to +45.8) in the SAE-feature basis and +17.2 points (95% interval +8.7 to +28.1) in the raw-neuron control.
Attribution minus the strongest cheap baseline, by taskDot-and-whisker chart, SAE-feature basis and raw-neuron control. On all seven single-token tasks, every estimate sits within 6.0 points of zero. On IOI, where the signal is distributed, the gap is +30.1 points (95% interval +16.0 to +45.8) in the SAE-feature basis and +17.2 points (95% interval +8.7 to +28.1) in the raw-neuron control.

Attribution minus the strongest cheap baseline (sufficiency, percentage points). Gemma-2-2B, layer 7, top 64 units, paired-bootstrap 95% intervals.
TaskSAE-feature basisRaw-neuron controlPrompt pairs
Capitals+6.0 (95% interval +2.3 to +10.1)+3.0 (95% interval +1.7 to +4.3)20
Country to language+0.3 (95% interval −2.8 to +3.8)+2.5 (95% interval +1.6 to +3.4)20
Past tense+2.1 (95% interval +0.3 to +4.2)+2.5 (95% interval −0.2 to +5.3)30
Comparative−1.3 (95% interval −3.9 to +0.6)0.0 (95% interval −1.8 to +1.5)24
Plural−2.1 (95% interval −5.8 to +1.9)+1.0 (95% interval −1.5 to +3.4)16
Antonyms−5.8 (95% interval −14.6 to +0.7)+1.2 (95% interval −1.6 to +3.9)24
Successor−0.5 (95% interval −5.8 to +6.6)−5.0 (95% interval −11.9 to +2.6)11
IOI (indirect object identification, distributed)+30.1 (95% interval +16.0 to +45.8)+17.2 (95% interval +8.7 to +28.1)18
Gemma-2-2B, layer 7: attribution minus the strongest cheap baseline, for SAE features and a raw-neuron control, with paired-bootstrap 95% intervals. Attribution pulls clear only on IOI (indirect-object identification), the distributed circuit.
Design choice
I grade every method against the strongest cheap baseline I can build, and measure the known answer instead of assuming it.

A lost reply. Not a lost result.

System sketch, not a live run.

create_guarded_issueAgenthas resultToolno new writeExecution recordSQLitecase-next-resultappliedreply lostget_operation_statusreturn_stored_resultIssues created1Still 1. No second issue.

Stored result returned

Recover the result.

After a restart, the agent calls get_operation_status with the same operation_id. The answer is return_stored_result, and nothing runs again.

The tool finished the work, and the agent never heard back. A blind retry would create the issue twice. I built a reference that asks the execution record what happened before doing anything again.

A lost reply means unknown, not failed. Ask the system of record before retrying anything with side effects.

Checkout, when the buyer is an agent.

System sketch, not a live run.

Policy decision
  1. Intent
  2. Confirmationconsent ledger, hash‑chained
  3. Authorisationpolicy hook
  4. Fulfilment
  5. Receiptprovenance manifest
Allow: Order fulfilled. It leaves with a receipt and a detached provenance manifest.

An order arrives that no person typed. I designed the path it has to take (consent recorded, a policy check, a receipt with provenance) and built it.

Settle what an agent may buy on its own, and what needs a person, while designing the checkout, not after an order ships.

A useful result
can be “no.”

No model. Full marks.

In No Model Floor, a fixed action behind a guard reached the benchmark’s ceiling. There was no headroom left for a model to show it could adapt.

configurations tested
328

If a fixed baseline already reaches the ceiling, the benchmark cannot credit a model.

Studies10 · 7 with a recorded result

Pick a question.

Each study shows its result and where the evidence stops.

  • Result recorded (7)
  • Status only, no measured result yet (3)

Left to rightTop to bottom: release order within each question.

GhostTrace

Research software

Does a behavioural signal survive repeated self-distillation?

Finding toy tier only

A behavioural result from a toy setting may not carry over to a real model. Test it where you deploy.

Read the study Discuss this: GhostTrace

Everything else on the bench.

The smaller builds, by area, newest first.

All public repositories on GitHub

About

I want to understand what a new capability can do. Building with it is how I find out.

Each wave, the same method.

  1. From 2004

    Mobile

    I started in mobile, when the consumer-tech frontier moved one new category at a time: operators first (O2, T-Mobile, EE and TalkTalk), then product and portfolio leadership at the device maker KAZAM.

  2. Connected devices and immersive computing

    IoT, wearables, tablets, connected home, smart audio, then VR and AR, including senior product management for VR, wearables and SmartThings at Samsung. Leadership teams across sectors, UK first, then the US.

  3. 2016 to 2021

    Founding Captivate

    I founded Captivate, an applied-innovation consultancy for XR, AI and emerging technology, and ran it for five years. Clients included Synthesia, Gucci, NBC Universal and Lloyd's of London; for Synthesia I was VP of global partnerships in 2018 and 2019. That is where AI caught me: early generative models for synthetic media, alongside immersive work. In parallel, from 2017 to 2018, I led brand partnerships across EMEA for the VR company Jaunt.

  4. 2017

    Pulled into AI

    AlphaGo, OpenAI's Dota work and AlphaGo Zero. I have read the field closely since.

  5. 2019 to 2022

    Emerging technology and XR strategy, PwC UK

    I led strategy and client work on XR and spatial computing where it met AI and data, with a specialist team building training simulations and spatial analysis.

  6. 2022 to 2024

    AI and emerging technology, PwC US

    An international assignment to PwC US in Austin. It began with immersive technology and grew into AI and generative AI: commercialising both for clients and co-developing solutions with technology alliance partners.

  7. 2024 to 2025

    Technology business partner, PwC UK

    Back in the UK, the link between PwC's central technology team and its consulting teams: shaping AI and generative-AI propositions for clients, and bringing together the people to deliver them.

  8. Sep 2025 to now

    Frontier AI strategy and enterprise innovation, PwC US

    A move to PwC US. First, agentic AI delivery for clients; now, with PwC Intelligent Enterprise, strategy, research and hands-on solution development: taking an idea from an early research question to a testable solution and a credible path to scale.

  9. 2026

    Building in public

    Specifications from February 2026, an interpretability study in June, then nine more studies in September.

More than twenty years across the UK and US.

The work today

At PwC Intelligent Enterprise I work across strategy, research and hands-on solution development, helping organisations rethink how work gets done when AI becomes part of the operating model. The studies and builds here are my own, independent of that role.

The pattern I keep seeing

The arc tends to repeat: loud enthusiasm, then a small number of teams quietly figuring out what is real and shipping it.

How I build

I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me. Several studies here end negative or inconclusive, and they say so.

What I care about now

The space between an emerging capability and a redesigned enterprise. That is where evidence, judgment and thoughtful implementation matter just as much as the technology.

Releases12 citable releases

Written down, so you can check it.

Specifications, a white paper and research software, each published with the reasoning beside the code, so someone else can reproduce it or build on it.

Full record on ORCID

Working on agent supply chains? Start with the Tool Bill of Materials (TBOM), Skill Bundle Attestation (SBA) or Tool Security Advisory (TSA).

A DOI records a release. It is not peer review.

Agent Boundary Assurance

An evidence discipline for what a local or enterprise agent accesses, remembers, transforms and sends.

White paper

Zenodo DOI 10.5281/zenodo.20815336

Code for Agent Boundary Assurance

Personal AI Sovereignty Lab (PAISL)

A local-first benchmark scaffold for testing whether a personal agent stays useful while respecting the user's data boundaries.

Benchmark scaffold

Zenodo DOI 10.5281/zenodo.20753227

Code for Personal AI Sovereignty Lab (PAISL)

The nano interpretability sequence

  • nanocircuits

    Ground-truth-calibrated circuit discovery on small transformers.

    DOI 10.5281/zenodo.20611793

  • nanofeatures

    The per-element attribution-cost boundary on Gemma-2-2B and GPT-2 with SAEs.

    DOI 10.5281/zenodo.20611795

  • nanoassembly

    Multi-layer SAE feature circuits and attribution-cost calibration.

    DOI 10.5281/zenodo.20611788

Research software, three releases

Zenodo

All releases8 more

ProofKern

A harness that counts an AI-generated Metal kernel as faster only after it passes a float64 correctness oracle and a bootstrap confidence-interval timing rule.

Research software

Code release, no DOI

nanoIM

A from-scratch lab for one question: if two interactions flatten to the same final chat transcript but require different next actions, what can a transcript-only model know?

Research software

Zenodo DOI 10.5281/zenodo.20492362

Code for nanoIM

nanoAWM

A tiny agent world model lab: compact models predict the consequences of an action before it runs, then plan against those predictions.

Research software

Zenodo DOI 10.5281/zenodo.20429920

Code for nanoAWM

Agent Assurance Case

A portable, signed evidence object for agent release decisions. The verifier recomputes the verdict offline.

Draft specification and verifier

Zenodo DOI 10.5281/zenodo.20185170

Code for Agent Assurance Case

Damn Vulnerable Agent Asset Corpus

Deliberately vulnerable and deliberately clean agent release fixtures, with expected findings and Agent Assurance Case templates, for testing agent-asset assurance tools.

Conformance corpus

Zenodo DOI 10.5281/zenodo.20186918

Code for Damn Vulnerable Agent Asset Corpus

Tool Security Advisory

Machine-readable security advisories for MCP tools.

Specification and preprint

TechRxiv DOI 10.36227/techrxiv.177155646.67434382/v1

Code for Tool Security Advisory

Skill Bundle Attestation

Deterministic bundle identity, content attestation and verification tooling for agent skills.

Specification

Zenodo DOI 10.5281/zenodo.18485741

Code for Skill Bundle Attestation

Tool Bill of Materials

Signed manifests that bind MCP server releases to immutable tool metadata, so a client can check what it is about to call.

Specification

Zenodo DOI 10.5281/zenodo.18458945

Code for Tool Bill of Materials

Notes between releases: LinkedIn

Contact

What are you working on?

Send the question you’re trying to answer and what you’ve tried so far. A short note is fine.

What is it about?

or email hello@jasonlovell.ai

Replies usually inside a working day.