Skip to content
All findings

About

Building AI-native systems for real use cases.

100+ small experiments on AI and human judgment.

A public record of testing what actually holds up: how AI behaves under different prompts, models, and conditions, and where it quietly fails. Each finding ships with what survived testing, what testing killed, and honest limits on scope.

The author

I run the experiments and publish the findings: where AI is dependable enough to put real work on, and where it isn't. Contact and longer-form writing on LinkedIn.

Method

Every finding here was tested before being published. The test stays visible as part of the finding, not hidden in methodology. Three sections run below each post:

  • What survived testing: hypotheses the data supported.
  • What didn't survive: hypotheses I went in with that the data killed. Part of the finding, not against it.
  • Honest limits: scope conditions the finding does not stretch to cover.

When a previously published claim is later corrected or withdrawn, the record tracks that separately. See the record.

Reading and citing

Each bullet under a finding has a stable anchor you can cite by URL fragment (/slug#s1 for the first surviving claim, #d1 for the first that didn't survive, #l1 for honest limits). Hover over any bullet to reveal the permalink.

Each post also has a “Cite this finding” button that copies a formatted citation to your clipboard.

Errata and corrections

If you find a problem with a finding's data, methodology, or reasoning, the corrections channel is LinkedIn DM. Posts with receipts publish the raw data; specific issues with the raw data, scoring, or analysis are the easiest to engage with.

Corrections that hold up get published on the record, with attribution.

26 findings published since Mar 23, 2026. The published findings are the curated subset of a larger experiment log.

New findings when they land.

No spam. Just what held up.