Writing
Long-form pieces on the AI Reliability Layer, the Veridi methodology, and the engineering practice the category requires. Citation-grounded, calibrated, dated.
Five pieces, currently. Start with the Reliability Layer manifesto for the four-concern framing of the runtime infrastructure category I am building toward. Read the Veridi methodology writeup for technical depth on calibration, gaming countermeasures, and pipeline architecture; it is the worked example that grounds the manifesto’s framing. Trust and trustworthiness (part one of two) is the accessible entry point: why putting a trust-inspiring face on an untrustworthy system is dangerous, and the three barriers that have to be faced first. AI agents that test accessibility announces two open-source Claude Code skills for agent-driven WCAG testing and makes the case that accessibility tooling and AI evaluation are one conversation. Accessibility is a reliability concern argues the structural version of that case: the manifesto’s four runtime concerns already cover accessibility engineering, and the two fields converged independently on the same findings.
- Accessibility is a reliability concern
Evaluation discipline and accessibility engineering are one posture applied to two substrates. The Reliability Layer's four runtime concerns already cover accessibility, once you notice that an interface is a system output and a user session is runtime.
- AI agents that test accessibility: two skills, released
Two Claude Code skills for agent-driven WCAG testing, released as open source with their test-contracts intact. What agent testing catches, what it cannot, and why the must-not-contain list is the piece worth defending.
- Trust and trustworthiness
AI agents are not trustworthy out of the box, and putting a trust-inspiring face on an untrustworthy system is dangerous. The three barriers to earning user trust - overconfidence, security, and reliability - and why they have to be faced before the face goes on. Part one of two.
- The AI Reliability Layer
A four-concern runtime taxonomy (verification, calibration, adversarial robustness, recovery) and a positioned framing of the territory it covers. Engages with the close adjacencies (AI TRiSM, NIST AI RMF, AI SRE).
- Veridi methodology: what the numbers actually mean
A deep-dive on the Veridi verification system and the Pragma + Praxis assessment frameworks that run on top of it. Calibration (selective Brier 0.0253, coverage 89%, abstention correctness 11/11, accuracy 99.5%), architecture, gaming countermeasures, limitations, and refinement backlog, grounded in published research.