Availability
I am a salaried full-time Lead Engineer at BI Worldwide. Here’s what I am open to outside that role, and the conditions that make a fit.
Hiring inquiries
Open to senior to staff individual-contributor roles in Applied AI Engineering, Forward Deployed Engineering, Solutions Architecture, and AI Evangelism. Preferred locations: Minneapolis/St. Paul, San Francisco Bay Area, New York City, Seattle, or anywhere in Canada. Other areas may be possible given an ideal role fit.
The strongest fitting signal, as I see it, is whether your team’s runtime concerns map to the four pillars in my Reliability Layer manifesto: verification, calibration, adversarial robustness, and recovery. A team working that gap gets someone who has built verification infrastructure under adversarial conditions and can translate it across product, engineering, and design. For relevant evidence, see work.
A more concrete fit signal: your team is deploying AI in production and has started accumulating incidents or quality drift that benchmark eval did not predict. If your runtime concerns are verification, calibration, adversarial robustness, or recovery, and you have a senior-IC role open, we should talk.
For hiring inquiries, email shawn@nettertech.com with the subject hiring inquiry. Include the role link, level or compensation band, and a sentence on team context.
Verification: LinkedIn · GitHub · AICred: Rank 1, 10.0/10 (2026-07-21) · References available on request.
Contract Reliability reviews
I take on one to three Reliability Layer engagements per quarter. Scope: an outside read on a production AI deployment’s verification, calibration, adversarial robustness, and recovery posture. Deliverable: a written assessment with calibrated findings and a prioritized remediation list. Engagement formats: a one-week onsite intensive (eval-harness bootstrap, agent reliability audit, or calibrated-uncertainty workshop), or a two-to-four-week project-scoped review. Pricing on inquiry; senior-IC consulting rates apply.
Accessibility sits inside that scope rather than beside it. Where a deployment’s interface is the thing failing its users, the review treats that as a reliability finding on the same four concerns, for the reasons set out in Accessibility is a reliability concern. I do not sell standalone conformance audits.
Adjacent surface: short-form expert consultations via established expert networks (AlphaSights, Third Bridge, EveryExpert). Different product, different time profile. Route those engagements through the network rather than this page.
Not a fit: general AI consulting, prompt-engineering coaching, training-data work, or engagements that require exclusive availability for a multi-month period.
For contract inquiries, email shawn@nettertech.com with the subject reliability review.
What I am not available for
- Cold outreach for roles outside the list above
- Recruiter outreach without a specific role link
- General consulting outside the Reliability scope
Routing inquiries with the subject conventions above is the fastest path to a real reply.
Methodological anchors
The four-concern taxonomy in the Reliability Layer manifesto is the public-facing frame. Underneath, the work draws on established disciplines outside the typical AI-tooling stack. For prospects who want the depth read before a discovery call, the anchor list is below.
See the full discipline anchor set
Psychometrics: Krippendorff α (inter-annotator agreement floors), MTMM (Campbell & Fiske 1959 multitrait-multimethod), Cronbach-Meehl nomological networks, Jacobs & Wallach 2021 seven-aspect measurement audit, McDonald ω.
Calibration science: Brier decomposition (Murphy 1973), ECE/MCE/SCE (expected, maximum, static calibration error), temperature scaling, reliability diagrams, Tetlock-style tournament Brier benchmarking.
Statistical-test discipline: Wilson score confidence intervals, McNemar paired test with Bowker symmetry generalization, Holm-Bonferroni FWER plus Benjamini-Hochberg FDR (multiple comparisons), Cohen 1988 power analysis, BCa bootstrap.
Safety engineering: PRA Levels 1/2/3 (NUREG-1855), HFMEA (VA NCPS Healthcare Failure Mode and Effects Analysis), FMEA (MIL-P-1629, 1949 origin), ALARP tier semantics (UK HSE 2001; Edwards v National Coal Board 1949).
High-reliability-organization (HRO) theory: WHO Surgical Safety Checklist (Haynes 2009 NEJM, mortality 1.5% to 0.8%), Hollnagel Safety-II, Hopkins constraint analysis, ASRS-analog confidential reporting (FAA AC 00-46F).
Intelligence-community probability calibration: ICD-203 verbal probability bands.
Software-quality engineering: ISO/IEC 25010:2023 (eight quality characteristics plus opt-in Safety), SWEBOK v4, ISO 14764:2022 (maintenance), Nygard 2011 ADRs, Ford+Parsons+Kua fitness functions (evolutionary architecture), Twelve-Factor.
Secure-software development: OWASP ASVS 5.0, NIST SSDF, SLSA v1.2, CVSS v4.0 (FIRST.org 2023), SSVC (Spring et al. 2021 IEEE S&P).
AI-native threat and failure taxonomies: MITRE ATLAS (AML.T#### technique mapping), MAST (NeurIPS 2025 multi-agent failure modes), Microsoft agentic AI taxonomy.
LLM-evaluation infrastructure: Inspect AI (Dataset/Solver/Scorer), FEVER-score, StrongREJECT, Spotlighting (Hines 2024), indirect prompt injection (IPI), PromptBench, Length-Controlled Win Rate.
Forecasting and expert-judgment research: Tetlock superforecaster Brier benchmarking, Klein 1998 naturalistic decision-making, Kahneman and Klein 2009 on expert-intuition conditions.
Normative philosophy (Pragma methodology only): Frankfurt sufficiency, Sen capability, Parfit priority, Rawls liberty, Raz-Chang incommensurability.
The instrumented form lives in the Veridi methodology writeup and across the Veridi, Praxis, and Pragma published frameworks. Veridi v1.2, Praxis v1.4, Pragma v1.6 as of May 2026.