Alan De Vaney

I build AI systems that can prove what they did, and refuse when they can’t.

Grounded generation, deterministic cores, eval-driven development, and a receipt for every action. Built end to end, run in production, operated solo.

19 years shipping software Orange County, CA Python · TypeScript · C++ LLM · RL · CV, in production

Interrogation desk, the thesis, live

Ask the record anything.

This bot answers only from a curated evidence corpus and must cite it. The server verifies every citation, an answer it can’t back gets stamped refused, not improvised. That constraint is the same architecture the systems below run on.

Grounded-or-refuse · citations verified server-side№ live

Ask about the systems, the incidents, the method, or the 19 years. If it isn’t in the record, I’ll say so.

source: E53, this bot is itself the demonstration

Verified

Ledger, systems built end to end

The work.

Every card links to a full account: problem, architecture, the hard decision, the numbers, and what broke. Code is proprietary; the systems run, and I walk through any of them live.

Record, 2007 to present

Nineteen years, in brief.

2025, present

Production AI systems, end to end

Designed, built, and operate the four systems above, solo. Product, architecture, implementation, deploys, backups, incident response, postmortems. LLM, RL, and CV workloads with evals, cost controls, and guardrails as first-class architecture.

2007, 2024

Client web development, every era of the stack

Paid web work from 2007 onward, through the PHP/jQuery years, the SPA turn, and into the modern TypeScript stack. The full account of clients and engagements is available on request; this public record keeps client work private by default.

2007

First paid build

Where the ledger opens.

Method, what holds across all of it

Four rules, no exceptions.

These aren’t aspirations; they’re load-bearing architecture in every system above, and in the bot at the top of this page.

Grounded or refused

Generated claims trace to evidence or don’t ship. Gaps get acknowledged, not papered over. Enforced in the system, server-side, not in the prompt.

Deterministic core, LLM at the edge

The model proposes; typed, testable machinery disposes. Numbers come from engines and validators, never from a language model’s mouth.

Eval-driven, or it isn’t engineering

Non-deterministic systems get golden sets, agreement measurement, and calibration loops that only ship a change when it provably helps.

Every action leaves a receipt

Immutable records of what was done, with what inputs, at what cost. Auditability is a feature users see, and the reason autonomy can be granted at all.

Writing, the reasoning, in public

Essays

Colophon

About

I’ve been paid to build software since 2007, client web work through every era of the stack, and for the last two years, production AI systems built end to end: retrieval pipelines, agentic workflows with guardrails, evaluation harnesses, cost controls, and the unglamorous operations underneath.

I work the full span: product decisions, architecture, implementation, and running the result in production with real users. The common thread is a distrust of unverifiable output, mine included, and systems designed so trust doesn’t have to be assumed.

I’m based in Orange County, California, and open to AI engineering leadership jobs.