Writing¶
Fewer, deeper pieces. One good implementation note on provenance-aware retrieval is worth more than thirty posts about AI news, and the second kind is now free to produce, which is exactly why it carries no signal.
Everything here is dated and status-labelled. Nothing is published without a human verification pass over its factual claims — drafting is assisted, verification is not.
In progress¶
Latent spaces and agentic revenue growth management — working paper, in editorial review.
A long-form piece on semantic architecture for agentic systems: what a latent space gives you that a taxonomy does not, and where that breaks down in an enterprise setting where meaning has to be governed rather than inferred.
The sourcing pass is done. It was more instructive than expected. Of 55 references, roughly two dozen turned out to be SEO articles, forum posts or aggregator pages — and they carried some of the most checkable claims in the paper. Four were substantively wrong, not merely weakly sourced: a wrong EU AI Act deadline, a margin figure attributed to a company that never reported it, a benchmark score cited to a Reddit thread when the vendor's own page had it, and two computer-vision papers cited as evidence about language-model agents because an aggregator had collapsed two literatures under one keyword.
Four numbers were deleted rather than re-sourced. The paper is shorter and considerably more defensible.
One thing still stands between the draft and publication: an original-contribution section. A well-sourced synthesis of other people's work is not worth publishing under my name. The contribution has to come from the platform build — what the evaluation harness actually measured, and where semantic ownership turned out to be the binding constraint.
Then editorial review, then publication. It is not linked from here until it clears both.
Planned¶
The model was never the difficult part. A field note from the engagement that moved me out of general software engineering and into data: we shipped an end-to-end ML product without a data engineer, it succeeded, and every hard problem we hit late turned out to be a data problem in disguise. The argument is that the pipeline — provenance, lineage, definitional stability — is what makes a system end-to-end, and that the generative-AI wave is reproducing the same pattern at larger scale. Sketched in About; the essay owes a worked example before it is worth publishing.
Field notes on the patterns already written up in Methods — specifically what the fail-loud pattern costs politically in the first month, which is the part of it nobody writes about.
Publishing standards¶
Applied to everything on this site, recorded here so the standard is visible:
- Primary sources over commentary. Papers, standards, product documentation, original talks. Where a claim rests on someone else's summary, that is stated.
- Opinions labelled as opinions, distinct from measured findings.
last verifieddates on anything touching fast-moving tooling.- Original contribution required. If a piece only restates existing material, it does not get published.
- AI assistance disclosed where it is material — where the reader's interpretation of the artefact should change because of it.