What the harness has shipped
End-to-end teardowns of products built with CodeMySpec. The code, the specs, the failure modes, and what changed in the harness as a result.
Broken Oaths
A persistent hex-globe 4X. Twelve days. The hardest thing the harness has built.
A massively multiplayer 4X strategy game where losing a war makes you a vassal rather than a corpse. Built in a 12-day sprint in July 2026: 108,734 lines of Elixir, 279 executable BDD spex, seven production releases, real playtesters by day six. Also the build where 20 of 52 stories carry no QA record at all, and where the case-study tooling was caught fabricating its own evidence.
EarWitness
Live build in progress. The requirements are public before the code exists.
A local-first Otter alternative: on-device transcription, speaker diarization, audio-tap meeting capture with no bot in the room, an Otter-style editor, and MCP access for your AI assistant. Being built live by the harness right now; 11 stories with full Three Amigos records, 4 already browser-QA passed. Follow along on Reddit.
Legwork
My manual outbound process, productized in a weekend. The agent is one of the users.
An outbound-process composer built with Claude Opus 4.8 over one weekend (July 5-6, 2026). Five-stage prose runbooks an LLM agent works via 16 OAuth-authenticated MCP tools, with a human launch gate. Gating John's real prospects on day two, overturning his own guesses 8 times out of 12.
Keel
My own books. 9,225 transactions. The agent found $227,736 of phantom income.
An agent-first, double-entry, cross-border personal-finance ledger John built on his real books over 26 days (June 15 to July 10, 2026). The collaborative end of the dial: 2.9% of the conversation typed by the human, who supplied the accounting judgment no fixture could encode.
Cleaner CRM
Two prompts. Two UAT passes. Three days. Fable did the rest.
A full cleaning-business CRM two-shotted with Claude Fable 5 (July 7-9, 2026). Owner back office, cleaner field app, public booking funnel with real Stripe payments, deployed to live UAT. 172 subagents, 0.4% of the conversation typed by John, and an honest accounting of everything that broke.
Market My Spec
Zero prompts. Zero lines of code. Four days. Same harness.
A small MCP server built by CodeMySpec across one weekend (May 3-6, 2026). Configurable per-component workflow, Three Amigos as a real agent task, BDD-spec boundary protection. The honest teardown of the experiments that made zero-prompts possible.
MetricFlow
The harness wrote the code. Here's what worked and what didn't.
A Phoenix LiveView marketing analytics platform built by CodeMySpec in 13 working days. OAuth integrations across five platforms, multi-tenant accounts with agency white-labeling, correlation analysis, AI insights. The first end-to-end teardown.
Want the methodology behind these?
The case studies are the receipts. The methodology is the process that produced them.