InVinetory
Point a phone at a wine label and it identifies the bottle, researches it once for every cellar, and files it in a shared household inventory. Describe dinner out loud and a voice sommelier picks from bottles you actually own.

I own two or three bottles of most wines and keep them in fridge drawers. I wanted to know what I have and where it is, and to get a pairing from my bottles rather than a shopping list.
AI engineering, measured
The interesting work was making model calls reliable and cheap, and proving it with numbers:
- Label identification: Haiku reads the label first; Sonnet retries only when the first read is unsure. On 14 real photos (seven bottles, front and back) it got 13 right, with a median of 2.7 s and about 0.4¢ per read.
- Model confidence is poorly calibrated. A misread came back at 0.85 confidence. So hard rules gate it: a producer is required, and a vintage too unless the label says NV. Back labels without a year go to a review screen instead of silently becoming the non-vintage wine.
- Research runs once per wine and is shared by every cellar: about 22 s and 7¢. The newest search tools were tested and rejected for this job: 113 s and 16¢ for one wine.
- Sommelier: meal parsing and picking run in parallel, cutting latency from 15.6 s to about 6 s. The model only ever sees short codes for bottles in the cellar, and the server drops anything else, so it can’t recommend a wine you don’t own.
- Voice: speech-to-text is primed with the cellar’s own wine names as key terms, which fixes misheard producers.
- Cost control: every AI call is logged and priced. Each cellar has a monthly allowance, with alerts at 80% and 100%; research pauses at the cap, but label reads never do.
Engineering underneath
- Two layers of access control: routes check cellar membership, and Postgres row-level security is the backstop. Every new table needs a policy and a test.
- Three containers instead of a dozen. The original plan used a hosted backend, a job service and a separate host. It became Postgres (with the job queue inside it), a Hono server and the app, self-hosted.
- Prompt caching on the system block only. Automatic caching cached through the photo and paid the cache-write premium on every scan.
- Web research is treated as untrusted. Every server-side fetch is checked against public addresses on every redirect, and a stock bottle photo is kept only after a vision check confirms it’s the right bottle. Share images lie.
Built with AI
The brief estimated 9–13 weeks at ten hours a week. Claude Code built all six phases in two days, each phase ending with something testable on my phone. CLAUDE.md holds the rules every session must follow (RLS on every table, secrets only in the server env, model IDs only from config), and the eval scripts score model changes against ground truth, not vibes.
Screenshots are from a local copy with an invented demo cellar.

