Recetaria
Live multimodal pipeline: spoken story in, structured recipe out
- Where
- Independent
- When
- 2026
- Tags
- Gemini API · Schema Enforcement · Supabase · React
- Try it
- recetaria.co

Recetaria isn't a recipe app with AI features bolted on. It's a multimodal generation pipeline running in production, with real users, taken from first concept through to monetization independently. The same class of system as EverGen, except this one you can open in a new tab and use before you finish reading this page.
The problem it solves
Family recipes live in the way someone tells them, not in the way a cookbook writes them — full of real information, but unstructured, and gone once the person who knows it stops telling it. The product's promise is in its tagline: life happens, the recipes remain. So the design constraint is unusual: the user shouldn't have to think in database fields. They talk the way they'd talk to family, and the system does the structuring.
The pipeline
One spoken recording becomes a structured, shareable recipe in a single flow:
- 01
Speak it
Live in-browser dictation, with a recorder fallback for iOS and the embedded WhatsApp browser, where the mic behaves differently. Someone just talks.
Web Speech API - 02
Transcribe
The audio becomes text. The same endpoint also reads a photo of a handwritten recipe card, because the input the family has is sometimes paper, not a voice.
Gemini · STT + OCR - 03
Structure
The transcript becomes a schema-constrained recipe: ingredients with estimated measures, ordered steps, the story around it — and a prompt describing the finished dish.
Gemini 3.1 Flash-Lite - 04
Paint the cover
That prompt paints the recipe's cover in a naïf folk-art style, re-encoded to WebP so a phone on a slow connection still loads it.
Nano Banana 2 Lite - 05
Keep and share
Recipe, cover and the original audio are stored together behind row-level security, and shared by a link a relative can open and answer without an account.
Supabase
Every stage runs on the Gemini API, chained into one user-facing flow. Chaining generation stages turns out to be a harder problem than any single stage on its own: each step's output has to be reliable enough to become the next step's input, with no human in the loop to catch a bad hand-off before the user sees the result.
Schema enforcement
A language model will describe a recipe in prose happily and inconsistently. A database needs a fixed shape: ingredients as a list, quantities as parsed values, steps as an ordered sequence. Getting from the first to the second reliably means constraining the model's output to a schema and handling the cases where it still doesn't comply: a missing field, a quantity given as 'a pinch' where the schema expects a number, a step folded into the ingredients list.
None of that is exotic engineering, but all of it has to work every time, because there's no editor in the loop to fix it before it ships. This is the same discipline that made EverGen usable at studio scale, applied to a consumer product where the tolerance for a broken result is even lower.
What breaks
- Transcription errors cascading downstream: a mis-heard ingredient doesn't stay a transcription problem, it becomes a wrong recipe.
- Ingredient hallucination: the model filling gaps with plausible-sounding but invented ingredients or quantities.
- Narration that isn't linear: people double back, correct themselves, and mention an ingredient three sentences after the step that uses it.
Each failure mode gets a guardrail at the stage where it originates rather than a cleanup pass at the end. That ordering matters: a mistake caught at transcription is cheap, and the same mistake caught after structuring has already contaminated everything downstream of it.
Designing for the person who actually has the recipe
The people whose recipes are most worth saving are usually the oldest people in the family, and the least likely to type a structured form into an app. That single fact decided the interface: voice is the whole entry point, not a feature bolted onto a form, because it asks nothing new of the person using it. A typed form would have given me clean fields for free; accepting speech meant taking on transcription error, non-linear narration and hallucination as my problems instead of the user's — the right trade when the alternative is that the recipe never gets recorded at all.
The rest follows the same logic: bilingual in Spanish and English from day one, since these families often aren't working in English, and sharing is built around family rather than a public feed — a shared recipe is an invitation.
Concept to monetization
I took Recetaria through the entire arc independently: product idea, interface, pipeline architecture, visual direction, deployment, and monetization. That last step is the one most side projects never reach.
What it proves
Recetaria is the public, NDA-free version of the same argument EverGen makes privately: that I design and ship multi-stage generative systems, not just prompts. Anyone evaluating the EverGen case study does not have to take it on faith, because this one is live and every architectural decision described above is inspectable in the running product.
It also closes the gap that the studio work cannot. EverGen proves I can build something a team adopts and that moves a business number. Recetaria proves I can take a product from an idea to something people use and pay for, without a studio, a brief, or anyone else's roadmap. Those are different claims, and most portfolios can only make one of them.