跳到正文
原文
Google AI:DEV 作者专属(RSS)· Piyush Anand·· 4 小时前AI 评分52

EchoBook:用 faster-whisper 与 Ollama 在本地把祖辈语音备忘录整理成家庭菜谱书

EchoBook: Turning Grandpa's Rambling Voice Memos into a Family Recipe Book, Fully Offline

AI 导读

作者为 Hacktoberfest 挑战开发了 EchoBook,用 faster-whisper(Whisper small.en)本地转录语音备忘录,再由 Ollama + qwen2.5:7b 通过 JSON schema 约束输出生成结构化菜谱卡,全部处理在自家电脑离线完成,不上传云端。

正文

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

What I Built

The person

I built this for my grandfather.

For years he has recorded voice memos about the food our family grew up on: the dal my grandmother made every Sunday, the banana bread our neighbour Mrs. Fernandes taught him in the seventies, the chai he picked up from a railway-station chaiwala over forty years of mornings before work. Whenever someone asks "how did you make that?", he records another memo.

The problem

None of it is written down. The memos ramble: "Okay, is this thing recording?", a story about the day my father was born, back to the lentils, "four whistles, not three, she was very strict about that." There are no ingredient lists and no numbered steps. The measurements are "a big pinch" and "the ugly black bananas nobody wants to eat." The recipes exist only in his voice, scattered across audio files on an old phone, and once those files are lost the recipes go with them.

Typing them up by hand means listening to hours of audio and pausing every few seconds. Sending them to a cloud transcription or AI service means uploading my grandfather's voice and our family stories to somebody else's servers. I didn't want either.

The solution: EchoBook

EchoBook turns raw voice memos into a clean, searchable, shareable family recipe book, with every bit of processing done on the family's own computer using open-weight models.

You drop in a recording, and EchoBook:

  1. Listens. It transcribes the memo locally with an open-weight Whisper model.
  2. Writes the recipe. A local open-weight LLM turns the ramble into a structured recipe card: title, servings, every ingredient with its quantity exactly as he said it, numbered steps in the order he does them, and his own tips and warnings.
  3. Keeps his voice. Every recipe carries an "In Grandpa's words" section: a verbatim quote of the memory behind the dish ("You know, we made this the day your father was born. The whole hostel floor could smell it."), right above an audio player with the original recording.
  4. Adds it to the book. It's stored with the full transcript and a link to the source audio.

A recipe page with Grandpa's quote and the original recording

Features

  • 🎙️ Drop in a voice memo (mp3, m4a, wav, ogg, webm…) or pick a sample, and watch each stage run live: saving → listening (progress bar plus live transcript) → writing the recipe (live token count) → adding to the book.
  • 📖 Recipe cards with tick-off ingredient checklists, numbered steps, "Grandpa's tips", and the original transcript one click away.
  • 🗣️ "In Grandpa's words": an authentic quote plus the original recording on every recipe.
  • 🔍 Search by dish or ingredient (ghee, cardamom milk), with matching ingredients highlighted on each card.
  • ✏️ Review & edit: small models sometimes mishear, so every recipe can be checked against the recording and corrected before the book is shared. The family gets the final say, not the AI.
  • 🖨️ Export: download one recipe or the whole book as Markdown, or open a print layout (cover, table of contents, one recipe per page) and save it as a PDF to print and bind.
  • 🌐 A shareable link on Render: relatives anywhere can browse the book on their phones and hear Grandpa tell the story, while recordings are processed only at home.
  • 📴 Works offline: once the models are downloaded, transcription and structuring run with the network unplugged.

The recipes are cleaned up. His voice is left as it is.

Demo

🔗 Live recipe book: https://echobook-8fpj.onrender.com/

(It's on Render's free tier, so if it has been idle the first load can take ~30–50 seconds to wake up.)

What to click:

  1. Open "Sunday Dal". Read the In Grandpa's words quote and press ▶ to hear the memo it came from. Notice the quantities are his ("a big pinch" of asafoetida, "4 whistles"). Scroll down and expand Original voice memo transcript to compare the raw ramble with the finished card.
  2. Search ghee, then cardamom. Only the matching recipes remain, with the matched ingredient highlighted.
  3. Open "Masala Chai" for the railway-station story and the "three times" trick.
  4. Click Print book, then Print / Save as PDF, to get the whole family cookbook as a PDF.
  5. Click Add a memo. On the public site this explains that processing happens at home. That's deliberate (see How I Built It).

The local pipeline in action (this is the part that runs on the family's computer):

EchoBook demo: a voice memo becomes a recipe card, then search and the printable book

A real, unedited run on a CPU-only machine. The two waiting stages are sped up (labelled in the GIF); the full memo-to-recipe run took about 2½ minutes. Note the raw output still says "asan" (Assam); that's what the Review & edit screen is for.

A note on the sample audio: to avoid publishing private family recordings, the demo uses three stand-in memos I wrote in the style of Grandpa's real ones (false starts, side stories, loose measurements, a bit of family history) and voiced with Piper, an open-source local text-to-speech engine. Real recordings go through exactly the same pipeline.

Code

GitHub logo Anandpiyush21 / EchoBook

Turn family voice memos into a living recipe book, entirely on-device with open-weight models (faster-whisper + Ollama). Built for the Hacktoberfest 'Build for a Friend' challenge.

EchoBook

Turn family voice memos into a living recipe book, entirely on your own computer.

My grandfather has years of voice memos describing family recipes from memory. They ramble, there are no ingredient lists and no steps, and the recipes only exist in his voice, scattered across audio files EchoBook takes those recordings and turns them into a clean, searchable, printable recipe book. It keeps the important part: his own words and the original recording, attached to every recipe.

All processing runs locally on open-weight models. No recording, transcript or family story is sent to a cloud API, and once the models are downloaded the pipeline works with no internet connection at all.

EchoBook demo

The recipe book

A recipe, with Grandpa's own words and the original recording The local pipeline, stage by stage
Recipe page Pipeline progress

Features

  • Drop in a voice memo (mp3, m4a, wav, ogg, webm…) or pick a sample, and watch each stage run live…

The repo includes setup instructions, an architecture write-up, render.yaml, .env.example, the sample memos, a prompt-evaluation script, and a step-by-step demo_script.md.

How I Built It

Architecture

voice memo ─► faster-whisper ─► Ollama + qwen2.5:7b ─► SQLite ─► web UI
  (.m4a)     (Whisper small.en,   (JSON-schema-        (recipe +     (browse, search,
              on-device)           constrained output)  transcript +  edit, export)
                                                        audio link)
└──────────────── all on the family's computer, works offline ───────────────┘
                                   │ git push (finished recipes only)
                                   ▼
                     Render: same app in read-only "viewer" mode
Layer Tool Why
Speech-to-text faster-whisper (open-weight Whisper small.en, CTranslate2, int8) Accurate, fast on a plain CPU, fully local
Structuring Ollama + qwen2.5:7b-instruct (Apache-2.0) Strong instruction-following at 7B; supports schema-constrained output
Storage SQLite, mirrored to data/recipes.json Zero setup; the JSON snapshot is diffable and deployable
Backend FastAPI + a background worker Long jobs don't block the UI; per-stage progress API
Frontend Plain HTML/CSS/JS No framework, no build step; mobile, dark mode and print styles
Hosting Render (free web service, render.yaml) A shareable link for the family, auto-deploy on push

Stage 1: Listening (faster-whisper)

faster-whisper runs OpenAI's open-weight Whisper model on CTranslate2: int8 on CPU, float16 on a GPU. Two settings made a real difference for this use case:

  • Voice activity detection skips the long pauses that old memos are full of.
  • An initial prompt seeded with kitchen vocabulary ("ghee, cumin, turmeric, cardamom, asafoetida, toor dal, Fahrenheit…") biases recognition toward the words that matter most in a recipe.

A one-minute memo transcribes in about 15 seconds on a 12-core CPU. The transcript is still imperfect ("hing pia saffitida", "the tor dal", "still worn" for "still warm"), which is exactly why the next stage has to be careful.

Stage 2: Writing the recipe (Ollama + Qwen 2.5)

Most of my time went here. The transcript goes to the local model with an "archivist" system prompt, and Ollama's structured outputs feature constrains decoding to a JSON schema:

{
  "title": "...", "description": "...", "servings": "...", "total_time": "...",
  "ingredients": [{ "quantity": "...", "item": "...", "note": "..." }],
  "steps": ["..."], "tips": ["..."], "story_quote": "...", "tags": ["..."]
}

Because generation is constrained to this schema, the response always parses, and I never have to scrape JSON out of chatty text. Getting it to parse was easy. Getting it to be faithful was the real work:

  • The 3B model invented things. I started with qwen2.5:3b for speed. It added a "30 minutes" total time nobody mentioned, gave the dal a "don't over-mix" tip that belonged to the banana bread, and turned "one teaspoon of baking soda" into "½ tsp". For a family recipe, a wrong quantity is worse than a missing one.
  • Stricter rules plus a bigger model fixed most of it. I rewrote the prompt around faithfulness: use only what is said; copy every amount exactly as spoken; never add a unit he didn't say ("four cardamom pods" → quantity 4, not 4 tsp); leave a field empty rather than guess; include ingredients mentioned in passing ("cinnamon if you like it", "if you have walnuts"). With those rules, qwen2.5:7b copied amounts exactly and left unknown fields blank.
  • Examples in a prompt leak. I gave one example of fixing a misheard word involving "Assam", and the model started naming unrelated dishes "Assam Dal". I replaced it with a neutral example.
  • "Verbatim" isn't always verbatim. Small models quietly paraphrase quotes, and the whole point of In Grandpa's words is that they are his words. So after generation, EchoBook fuzzy-matches the model's quote against the transcript and snaps it to the closest real span of one to three sentences.
  • Humans get the last word. Even the 7B model occasionally slips (one run wrote "4 tsp" of cardamom pods). Rather than pretend the AI is perfect, I built a Review & edit screen with the recording and transcript right next to the form, so a family member can check each recipe before the book is shared.

To iterate quickly I wrote scripts/eval_structuring.py. It caches transcripts, re-runs only the LLM step, and flags any ingredient quantity whose numbers never appear in the transcript, plus any quote that isn't verbatim. That made comparing prompts and models a two-minute loop instead of a re-recording session.

A performance surprise: my development machine has only a small, unsupported GPU, and Ollama was quietly offloading a sliver of the model onto it. Prompt processing crawled at ~4 tokens/s. Forcing pure CPU (OLLAMA_NUM_GPU=0) made it 6x faster (~27 tokens/s) and brought a full one-minute memo down to about 2 minutes end to end on CPU. On a proper GPU the whole pipeline takes seconds, and you can step up to large-v3 and a 14B model through .env.

Stage 3: Storage

SQLite stores each recipe alongside its full original transcript, the source audio file, and which models produced it, so provenance is never lost. Every write is also mirrored to data/recipes.json, a human-readable, diffable snapshot that doubles as the deployment artifact.

The web app

FastAPI serves a small JSON API and the static frontend. Processing runs in a single background worker (the models are large, so one job at a time keeps memory predictable). The UI polls /api/jobs/{id} for stage, transcription progress, the live transcript and the token count. The frontend is one page of plain JavaScript with hash routing, works on phones, follows dark mode, and has a print stylesheet that turns the book into a clean PDF with a cover, a table of contents and page breaks between recipes.

Deploying with Render: inference stays local, viewing goes to the cloud

Whisper plus a 7B LLM won't run on a free web instance. More importantly, family recordings shouldn't be processed in the cloud at all, because that's the whole point. So the same codebase runs in two modes:

At home (ECHOBOOK_MODE=local) On Render (viewer)
Upload & process new memos ✅ faster-whisper + Ollama ❌ disabled (read-only)
Browse, search, listen, export ✅ ✅
Dependencies requirements.txt (includes faster-whisper) requirements-web.txt (FastAPI + uvicorn only)
Data SQLite, mirrored to recipes.json SQLite rebuilt from recipes.json on boot

The publishing workflow:

  1. Process memos at home and review or fix them in the UI.
  2. git add data && git commit && git push.
  3. Render auto-deploys and rebuilds the book from the snapshot.

The viewer installs no ML dependencies, so it builds in seconds and fits the free tier. The repo includes a render.yaml blueprint, and the app automatically falls back to read-only mode when it detects it's running on Render, so a missing setting can never expose an upload endpoint. Relatives get a link they can open on their phones, and the cloud only ever sees the finished recipes the family chose to publish.

What's next

  • Batch-import a whole folder of old memos overnight.
  • Multilingual memos: Whisper's multilingual models can handle Hindi and mixed Hindi-English recordings, with the recipe written out in English.
  • Let family members add photos of the finished dish to each recipe.

Why Does Open Innovation Matter?

These recordings are my grandfather's voice and our family's history, and uploading them to a cloud API to save some typing never felt right. Because Whisper and Qwen are open-weight and run locally, no audio, transcript or family story ever leaves the machine during processing, and the whole pipeline works with the network cable unplugged. There's no per-recording API cost, so digitizing years of memos costs nothing but electricity. Because the model runs on our own machine, I could tune the structuring prompt freely, compare models side by side, and swap in a bigger one on a better GPU without anyone's permission or a pricing page. And a family archive should outlive any single product: in twenty years, when today's cloud APIs have changed or shut down, these model weights and this code will still turn Grandpa's voice into recipes.

Prize Categories

  • Best Use of Render: the live, shareable family recipe book runs on Render (https://echobook-8fpj.onrender.com/), deployed from a render.yaml blueprint with auto-deploy on every push. Inference is kept local on purpose, and only finished, family-approved recipes are deployed: a lightweight viewer with no ML dependencies that builds in seconds on the free tier.

Thanks for reading. If you have a grandparent with stories in their voice memos, record a few more this weekend. 🍲

来源:Google AI:DEV 作者专属(RSS) · dev.to