本地 Gemma 3 4B 把化验单高值判为"正常范围",开发者改用代码做数值比较
"My local model called a flagged result 'within range', so I stopped letting it do arithmetic"
开发者用本地 Gemma 3 4B(Ollama、4-bit 量化、约 12.7 tokens/s)构建化验单解读工具 Plain-Words 时,发现模型把报告标记为偏高的结果说成"在参考范围内",且同一提示词两次回答不一致。
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
The first time I asked a local model to explain a lab report, it told me a result the report had flagged as high was "within the reference range". I ran the same prompt again and got a different answer.
I was building Plain-Words for Bhavya, a classmate. When a confusing document lands in their family, they ask the family doctor if it's medical, and upload it to Google and try to work it out themselves if it isn't, which they described as very stressful. A tool meant to make that less stressful can't be wrong about the one thing that matters. So most of this project is about what the model isn't allowed to decide.
What I Built
Plain-Words takes a photo, PDF or pasted text of a document and gives back a plain-language explanation. For a lab report, that's:
- a short summary of which results fall outside the range printed on the report
- one key point per result, tagged "Above listed range" or "Within listed range"
- a glossary of the terms and units
- specific questions to ask a doctor
- all of it again in Hindi, because that's what an older relative might read
There's a history with a delete button and a thumbs up/down on every explanation.
The version I built and tested runs entirely on my laptop with a local model, and nothing leaves the machine. I also put a hosted demo online, which works differently (see below).
Demo
Live demo: https://plain-words-ox04ps1vb-rehan-6df9.vercel.app?_vercel_share=9PNLhyXGNTZ4cv3g2bGr6V9km6pcsXW2
The hosted demo sends text to a server and to an AI provider, so use it with the fictional sample report only. There's a "Try the fictional sample report" button on the first screen. Photo reading may not work there.
The first version, which Bhavya tested (fictional sample report):
No real patient data is in the repo, the screenshots or this post.
Code
https://github.com/Rehan1604/plain-words
MIT-licensed. The Gemma model keeps its own terms.
How I Built It
The front end is React and Vite. Behind it is FastAPI, with Tesseract for photos, SQLite for history and a model behind a small LLMProvider interface.
phone browser → FastAPI → PDF text / Tesseract OCR
├→ range parser (plain code): above / below / within
└→ LLMProvider → Gemma 3 4B (local, Ollama): glossary + Hindi glossary
→ SQLite
The decision that matters most is what the model doesn't get to do. My first attempts simply asked Gemma to explain a sample HbA1c report. On a report that flagged the result high, it told me it was within the reference range, and in a bare prompt it drifted into diet advice. A small model on a CPU isn't a reliable number comparer.
So the app doesn't ask it to. A parser reads each result and the reference range printed next to it, and plain code decides above, below or within. That output becomes the summary, the key points and the doctor questions, with no model involved. The model only writes the glossary, and the Hindi version of it. For documents with no numbers, like a notice or a form, the model writes the summary instead. That path is less reliable, and I haven't tested it much.
Each glossary entry quotes a line from the document, and the code checks the line really appears there. That proves the quote exists, not that the explanation is right. The meanings are the model's general knowledge and can be wrong.
The local model. Gemma 3 4B through Ollama, on the CPU. I'd planned around having a GPU, but the laptop I built this on has only Intel UHD graphics (i5-13420H, 16 GB RAM). The 4B model at 4-bit quantization ran at about 12.7 tokens per second, which is slow but usable. I compared it with Qwen 2.5 3B, which was faster at about 15.9 tokens per second. Its Hindi was gibberish: it talked about an injection and "five to ten months" for a blood sugar value that had neither in it. For a tool that's supposed to calm someone down, that settled it.
Gemma comes under Google's own Gemma Terms of Use, not a standard open-source license, so I call it open-weight, not open-source.
The hosted demo. Because everything sits behind LLMProvider, the hosted option was a new provider class plus a few config values, selected with one environment variable (LLM_PROVIDER=groq). The backend runs on Render and calls openai/gpt-oss-20b, an open-weight model, through Groq. I haven't measured speed or Hindi quality on that model, so every timing and Hindi claim in this post is about local Gemma.
Speed (local). After I added the guardrails, a full run took 262 seconds. I'd made the model write a summary, key points and doctor questions, then replaced all of it with code anyway. Leaving only the glossary to the model brought it to about 35 seconds. The Hindi step fell from 52 to about 19 seconds for the same reason.
There are 10 automated tests: the range parser, the Hindi templates and the API round trip with a fake model. I checked the photo path, the real model output and the Hindi quality by hand.
Why Does Open Innovation Matter?
- Privacy, when you run it locally. A lab report is the kind of document I'd hesitate to paste into a cloud service. In the local version the photo, the extracted text and the history stay on one machine. The hosted demo gives that up, which is why it carries a warning and I only use the fictional report there.
- Cost. The local version has no API bill and no account. A student laptop is enough.
- Picking and swapping the model. I put Gemma and Qwen side by side on the same task and chose the one that handled Hindi. Later I added a hosted open-weight model behind the same interface without touching the rest of the app.
- Designing around the weaknesses. The model is swappable and runs where I choose, so I could build around what it gets wrong instead of waiting for a better release.
Testing It With Bhavya
Bhavya tried the local version on an Android phone over Wi-Fi, using the sample report (fictional data) and the in-app camera. What follows is paraphrased from their written feedback, which they cleaned up afterwards, so none of it is a direct quote:
- It worked. The OCR read most of the text, though a couple of numbers and formatting spots were harder.
- They waited and wondered whether it was still processing.
- They found the result much easier to understand than the original report.
- The Hindi sounded natural rather than word-for-word, and leaving some medical words in English was fine. They thought it would help parents and grandparents.
- They had to ask what the "ask your doctor" lines meant, and I explained they weren't a diagnosis. They also weren't sure at first whether the numbers were normal until they read below.
- They'd use it again for blood-test reports, prescriptions and relatives' documents, to get the meaning without looking up every word.
- They suggested a clearer progress state, and said to keep the warning to check the original document.
One tester and one sample document, so I wouldn't call this validation.
What Changed After That
| Bhavya's feedback (paraphrased) | What changed |
|---|---|
| Wondered whether it was frozen during the wait | Spinner and a "still working, not frozen" message, plus a note that photos take longer |
| Wasn't sure if numbers were normal until reading below | An "Above / Within listed range" badge on every result, instead of color alone |
| Had to ask what the doctor lines meant | A one-line "not a diagnosis" note under that heading, in English and Hindi |
| Photos can misread numbers; keep the check-the-original warning | A panel showing the exact text read from the photo, with a note to compare it against the paper |
Afterwards I also redesigned the interface (range bars, summary counts, a one-tap sample button, dark mode). That wasn't based on Bhavya's feedback, and Bhavya hasn't seen these changes yet.
What I Got Wrong
- I let the model write the whole explanation, then replaced most of it with code anyway. That was the 262-second version.
- I asked the model to write Hindi for sentences the code had already computed. It returned English. I switched to fixed Hindi templates for those parts.
- Gemma sometimes fell into a repetition loop while writing JSON, and Ollama killed the request with a "token repeat limit reached" error. I capped the output length and retry when that happens.
- One glossary entry in Hindi contradicted itself, saying "after eating" and "after not eating for 8 hours". The 8 hours wasn't in the report at all. I tightened the prompt against adding numbers, but it can still happen.
- My test script looked dead because it printed nothing for over a minute, so I added a heartbeat line. Bhavya hit the same wall in the app.
What I Didn't Do
- I haven't tested it on real reports, only a fictional sample, and only with one tester.
- The parser handles two line formats. Other layouts fall back to the model alone, which is less reliable.
- Bhavya was the only Hindi reader.
- The glossary meanings are the model's general knowledge. Nobody has checked them against a medical source.
- The hosted backend has no login or rate limiting, and its CORS setting is wide open. Anyone with the URL could call it, so the Groq key's usage is capped on their side.
- I haven't measured the hosted model's speed or Hindi, and photo reading there is untested.
- No staged progress bar yet (reading the photo, then explaining it).
Prize Categories
Best Use of Gemma. Gemma 3 4B, an open-weight model, runs locally through Ollama and writes the glossary, the Hindi glossary and the summary for documents without numbers. Code, not the model, compares the numbers, because Gemma got that wrong. (The hosted demo uses a different model, so to see Gemma in action, run the local version.)
来源:Google AI:DEV 作者专属(RSS) · dev.to
