Feedback Compiler:为朋友打造的隐私优先本地垂直切片
Feedback Compiler: a privacy-first local vertical slice for a friend
Feedback Compiler 是一个隐私优先的本地工作区,通过 Ollama 在本地运行开源 Gemma 模型 gemma4:e2b-it-qat,把 Slack 消息、邮件、会议记录等异构反馈编译为含待办行动、已确认决策、冲突与未决问题、截止日期和来源引用的结构化审阅面板。
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
Feedback Compiler is a small privacy-first local workspace built for a friend who receives feedback from a team through multiple communication channels.
The information does not arrive as a clean task list. It arrives as:
- Slack-style messages;
- emails;
- meeting notes;
- support tickets;
- multilingual chats;
- transcript excerpts.
The problem is not simply summarising these messages.
The useful questions are:
- What should become an action?
- What has already been decided?
- What is still unresolved?
- Are two requests duplicates or separate?
- Is there a genuine conflict?
- Is there a deadline?
- Which original messages support each conclusion?
Feedback Compiler transforms these fragments into a structured review surface containing:
- proposed actions;
- confirmed decisions;
- conflicts and open questions;
- deadlines;
- duplicate relationships;
- source references.
The person remains responsible for reviewing and approving the result.
This is intentionally a vertical slice. It is not an autonomous project manager, an enterprise integration platform or a claim of production-ready semantic accuracy.
Demo
Watch the 1:37 Loom walkthrough.
The recording shows the local product flow:
- heterogeneous feedback enters the workspace;
- the input keeps its source identity;
- the local model compiles the batch;
- the user reviews actions, decisions, conflicts and deadlines;
- the result is prepared for a manual handoff.
The recording uses synthetic feedback. It shows the local application, not a hosted model endpoint.
There is also an interactive public replay demo.
The public page allows reviewers to browse:
- 30 benchmark cases;
- 8 heterogeneous casebook examples;
- a six-input mixed set;
- paired structured outputs.
The public replay uses preserved synthetic fixture outputs. It never contacts Ollama.
Why I Started With This Problem
I started with a person and a repeated workflow rather than with a model.
A general-purpose summariser can produce fluent text while still making consequential mistakes:
- turning an opinion into a decision;
- inventing a deadline;
- merging requests with different scopes;
- losing the source of a conclusion;
- hiding a conflict inside a short summary;
- treating an unresolved question as completed work.
The core product rule became:
Every proposed interpretation should remain connected to the feedback that supports it, and the person should be able to review it before it becomes work.
That rule shaped both the interface and the evaluation process.
From Idea to Vertical Slice
The initial idea was much broader: understand workplace communication and connect the result directly to tools such as Slack, Linear, Jira or Notion.
That scope was too large for a trustworthy first version.
I reduced the problem to one complete loop:
feedback fragments
↓
local compilation
↓
structured, source-linked output
↓
human review
↓
prepared handoff
The current project proves this loop end to end.
It does not attempt to solve authentication, real-time integrations, enterprise retention, compliance or automatic task creation. Those are separate engineering and product problems.
Starting with one complete loop made it possible to evaluate the semantic task before building a large integration layer around uncertain behaviour.
How I Built It
Feedback Compiler is a local React/Vite web application connected to Ollama on the same computer.
The project uses the open-weight Gemma model gemma4:e2b-it-qat. The model receives heterogeneous feedback and returns structured output containing actions, decisions, conflicts, duplicates, open questions and deadlines.
The application adds a review layer around the model:
- source IDs remain attached to every result;
- structured output is validated before being shown;
- deterministic post-processing normalises relations and categories;
- the user reviews the result before any handoff;
- destination buttons prepare text but do not write to external services.
The model is not fine-tuned for this project. The initial specialisation comes from the input contract, the conservative extraction rules, the output schema and the evaluation process.
The repository contains the local setup commands, the application, the datasets, the evaluation artifacts and the privacy documentation.
How the Product Works
The application has four main review areas.
Do
Proposed actions that may become work.
Decide
Choices or commitments that appear to have been explicitly confirmed.
Clarify
Conflicts, ambiguities, unresolved questions and missing information.
Trace & timing
Source references, duplicate relationships and deadlines.
The user can:
- load the mixed example set;
- add or edit feedback;
- compile the batch locally;
- inspect the structured review board;
- trace each item back to its source;
- edit, accept or reject the interpretation;
- prepare text for another tool.
The handoff buttons prepare text for Slack, email, Linear, Jira or Notion. They do not log in, publish or write automatically.
Privacy as a Product Boundary
The main architectural decision was local inference.
The model runs on the same computer as the workspace through Ollama. After the one-time model download, inference can run without sending the feedback to a hosted inference API.
The current application does not:
- authenticate with Slack, email, Linear, Jira or Notion;
- publish content automatically;
- modify a remote work system;
- store private feedback in a hosted database;
- claim that local processing automatically makes sensitive information safe.
The active browser session is temporary. Input and output remain in the active session and are not retained after the page is closed, unless the user explicitly exports them.
“Local” is a boundary, not a guarantee.
Real use would still require:
- redaction of unnecessary personal or client information;
- retention and deletion rules;
- protection of exported files;
- network-boundary verification;
- access controls;
- backup and synchronisation review;
- threat modelling;
- security review.
The project is privacy-first because the data path is deliberately constrained. It is not privacy-proof by slogan.
Choosing the Model
The project uses the local Gemma model gemma4:e2b-it-qat through Ollama.
The goal was not to select the largest model available. The goal was to use a compact model that could run on a personal computer while keeping the inference path inspectable and replaceable.
The model has not been fine-tuned on private client material.
The current behaviour comes from:
- a conservative system instruction;
- a structured output contract;
- strict source-ID preservation;
- explicit category rules;
- deterministic validation;
- deterministic post-processing;
- human review.
Fine-tuning may become useful later, but only after building a carefully reviewed, privacy-safe and error-driven dataset.
Making the Inputs More Realistic
The first examples were short and strongly oriented toward one design workflow.
That was useful for prototyping, but it did not represent the intended daily work.
I expanded the input surface to include:
- short direct requests;
- Slack-style threads;
- email feedback;
- meeting notes;
- support tickets;
- multilingual exchanges;
- ambiguous feedback;
- prompt-injection-like content;
- transcript-style excerpts.
The interface preserves input type as metadata, but the label itself is not treated as evidence.
A message labelled “meeting notes” does not automatically become a decision. The text must still support that interpretation.
Transcript support is intentionally limited. A production transcript pipeline would require first-class speaker attribution, timestamps, chunking and cross-chunk relation handling.
Building the Evaluation Before Polishing the Product
I created a controlled synthetic benchmark of 30 cases.
The cases cover failure modes such as:
- missed requests;
- opinions incorrectly classified as decisions;
- false duplicates;
- missed duplicates;
- false conflicts;
- missed conflicts;
- invented details;
- date hallucination;
- negation failures;
- prompt injection;
- multilingual duplicates;
- unsupported coreference;
- missing priorities;
- deadline conflicts;
- stale or superseded requests.
I then created a separate heterogeneous casebook of 8 cases to test input diversity.
Finally, I added a small attributed public-data input pack shaped from:
- PolyAI Banking77;
- OpenAssistant/oasst1;
- QMSum.
The public-data pack was not used as training data and was not treated as semantic gold. Its purpose was to test whether the input pipeline could handle different communication shapes while preserving structure and provenance.
The evaluation layers have different purposes:
- the controlled benchmark tests explicit semantic rules;
- the heterogeneous casebook tests input diversity;
- the public-data pack tests transport and shape tolerance;
- the manual review tests whether structurally valid output is actually useful.
None of these proves that the system understands all real workplace communication.
Testing and Iteration
The final controlled run used the local Gemma model with deterministic post-processing.
Results:
- 30/30 cases completed;
- 30/30 outputs were schema-valid;
- 30/30 outputs preserved valid provenance;
- 0 runtime failures;
- 28 heuristic passes;
- 2 heuristic partial results;
- 0 heuristic failures.
The heterogeneous run produced:
- 8/8 schema-valid outputs;
- 8/8 provenance-valid outputs;
- 0 runtime failures;
- 6 heuristic passes;
- 2 heuristic partial results.
The public-data input pack produced:
- 8/8 schema-valid outputs;
- 8/8 provenance-valid outputs;
- 0 runtime failures.
These are contract and workflow results. They are not semantic accuracy scores.
A valid JSON response can still contain an incorrect interpretation.
For that reason, I manually reviewed the controlled result:
- 13 cases were accepted;
- 16 were partial;
- 1 was rejected.
The main remaining quality problems were:
- missing structured scope;
- lost conditions;
- priority handling;
- deadline normalisation;
- dense multi-relation reasoning.
This changed the product direction. The system should not pretend to be autonomous while meaningful interpretation debt remains.
What Changed During Development
The first local baseline exposed several problems:
- weak provenance preservation;
- inconsistent relation handling;
- over-production of requests;
- confusion between decisions, requests and open questions;
- slow or unstable reasoning in some configurations.
I improved the system incrementally by adding:
- stricter source-ID constraints;
- clearer category precedence;
- explicit opinion-versus-decision rules;
- conservative duplicate and conflict handling;
- relative-deadline preservation;
- protection against instructions embedded inside feedback;
- deterministic post-processing;
- schema and provenance validation;
- manual contract-aware review.
The most important improvement was not simply increasing a score.
It was making the failure surface visible.
The system now distinguishes between:
- runtime failure;
- invalid structure;
- invalid provenance;
- heuristic mismatch;
- semantic review debt.
That separation makes future improvements measurable.
Deployment and Local Use
The public replay is only a review surface.
To run the actual compiler:
git clone https://github.com/MathCat0000/feedback-compiler-week1.git
cd feedback-compiler-week1
# Install Ollama from https://ollama.com/download
npm install --prefix demo
npm run model:pull -- --model gemma4:e2b-it-qat
npm run app
Then open:
http://localhost:5173/#workspace
The first model download requires internet access. After installation, the compilation path remains local through Ollama.
The repository contains:
- the React workspace;
- the local Ollama setup flow;
- the structured output contract;
- the evaluation datasets;
- the heterogeneous input cases;
- attributed public-data input shapes;
- selected evaluation artifacts;
- privacy documentation;
- reproducible tests.
It does not contain private client feedback or team transcripts.
How a Company Version Could Be Extended
The current repository does not contain enterprise integrations.
However, the vertical slice establishes a contract that could support them:
company input adapters
↓
controlled local or private inference boundary
↓
source-linked structured output
↓
human approval
↓
company-approved connectors
A future deployment could add adapters for:
- Slack;
- Microsoft Teams;
- email;
- customer-support systems;
- ticketing tools;
- meeting transcription systems;
- internal knowledge platforms.
Those connectors could run inside a company-controlled environment, with outbound actions limited to reviewed and approved results.
That would still require:
- authentication and authorisation;
- tenant isolation;
- audit logs;
- retention policies;
- redaction;
- encryption;
- network controls;
- security review;
- explicit approval before external actions.
The current project demonstrates that the local compilation and review layer can be built first. It does not claim that the enterprise integration layer already exists.
Why Does Open Innovation Matter?
Open innovation made it possible to build around a local model instead of making a hosted API the centre of the product.
That matters because privacy is part of the user problem.
With a local model, the project can inspect:
- which model is running;
- where inference happens;
- what is sent to the model;
- what output is retained;
- where human review is required.
The trade-offs are real:
- setup takes time;
- local inference can be slower;
- compact models can be less reliable;
- the user must still review the result.
Open innovation makes those trade-offs visible rather than hiding them behind a remote API.
What This Project Does Not Claim
Feedback Compiler does not currently claim:
- perfect semantic accuracy;
- autonomous task creation;
- production-ready transcript processing;
- enterprise compliance;
- automatic privacy protection;
- authenticated integrations;
- successful validation with real client data;
- fine-tuning for a specific organisation.
It is a working local vertical slice that demonstrates a narrower claim:
A compact local model can help transform heterogeneous feedback into a structured, source-linked review surface while keeping human approval in the loop.
That claim is supported by the preserved local runs, but broader real-world validation is still required.
My Agent Session
I am not publishing an agent session link.
The working session could expose private project context and development details that are not necessary to evaluate the product. The public repository, evaluation artifacts and Loom walkthrough provide the relevant evidence.
Code
The complete repository is available here:
Prize Categories
Best Use of Gemma
来源:Google AI:DEV 作者专属(RSS) · dev.to