跳到正文
原文
Google AI:DEV 作者专属(RSS)· Aizen-Kun·· 4 小时前AI 评分29

Proofdesk:在批准 AI 草稿前先审查证据

"Proofdesk: review the evidence before you approve the draft"

AI 导读

Proofdesk 是一个面向 AI 辅助草稿的编辑审查队列,将原文与修改稿并排展示,列出作者标注的主张及对应来源摘录,审稿人可批准或附理由退回,两种操作都不会直接发布内容。

正文

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange.

What I built

One of Proofdesk's sample drafts promises a 90% reduction in editing time for every team. There is no supporting source. It is an intentionally bad claim, and the app refuses to approve it.

Proofdesk is a small editorial review queue for AI-assisted drafts. It puts the original text beside a proposed revision, lists the claims the author has identified, and shows the source excerpts attached to them. A reviewer can approve the revision or send it back with a reason. Neither action publishes anything.

I built this with Hermes Agent, which generated the implementation and test scripts. This write-up was also prepared with AI assistance. The app itself does not call a paid model: it reviews supplied proposals rather than generating a new answer whenever someone opens a page.

That distinction kept the scope manageable. The part I wanted to explore was the handoff from a plausible draft to an explicit editorial decision.

Demo

Watch the 45-second guided walkthrough (MP4).

The video is a sequence of screenshots captured from the working app, not a continuous screen recording. The first four scenes demonstrate browser-local review decisions. The last scene shows the app reading the actual Sanity dataset. Each mode is labelled on screen.

The walkthrough shows an unchecked source blocking approval, a checked source allowing a recorded decision, and the unsupported performance claim going back for changes.

Sanity project ID: c033ya5z

Dataset: production

The dataset contains three demonstration briefs, not customer data. The screenshot below is the connected view:

Proofdesk reading its review queue from Sanity Content Lake

Code

Download the runnable source archive (ZIP).

The archive contains the Astro app, Node server routes, Sanity schema declarations, fixture seeder, and tests. It excludes credentials, dependencies, and compiled server output. It is a source download, not a Git repository.

To try the local demo after extracting it:

npm ci
npm test
npm run dev

Open http://127.0.0.1:4321. No account or API key is needed for that mode. The README explains how to connect a separate Sanity project using server-side credentials. Review access credentials for my dataset are deliberately not public.

How I used Sanity

A proofBrief document keeps the original copy, proposed copy, status, claims, source excerpts, and review history together. Claims refer to stable source IDs within that brief. The schema also records which source IDs a reviewer checked and the text of the proposal they reviewed.

The approval process uses Sanity's revision checks:

  1. The server authenticates the review request and fetches the latest document.
  2. It checks the revision the reviewer saw, the allowed transition, and the evidence attestations.
  3. It patches the status and history with ifRevisionID.
  4. It reads the document back and checks that the expected review event was saved.

The mutation is deliberately narrow:

{
  mutations: [{
    patch: {
      id: doc._id,
      ifRevisionID: doc._rev,
      set: { status: doc.status, history: doc.history }
    }
  }]
}

If another editor changes the document between the read and the write, Sanity rejects the patch. The app asks the reviewer to reload instead of overwriting the newer work.

This is a custom workflow built with Content Lake and the HTTP API. It does not use Sanity's Workflows product, App SDK, or Knowledge Bases.

What the build got wrong

The first review implementation handled approval but not rejection. A failing test exposed that: a brief without evidence could not even be sent back for changes. Rejection now requires a reason, while approval requires the evidence checks.

A browser test found another mistake. When browser storage failed, the interface briefly produced a warning and then overwrote it with a success message. The corrected version tells the reviewer that the decision exists only in the current tab.

There was also a test assumption to fix. Astro rejects some cross-origin form requests before the endpoint's JSON validator runs. The HTTP checks now distinguish the framework's rejection from the route's content-type validation rather than treating both as the same response.

These were useful corrections to AI-generated code. A clean-looking screen did not reveal them.

What was actually tested

Seven automated tests passed, along with HTTP checks, browser interactions, a production build, and a dependency audit that reported no vulnerabilities at the time of the build. The interface was inspected at desktop and mobile widths.

For the Sanity integration, a separate live test created a temporary document, approved it, and verified the stored event by reading it back. It then forced a concurrent update between the application's read and write. The stale patch was rejected, and the test confirmed that the document remained in review. Finally, the test removed its temporary document and verified the removal. The source archive includes qa/live.mjs and its recorded results.

The boundary I would keep

Proofdesk checks the structure of the evidence and records a reviewer's attestation. It cannot tell whether a quotation is accurate, whether a source is trustworthy, or whether the author left a factual claim out of the claim list. The intentionally unsupported sample has an empty citation list; the app is not doing semantic fact-checking.

The reviewer label is self-entered, and connected writes use a shared access token. That is adequate for a controlled prototype, not a production identity system. Browser history can be edited locally, and Sanity administrators can edit dataset records, so the log is not tamper-proof.

A production version would need individual reviewer identities and a controlled path for submitting and revising proposals. For this entry, the narrower workflow is working: a proposal stays in review until someone makes a recorded decision about that revision.

来源:Google AI:DEV 作者专属(RSS) · dev.to