跳到正文
原文
Google AI:DEV 作者专属(RSS)· Shreyas·· 4 小时前AI 评分42

Aloud:用眨眼控制的浏览器 AAC 通信应用,为 ALS 患者发声

Aloud-Eye communication System

AI 导读

Aloud 是一款基于浏览器的眼控 AAC 通信应用,让无法用手或说话的人仅靠长眨眼、抬眉或握拳即可组句并朗读。它用 MediaPipe 的 FaceLandmarker 和 HandLandmarker 模型在浏览器本地完成手势检测,视频不出设备,配合 Web Speech API 语音输出,无需专用眼动硬件。

正文

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Aloud is a browser-based communication app for people who cannot use their hands or voice. People with ALS or locked-in syndrome often keep control of their eyes and face. Aloud lets them build sentences and speak them aloud using only a long blink, an eyebrow raise, or a closed fist held to the camera.

Items on screen highlight one at a time in a scanning loop. A deliberate long blink selects the highlighted item. Quick, natural blinks never select anything. It has a category-based phrase grid, a spell-it-out keyboard, full-screen spoken-message output, and a profile page showing real usage data (most-used words and phrases, message counts).

Who it's for: My family friend who lost is ablity to speak due to ALS.

The problem: Dedicated eye-tracking hardware for AAC costs thousands of dollars and is out of reach for most families. Aloud runs on a laptop webcam in a browser tab, with nothing to install and no special hardware.
**
Demo**
Live: https://hacktoberr.vercel.app/

Code

Aloud — Eye-controlled AAC

Aloud is an eye-controlled augmentative and alternative communication (AAC) app. It lets users who can only move their eyes build sentences and have them spoken aloud. The project focuses on clarity, accessibility, and minimal friction for low-mobility users.

Key principles

  • Design for eye- and switch-only input: large clear targets, one focal point per screen.
  • Accessibility-first: accessible names, screen-reader support, contrast and reduced-motion respect.
  • Minimal, deliberate visual language: all tokens come from styles/tokens.css.
  • Keep implementations small, clear, and well-justified — prefer short, readable code.

Tech stack

  • React 19
  • Next.js (App Router)
  • Plain CSS with design tokens (no Tailwind or CSS-in-JS)
  • @mediapipe/tasks-vision for eye/blink tracking
  • Web Speech API for text-to-speech
  • Gemini 2.5 Flash (lib/gemini.js, server-side REST API) for next-word suggestions

Quick start

  1. Install dependencies:
npm install
  1. Run the development server:
npm run dev

Notes: API keys (GEMINI_API_KEY) must be provided via .env.local and…

How I Built It

Stack: React 19, Next.js (App Router), @mediapipe/tasks-vision, Web Speech API, deployed on Vercel.

Open models at the core. All gesture detection runs on MediaPipe's open FaceLandmarker and HandLandmarker models, entirely in the browser:

Blink and eyebrow control use the FaceLandmarker blendshape scores (eyeBlinkLeft/Right, browOuterUp*), not raw landmark distances. A hysteresis threshold plus a sustained-duration check separates a deliberate hold from a natural blink.
Palm control uses the HandLandmarker model, loaded only when that mode is active. It is deliberately binary (open hand = idle, closed fist held briefly = select) so it's predictable for users with limited control.
One selection path. Every input mode calls the same select() function, and only one detection model runs at a time.

Privacy and offline behavior. Video never leaves the device. Detection, scanning, the keyboard, and speech via on-device OS voices all work without a connection once the page has loaded.

Hard problems:

Quick blinks selecting items. This took a full debugging cycle (hysteresis, sustained duration, head-movement suppression, adaptive baseline for lighting changes).
Stale-item bugs, where the item selected was the one highlighted at the moment of confirmation rather than at blink onset.
Slow startup from synchronous model loading, fixed with preloading, async init, and dynamic import.

I built this with an agentic coding tool (Antigravity), driven by staged single-feature prompts and an AGENTS.md file as the source of truth for design tokens, structure, and rules.

Honest status: [STATE WHAT IS ACTUALLY WORKING. E.g. "AI sentence composition via Gemini is scaffolded but not yet live; the app works fully without it."]

Why Does Open Innovation Matter?

Commercial eye-gaze systems are expensive because they are closed hardware plus closed software. Aloud is possible because Google released the landmark models openly and they run client-side:

Cost: a webcam replaces specialized hardware.
Privacy: a person's face video never has to be sent to a server. With a closed cloud vision API, it would have to be.
Offline use: a hosted API would make a person's ability to communicate depend on a connection and on someone else's pricing and uptime.

Tunability: because I could read how the blendshape outputs behave, I could build custom hysteresis and anti-false-positive logic on top, which matters when a false selection can say the wrong thing for someone who can't easily correct it.

来源:Google AI:DEV 作者专属(RSS) · dev.to