Deskferry 如何为 AI 智能体构建实时看板 Tasks
How we built a live Kanban board for AI agents
Deskferry 为在 Gmail、Slack、HubSpot 中工作的 AI 智能体构建了实时看板 Tasks,用户可把任务分配给智能体并观看卡片从 Backlog 移到 Done。
At Deskferry we run AI agents that work inside tools like Gmail, Slack and HubSpot. For a long time, the only way to see what an agent was doing was a run log. Useful for debugging, useless for actually managing work.
So we built Tasks: a Kanban board where users plan work, assign it to agents, and watch cards move from Backlog to Done in real time. This post covers the three design problems that mattered most: modelling task state, pausing for human approval, and keeping the board live.
1. A task is a state machine, not a status field
The first version had a status column and agents wrote to it whenever they felt like it. That fell apart fast: tasks got stuck, two updates raced each other, and the UI showed states that made no sense.
We replaced it with an explicit state machine. Every task is in exactly one state, and only defined transitions are allowed:
backlog -> in_flight (agent assigned and picks it up)
in_flight -> needs_you (agent hits an action that requires approval)
needs_you -> in_flight (user approves or answers)
needs_you -> backlog (user rejects, task returns for rework)
in_flight -> done (agent completes all steps)
in_flight -> backlog (agent fails; error attached to the card)
The four board columns map directly to these states, so the UI can never show something the backend doesn't consider valid. Every transition is written as an event with a timestamp and the actor (agent or user), which gives us an audit trail for free.
Lesson: if agents and humans both change the same object, make illegal states impossible at the data layer, not in the UI.
2. Human-in-the-loop as a first-class state
Most human in the loop AI setups we've seen treat approval as a blocking prompt: the agent waits on a promise until someone clicks a button. That's fine in a demo and terrible in production. Users approve things hours later, servers restart, and you can't hold a process open all afternoon.
We treat "waiting for a human" as a durable state instead:
Before any action tagged as sensitive (send, spend, write to a record), the agent stops and serialises what it was about to do, including the full proposed action and the context it used to decide.
The task transitions to needs_you and an approval item appears on the user's Desk.
The agent process ends. Nothing is held in memory.
When the user approves, edits or rejects, we rehydrate the agent from the saved checkpoint and continue from the exact step it paused on.
This makes approvals cheap. A task can sit in Needs You for a minute or a weekend with no cost. It also makes the approval screen better, because the user sees precisely what will happen, not a vague "Agent wants to continue."
Lesson: approval is a pause in a workflow, not a modal dialog. Persist it.
3. Keeping the board live
A board where cards move on their own only works if the user actually sees them move. Polling every few seconds felt sluggish and wasted requests on idle boards.
Instead, every state transition from point 1 publishes an event, and the board subscribes to events for the current workspace over a persistent connection. The client applies each event to its local copy of the board, so a card slides between columns the moment the backend commits the change.
Two details made this reliable:
Events carry a version number per task. The client ignores any event older than what it already has, which handles out-of-order delivery after a reconnect.
On reconnect, the client refetches the board once, then resumes applying events. Simple, and it avoids building a replay system.
Lesson: for AI agent monitoring, the event log you already need for auditing is also the best feed for your UI.
- What changed for users
The technical work was in service of one product shift. Before Tasks, people had to design an AI agent workflow before getting any value. With Tasks, they write a sentence and assign it. The agent works out the steps, the board shows progress, and the Desk collects every decision that needs a human.
That turned out to be the right abstraction for AI agent orchestration from the user's side: not a graph of nodes, but a list of tasks with owners.
Takeaways if you're building something similar
Model agent work as an explicit state machine with logged transitions.
Make human approval a persisted state with checkpoints, not a blocking call.
Drive your live UI from the same events you use for auditing.
Show users tasks and owners, not pipelines. They already know how to manage those.
Tasks is live in Deskferry now. If you want to see the board in action, you can try it free at deskferry.app. Happy to answer questions about any of this in the comments.
来源:Google AI:DEV 作者专属(RSS) · dev.to