Don't type,
just point and speak.
Works with the agents you already use
How it works
Three steps. No prompt draft.
- Step 1
Record your screen
Capture one window or every display in a single session.
- Step 2
Point and speak
Circle the UI and say what you want changed. Skip the written brief.
- Step 3
Hand off to your agent
Open the session in Cursor, Claude, or Codex and let it take the next step.
Built for agents that read folders, not guess from screenshots
Mark the exact UI
Pen, rectangle, and text — so the agent knows which button you mean.
Fewer tokens, clearer intent
You get a folder the agent can skim — frames and speech, not a giant video dump.
Multi-monitor in one take
App on one display, editor on the other — captured together.
Talk through the change
Your voice is transcribed and lined up with what was on screen when you said it.
“Make this button match the sidebar colour…”
Handoff where you already work
Open the session in Cursor, Claude Code, or Codex — no new chat app.
Screenshots vs recordings
A screenshot can't show what changed.
Static shots lose clicks, timing, and what you said. Annotate keeps the motion so your agent sees the full story.
One frozen frame. No clicks, no timing, no voice — the agent guesses the rest.
“Make this button match the sidebar colour…”
Frames, gestures, and speech stay in sync — so the agent follows what you meant, not what one still frame happened to show.
Stays on your Mac. Never uploaded.
Frames and transcripts live on your machine. There's no cloud upload for your screen.
Local by design — nothing leaves the device.
Next time the task is hard to explain, show it.
macOS · Apple Silicon (arm64)