The open, agent-agnostic version of macOS computer use. A single signed Swift binary that exposes your Mac as a standard MCP server. Point Claude Code, Cursor, Codex CLI, Gemini CLI, or your own agent at it and the agent can see and operate the apps on your Mac — in the background, without hijacking your cursor or stealing focus.
Status: pre-1.0, used in production by the authors. macOS only. Runs while the Mac is unlocked (the system is kept from idle-sleeping during active sessions; if the screen locks, mutating tools pause with a recoverable error until you unlock).
|
Placeholder for The agent drives an occluded background app — visible agent cursor gliding to each target, "Agent working" chip on screen — while the user keeps typing in a different app. No focus fight, no cursor hijack. |
- Why it's different
- Install
- Quickstart per client
- How it works
- Architecture
- Tools
- What's new this cycle
- Requirements
- Safety
- Configuration
- Comparison
- Distribution notes
- Known limitations
- Development
- Contributing
- License
- Universal — works on every app. Accessibility-first for precision, with an automatic pixel-coordinate fallback (z-order hit-testing) for apps with poor or absent accessibility trees. The agent's own vision provides the grounding — no bundled ML model. Web apps included: Chromium/Electron web-content accessibility is enabled on demand, and structural wrapper nodes are collapsed so deeply nested page content actually reaches the agent.
- Background-safe. A layered input ladder (AX action → per-window event → per-PID event, with explicit global fallback for clicks) delivers actions to the target app without moving the real cursor or changing focus by default. You keep working while the agent works.
- Agent-agnostic. Standard MCP over stdio. Any compliant client connects with one line of config — no lock-in.
- Observable. A smooth self-drawn agent cursor (separate from your real pointer) glides to each target so you can watch what the agent does.
- Verified outcomes. Every mutating action re-reads the target and reports whether the effect actually happened — not just whether the call returned without throwing. See What's new this cycle.
- Multi-session safe. Sessions are thin shims over one shared engine daemon (spawned on demand, retired on version changes, self-reaping when idle), so any number of concurrent agents go through a single process that owns capture, accessibility, input, and the cursor — and per-app leases keep two agents from interleaving actions inside the same app.
- Reliable. Every action returns fresh app state (screenshot + accessibility tree). Element IDs retain exact live AX handles and fail stale if the app recreates or detaches them; destructive actions pass a confirmation policy.
- One native binary. Zero runtime dependencies, frictionless install.
macOS 14 or newer. On first run the binary asks for Accessibility and Screen Recording permission (Input Monitoring is not required).
Heads-up: the first tagged release may not be published yet. If the Releases page is empty, use build from source below. These steps are correct as soon as the first release is cut.
Download the notarized .app (or standalone binary) from the
latest release,
move it to /Applications (or anywhere on your PATH), and confirm it runs:
computer-use-mcp version
computer-use-mcp doctor # check Accessibility / Screen Recording grantsA notarized, stably-signed build matters here: macOS ties Accessibility and Screen Recording grants to a signing identity, so a notarized bundle keeps its permissions across upgrades (see Distribution notes).
git clone https://github.com/minghinmatthewlam/computer-use-mcp.git
cd computer-use-mcp
swift build -c release
.build/release/computer-use-mcp versionThe binary is .build/release/computer-use-mcp. Put it on your PATH, or point
your MCP client's command at the absolute path. To build the stable-identity
.app wrapper locally:
python3 scripts/build_app_bundle.py # local .app wrapper
python3 scripts/deploy_app_bundle.py # release build + install to ~/ApplicationsA Homebrew formula is drafted under
packaging/homebrew/computer-use-mcp.rb
as groundwork. It is not tapped or published yet; once the first release ships
with a stable tarball, the intended flow is:
# Not live yet — planned:
brew install minghinmatthewlam/tap/computer-use-mcpThe server speaks MCP over stdio; every client launches it the same way —
computer-use-mcp serve. Pick your client below.
Claude Code
claude mcp add computer-use -- computer-use-mcp serveThat registers the server for the current project. Use -s user to register it
globally for all projects. Verify with claude mcp list.
Cursor
Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (per-project):
{
"mcpServers": {
"computer-use": { "command": "computer-use-mcp", "args": ["serve"] }
}
}Codex CLI
Add to ~/.codex/config.toml:
[mcp_servers.computer-use]
command = "computer-use-mcp"
args = ["serve"]Gemini CLI
Add to ~/.gemini/settings.json:
{
"mcpServers": {
"computer-use": { "command": "computer-use-mcp", "args": ["serve"] }
}
}Any other MCP client
Any MCP-compliant client that launches a stdio server works. The command is
computer-use-mcp with a single argument serve:
{
"mcpServers": {
"computer-use": { "command": "computer-use-mcp", "args": ["serve"] }
}
}If the binary is not on your PATH, use its absolute path as command.
Each MCP client spawns serve, a thin stdio shim; tool calls are forwarded to a
shared engine daemon (one per user, spawned on demand over a unix socket) that
owns accessibility, screen capture, input delivery, and the agent cursor. One
engine process means concurrent agent sessions cannot collide on shared system
services, and short per-app leases keep two sessions from interleaving actions
inside the same app. Tool calls fail fast with structured DAEMON_UNAVAILABLE
metadata if the daemon cannot be reached; they never fall back in-process.
The daemon also owns privacy-filtered runtime metrics: events are buffered and
flushed in bounded batches to rotating JSONL plus an aggregate summary.
Click interactions first resolve to an accessibility element and a screen point, then descend a delivery ladder, stopping at the first tier that works:
- Accessibility action (
AXPress, etc.) — precise, background, no event posted. - Per-window event — a
windowNumber-routed event delivered to the target process, so the action lands without activating the app or moving the cursor. - Per-pid event — delivered to the process when no window id resolves.
- Global cursor — opt-in last resort only (
allow_global_cursor: true); it moves the real pointer, then restores it.
State is re-perceived after every action and returned to the caller, so the agent always acts on current ground truth. Element ids carry a snapshot generation, so reusing a stale id fails loudly instead of mis-clicking.
For the precise production contract across observation, dispatch, coordinate spaces, foreground/background guarantees, TCC requirements, stale snapshots, and failure recovery, see Modality Contract. For strict background focus/cursor behavior, see Background Control Contract.
flowchart LR
subgraph Clients["MCP clients"]
C1["Claude Code"]
C2["Cursor"]
C3["Codex CLI"]
C4["Gemini CLI / your agent"]
end
C1 & C2 & C3 & C4 -->|stdio MCP| SHIM["serve (thin stdio shim)"]
SHIM -->|unix socket| DAEMON["shared engine daemon<br/>(one per user, per-app leases)"]
subgraph ENGINE["Engine"]
AX["Accessibility<br/>(perceive + AX actions)"]
SCK["ScreenCaptureKit<br/>(background-safe capture)"]
LADDER["Input ladder<br/>AX → per-window → per-pid → global"]
OVERLAY["Agent cursor + status chip overlay"]
end
DAEMON --> AX & SCK & LADDER & OVERLAY
AX & SCK & LADDER & OVERLAY -->|drive / observe| APPS["Target macOS apps<br/>(foreground or occluded)"]
A longer walkthrough of the process model, dispatch ladder, outcomes, and source map lives in docs/architecture-overview.html.
Perceive get_app_state (with scope_element_id/max_elements for huge
windows, skeleton: true for a shallow overview of a large tree, and ocr: true
for apps that draw their own UI) · list_apps · list_windows
Act click · type_text · press_key · scroll · drag · set_value ·
select_text · perform_secondary_action
System open_app · open_url · manage_window · read_clipboard ·
write_clipboard · health_report
Element actions use ids from get_app_state; coordinate-capable tools also
accept raw screenshot coordinates. The intentionally small surface keeps
perception, delivery, and verification in one inspectable path.
Action results return a reduced-resolution screenshot to keep the agent loop fast,
and skip resending the element tree when the action changed nothing (existing ids
stay valid). When the UI did change, results carry a compact diff of what changed,
appeared, or disappeared — elements that survive a change keep their ids, so
everything the agent holds stays valid. Pass include_screenshot: false for
tree-only results, include_state: false for a bare confirmation (fastest), and
call get_app_state whenever full-resolution pixels are needed.
Three capabilities landed on main this cycle; the README documents them so it
stays truthful about current behavior.
Mutating tools no longer report "success = the call didn't throw." Each action
now does a read → act → re-read of the target and classifies whether the intended
effect actually occurred, surfaced in a computer-use-mcp/outcome block in the
result's _meta. The classification is one of four values:
| Classification | Meaning |
|---|---|
success |
The effect was observed, or the target was already in the requested state (an idempotent no-op is a success). |
unsupported |
The target cannot perform this action (disabled control, no settable value). Retrying won't help. |
effect_not_verified |
Dispatched without error, but no confirming change was observed — a failure_domain distinguishes a likely-dropped background event (transport, retry at a higher tier may help) from a control that lied about acting (verification). |
verifier_ambiguous |
The action may well have worked, but the verifier couldn't read enough state to prove it (secure field, unobservable menu item). Never a false failure. |
isError is unchanged — a thrown exception still sets it, and everything else
stays false — so agents that ignore _meta behave exactly as before. Agents
that read the outcome block get an honest verdict plus a plain-language sentence in
the text body for non-success results. The full design, including the per-tool
false-success trap matrix, is in
docs/outcome-contract.md.
The input ladder silently falls from tier to tier; results previously reported
only the tier that finally landed the event. A computer-use-mcp/delivery block
in _meta now carries delivery_tier, a fallback_reasons array explaining why
each higher tier was skipped (e.g. axActionUnsupported, windowNumberUnresolved,
eventBridgeFailed, globalCursorRequested), and ui_changed. This is pure
telemetry — it does not change which tier is attempted — and it lets an agent
reason about whether escalating the delivery tier is worth trying.
Two ways to keep large windows from blowing up the tree:
skeleton: truereturns a shallow overview: the outline recurses a few levels, then a deeper container is emitted with achildren_countannotation instead of its subtree, and stays a drill target. Pass that container's id asscope_element_idfor a full scoped re-query. Skeleton is the overview,scope_element_idis the drill-in, andmax_elementsstill bounds either.- Dense-collection viewport windowing (automatic): virtualized collections
(lists, tables, outlines, grids) with many children are windowed to the
on-screen slice — preferring the app's own visible-rows attributes, else a
viewport-frame intersection — instead of a blind first-N prefix that ignores
scroll position. Off-window items are summarised as a count on the container's
line (never silently dropped), and materialised rows retain live-handle identity
while the app keeps the same AX objects. Use
scope_element_idto inspect a specific retained container with the full element budget.
- macOS 14+
- Permissions granted on first run: Accessibility and Screen Recording
(Input Monitoring is not required). Run
computer-use-mcp health_reportto inspect current identity/permission state, orcomputer-use-mcp doctor --promptwhen you intentionally want macOS prompts.
The server gates risky actions itself (it does not trust the calling agent).
Destructive/irreversible button clicks (Delete, Erase, Reset, …), typing into
secure password fields, and actions against apps on a confirmation list return a
recoverable Confirmation required: … error until the caller retries with
"confirm": true.
Browser pages get their own gate: before acting in a known browser the server
reads the current URL from the accessibility tree and applies the URL policy —
url_deny patterns block the action outright (confirm does not override),
url_confirm patterns (plus built-in payment-page defaults) require confirm
per action. The server also yields to the human: when real hardware input was seen
in the last second and the target app is the one the user is working in (or the
action uses the global cursor), the call returns a recoverable error instead of
interleaving with the user (see interference_idle_seconds).
See SECURITY.md for the TCC permission model, threat model, and private disclosure process.
Every option is settable as an environment variable (COMPUTER_USE_MCP_<KEY>) or a
key in ~/.config/computer-use-mcp.json (env wins):
| Key (file) / variable | Effect |
|---|---|
cursor / COMPUTER_USE_MCP_CURSOR=0 |
Hide the animated agent-cursor overlay (on by default; set 0 for headless/CI). |
cursor_idle_fade |
Seconds of quiet before the agent cursor fades (default 12). |
cursor_topmost / COMPUTER_USE_MCP_CURSOR_TOPMOST=1 |
Keep the agent cursor unconditionally above every window (escape hatch from target-relative z-order). |
status_chip / COMPUTER_USE_MCP_STATUS_CHIP=0 |
Hide the "Agent working" pill shown on every display during activity (on by default). |
no_safety / COMPUTER_USE_MCP_NO_SAFETY=1 |
Disable the safety policy entirely. |
confirm_apps |
Apps (name or bundle id) where every action needs confirm. |
destructive |
Extra destructive label substrings to gate. |
url_deny |
URL substrings where browser actions are blocked outright (confirm does not override). |
url_confirm |
Extra URL substrings where browser actions need confirm (defaults cover payment pages). |
ax_timeout |
Per-call accessibility timeout in seconds (default 2). |
ax_element_timeout |
Per-element AX messaging timeout during tree traversal (default 0.25s). |
viewport_probe_timeout |
Tighter AX timeout for dense-collection viewport frame probes (default 0.05s). |
no_app_lease |
Disable per-app session arbitration. |
app_lease_seconds |
How long an app stays leased to a session after its last action (default 10). |
no_interference_yield |
Disable yielding to real user input (yield is on by default). |
interference_idle_seconds |
Hardware quiet time required before acting in the app the user is working in, or via the global cursor (default 1; 0 disables). |
no_sleep_assertion |
Do not hold a prevent-idle-sleep assertion while tool calls are flowing. |
no_telemetry / COMPUTER_USE_MCP_NO_TELEMETRY=1 |
Disable funnel/telemetry recording. |
show_meta / COMPUTER_USE_MCP_SHOW_META=1 |
Dump _meta (focus / delivery / outcome) on call harness output. |
log / COMPUTER_USE_MCP_LOG=1 |
Per-tool-call stderr log lines (name, ok/error, duration). |
max_actions_per_sec |
Optional global throttle on tool calls (off by default). |
COMPUTER_USE_MCP_SKYLIGHT=1 |
Env-only: enable the opt-in SkyLight SLEventPostToPid input rung (not a config-file key). |
How computer-use-mcp relates to the closest projects. This is a factual snapshot; cells marked "Not public" mean the project is closed-source and the behavior isn't documented, not that it's absent.
| computer-use-mcp | OpenAI Codex computer use | actuallyepic/background-computer-use | lahfir/agent-desktop | |
|---|---|---|---|---|
| Protocol | MCP over stdio (agent-agnostic) | Closed, Codex-only | HTTP loopback (no auth) | Rust CLI |
| Background input | AX action → per-window → per-pid event ladder, opt-in global cursor | Not public | Weaker; only physical path is an experimental WindowServer route | Global HID tap (raises windows) |
| Capture API | ScreenCaptureKit | Not public | CGWindowListCreateImage (deprecated) |
screencapture shell-out |
| Verified outcomes | Yes (classification enum in _meta) |
Not public | Yes | Yes |
| Visible agent cursor | Yes (self-drawn overlay) | Not public | No | No |
| Multi-session | Yes (shared daemon + per-app leases) | N/A (hosted) | No | No |
| Safety gates | Yes (destructive / URL / interference / screen-lock) | Provider-side | Minimal (no auth) | Not documented |
| License | MIT (open source) | Closed | Open source | Open source |
The two independent open-source efforts above both converged on verified outcomes ("don't trust an AX success — re-read and confirm the effect"), which is the model computer-use-mcp adopts in docs/outcome-contract.md. computer-use-mcp's distinguishing bets are the standard MCP surface, the shared-daemon multi-session model, and the visible agent cursor.
The binary needs Accessibility and Screen Recording permission, and macOS
ties those grants to the host process that spawns the server (your terminal or
agent app). A rebuilt binary keeps its grants; a different host needs its own. Use
computer-use-mcp health_report --json to record the current executable, bundle id
(if any), parent process, permission state, and daemon socket/secret paths without
revealing daemon secrets. Add --probe-capture when you intentionally want a
bounded ScreenCaptureKit/replayd responsiveness probe. For redistribution, codesign
with a Developer ID and notarize
(codesign --sign "Developer ID Application: ..." && xcrun notarytool submit ...)
so TCC grants attach to a stable identity. See
Permissions and app identity for the first
productionization checklist.
- Background delivery uses macOS per-process event posting, which is
app-dependent: a few apps that require real keyboard focus (e.g. some pro audio
apps, secure input fields) may ignore background events. For clicks, use
allow_global_cursor: trueas an explicit fallback. - Menu key-equivalents (e.g.
cmd+a) are reliable when the app is the key window; some apps ignore them when targeted purely in the background. - macOS only (the engine is built on Accessibility, ScreenCaptureKit, and CoreGraphics). The protocol/tool layer is OS-agnostic.
swift build
swift test
.build/debug/computer-use-mcp serve # run the stdio MCP server
.build/debug/computer-use-mcp call get_app_state '{"app":"Calculator"}' # drive one tool
.build/debug/computer-use-mcp health_report --json # non-mutating diagnostics
.build/debug/computer-use-mcp health_report --probe-capture # bounded capture-service probe
.build/debug/computer-use-mcp doctor # check permissions
python3 scripts/preflight.py # CI-safe release preflight
python3 scripts/build_app_bundle.py # local .app wrapper build
python3 scripts/deploy_app_bundle.py # release build + bundle + install to ~/Applications + daemon handover
python3 scripts/deploy_app_bundle.py --check # exit 1 if the installed bundle is older than source/build
python3 scripts/preflight.py --use-app-bundle # non-live checks through the .app executable
python3 scripts/e2e_demo.py # safe structured smoke artifact; no GUI mutationDefault CI covers the non-mutating path: package build, pure unit tests, and CLI
version/help/health_report --json smoke checks. Live app-control checks are
local-only because they require a logged-in macOS desktop plus Accessibility/Screen
Recording permission and can operate real apps. See
Testing and Preflight for the testing tiers, local command loop,
and release preflight expectations. The architecture-level behavior target for
these tiers is captured in
Modality Contract.
Use deterministic checks for hosted CI and live GUI checks only on a local Mac or a future self-hosted macOS runner with explicit permissions.
| Tier | What it proves | Safe for hosted CI? | Command |
|---|---|---|---|
| Deterministic unit tests | Pure Swift behavior such as parsing, safety policy, coordinates, and tree shaping. | Yes | swift test |
| Structured dry-run smoke | The benchmark entrypoint, schema, git/macOS metadata collection, and opt-in gate. It does not start the MCP server or open apps. | Yes | python3 scripts/e2e_demo.py |
| Release preflight | Build, unit tests, CLI smoke, health report, and dry-run background eval in one JSON report. | Yes | python3 scripts/preflight.py |
| Local app-bundle runtime | Produces an ad-hoc signed .app wrapper and runs non-live CLI/dry-run checks through Contents/MacOS/computer-use-mcp. |
No | python3 scripts/preflight.py --use-app-bundle |
| Deterministic background eval | The fixture-app path mutates a stable AX text field while preserving the current frontmost app. | No | python3 scripts/live_background_eval.py --live |
| Real-app compatibility smoke | Lightweight live matrix for Finder read-only discovery and TextEdit background stdio behavior. | No | python3 scripts/real_app_smoke.py --live |
| Live GUI smoke | The real MCP stdio path against TextEdit in the background, including perceive, type, select, and focus-stability checks. Setup launches TextEdit without activation and fails if the current frontmost app changes. | No | python3 scripts/e2e_demo.py --live |
Live GUI mutation is opt-in. Use --live for local manual runs. Do not add the
live tier to hosted CI — it depends on an unlocked macOS desktop, TCC permissions,
Finder/TextEdit behavior, and user-visible app state. Before running the live tier,
build the binary and make sure the spawning terminal has Accessibility and Screen
Recording permission:
swift build
.build/debug/computer-use-mcp doctor --prompt
python3 scripts/e2e_demo.py --liveContributions welcome. See CONTRIBUTING.md for build/test conventions and PR expectations, SECURITY.md for the security model and private disclosure process, and CHANGELOG.md for release history.
MIT — see LICENSE.