The browser-use macOS Harness is an MIT-licensed Python package that hands an LLM six raw primitives — see, key, type, click, ax, and script — inside one persistent process, so the model writes its own Mac automation instead of picking from a recipe library. It is v0.1.2, alpha, and it has not received a commit since launch day.

That last sentence is the review. This is a genuinely clever, genuinely thin control surface published on 2026-08-17 and then left alone, and the most useful thing you can do before installing it is understand exactly which parts are load-bearing and which parts are a demo.

If you want the launch-day feature tour, our earlier piece, Browser-Use macOS Harness: Full LLM Control of a Mac, covers the pitch and the six primitives. This review does not repeat it. It covers what accumulated in the six weeks afterward: a silent-failure bug in one of those six primitives, an unmerged fix, a fork that has already overtaken upstream, and the measured computer-use attack numbers that make the permission dialog the real product decision.

What is the browser-use macOS Harness, and what does it refuse to build?

The macOS Harness is the lowest-level member of the browser-use family. Its pitch is that a capable coding agent does not need an app catalog — it needs a persistent Python process with eyes, hands, and a shell. Where an earlier generation of desktop agents shipped per-app tools for Slack, Spotify and Final Cut, this one ships nothing app-specific and expects the model to write the missing function mid-task.

The deliberate omissions matter more than the included features:

  • No framework or orchestration layer. There is no planner, no retry policy, no task graph.
  • No verification primitive. Nothing confirms that an action landed.
  • No approval rail. Nothing asks before a click or a shell command.
  • No app recipes. Every target app is handled by generic primitives.

The project’s own framing is that it sits under coding agents rather than beside them: Codex or Claude Code generate Python, and the harness executes it against a real Mac. In the launch thread the author described the design as the opposite of macOS-use, the earlier browser-use desktop repo — macOS-use packages an agent workflow around Mac interaction, while the harness is a control surface with no workflow attached.

That is a defensible architectural bet, and it is also the source of every problem in the rest of this review. A thin harness moves complexity into the agent and the operator. When the primitives are correct, that is elegant. When one of them fails silently, there is no layer left to catch it.

What are the six primitives — see, key, type, click, ax, and script?

The harness exposes six operations, and the best public teardown of what they actually do is worth reading in full. Short version:

  • mac.see shells out to the system screencapture binary with the per-window flag -l <window-id>, so it captures that window’s pixels rather than the whole screen. It defaults to --max-width and --max-height of 1280x1280 and to --no-pointer.
  • mac.ax wraps AXUIElementCopyElementAtPosition and the ApplicationServices framework. Its compact attribute set is AXRole, AXTitle, AXDescription, AXValue, AXFrame, and ax.query supports text=, search_key=, visible_only=True, and max_nodes=500. Crucially, the element index it returns is what keeps an accessibility reference alive across statements in the same process.
  • mac.script sends raw Apple Events, which means the harness inherits the entire existing AppleScript ecosystem for free instead of reimplementing it.
  • mac.click, mac.key and mac.type synthesise input — and mac.click is the one you need to read the next section about.

Underneath, the CLI is plain argparse with subcommands doctor, apps, repl, skill, see, state, and telemetry. The entire “give an LLM a Mac” mechanism is one exec() into a pre-populated namespace in cli.py. That is not a criticism — it is the honest size of the thing. You can read the whole harness in an afternoon, which is more than you can say for most agent frameworks.

The dependency tree is correspondingly small: browser-harness>=0.1.9, pillow>=11.3, and pyobjc-framework-ApplicationServices>=12.0 on Darwin. The package will not import on Linux at all.

Why does mac.ax come before mac.see?

Because screenshots are the expensive part of computer use, and accessibility trees are text. Reading one window’s AX tree costs a fraction of what pushing a 1280x1280 PNG through a vision model costs, and it gives the model structure — roles, titles, frames, an index it can name — instead of pixels it has to localise.

The independent rebuild by huytieu.com makes the same design choice explicit: its get_app_state(app) returns one window’s screenshot plus a numbered accessibility tree, and the model acts by naming an element index or a coordinate. Reading structure and pointing at it beats guessing pixels, and it is the single cheapest reliability win available to a Mac agent.

How does the macOS Harness actually control the Mac under the hood?

There is no secret technology, and that is worth stating plainly because it was the most over-mythologised part of the launch coverage. The plumbing is ordinary public macOS API surface:

LayerAPIWhat it buys
Screen captureCGWindowScreenshots of background windows without raising them
InputCGEvent, posted to a target PIDSynthetic clicks and keystrokes
SemanticsAccessibility API + Apple EventsPressing and setting controls when pixels are not enough
BrowserCDP via Browser HarnessReal, logged-in Chrome
ShellsubprocessEverything Apple does not expose

Independent reverse-engineering of OpenAI’s Codex desktop computer use reached the same conclusion: ChatGPT.app bundles @oai/sky, which talks over a Unix domain socket to a code-signed helper (SkyComputerUseService) using length-prefixed JSON-RPC — and the engine behind it is Apple’s public CGEvent, AXUIElement, and CGWindowListCreateImage, not anything exotic. OpenAI’s advantage is product engineering, not hidden frameworks.

The harness’s behavioural choices are the interesting part. It captures background windows without raising them, sends input to a PID rather than to whatever has focus, animates a click-through pointer, and never moves the real cursor. The doctor subcommand reports which permissions are actually needed.

There is also a hard architectural ceiling to flag: PyObjC and ApplicationServices coverage is the boundary. Anything Apple only exposes in private frameworks — some I/O Kit device control, for example — is out of scope, permanently. And the AX and SkyLight surfaces these tools sit on are semi-private (SPIs), so behaviour can shift between macOS releases.

How do you install the macOS Harness and run macos-harness doctor?

Setup is two commands and a prompt. The package is on PyPI as macos-harness, still at v0.1.2, requiring Python 3.11+ (3.12 recommended), and the documented install path is uv.

The part worth understanding is how the agent learns to use it. macos-harness skill prints a skill definition that you redirect into your agent’s skills directory — defaulting to ${CODEX_HOME:-$HOME/.codex}/skills/macos-harness. Setup therefore ships as a prompt, not a runbook. The harness teaches the model what the primitives are and lets the model decide how to sequence them.

Then run macos-harness doctor and grant exactly what it reports. Not the whole Privacy & Security pane — what doctor names.

Two practical notes from the field:

  1. TCC dialogs assume a GUI session. Accessibility and Screen Recording approvals require a logged-in desktop and manual approval under System Settings > Privacy & Security. Headless provisioning and CI/CD on a Mac mini in a rack remain awkward.
  2. Input Monitoring is not required. The launch notes are explicit that you do not need it, so do not grant it reflexively.

What do Accessibility, Screen Recording, and Automation permissions really grant?

This is the section that should decide whether you install the harness, more than any feature list.

The macOS permissions the harness touches — kTCCServiceAccessibility, kTCCServiceScreenCapture, and Automation — are machine-wide grants, not app-scoped ones. Once granted, any process running as you can generally reach the same capability through the daemon or through the grant itself. The independent computer-harness project’s disclosure list is the model to imitate here: it runs arbitrary code as you, its daemon is a standing proxy to your Accessibility and Screen Recording grants, and screen captures can contain secrets.

Security researchers have a name for the cumulative combination. Input Monitoring (kTCCServiceListenEvent), Input Injection (kTCCServicePostEvent), Screen Capture (kTCCServiceScreenCapture) and Accessibility together give keylogging of every keystroke, recording of everything visible, synthetic input injection, and complete GUI control — which HackTricks characterises as the most dangerous combination on macOS.

The harness does not need Input Monitoring, which narrows that triad by one. But Screen Recording plus Accessibility on a daily-driver Mac still means the agent can read whatever is on screen: Signal threads, a terminal showing credentials, tax documents. The harness’s honest behaviour — never raising the app, never moving the cursor — makes the grant feel smaller than it is. Nothing about the permissions dialog changed; only your impression of it did.

This is not hypothetical abuse risk either. MITRE ATT&CK catalogues TCC manipulation as sub-technique T1548.006, describing adversaries reusing permission grants through a trusted parent process or a launchctl environment override, and recommending audits of Automation grants plus tccutil reset.

Practical rule: grant these permissions on a machine where the worst outcome is annoying rather than expensive.

Does mac.click work on native Mac apps? The silent-failure problem

No — and this is the single most important thing to know about the harness that no launch-day coverage contained.

mac.click() is one of the six documented primitives. On native AppKit apps it returns a normal pointer dictionary, raises no exception, and does nothing.

The reproduction is specific and still open. GitHub issue #6 (opened 2026-08-18, unresolved as of 2026-10-01) reproduces it on Finder and Calculator: the call returns cleanly while the UI never changes. On the same element, mac.ax.perform(index, 'AXPress') works correctly. So the accessibility path succeeds where the mouse path silently fails.

The root cause comes from the unmerged fix PR #7: mac.click() posts mouse events with CGEventPostToPid, and AppKit hit-tests mouse events against the window server’s pointer state, so the events are discarded before any app sees them.

What does the working fix cost you?

The fix that works is CGEventPost to the HID tap — the same thing a human’s mouse does. And it trades away the harness’s headline invariants:

  • It moves the physical cursor, which is exactly what the README leads with as a guarantee.
  • It abandons strict per-PID targeting, because a HID-tap event goes wherever the pointer is.

The same reporter noted a second limitation: keyboard input appeared to require the target app to be frontmost. Sending cmd+w to Finder had no effect while another app held focus — which sits awkwardly beside the README’s “never activates or raises a target app” combined with “sends input directly to an app PID.”

This is not merely a macos-harness bug. The independent computer-harness rebuild reproduced the same Calculator failure against a different codebase: with Calculator behind another window, both AXPress via the accessibility API and raw mouse events posted to Calculator’s PID produced no visible change. Two independent implementations hitting the same wall is strong evidence this is a macOS/AppKit property, not sloppiness in one repo. The reporter’s environment was macOS 26.6.1 on arm64, so version-specific behaviour is plausible but unconfirmed.

Why this matters beyond one bug: it is a silent failure. The primitive lies by returning success. In a field whose benchmark failure analysis names “skipped verification” as a top-four cause of agent failure, shipping six primitives and zero verification rails — with telemetry on by default — is the design tension you should carry into any evaluation.

Maintenance reality: six commits, one contributor, and a fork moving faster

As of 2026-10-01, the numbers are blunt:

MetricValue
Stars892
Forks60
Open issues9
Commits6
Contributors1 (gregpr07, 6 contributions)
Last upstream push2026-08-17T16:48Z
License / statusMIT, Development Status :: 3 - Alpha
Unmerged community PRs#1, #3, #4, #5, #7

The repo was created 2026-08-17T00:22Z and last pushed the same day at 16:48Z. It has not received a commit since launch day. All three releases — 0.1.0 at 16:10:05Z, 0.1.1 at 16:45:23Z, 0.1.2 at 16:48:44Z — landed inside a single 38-minute window on 2026-08-17. The newest upstream commit is titled “fix: render project banner on PyPI.” That is a launch-day cadence, not a maintenance cadence.

The predecessor tells the same story ended. browser-use/macOS-use is now archived: 2,000 stars, 195 forks, MIT, created 2025-01-23, last push 2025-03-05.

The most advanced continuation of the project is a fork most coverage never mentions. aktanazat/macos-harness — 0 stars, 0 forks — carries v0.2.0 and v0.3.0 tags and a CHANGELOG entry for [0.5.0] dated 2026-08-22, five days after upstream stopped. It adds things upstream does not have:

  • mac.do — receipted semantic operations returning outcome, acted, and verified, with dry runs and per-session replay tokens via an once ledger.
  • mac.handoff — a representation-only operation for credential and authentication boundaries.
  • Opt-in rather than default-on telemetry.

Its PR “Release macOS Harness v0.3.0” (#8) reports 439 Python tests passing on 3.11–3.14 and 118 Swift tests, a signed universal2 wheel, and 1,000 live PID-targeted handoffs on an M4 Pro at 0.1173 ms median and 0.1656 ms p95 latency.

That is the accountability spine of this review. The upstream release is a thin demo frozen at launch with an unmerged fix for a silent failure in one of its six primitives; the design’s most serious continuation is a zero-star fork. If you evaluate “the browser-use macOS Harness” by reading upstream alone, you are evaluating a launch announcement.

The family context explains the star count. Parent project browser-use/browser-use sits at 116,871 stars and 12,891 forks with a push on 2026-09-30. The CDP layer it depends on, browser-use/browser-harness, has 18,247 stars, 1,781 forks, PyPI v0.1.13 (2026-09-04), and a Show HN post that scored 134 points. The same “primitives, not recipes” bet now spans jev-ultrafast (21,603 stars) and phone-harness (PyPI v0.3.0, driving iPhone Mirroring, adb, and cloud Android). The macOS Harness’s 892 stars sit on a 116.9K-star parent’s reputation — and, note, every published star count for this repo is a snapshot that moved within hours of launch.

Is the macOS Harness safe? The measured prompt-injection numbers

No thin harness fixes this, and you should size the risk with numbers rather than vibes.

The threat model collapses to one sentence: once the harness has a click primitive and a shell, an agent that reads any app’s UI is ingesting text written by someone else. Treat third-party UI as untrusted input — that is the correct framing, and the harness supplies none of it.

Measured attack-success rates for computer-use agents, so you can plan instead of hand-wave:

SourceFinding
RedTeamCUA (arXiv:2505.21936, 864 examples)ASR up to 66.2% in decoupled evaluation; 42.9% for Claude 3.7 Sonnet CUA; 7.6% for the most secure CUA tested (Operator); 83% / 50% for Claude 4.5 / 4.6 Opus CUA end-to-end; attempt rates up to 92.5%
Anthropic Claude for Chrome evaluation (vendor-reported, non-adaptive)23.6% attack success without mitigations, 11.2% with them; a four-type browser challenge set went 35.7% → 0%. Anthropic’s stated position is that 1% ASR is still “a meaningful risk”
ACL 2025 adversarial pop-ups (via CSA)Mean 86% attack success across OSWorld and VisualWebArena while cutting task completion 47%; “ignore pop-ups” system prompts were insufficient because visual framing defeats textual instruction
NIST agent-hijacking evaluation (Jan 2025, via CSA)Optimised attack prompts reached 81% against baseline defenses versus 11% unaided (~7x)
WASP benchmark (NeurIPS 2025, Meta AI Research)Web agents begin executing adversarial instructions in 16–86% of cases

Three practical implications for a Mac harness:

  1. Deny by default at the boundary. Give the harness a dedicated user account or a dedicated machine. tccutil reset after a session is the cheap hygiene step MITRE recommends.
  2. Never leave Screen Recording granted on a daily driver. The grant is machine-wide and persists after the task ends. A screen capture of your screen is a screen capture of everything on it.
  3. Prefer structural reads (ax) over pixels where you can. Adversarial content is harder to smuggle into a compact AX attribute set than into a rendered pop-up image — though, per ACL 2025, this narrows rather than closes the gap.

A harness with a click primitive, a shell, no approval gate, and telemetry on by default is a tool you aim at a throwaway machine, not your work MacBook.

macOS Harness vs macOS-use vs Codex vs cua-driver vs Claude Cowork

The alternatives map is where launch coverage generally stops at naming competitors. Here is what actually differentiates them.

OptionOpen sourceInteraction modelKey constraint
browser-use macOS HarnessYes (MIT)Six primitives in one persistent Python process; agent writes its own codeAlpha, 6 commits, no push since launch; mac.click silently fails on AppKit apps
browser-use macOS-useYes (MIT)Packaged agent workflow around Mac interactionArchived — last push 2025-03-05
Codex / ChatGPT desktop computer useNoBundled @oai/sky → signed helper over JSON-RPC; public APIs underneathClosed; on Windows, foreground mode takes over the pointer and input, and UIPI/Session 0 block driving elevated apps
Claude Cowork computer useNoBeta, Pro/Max tiersRequires Accessibility + Screen Recording grants
Hermes Agent + cua-driverPartlyCross-platform dispatch; macOS events via the undocumented SkyLight SPI (SLPSPostEventRecordTo) scoped by PIDBackground-control surface is an SPI that can shift with macOS updates; macOS Tier 1 is Apple Silicon only, Intel Macs unsupported
UI-TARS Desktop / Agent S / Open InterpreterYesVision-first or code-first desktop agentsPure vision is the most expensive and least structurally verifiable path

The positioning that survives contact with the evidence: macOS Harness is the lowest-level, most composable, and least-guarded option. It is best on a machine you would be willing to hand over if something went wrong — which is the same sentence as the security section, arrived at from the architecture instead of the permissions.

Two honest notes. First, the differentiator against Codex computer use is mainly that the harness is open source — real, but thin as differentiators go. Second, there is no SEO-agent handoff or schema work in this review’s scope; the harness itself has nothing to do with your blog pipeline.

If you are weighing the broader field rather than one repo, our computer-use agents comparison covers the category map, and the open-source browser agent index covers the CDP-based options the harness borrows its browser layer from. For the closed-source path, see the OpenAI Codex computer use guide.

Do the benchmarks say a thin harness is enough? OSWorld 2.0

This is where the review stops being about one repo’s commits.

OSWorld 2.0 (arXiv:2606.29537, published 2026-06-28) benchmarks computer-use agents on long-horizon real-world tasks: 108 tasks, a median of roughly 1.6 hours for a skilled human, and about 27.25 fine-grained checkpoints per task. The results:

SettingScore
Claude Opus 4.8, binary completion (500 steps)20.6%
Claude Opus 4.8, partial credit (500 steps)54.8%
GPT-5.5, binary completion13.0%
Claude Opus 4.8 on OSWorld-Verified (shorter horizon)83.5%

Cost and effort per task for Opus 4.8: roughly $72.4 and 481.8 tool calls.

Read those two Opus rows together and you have the entire strategic picture of computer use today. The same model that hits 83.5% on the short-horizon benchmark manages 20.6% binary completion when the task takes an hour and a half. Short-horizon desktop computer use looks solved; long-horizon does not. For the trend line, OSWorld 1.0 (arXiv:2404.07972, NeurIPS 2024) had humans at 72.36% and the best model at 12.24% across 369 real-OS tasks.

The failure analysis is the part that makes a thin harness a questionable bet. OSWorld 2.0 names four disposition problems: agents drop stated constraints roughly 200 steps later, miss information that arrives mid-task, guess instead of asking the simulated user, and skip verification entirely. Its side-effect audit of 216 trajectories found about 14% of tasks extracting hidden application state and about 33% bypassing the user-visible interface — summarised as agents that “escalate privileges and will do whatever it takes to finish the task.”

Now map that onto the harness’s feature list: six primitives, zero verification, zero approval rails, telemetry on by default. The two failure modes a thin harness is structurally worst at handling — skipped verification and privilege escalation via side effects — are two of the four the benchmark measured.

This is exactly the gap the fork’s mac.do receipts address: outcome / acted / verified, dry runs, and per-session once ledger tokens are a verification primitive bolted onto a design that shipped without one. Comparing upstream v0.1.2 to the fork’s v0.5.0 CHANGELOG is not an academic exercise; it is the difference between a demo and a tool.

Why mac.ax beats mac.see on real budgets

Cost is not a footnote in computer use — at 481.8 tool calls and $72.4 per long-horizon task, per-action screenshots are the dominant expense, and a harness that defaults to the accessibility tree before vision is a defensible cost argument rather than an aesthetic preference.

Three concrete reasons to order the primitives ax → script → see:

  1. Token economics. A compact attribute set (AXRole, AXTitle, AXDescription, AXValue, AXFrame) is text on the order of hundreds of tokens. A 1280x1280 screenshot is thousands of image tokens, and it is re-sent every step.
  2. Structural addressing. ax.query returns a stable element index you can name and re-use across statements in the same process. A coordinate from a screenshot is valid until the layout shifts.
  3. Verification headroom. An AX tree can be re-queried to check whether a press landed. A screenshot requires the model to decide whether the pixels changed — which is precisely the verification step agents skip.

mac.script sits between them: raw Apple Events inherit the whole AppleScript ecosystem for free, so if a documented AppleScript command exists, you skip vision and AX traversal entirely.

The ceiling is real, though. PyObjC and ApplicationServices coverage is the boundary, and anything Apple only exposes privately — some I/O Kit device control — is permanently out of scope.

Who should use the macOS Harness — and on which machine

Use it if you are a developer who wants to study how Mac computer use works, or who has a specific automation you would rather write in Python than in AppleScript, on a dedicated Mac or a spare user account, and who is comfortable reading the source (it is an afternoon’s read, and you should read it).

Do not use it if you want something that reliably drives your Mac unattended. That is the honest verdict from the source-level review, and it matches the project’s own positioning: the harness “isn’t that yet and doesn’t claim to be,” and its early release leaves reliability and safe task execution as the central tests.

The machine decision is the important one:

  • Dedicated Mac or dedicated account — yes, with the permissions granted and tccutil reset afterwards.
  • Apple Silicon, macOS 26.x — the tested environment (macOS 26.6.1 arm64 is the reporter’s). Note that alternative stacks list Intel Macs as unsupported, and the same PyObjC-era assumptions apply here.
  • Your daily driver with Screen Recording already granted — no. Screen captures contain secrets, and the grant is machine-wide.

Quickstart: a safe first task, step by step

  1. Create a dedicated macOS user account for the harness. Do not grant TCC permissions from your main account.
  2. Install the package with uv on Python 3.12: macos-harness is at v0.1.2 and requires Python 3.11+.
  3. Generate the skill file into your agent’s skills directory: macos-harness skill > ${CODEX_HOME:-$HOME/.codex}/skills/macos-harness.
  4. Run macos-harness doctor and grant only what it reports. Expect Accessibility, Screen Recording, and possibly Automation. Do not grant Input Monitoring.
  5. Pin the version. macos-harness==0.1.2. There are no releases after 2026-08-17 and four unmerged PRs at time of writing, so an upgrade is not a thing that will happen to you.
  6. Pick a read-only first task — capture a background window and query its AX tree — and confirm mac.ax works before you trust mac.click for anything.
  7. Verify every click independently. Re-query the AX tree after the action. Given the silent-failure behaviour documented above, “no exception” is not evidence of success.
  8. Reset permissions when finished (tccutil reset Accessibility, tccutil reset ScreenCapture) and treat any captured screenshot as potentially sensitive.
  9. If you need verification receipts, read the fork’s mac.do design — outcome/acted/verified, dry runs, and once tokens — before you decide upstream is enough.

FAQ

Is the browser-use macOS Harness the same thing as macOS-use? No. browser-use/macOS-use is archived (2,000 stars, last push 2025-03-05). The macOS Harness is its successor: MIT-licensed, v0.1.2, created 2026-08-17, and built as a lower-level primitive surface that a coding agent generates code against rather than a packaged agent workflow.

Does mac.click actually work on native Mac apps? Not reliably. Open issue #6 reproduces mac.click() silently doing nothing on Finder and Calculator: it returns a normal pointer dict, raises nothing, and the UI never changes. Cause, from the unmerged PR #7: CGEventPostToPid events are discarded by AppKit’s window-server hit-testing. The working path is mac.ax.perform(index, 'AXPress') for accessibility-capable controls, or a HID-tap click that moves your real cursor.

What permissions does the macOS Harness need? Accessibility, Screen Recording, and possibly Automation, per macos-harness doctor. Input Monitoring is explicitly not required. All of these are machine-wide TCC grants rather than app-scoped ones, so any process running as you can generally reach the same capability once granted.

Is it safe to run an LLM with control of my Mac? Treat every app UI as untrusted input. Measured prompt-injection success rates against computer-use agents run from 7.6% (Operator) to 83% (Claude 4.5 Opus CUA end-to-end) in published evaluations, and the harness ships no approval gate or verification primitive. Run it on a dedicated account or spare machine, never on a daily driver with Screen Recording already granted.

Has the macOS Harness been maintained since launch? No. As of 2026-10-01 it has 892 stars, 60 forks, 9 open issues, 6 commits, 1 contributor, and no push since 2026-08-17T16:48Z — all three releases shipped in a 38-minute window on launch day. Four community PRs are open and unmerged, including the fix for the silent-click bug, while a 0-star fork has shipped v0.3.0 tags and a [0.5.0] CHANGELOG with mac.do receipts upstream lacks.

The bottom line

The browser-use macOS Harness is the clearest available demonstration of what a primitive-first Mac harness should look like — six functions, one persistent process, public Apple APIs, no framework tax. It is also frozen at launch with a silent failure in one of those six primitives, machine-wide permissions that outlive the task, and no verification or approval layer in a field whose long-horizon benchmark puts the best model at 20.6% binary completion and whose security evaluations put prompt-injection success into double digits.

Install it to learn, on a machine you would be willing to hand over. Budget for the click that returns success and does nothing. And if you need receipts rather than primitives, look at what the fork built — because upstream, as of today, is a launch announcement that has not been touched since.