Building FynPDF on macOS, Maheep Kumar walks through failed AI UI-testing approaches (screenshots, VNC, XCUITest, generic computer-use) and why a small AXUIElement test API finally let agents drive the app in the background.
All you need is accessibility
Maheep Kumar — 27 September 2026
I’ve been building FynPDF, a PDF reader and editor for macOS, with its own PDF engine written from scratch. The engine is easy to test headlessly. The user-facing viewer/editor is not: stale thumbnails, 200 ms blank saves, stuck popovers, mis-hit clicks — none of that shows up in a unit test.
I wanted AI agents to exercise the application: open, scroll, zoom, annotate, save, reopen — and tell me when something went wrong. The interesting part wasn’t deciding accessibility was the answer up front; the requirements emerged one failed approach at a time.
Attempt 1: Screenshots + System Events
A skill file had the agent drive FynPDF via menus/shortcuts, screenshot, and repeat. Problems: the app stole focus; clicks landed in the wrong window; screenshots burn tokens; every step was slow. A runaway render needed a watchdog. Deleted the skill. Lesson: the test must not interfere with me.
Attempt 2: Separate macOS account + VNC
Isolation fixed desktop risk but still forced pixel reasoning and expensive observations. Lesson: talk to the application, don’t continuously stare at a desktop; keep the model out of every micro-interaction.
Attempt 3: XCUITest
Proper identifiers and a regression suite — but synthesized HID events take over the machine, fight the Debug build, and couple tightly to UI architecture (title-bar tabs broke “the window” assumptions).
Attempt 4: General computer-use agent
Background accessibility-tree interaction was closer, but pixel clicks failed under overlays, snapshots were expensive, focusing fields activated the app, and a person-like agent is the wrong abstraction for a deterministic test runner.
What finally worked: accessibility as a test API
A ~900-line Swift CLI, fynpdf-ax, drives a Debug build that never comes to the front via AXUIElement: press actions (no pointer), set values (no focus/activation), find by identifier, walk menus. First day found real bugs (popover never dismissed; text box too small).
Tests are JSON flows — fixture PDF + steps (menu, action, set, press, capture, same/differ pixel assertions, parity, relaunch). The agent writes the flow; the runner executes hundreds of actions without burning model tokens per click. Making the app testable also made it more usable with VoiceOver and Switch Control.
Where accessibility isn’t enough
Visual checks still need pixels; some bugs need in-process tile inspection; real drags/ink still matter. The point isn’t that accessibility replaces every UI test — it’s a semantic control surface that doesn’t require owning the desktop.
Advice
Give every control an accessibility identifier.
Add accessibility actions for anything that currently needs a mouse.
Keep the flows in the repo — a repro that runs once is a demo; a checked-in flow is a test.