A device toolkit for interaction, repeatable QA, and native or React Native diagnostics.
Compare mobile automation tools.
See what’s supported—and the evidence behind it.
| Select | Agent | Platforms | MCP | Flows | Debugging | Profiling | Skills | Best suited for | Details |
|---|---|---|---|---|---|---|---|---|---|
| iOSAndroid | Deep app diagnostics | ||||||||
| iOSAndroid | Coding-agent feedback loops for mobile, TV, and web app verification | ||||||||
| iOS | Streaming Apple simulators into an agent or browser | ||||||||
| iOSAndroid | Direct device control via MCP tools | ||||||||
| iOSAndroid | Repeatable end-to-end tests | ||||||||
| iOS | Native Apple development with integrated Xcode build, test, and simulator tooling | ||||||||
| iOS | Lightweight iOS simulator UI automation for MCP coding agents | ||||||||
| iOSAndroid | Appium-based mobile, tvOS, and remote-grid automation | ||||||||
| iOSAndroid | Vision-driven UI testing across web, mobile, and desktop | ||||||||
| iOSAndroid | Natural-language phone task automation | ||||||||
| iOSAndroid | Multi-step task execution | ||||||||
| Android | Self-hosted vision-model phone agent for Android, HarmonyOS, iPhone | ||||||||
| Android | Multi-agent GUI automation research on Android, desktop, and web | ||||||||
| Android | Android app exploration and documentation-driven automation research |
A shared CLI, MCP server, and typed Node.js runtime that lets coding agents inspect, control, and verify iOS, Android, HarmonyOS, TV, web, macOS, and Linux apps, and preserves evidence for review.
Apple simulator streaming and control through a browser preview, CLI, and portable agent skill. Includes accessibility inspection, camera injection, and rendering diagnostics.
An MCP server for automating iOS and Android apps via accessibility trees or screenshot-based taps, covering simulators, emulators, real devices, and optional Mobile Next Cloud remote devices.
A UI testing framework with agent access through its bundled MCP server.
An MCP server and CLI for building, running, testing, and debugging Xcode projects on iOS simulators, physical Apple devices, and macOS.
A Model Context Protocol server that lets an MCP-integrated AI assistant inspect and control an iOS simulator: accessibility tree queries, tap/type/swipe, screenshots, video recording, app install/launch/terminate, and deep links.
An MCP interface to the Appium ecosystem, driving embedded local or remote Android, iOS, and tvOS sessions with AI-assisted locators and test generation.
An AI-powered GUI agent and testing kit for writing, running, and debugging visual UI workflows across web, Android, iOS, and desktop interfaces.
An open-source framework for controlling Android and iOS devices with LLM agents, offering a CLI, Python API, and an optional Mobilerun Cloud service for hosted or managed devices.
A phone agent that decomposes natural-language tasks and acts on UI state.
An open phone-agent vision-language model and framework that automates Android, HarmonyOS, and iPhone (via WebDriverAgent) apps from natural-language instructions, using a self-hosted or third-party OpenAI-compatible model endpoint.
A family of visual GUI agents and models (Mobile-Agent v1-v3.5, GUI-Owl, PC-Agent, UI-S1, GUI-Critic-R1) from Alibaba's Tongyi Lab. Mobile-Agent-v3 deploys a multi-agent planning framework on real Android/HarmonyOS phones via ADB/HDC, while the GUI-Owl-1.5 models are described as supporting desktop, mobile, and browser automation.
A CHI 2025 research framework where a multimodal LLM (e.g. GPT-4V or Qwen-VL-max) learns to operate Android apps through autonomous exploration or human demonstration, building a per-app documentation knowledge base of UI elements that it then uses to complete new tasks via tap/swipe/type actions over ADB.
Different tools, different jobs. These are documented capabilities, not benchmark scores. “Unverified” means we haven’t established support from the sources.