Curated list of papers + libraries related to computer GUI use via LLMs.
Highly opinionated, focus on quality vs quantity.
- Clickyy - Shake your cursor to summon an AI agent that sees your screen and clicks, types, drags, and acts for you on macOS.
- Try computer use on your Mac in one click.
- Agent QA - Self-improving QA agent for natural-language web and mobile regression tests, with execution memory and UI-change adaptation.
- Fazm - MIT-licensed, open-source voice-controlled AI agent for macOS using accessibility APIs and ScreenCaptureKit.
- Openwork - MIT-licensed, open alternative to Anthropic's Cowork with multi-LLM support for browser automation.
- optics-framework - Apache-2.0 framework for LLM-driven GUI automation of mobile, web and Smart TV apps: a natural-language ReAct mode operates live apps through screenshot → LLM → validated action loops, and an MCP server exposes tap/swipe/verify keywords so any AI agent can drive real devices.
- ClawBench: Can AI Agents Complete Everyday Online Tasks? (code) (project) (UBC, Vector Institute, UWaterloo) (04/26)
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning (Tsinghua U) (11/24)
- Anthropic Claude Computer Use API (Anthropic) (10/24)
- OmniParser for Pure Vision Based GUI Agent (code) (Microsoft) (08/24)
- ECLAIR: Enterprise sCaLe AI for woRkflows(code) (Stanford U) (05/24)
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (code) (HKU) (05/24)
- Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs (code) (Apple) (04/24)
- SeeAct: GPT-4V(ision) is a Generalist Web Agent, if Grounded (code) (OSU) (01/24)
- CogAgent: A Visual Language Model for GUI Agents (ZhiPu)(12/23)
- AppAgent: Multimodal Agents as Smartphone Users (code) (TenCent) (12/23)
- SoM : Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V (code) (Microsoft) (10/23)
