zuchenli@umich:~/projects$

~/projects/gui-agent-fragility

How Fragile Are GUI Agents?

2026 winter Project lead University of Michigan

[GUI-agents][benchmark][LLM][evaluation]

GUI agents are usually evaluated on frozen snapshots of interfaces — but real interfaces evolve. We built a benchmark grounded in five recurring interface-evolution categories — layout, semantic, visual, structural, and workflow shifts — and generated controlled synthetic websites so each perturbation can be isolated and measured.

I led the project organization and ran the agent experiments using the VisualWebArena harness.