Zuchen Li

projects / gui-agent-fragility

How Fragile Are GUI Agents?

2026 winter Project lead University of Michigan

GUI-agentsbenchmarkLLMevaluation

GUI agents are usually evaluated on frozen snapshots of interfaces — but real interfaces evolve. We built a benchmark grounded in five recurring interface-evolution categories — layout, semantic, visual, structural, and workflow shifts — and generated controlled synthetic websites so each perturbation can be isolated and measured.

I led the project organization and ran the agent experiments using the VisualWebArena harness.