An agent can read a receipt, enter the right number, and still fail to update a spreadsheet. It might miss a second receipt, edit the wrong cell, or close the window before saving. The difficult part is connecting perception to a sequence of actions whose result can be checked.

Caiming Xiong’s seventh lecture in Berkeley’s Advanced Large Language Model Agents connects three parts of that problem: an executable environment, useful training trajectories, and a model that can both choose an action and locate its target. OSWorld, AgentTrek, TACO, and Aguvis address different parts of this chain.

An original teaching example, not an OSWorld benchmark run. The animation separates an action, the next observation, and evidence in the saved artifact.