An MCP tool can pass normal tests and still fail when an agent tries to use it.

The implementation may be correct, but the model may choose the wrong tool. It may send the wrong arguments. The tool description may be too vague. Or the result may contain so much data that the model gets lost.

MCP Apps add a few more places where things can break. Now there is a UI resource, an iframe, and a bridge between the app and the host.

I think the easiest way to test all of this is to split it into layers. Each layer should answer a different question.