# When a WebMCP tool returns an MCP App

> WebMCP lets a website hand tools to the agent. MCP Apps lets a server hand UI to the human. We built an experiment where one hands off to the other.

*Published 2026-09-30 · 8 min read*

---

WebMCP gives agents tools from a website. MCP Apps gives humans UI from an MCP server.

So what happens when a WebMCP tool returns an MCP App? 🤔

Picture an agent working with a website you already have open. Most steps are plain tool calls, but then it needs *you* to compare a few options, move a slider or check a preview.

Could the website open a small UI inside the agent's side panel for exactly that step? We built an experiment to find out.

<Note>
  **The proposals below are our own**, not positions of the MCP Apps working group or the WebMCP community group. Take them as starting points for a chat 😅
</Note>

## Two directions

You probably know [WebMCP](https://github.com/webmachinelearning/webmcp) already - a page registers tools, then an agent in the browser calls them. One detail matters here: **WebMCP defines tools only**. No MCP resources and no prompts.

[MCP Apps](https://modelcontextprotocol.io/extensions/apps/overview) goes in the other direction. An MCP tool points to a `ui://` resource, and the host renders it as a sandboxed **View**: a chart, map, form or whatever the human needs for the next step. Think of it as the **last mile of interaction**: the agent does the heavy lifting, and the View gives the human the right controls for the decision that remains.

So which one should you use? For us, one question makes the choice much simpler:

> Where is the human looking?

<AttentionSplit />

The **split** case is the one we wanted to test. The website remains the main surface, but one step works better as a focused UI inside the agent's side panel.

## Two ways to combine them

There are two bridges between WebMCP and MCP Apps, and they run in opposite directions:

<TwoBridges />

1️⃣ **A View exposes tools to the host.** The MCP Apps draft spec proposes this as *App-provided tools*: the host can operate the View it rendered. The tools are shaped like WebMCP tools, but they travel over the MCP Apps channel, not WebMCP.

2️⃣ **A WebMCP tool returns a View.** When the agent calls the tool, the website gives the human an interface.

We tested the second one 👇

## Demo - Form / Factor

We built **Form / Factor**, a home-gym equipment page. It recommends equipment based on your measurements, the space you have and your preferred balance between strength, variety and convenience.

The whole demo is on GitHub if you want to run it yourself: <RepoLink href="https://github.com/Kulikowski/webmcp-apps">Kulikowski/webmcp-apps</RepoLink> - a local page plus an unpacked Chrome extension. You'll need to enable `chrome://flags/#enable-webmcp-testing`.

<PostImage
  src="/form-factor.png"
  alt="The Form / Factor demo in three panels: WebMCP tools in Chrome DevTools, the live home-gym equipment page and an Equipment fit controls View in the agent side panel."
  width={3780}
  height={1957}
  quality={90}
  caption="The page's WebMCP tools in DevTools (left), the live website (middle) and the fit-controls View the agent opened in the side panel (right)."
/>

The human asks the side-panel agent for help choosing equipment. The agent is scripted, with no LLM - we're testing the bridge, not the model.

The agent calls `gym_open_fit_sidecar`, which returns **text for the model**, the page state as `structuredContent`, and **the View itself** as an embedded `ui://` resource, plus a `_meta.ui.resourceUri` link telling the host which View to open for the human.

```js
document.modelContext.registerTool({
  name: "gym_open_fit_sidecar",
  description: "Open interactive equipment-fit controls for the user. Returns an MCP Apps UI resource.",
  inputSchema: { type: "object", properties: {}, additionalProperties: false },
  annotations: { readOnlyHint: true },
  async execute() {
    return {
      content: [
        { type: "text", text: "Interactive equipment fit controls opened for the user." },
        // The View itself, as a standard MCP embedded resource
        { type: "resource", resource: {
          uri: "ui://form-factor/equipment-fit",
          mimeType: "text/html;profile=mcp-app",
          text: FIT_SIDECAR_HTML,
        } },
      ],
      structuredContent: snapshot(),
      _meta: {
        // MCP Apps puts this on the tool; registerTool() has no _meta
        ui: { resourceUri: "ui://form-factor/equipment-fit" },
        // Not standard: stands in for MCP Apps' visibility: ["app"]
        "webmcp-apps": {
          allowedPageTools: ["gym_update_profile", "gym_set_preferences"],
        },
      },
    };
  },
});
```

The extension renders that View in the side panel through a sandbox proxy - controls for measurements, room size and the training mix. As the human changes them, the View calls two more page tools, `gym_update_profile` and `gym_set_preferences`, and the recommendation updates on the live page 😎

<McpWebLoop />

The agent finds the tool and opens the View, the human makes the decision, and the website remains the **source of truth**.

### What's standard, and what isn't

Inside the side panel, the View follows the [MCP Apps](https://github.com/modelcontextprotocol/ext-apps/blob/main/specification/2026-01-26/apps.mdx) protocol. It's a real `ui://` resource with the MCP Apps MIME type, and the extension renders it through a sandbox proxy, completes the `ui/initialize` handshake, then exchanges JSON-RPC `tools/call` requests and tool-result notifications over `postMessage`.

WebMCP has tools only, so three pieces of MCP Apps have to move:

- **The link moves to the result.** MCP Apps puts `_meta.ui.resourceUri` on the tool definition, but WebMCP's `registerTool()` has no `_meta`, so the host can't know about the View in advance.
- **The View is embedded in the result.** An MCP Apps host fetches it with `resources/read`. WebMCP has no resources, so the same resource contents ride in the result's `content`. MCP Apps deferred embedded resources; MCP-UI uses them.
- **App-only tools become an allowlist.** WebMCP can't mark tools with MCP Apps' `visibility: ["app"]`, so the page names them in `allowedPageTools` under its own `webmcp-apps` key in `_meta`.

Rendering the View at all is still a private convention between our page and our extension: WebMCP has no rule that a host should render UI found in a tool result. A client that doesn't know the convention may pass the whole result to the model as text, HTML included. This experiment shows that the two can *meet*, not that they already interoperate.

That also means our extension has to provide the guardrails on its own 🔑

- Render only a known `ui://` URI, and only when the result embeds that exact resource with the MCP Apps MIME type.
- Talk only to the expected page origin, and render the View through a sandbox proxy - the double-iframe pattern of the MCP Apps reference host - under a deny-by-default CSP. The View gets an opaque origin, so it can't read the site's cookies or storage or reach the extension's privileges.
- Keep an allowlist of page tools the View can call. `gym_update_profile` and `gym_set_preferences` behave like MCP Apps *app-only* tools, but WebMCP cannot mark them that way, so any other agent on the page still sees them.

## Open questions

The experiment shows the pieces *can* fit together. It doesn't settle how they *should*. These are the questions we kept asking ourselves, and where we currently lean.

### 1. Should the tool return the UI, or point to it up front?

In MCP Apps, a tool points to its UI before anyone calls it - that `ui://` pointer from earlier. The host can review and prefetch the UI.

Our `gym_open_fit_sidecar` does the opposite. Nothing tells the agent that UI is coming until the result arrives with the `_meta.ui.resourceUri` link and the View inside. WebMCP's `registerTool()` has no `_meta`, so there's nowhere to put the link up front.

We lean towards **both**: point to the UI when the tool is registered, so the agent can prepare for it, then include it in the result.

### 2. What if nobody can see the UI?

WebMCP also describes a [headless browsing scenario](https://github.com/webmachinelearning/webmcp/blob/main/README.md#:~:text=Headless%20browsing%20scenarios%3A%20Tools%20exposed%20for%20human%2Din%2Dthe%2Dloop%20can%20also%20be%20used%20for%20task%20completion%20in%20headless%20scenarios%2C%20and%20is%20particularly%20useful%20when%20switching%20between%20human%2Din%2Dthe%2Dloop%20and%20headless%20experiences.): the agent works on its own, then pulls in a human when it needs one.

A tool that opens UI sounds especially useful here. But what if nobody is around to see it? 🙃

Two different things matter:

- **Presence:** is a human looking at the page, looking at the agent, or not present at all?
- **Rendering:** can the agent show an MCP App at all, or only text?

Presence decides *whether and where* to show UI. Rendering decides the *format*. If nobody is there, the tool should say "needs a human" instead of opening UI for an empty room.

The agent could simply tell the page what it supports, but that gives sites a new way to fingerprint agents. A safer option might be to return text *plus* optional UI, like email's `multipart/alternative`, and let the agent pick what it can show. That's what Form / Factor does: a text block for any agent, and the View for one that can render it.

### 3. Where should a confirmation happen, and who should confirm it?

Deleting a GitHub repository shows a confirmation dialog. If an agent calls a "delete repo" tool while you watch the page, do you need the same confirmation again in the side panel?

We lean towards the **page UI** - it's the dialog users already know and trust. The side panel is for steps the page has no good UI for, like our training-mix triangle.

Browser agents can also click the page. If the agent can press "Delete" itself, showing the dialog doesn't prove that a human approved the action. The same problem applies to a "confirm" button inside the View.

So the real question isn't only *where* the dialog lives, but *who is allowed to answer it*. We think a tool needs a way to say *"this step is for the human"*: pause, show the UI, and let the browser verify that the answer came from a person, not the agent. WebMCP is already discussing something close to this: [`requestUserInteraction() / requestUserInput()`](https://github.com/webmachinelearning/webmcp/issues/165).

A page could also draw something that imitates the agent's own interface. At minimum, page-originated UI should show its origin, the way browsers label permission prompts. We don't know yet whether a click inside that View should approve the next tool call.

### 4. Can the agent operate the View it just opened?

*"Can you fill in that form?"* MCP Apps lets a View expose tools to the host - the *App-provided tools* from bridge 1️⃣.

They look like WebMCP tools, but they aren't. To the browser agent, the View is a sandboxed iframe, not a page.

Our preference is **one authoring API with two transports**. The View could use an API shaped like `document.modelContext.registerTool`, while the SDK carries those calls over the MCP Apps channel. Developers learn one API, and the protocols stay separate.

### 5. Who should move: WebMCP or MCP Apps?

Our bridge needed three substitutions because the two specs don't meet in the middle. There are two ways to close that gap.

**WebMCP adds resources.** Then the MCP Apps flow works almost unchanged: if tools can also carry `_meta` to link a `ui://` resource, the host reads it before the call, so it can review and prefetch the UI. The cost is a new primitive for a standard that has [stayed tools-only so far](https://github.com/webmachinelearning/webmcp/issues/25) - though there's already an [open proposal for resources](https://github.com/webmachinelearning/webmcp/issues/151).

**MCP Apps accepts embedded resources.** Then WebMCP stays tools-only, and a tool result can carry its View directly. MCP already has the content block for this, and MCP-UI shipped it. The cost is losing up-front review, unless tools can also point to their UI at registration.

The second is the smaller change for both specs. The first keeps MCP Apps' security model intact. We don't think the answer is obvious.

It touches a bigger question: **how MCP-compatible should WebMCP actually be**?

## What we'd propose

Starting points for discussion, not positions of either group:

- **For WebMCP:** let tool results carry UI next to text, let `registerTool()` carry `_meta` so tools can point to their UI up front, support *app-only* tools, and provide a standard way to hand a step to the human and wait.
- **For MCP Apps:** define a **web host profile** for displaying Views supplied by web pages, labelling their origin and routing calls back to the page.
- **For both:** share a model of presence and rendering.

## One loop

Not every website should become an MCP App, and an MCP App shouldn't rebuild a full website inside chat.

But the two can meet for one useful moment: the agent works with the website, the website opens the right controls, the human decides, and the agent carries on 🤝

Today that takes a custom extension. If you have opinions on the questions above, bring them to the specs: [WebMCP issues](https://github.com/webmachinelearning/webmcp/issues) for WebMCP and [ext-apps issues](https://github.com/modelcontextprotocol/ext-apps/issues) for MCP Apps - we'd love to hear them 💪

The demo code is in <RepoLink href="https://github.com/Kulikowski/webmcp-apps">Kulikowski/webmcp-apps</RepoLink>, and issues about the demo itself are welcome there.

More soon. ✨
